ClinicalMC: A Benchmark for Multi-Course Clinical Decision-Making with Large Language Models
The article introduces ClinicalMC, a benchmark designed for evaluating large language models in multi-course clinical…
Recent ai-research headlines from arXiv cs.AI.

The article introduces ClinicalMC, a benchmark designed for evaluating large language models in multi-course clinical…

MedCUA-Bench is a newly introduced benchmark designed specifically for clinical computer-use agents. It aims to…

The study investigates the impact of demographic bias on skin lesion classification using ResNet-based models. It…

The article discusses a new framework called the Pre-Reasoning Perception Framework (PRPF) designed to enhance…

A recent paper argues that superintelligence developed from a solipsistic approach to AI design is unlikely to be…

The study investigates whether real-world datasets contain natural experiments, which are implicit interventions…

The article discusses a novel approach for enhancing Visual Question Answering (VQA) by distilling rules from Large…

The paper presents a negative result regarding cross-model activation transfer in a multi-hop reasoning setting using…

The article discusses LEAP, a new framework designed to enhance the capabilities of Large Language Models (LLMs) in…

The article discusses the challenges of benchmark auditing in artificial intelligence, particularly regarding…

The Violation Situation Pattern (VSP) is a new knowledge-graph pattern designed to improve compliance violation…

The paper introduces InfoMem, a new reward mechanism designed for training long-context memory agents in artificial…

The article introduces CP-Agent, a multimodal large language model designed for cellular morphological profiling under…

The paper investigates the effectiveness of interaction trajectories in training terminal agents. It reveals that…

The Deterministic Memory Framework (DMF) aims to enhance memory systems for conversational AI agents. It replaces…

StepFinder is a new framework designed for failure attribution in multi-agent systems. It aims to improve the…

The paper presents a formal definition and meta-model for the Machine Theory of Mind. It integrates insights from…

The paper introduces ThoughtFold, a framework designed to improve the efficiency of Large Reasoning Models (LRMs) by…

The paper presents a new compositional authorization framework for managing delegation and scope in agentic AI…

The paper introduces SAGE, a framework for evaluating socialized evolution in agent ecosystems. It compares two…

The paper introduces an SLM-based Agent Orchestration Gateway designed for AI-driven virtual worlds. This gateway…

The paper discusses a new approach to optimize coding agents by reducing input-token costs. It introduces a middleware…

The paper introduces a framework to improve instruction following in Large Reasoning Models (LRMs) by addressing the…

The paper introduces TSQAgent, a framework designed to improve the assessment of time series data quality using large…

A recent study investigates gender-dependent disparities in medical triage recommendations made by large language…

The paper discusses advancements in propositional defeasible standpoint logic, focusing on non-monotonic entailment.…

The paper introduces NovelAPIBench, a dynamic benchmark designed to evaluate large language models' ability to use…

The paper introduces ChemCoTBench-V2, a benchmark designed for evaluating chemical reasoning in large language models.…

EvoDrive is a new framework designed for generating safety-critical scenarios in autonomous driving systems. It…

The DeepSpeak-Agentic dataset consists of over 37 hours of semi-structured conversations between humans and AI agents.…

The article introduces SkillPyramid, a framework designed to enhance the skill consolidation of self-evolving AI…

The paper introduces Dynamic Objective Selection with Safeguards (DOSS) for financial decision-making. DOSS aims to…

The paper introduces Code-on-Graph (CoG), a new framework for integrating Large Language Models (LLMs) with Knowledge…

The paper introduces derivation graphs to enhance the understanding of do-calculus reasoning. These graphs help in…

The paper titled 'When to Re-Plan: Subgoal Persistence in Hierarchical Latent Reasoning' explores the balance between…

The paper titled 'Proof-Refactor' addresses the challenges in generating formal proofs using Large Language Models…

The article introduces the Lab Agent Protocol (LAP), designed to enhance the interaction between autonomous agents and…

The article presents BrickAnything, a new framework for generating buildable brick structures from 3D shapes. This…

The paper titled 'Can LLMs Introspect? A Reality Check' questions the ability of large language models (LLMs) to…

The paper discusses the need for persistent memory in long-running AI agents. It critiques current memory systems and…

The paper presents POLAR, a framework designed for personalizing embodied multimodal large language model agents…

The paper discusses the need for improved benchmarks in Constraint Acquisition (CA) research. Current benchmarks are…

The paper discusses the aging of AI agents deployed in operational systems and introduces a new benchmark called…

The paper discusses two innovative frameworks for creating autonomous AI systems to enhance scientific workflows.…

The paper introduces Anchor, a task-generation pipeline designed to address artifact drift in AI agent benchmark…

The paper introduces OmniToM, a benchmark designed to evaluate the Theory of Mind capabilities in large language…

The paper introduces JobBench, a new benchmark for evaluating AI agents based on human needs rather than economic…

The paper discusses a framework for managing uncertainty in procedural knowledge generated by large language models…

The paper titled 'ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence' presents a new…

A new study proposes an automated method for selecting layers in large language models to improve hallucination…
WeSearch's declared handling of arXiv cs.AI's content. Indexing, snippets, summaries, retrieval and training are separate questions — see the rights registry or read this source's machine-readable record.