Stop Comparing LLM Agents Without Disclosing the Harness
The paper titled 'Stop Comparing LLM Agents Without Disclosing the Harness' argues that the performance of language…
Recent ai-research headlines from arXiv cs.AI.

The paper titled 'Stop Comparing LLM Agents Without Disclosing the Harness' argues that the performance of language…

The paper presents methods for formal verification of agent skills, addressing a gap in the verification process. It…

The paper titled 'Machine Psychometrics' explores a new approach to understanding artificial intelligence through…

The paper titled 'From Accuracy to Auditability: A Survey of Determinism in Financial AI Systems' explores the…

The article introduces QUIVER, a formal framework designed to quantify perturbation propagation and bifurcation in…

The article discusses a new approach to job shop scheduling using rollout-calibrated hyper-heuristics. This method…

The paper introduces LGMT, a new framework for evaluating the reasoning reliability of large language models (LLMs).…

The article discusses the limitations of large language models (LLMs) in tasks requiring causal reasoning and…

The article discusses growth dynamics in deterministic equational discovery substrates across three toy domains. It…

A new model for autonomous robot learning has been proposed, focusing on a thinking-learning interaction approach.…

A new survey examines the trustworthiness of agentic AI systems, focusing on safety, robustness, privacy, and system…

The paper presents a new framework called Reason--Imagine--Act (RIA) for enhancing decision-making in autonomous…

The paper introduces LC-ERD, a framework designed to enhance self-evolving reasoning in Large Language Models. It…

EvoSci is a proposed multi-agent framework designed to enhance scientific discovery through bio-inspired evolution and…

A new framework utilizing Neutrosophic Logic has been proposed to address epistemic uncertainty in Large Language…

EvoCode-Bench is a newly introduced benchmark designed to evaluate coding agents in multi-turn iterative interactions.…

The paper introduces SkillEvolBench, a benchmark designed to evaluate the transition from episodic experience to…

The paper introduces MAPLE, a new method for evaluating policies in imperfect-information games using a tree search…

The paper titled 'HyperGuide' presents a novel approach to enhance multi-step reasoning in large language models. It…

The article presents a neuro-inspired framework for planning and control in artificial intelligence. It introduces…

The article discusses a new framework called Palette designed for safety alignment in large language models (LLMs).…

The paper discusses the role of context sparsity in large language model (LLM) efficiency. It argues that the…

The paper presents EPPC-OASIS, a framework designed for mining electronic patient-provider communication in secure…

The paper discusses agentic misalignment in multi-agent systems, particularly in automated workflows. It defines this…

The paper explores the impact of multi-agent reinforcement learning (RL) on large language model (LLM) workflows. It…

The paper discusses the challenges of measuring performance in Large Language Models (LLMs) as they move into…

The paper discusses the limitations of current hallucination benchmarks for Large Language Models (LLMs) in…

The paper examines how well AI models adhere to their specified behavioral guidelines. It introduces a multi-method…

The paper titled 'Toward Enactive Artificial Intelligence' advocates for integrating enactive approaches to perception…

The paper analyzes the routing behavior of the Mixtral 8x7B-Instruct model under different prompt conditions. It finds…

The study investigates the effectiveness of LLM-generated synthetic data in low-resource multi-label patent…

The paper presents a new framework for adaptive human-AI coordination called Intrinsic Action Disentanglement (IAD).…

The article discusses a new framework called Partner-Aware Skill Discovery (PASD) designed to enhance human-AI…

The paper discusses the generation of Game Code World Models (GameCWMs) using Large Language Models (LLMs). It…

The paper discusses ethical-use constraints in open-weight AI models and their implications for governance policy. It…

The paper discusses the issue of premature confidence in language models, which leads to flawed reasoning. It…

The article introduces ConceptM$^3$oE, a new framework for computational pathology that integrates multimodal…

The article discusses advancements in graph few-shot learning through a novel model called VISION. This model…

The paper introduces Psych LM, an iOS application designed for psychological coaching using a local-first…

The article introduces JT-Safe-V2, a large language model aimed at enhancing the safety and trustworthiness of…

The paper explores the limitations of In-Context Reinforcement Learning (ICRL) in the context of Ad-Hoc Teamwork…

The article introduces State-Adaptive Memory (SAM), a framework designed for long-horizon reasoning in artificial…

The paper titled 'SPACE: Unifying Symmetric and Asymmetric Routing Problems for Generalist Neural Solver' presents a…

The paper introduces AgentFugue, a framework designed for scaling agent capabilities in long-horizon tasks through…

The article presents TIGER, a framework designed for enzyme-reaction retrieval in computational biology. It addresses…

The article presents a new cooperative multi-agent decision system called Market Regime Council (MRC) for portfolio…

The paper discusses the vulnerabilities of Large Reasoning Models (LRMs) to jailbreak attacks due to their…

The study explores hypothesis generation and inductive inference in children and language models. It compares how both…

The paper introduces DemoEvolve, a method for enhancing agent harness evolution using demonstrations. This approach…

A new study proposes an emission-aware reinforcement learning strategy for electric vehicle charging. This approach…
WeSearch's declared handling of arXiv cs.AI's content. Indexing, snippets, summaries, retrieval and training are separate questions — see the rights registry or read this source's machine-readable record.