WeSearch
Hub / ai-research / arXiv cs.AI
ai-research · source

arXiv cs.AI on WeSearch

Recent ai-research headlines from arXiv cs.AI.

Stop Comparing LLM Agents Without Disclosing the Harness
arXiv cs.AI

Stop Comparing LLM Agents Without Disclosing the Harness

The paper titled 'Stop Comparing LLM Agents Without Disclosing the Harness' argues that the performance of language…

5/26/2026 · 3 min read · 44 views
Methods for Formal Verification of Agent Skills: Three Layers Toward a Mechanically Checkable Capability-Containment Proof
arXiv cs.AI

Methods for Formal Verification of Agent Skills: Three Layers Toward a Mechanically Checkable Capability-Containment Proof

The paper presents methods for formal verification of agent skills, addressing a gap in the verification process. It…

5/26/2026 · 3 min read · 38 views
Machine Psychometrics: A Mathematical Psychology of Artificial Intelligence
arXiv cs.AI

Machine Psychometrics: A Mathematical Psychology of Artificial Intelligence

The paper titled 'Machine Psychometrics' explores a new approach to understanding artificial intelligence through…

5/26/2026 · 3 min read · 38 views
From Accuracy to Auditability: A Survey of Determinism in Financial AI Systems
arXiv cs.AI

From Accuracy to Auditability: A Survey of Determinism in Financial AI Systems

The paper titled 'From Accuracy to Auditability: A Survey of Determinism in Financial AI Systems' explores the…

5/26/2026 · 3 min read · 34 views
QUIVER: A Formal Framework for Quantifying Perturbation Propagation and Bifurcation in Compound AI Systems
arXiv cs.AI

QUIVER: A Formal Framework for Quantifying Perturbation Propagation and Bifurcation in Compound AI Systems

The article introduces QUIVER, a formal framework designed to quantify perturbation propagation and bifurcation in…

5/26/2026 · 3 min read · 44 views
Low-Cost Labels, Reliable Choices: Rollout-Calibrated Hyper-Heuristics for Job Shop Scheduling
arXiv cs.AI

Low-Cost Labels, Reliable Choices: Rollout-Calibrated Hyper-Heuristics for Job Shop Scheduling

The article discusses a new approach to job shop scheduling using rollout-calibrated hyper-heuristics. This method…

5/26/2026 · 2 min read · 37 views
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs
arXiv cs.AI

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs

The paper introduces LGMT, a new framework for evaluating the reasoning reliability of large language models (LLMs).…

5/26/2026 · 2 min read · 39 views
Why We Need World Models for AGI: Where LLMs Fail and How World Models May Outperform
arXiv cs.AI

Why We Need World Models for AGI: Where LLMs Fail and How World Models May Outperform

The article discusses the limitations of large language models (LLMs) in tasks requiring causal reasoning and…

5/26/2026 · 3 min read · 31 views
Saturating Scaling Laws for Equational Discovery: A Phenomenology of Growth Dynamics in Three Toy Substrates with Two Real-World Replications
arXiv cs.AI

Saturating Scaling Laws for Equational Discovery: A Phenomenology of Growth Dynamics in Three Toy Substrates with Two Real-World Replications

The article discusses growth dynamics in deterministic equational discovery substrates across three toy domains. It…

5/26/2026 · 3 min read · 36 views
Beyond Predefined Learning Objects: A Thinking-Learning Interaction Model for Up-to-Date Autonomous Robot Learning
arXiv cs.AI

Beyond Predefined Learning Objects: A Thinking-Learning Interaction Model for Up-to-Date Autonomous Robot Learning

A new model for autonomous robot learning has been proposed, focusing on a thinking-learning interaction approach.…

5/26/2026 · 3 min read · 33 views
Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security
arXiv cs.AI

Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security

A new survey examines the trustworthiness of agentic AI systems, focusing on safety, robustness, privacy, and system…

5/26/2026 · 3 min read · 42 views
Reason--Imagine--Act: Closed-Loop LLM Decision Making with World Models for Autonomous Driving
arXiv cs.AI

Reason--Imagine--Act: Closed-Loop LLM Decision Making with World Models for Autonomous Driving

The paper presents a new framework called Reason--Imagine--Act (RIA) for enhancing decision-making in autonomous…

5/26/2026 · 3 min read · 39 views
LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition
arXiv cs.AI

LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition

The paper introduces LC-ERD, a framework designed to enhance self-evolving reasoning in Large Language Models. It…

5/26/2026 · 3 min read · 42 views
EvoSci: A Bio-Inspired Multi-Agent Framework for the Evolution of Scientific Discovery
arXiv cs.AI

EvoSci: A Bio-Inspired Multi-Agent Framework for the Evolution of Scientific Discovery

EvoSci is a proposed multi-agent framework designed to enhance scientific discovery through bio-inspired evolution and…

5/26/2026 · 2 min read · 35 views
Breaking the Chains of Probability: Neutrosophic Logic as a New Framework for Epistemic Uncertainty in Large Language Models
arXiv cs.AI

Breaking the Chains of Probability: Neutrosophic Logic as a New Framework for Epistemic Uncertainty in Large Language Models

A new framework utilizing Neutrosophic Logic has been proposed to address epistemic uncertainty in Large Language…

5/26/2026 · 3 min read · 41 views
EvoCode-Bench: Evaluating Coding Agents in Multi-Turn Iterative Interactions
arXiv cs.AI

EvoCode-Bench: Evaluating Coding Agents in Multi-Turn Iterative Interactions

EvoCode-Bench is a newly introduced benchmark designed to evaluate coding agents in multi-turn iterative interactions.…

5/26/2026 · 3 min read · 36 views
SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills
arXiv cs.AI

SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills

The paper introduces SkillEvolBench, a benchmark designed to evaluate the transition from episodic experience to…

5/26/2026 · 3 min read · 32 views
MAPLE: Multi-State Aggregated Policy Evaluation for AlphaZero in Imperfect-Information Games
arXiv cs.AI

MAPLE: Multi-State Aggregated Policy Evaluation for AlphaZero in Imperfect-Information Games

The paper introduces MAPLE, a new method for evaluating policies in imperfect-information games using a tree search…

5/26/2026 · 3 min read · 37 views
HyperGuide: Hyperbolic Guidance for Efficient Multi-Step Reasoning in Large Language Models
arXiv cs.AI

HyperGuide: Hyperbolic Guidance for Efficient Multi-Step Reasoning in Large Language Models

The paper titled 'HyperGuide' presents a novel approach to enhance multi-step reasoning in large language models. It…

5/26/2026 · 2 min read · 30 views
Neuro-Inspired Inverse Learning for Planning and Control
arXiv cs.AI

Neuro-Inspired Inverse Learning for Planning and Control

The article presents a neuro-inspired framework for planning and control in artificial intelligence. It introduces…

5/26/2026 · 3 min read · 35 views
Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs
arXiv cs.AI

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs

The article discusses a new framework called Palette designed for safety alignment in large language models (LLMs).…

5/26/2026 · 3 min read · 33 views
Inference Time Context Sparsity: Illusion or Opportunity?
arXiv cs.AI

Inference Time Context Sparsity: Illusion or Opportunity?

The paper discusses the role of context sparsity in large language model (LLM) efficiency. It argues that the…

5/26/2026 · 3 min read · 30 views
EPPC-OASIS: Ontology-Aware Adaptation and Structured Inference Refinement for Electronic Patient-Provider Communication Mining in Secure Messages
arXiv cs.AI

EPPC-OASIS: Ontology-Aware Adaptation and Structured Inference Refinement for Electronic Patient-Provider Communication Mining in Secure Messages

The paper presents EPPC-OASIS, a framework designed for mining electronic patient-provider communication in secure…

5/26/2026 · 3 min read · 29 views
A Sober Look at Agentic Misalignment in Automated Workflows
arXiv cs.AI

A Sober Look at Agentic Misalignment in Automated Workflows

The paper discusses agentic misalignment in multi-agent systems, particularly in automated workflows. It defines this…

5/26/2026 · 3 min read · 34 views
When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs
arXiv cs.AI

When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs

The paper explores the impact of multi-agent reinforcement learning (RL) on large language model (LLM) workflows. It…

5/26/2026 · 3 min read · 29 views
Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks
arXiv cs.AI

Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks

The paper discusses the challenges of measuring performance in Large Language Models (LLMs) as they move into…

5/26/2026 · 3 min read · 44 views
Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows
arXiv cs.AI

Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows

The paper discusses the limitations of current hallucination benchmarks for Large Language Models (LLMs) in…

5/26/2026 · 2 min read · 38 views
How Well Do Models Follow Their Constitutions?
arXiv cs.AI

How Well Do Models Follow Their Constitutions?

The paper examines how well AI models adhere to their specified behavioral guidelines. It introduces a multi-method…

5/26/2026 · 3 min read · 32 views
Toward Enactive Artificial Intelligence
arXiv cs.AI

Toward Enactive Artificial Intelligence

The paper titled 'Toward Enactive Artificial Intelligence' advocates for integrating enactive approaches to perception…

5/26/2026 · 3 min read · 36 views
Safety-Oriented Routing Analysis of Mixtral MoE Under Benign and Harmful Prompts
arXiv cs.AI

Safety-Oriented Routing Analysis of Mixtral MoE Under Benign and Harmful Prompts

The paper analyzes the routing behavior of the Mixtral 8x7B-Instruct model under different prompt conditions. It finds…

5/26/2026 · 3 min read · 30 views
When Does Synthetic Patent Data Help? Volume-Fidelity Trade-offs in Low-Resource Multi-Label Classification
arXiv cs.AI

When Does Synthetic Patent Data Help? Volume-Fidelity Trade-offs in Low-Resource Multi-Label Classification

The study investigates the effectiveness of LLM-generated synthetic data in low-resource multi-label patent…

5/26/2026 · 3 min read · 28 views
Adaptive Human-AI Coordination via Hierarchical Action Disentanglement
arXiv cs.AI

Adaptive Human-AI Coordination via Hierarchical Action Disentanglement

The paper presents a new framework for adaptive human-AI coordination called Intrinsic Action Disentanglement (IAD).…

5/26/2026 · 2 min read · 28 views
Partner-Aware Hierarchical Skill Discovery for Robust Human-AI Collaboration
arXiv cs.AI

Partner-Aware Hierarchical Skill Discovery for Robust Human-AI Collaboration

The article discusses a new framework called Partner-Aware Skill Discovery (PASD) designed to enhance human-AI…

5/26/2026 · 3 min read · 35 views
Distilling Game Code World Model Generation into Lightweight Large Language Models
arXiv cs.AI

Distilling Game Code World Model Generation into Lightweight Large Language Models

The paper discusses the generation of Game Code World Models (GameCWMs) using Large Language Models (LLMs). It…

5/26/2026 · 3 min read · 34 views
A governance horizon for ethical-use constraints in open-weight AI models
arXiv cs.AI

A governance horizon for ethical-use constraints in open-weight AI models

The paper discusses ethical-use constraints in open-weight AI models and their implications for governance policy. It…

5/26/2026 · 3 min read · 30 views
Understanding and Mitigating Premature Confidence for Better LLM Reasoning
arXiv cs.AI

Understanding and Mitigating Premature Confidence for Better LLM Reasoning

The paper discusses the issue of premature confidence in language models, which leads to flawed reasoning. It…

5/26/2026 · 3 min read · 34 views
ConceptM$^3$oE: Concept-Guided Multimodal Mixture of Experts for Interpretable Computational Pathology
arXiv cs.AI

ConceptM$^3$oE: Concept-Guided Multimodal Mixture of Experts for Interpretable Computational Pathology

The article introduces ConceptM$^3$oE, a new framework for computational pathology that integrates multimodal…

5/26/2026 · 3 min read · 34 views
Advancing Graph Few-Shot Learning via In-Context Learning
arXiv cs.AI

Advancing Graph Few-Shot Learning via In-Context Learning

The article discusses advancements in graph few-shot learning through a novel model called VISION. This model…

5/26/2026 · 3 min read · 26 views
The Model Is Not the Product: A Dual-Pillar Architecture for Local-First Psychological Coaching
arXiv cs.AI

The Model Is Not the Product: A Dual-Pillar Architecture for Local-First Psychological Coaching

The paper introduces Psych LM, an iOS application designed for psychological coaching using a local-first…

5/26/2026 · 3 min read · 24 views
JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data
arXiv cs.AI

JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data

The article introduces JT-Safe-V2, a large language model aimed at enhancing the safety and trustworthiness of…

5/26/2026 · 2 min read · 20 views
Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork
arXiv cs.AI

Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork

The paper explores the limitations of In-Context Reinforcement Learning (ICRL) in the context of Ad-Hoc Teamwork…

5/26/2026 · 3 min read · 34 views
SAM: State-Adaptive Memory for Long-Horizon Reasoning Agent
arXiv cs.AI

SAM: State-Adaptive Memory for Long-Horizon Reasoning Agent

The article introduces State-Adaptive Memory (SAM), a framework designed for long-horizon reasoning in artificial…

5/26/2026 · 3 min read · 35 views
SPACE: Unifying Symmetric and Asymmetric Routing Problems for Generalist Neural Solver
arXiv cs.AI

SPACE: Unifying Symmetric and Asymmetric Routing Problems for Generalist Neural Solver

The paper titled 'SPACE: Unifying Symmetric and Asymmetric Routing Problems for Generalist Neural Solver' presents a…

5/26/2026 · 3 min read · 27 views
AgentFugue: Agent Scaling for Long-Horizon Tasks through Collective Reasoning
arXiv cs.AI

AgentFugue: Agent Scaling for Long-Horizon Tasks through Collective Reasoning

The paper introduces AgentFugue, a framework designed for scaling agent capabilities in long-horizon tasks through…

5/26/2026 · 3 min read · 33 views
TIGER: Text-Informed Generalized Enzyme-Reaction Retrieval
arXiv cs.AI

TIGER: Text-Informed Generalized Enzyme-Reaction Retrieval

The article presents TIGER, a framework designed for enzyme-reaction retrieval in computational biology. It addresses…

5/26/2026 · 2 min read · 34 views
Market Regime Council for Dynamic Credit Assignment in Multi-Agent LLM Decision Systems
arXiv cs.AI

Market Regime Council for Dynamic Credit Assignment in Multi-Agent LLM Decision Systems

The article presents a new cooperative multi-agent decision system called Market Regime Council (MRC) for portfolio…

5/26/2026 · 3 min read · 36 views
Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs
arXiv cs.AI

Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs

The paper discusses the vulnerabilities of Large Reasoning Models (LRMs) to jailbreak attacks due to their…

5/26/2026 · 3 min read · 33 views
Hypothesis Generation and Inductive Inference in Children and Language Models
arXiv cs.AI

Hypothesis Generation and Inductive Inference in Children and Language Models

The study explores hypothesis generation and inductive inference in children and language models. It compares how both…

5/26/2026 · 3 min read · 32 views
DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations
arXiv cs.AI

DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations

The paper introduces DemoEvolve, a method for enhancing agent harness evolution using demonstrations. This approach…

5/26/2026 · 3 min read · 30 views
Emission-Aware Reinforcement Learning for Sustainable Electric Vehicle Charging and Carbon Dioxide Reduction Under Varying Renewable Penetration
arXiv cs.AI

Emission-Aware Reinforcement Learning for Sustainable Electric Vehicle Charging and Carbon Dioxide Reduction Under Varying Renewable Penetration

A new study proposes an emission-aware reinforcement learning strategy for electric vehicle charging. This approach…

5/26/2026 · 3 min read · 44 views

How WeSearch handles this source

WeSearch's declared handling of arXiv cs.AI's content. Indexing, snippets, summaries, retrieval and training are separate questions — see the rights registry or read this source's machine-readable record.

Indexing: Allowed Snippet: Allowed AI summary: Limited Retrieval / RAG: Not asserted Model training: Not asserted Commercial reuse: Not permitted

More ai-research sources

Visit arXiv cs.AI directly →