WeSearch
Hub / ai-research / arXiv cs.AI
ai-research · source

arXiv cs.AI on WeSearch

Recent ai-research headlines from arXiv cs.AI.

Understanding Rollout Error in Graph World Models
arXiv.org

Understanding Rollout Error in Graph World Models

Many planning environments, however, are not vectors or images; they are graphs of agents, tools, skills, routes, and…

6/29/2026 · 3 min read · 35 views
AlgoEvolve: LLM-driven Meta-evolution of Algorithmic Trading Programs
arXiv.org

AlgoEvolve: LLM-driven Meta-evolution of Algorithmic Trading Programs

Most current applications focus on static coding benchmarks. We extend this paradigm to algorithmic trading. This…

6/26/2026 · 2 min read · 33 views
Accelerating Returns and the Qualitative Engine for Science
arXiv.org

Accelerating Returns and the Qualitative Engine for Science

The paper discusses the concept of accelerating returns, which suggests that technological progress becomes…

6/26/2026 · 3 min read · 38 views
COrigami: An AI Pipeline for Co-Designing Flat-Foldable Visually Recognisable Origami
arXiv.org

COrigami: An AI Pipeline for Co-Designing Flat-Foldable Visually Recognisable Origami

Researchers have developed COrigami, an AI pipeline for co-designing flat-foldable visually recognizable origami. The…

6/26/2026 · 3 min read · 30 views
OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents
arXiv.org

OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents

The paper introduces OpenFinGym, a unified gym environment designed for evaluating quantitative‑finance agents across…

6/26/2026 · 3 min read · 32 views
What We are Missing in Multimodal LLM Evaluation?
arXiv.org

What We are Missing in Multimodal LLM Evaluation?

Computer Science > Artificial Intelligence arXiv:2606.26348 (cs) [Submitted on 24 Jun 2026] Title:What We are Missing…

6/26/2026 · 2 min read · 36 views
Geometry-Aware MCTS for Extremal Problems in Combinatorial Geometry
arXiv.org

Geometry-Aware MCTS for Extremal Problems in Combinatorial Geometry

Classical exact solvers suffer from combinatorial explosion for these types of problems, and standard reinforcement…

6/26/2026 · 3 min read · 30 views
Narration-of-Thought: Inference-Time Scaffolding for Defeasible Ethical Reasoning in Large Language Models
arXiv.org

Narration-of-Thought: Inference-Time Scaffolding for Defeasible Ethical Reasoning in Large Language Models

We introduce narration-of-thought (NoT), a system prompt that structures chain-of-thought into five sections:…

6/26/2026 · 3 min read · 34 views
When Agents Meet Electric Bus Fleet Operations: Pricing Behavior, Trade-offs, and Policy Implications in an Aggregator Framework
arXiv.org

When Agents Meet Electric Bus Fleet Operations: Pricing Behavior, Trade-offs, and Policy Implications in an Aggregator Framework

Electric bus fleets provide a relevant test case. Their operation requires continuous coordination between service…

6/26/2026 · 3 min read · 50 views
How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks?
arXiv.org

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks?

Computer Science > Artificial Intelligence arXiv:2606.26346 (cs) [Submitted on 24 Jun 2026] Title:How Do…

6/26/2026 · 3 min read · 32 views
The Verification Horizon: No Silver Bullet for Coding Agent Rewards
arXiv.org

The Verification Horizon: No Silver Bullet for Coding Agent Rewards

For today's coding agents, this intuition is being inverted: as foundation models develop stronger reasoning…

6/26/2026 · 3 min read · 35 views
Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems
arXiv.org

Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems

We formalize this as compositional behavioral leakage (CBL): interference between modules sharing a context window.…

6/26/2026 · 3 min read · 38 views
Governing Actions, Not Agents: Institutional Attestation as a Governance Model for Autonomous AI Systems
arXiv.org

Governing Actions, Not Agents: Institutional Attestation as a Governance Model for Autonomous AI Systems

This paper observes that human institutions have governed powerful autonomous actors not by monitoring their reasoning…

6/26/2026 · 2 min read · 34 views
Life After Benchmark Saturation: A Case Study of CORE-Bench
arXiv.org

Life After Benchmark Saturation: A Case Study of CORE-Bench

Siegel, Arvind Narayanan View a PDF of the paper titled Life After Benchmark Saturation: A Case Study of CORE-Bench,…

6/26/2026 · 3 min read · 30 views
Knowledge-augmented Agentic AI for Mental Health Medication Information Seeking
arXiv.org

Knowledge-augmented Agentic AI for Mental Health Medication Information Seeking

Integrating them without conflating evidence and anecdote is especially consequential in psychiatry, where poorly…

6/26/2026 · 3 min read · 36 views
Detecting and Controlling Sycophancy with Cascading Linear Features
arXiv.org

Detecting and Controlling Sycophancy with Cascading Linear Features

These data pairs determine the degree to which interpretability frameworks can reliably detect model features…

6/26/2026 · 3 min read · 40 views
Accelerating Skill Assessment in Chess: A Drift-Diffusion-Enhanced Elo Rating System
arXiv.org

Accelerating Skill Assessment in Chess: A Drift-Diffusion-Enhanced Elo Rating System

However, they inherently suffer from response lag due to their exclusive reliance on match outcomes, neglecting the…

6/26/2026 · 3 min read · 31 views
Agentic Analysis for Agentic Infrastructure: An LLM-Powered Pipeline for Comparative Governance of DAO and Corporate AI Protocols
arXiv.org

Agentic Analysis for Agentic Infrastructure: An LLM-Powered Pipeline for Comparative Governance of DAO and Corporate AI Protocols

We introduce an LLM-powered comparative pipeline for large-scale governance discourse analysis, integrating automated…

6/26/2026 · 3 min read · 42 views
Refusal Lives Downstream of Persona in Chat Models
arXiv.org

Refusal Lives Downstream of Persona in Chat Models

Researchers have found that refusal and persona traits in chat models are interconnected, with a compliant persona…

6/26/2026 · 2 min read · 34 views
Visual Graph Scaffolds for Structural Reasoning in Large Language Models
arXiv cs.AI

Visual Graph Scaffolds for Structural Reasoning in Large Language Models

The paper discusses the use of visual graph scaffolds to enhance structural reasoning in large language models (LLMs).…

6/3/2026 · 3 min read · 44 views
AURA: Action-Gated Memory for Robot Policies at Constant VRAM
arXiv cs.AI

AURA: Action-Gated Memory for Robot Policies at Constant VRAM

The paper presents AURA-Mem, a novel memory architecture designed for robotic policies that operates with constant…

6/3/2026 · 3 min read · 43 views
Evaluating Transformer and LSTM Frameworks for Prediction in Ungauged Basins
arXiv cs.AI

Evaluating Transformer and LSTM Frameworks for Prediction in Ungauged Basins

This study evaluates the effectiveness of Transformer and LSTM frameworks for predicting streamflow in ungauged…

6/3/2026 · 2 min read · 47 views
BehaviorBench: Modeling Real-World User Decisions from Behavioral Traces
arXiv cs.AI

BehaviorBench: Modeling Real-World User Decisions from Behavioral Traces

The paper introduces BehaviorBench, a benchmark designed to evaluate personalized decision modeling using real-world…

6/3/2026 · 3 min read · 45 views
ChatHealthAI: Aligning Electronic Health Record Representations with Large Language Models for Grounded Clinical Reasoning
arXiv cs.AI

ChatHealthAI: Aligning Electronic Health Record Representations with Large Language Models for Grounded Clinical Reasoning

ChatHealthAI is a proposed multimodal reasoning framework that aligns electronic health record (EHR) representations…

6/3/2026 · 2 min read · 39 views
Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection
arXiv cs.AI

Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection

Traj-Evolve is a self-evolving multi-agent system designed for modeling patient trajectories in lung cancer early…

6/3/2026 · 3 min read · 48 views
An Exploration of Collision-based Enemy Morphology Generation
arXiv cs.AI

An Exploration of Collision-based Enemy Morphology Generation

The paper explores novel methods for generating enemy morphologies in video games using player collision information.…

6/3/2026 · 2 min read · 44 views
Thinking Past the Answer: Evaluating Harmful Overthinking in Large Reasoning Models
arXiv cs.AI

Thinking Past the Answer: Evaluating Harmful Overthinking in Large Reasoning Models

The paper evaluates the phenomenon of harmful overthinking in Large Reasoning Models (LRMs). It introduces a new…

6/3/2026 · 3 min read · 44 views
Toward a Modular Architecture for Embedded AI Agent Systems at the Edge
arXiv cs.AI

Toward a Modular Architecture for Embedded AI Agent Systems at the Edge

The paper discusses a modular architecture for embedded AI agent systems designed to operate within the constraints of…

6/3/2026 · 2 min read · 50 views
Don't Gamble, GAMBLe: An Analytical Framework for AI-Driven Research Systems
arXiv cs.AI

Don't Gamble, GAMBLe: An Analytical Framework for AI-Driven Research Systems

The paper titled 'Don't Gamble, GAMBLe' introduces a framework for analyzing AI-Driven Research Systems (ADRS). It…

6/3/2026 · 3 min read · 43 views
When Helping Hurts and How to Fix It: Multi-Agent Debate for Data Cleaning
arXiv cs.AI

When Helping Hurts and How to Fix It: Multi-Agent Debate for Data Cleaning

The paper explores the impact of multi-agent debate on data cleaning processes. It finds that while debate can lead to…

6/3/2026 · 3 min read · 43 views
Handoff Debt: The Rediscovery Cost When Coding Agents Take Over Interrupted Tasks
arXiv cs.AI

Handoff Debt: The Rediscovery Cost When Coding Agents Take Over Interrupted Tasks

The paper discusses the concept of 'handoff debt' in coding tasks where agents take over interrupted work. It…

6/3/2026 · 2 min read · 44 views
Large AI Models in Dental Healthcare: From General-Purpose Systems to Domain-Specific Foundation Models
arXiv cs.AI

Large AI Models in Dental Healthcare: From General-Purpose Systems to Domain-Specific Foundation Models

A recent study explores the potential of large AI models in dental healthcare, highlighting the need for a unified…

6/3/2026 · 3 min read · 42 views
What Benchmarks Don't Measure: The Case for Evaluating Abstention Competence in Autonomous Agents
arXiv cs.AI

What Benchmarks Don't Measure: The Case for Evaluating Abstention Competence in Autonomous Agents

The paper discusses the limitations of current benchmarks for evaluating autonomous agents, particularly their failure…

6/3/2026 · 3 min read · 42 views
WISE-HAR: A Generalizable Ensemble Deep Learning Framework for WiFi-Based Human Activity Recognition
arXiv cs.AI

WISE-HAR: A Generalizable Ensemble Deep Learning Framework for WiFi-Based Human Activity Recognition

The paper presents WISE-HAR, an ensemble deep learning framework for recognizing human activities using WiFi signals.…

6/3/2026 · 3 min read · 52 views
Inducing Reasoning Primitives from Agent Traces
arXiv cs.AI

Inducing Reasoning Primitives from Agent Traces

The paper introduces a method called Reasoning Primitive Induction, which aims to enhance the performance of…

6/3/2026 · 2 min read · 41 views
AUDITFLOW: Executable Symbolic Environments for Structured Financial Reporting Verification
arXiv cs.AI

AUDITFLOW: Executable Symbolic Environments for Structured Financial Reporting Verification

The article discusses AuditFlow, a new framework designed for structured financial reporting verification. It utilizes…

6/3/2026 · 3 min read · 41 views
TriEval: A Resource-Efficient Pipeline for LLM Bias, Toxicity, and Truthfulness Assessment
arXiv cs.AI

TriEval: A Resource-Efficient Pipeline for LLM Bias, Toxicity, and Truthfulness Assessment

TriEval is a new pipeline designed to assess bias, toxicity, and truthfulness in large language models (LLMs)…

6/3/2026 · 3 min read · 50 views
RelGT-AC: A Relational Graph Transformer for Autocomplete Tasks in Relational Databases
arXiv cs.AI

RelGT-AC: A Relational Graph Transformer for Autocomplete Tasks in Relational Databases

The paper introduces RelGT-AC, a new model designed for autocomplete tasks in relational databases. It enhances the…

6/3/2026 · 3 min read · 46 views
ToolGate: Token-Efficient Pre-Call Control for Tool-Augmented Vision-Language Agents
arXiv cs.AI

ToolGate: Token-Efficient Pre-Call Control for Tool-Augmented Vision-Language Agents

The paper introduces ToolGate, a system designed to improve the efficiency of tool-augmented vision-language agents.…

6/3/2026 · 3 min read · 39 views
SkillDAG: Self-Evolving Typed Skill Graphs for LLM Skill Selection at Scale
arXiv cs.AI

SkillDAG: Self-Evolving Typed Skill Graphs for LLM Skill Selection at Scale

The paper introduces SkillDAG, a novel approach for selecting skills in large language models (LLMs) by modeling…

6/3/2026 · 3 min read · 40 views
CORE: Conflict-Oriented Reasoning for General Multimodal Manipulation Detection
arXiv cs.AI

CORE: Conflict-Oriented Reasoning for General Multimodal Manipulation Detection

The article discusses a new framework called CORE, which stands for Conflict-Oriented Reasoning, designed to enhance…

6/3/2026 · 3 min read · 49 views
DELTAMEM: Incremental Experience Memory for LLM Agents via Residual Trees
arXiv cs.AI

DELTAMEM: Incremental Experience Memory for LLM Agents via Residual Trees

The paper introduces DeltaMem, a framework designed to enhance memory management in Large Language Model (LLM) agents.…

6/3/2026 · 3 min read · 47 views
The Shadow Price of Reasoning: Economic Perspective on Optimal Budget Allocation for LLMs
arXiv cs.AI

The Shadow Price of Reasoning: Economic Perspective on Optimal Budget Allocation for LLMs

The article discusses a new approach to budget allocation for Large Language Models (LLMs) based on economic…

6/3/2026 · 2 min read · 42 views
Decomposing how prompting steers behavior
arXiv cs.AI

Decomposing how prompting steers behavior

The paper titled 'Decomposing how prompting steers behavior' explores how prompting influences the internal…

6/3/2026 · 3 min read · 50 views
From Long News to Accurate Forecast: Importance-Aware Fusion and PRM-Guided Reflection for Time Series Forecasting
arXiv cs.AI

From Long News to Accurate Forecast: Importance-Aware Fusion and PRM-Guided Reflection for Time Series Forecasting

A new framework has been developed to enhance time series forecasting by integrating news articles. This approach…

6/3/2026 · 3 min read · 50 views
DeskCraft: Benchmarking Desktop Agents on Professional Workflows and Human-in-the-Loop Collaboration
arXiv cs.AI

DeskCraft: Benchmarking Desktop Agents on Professional Workflows and Human-in-the-Loop Collaboration

The paper titled 'DeskCraft' introduces a new benchmark for evaluating desktop agents in professional workflows that…

6/3/2026 · 3 min read · 38 views
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning
arXiv cs.AI

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning

EvoTrainer is a new autonomous training framework designed for co-evolving LLM policies and training harnesses. It…

6/3/2026 · 2 min read · 45 views
Uncertainty-Aware Clarification in LLM Agents with Information Gain
arXiv cs.AI

Uncertainty-Aware Clarification in LLM Agents with Information Gain

The article discusses a new framework for Large Language Model (LLM) agents that aims to improve their performance in…

6/3/2026 · 2 min read · 43 views
Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation
arXiv cs.AI

Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation

The paper introduces a new framework called Think-Before-Speak (TBS) for multi-agent social simulation. TBS separates…

6/3/2026 · 3 min read · 44 views
GTBench: A Curriculum-Grounded Benchmark for Evaluating LLMs as Mathematical Research Assistants in Graph Theory
arXiv cs.AI

GTBench: A Curriculum-Grounded Benchmark for Evaluating LLMs as Mathematical Research Assistants in Graph Theory

The article introduces GTBench, a benchmark designed to evaluate large language models (LLMs) as mathematical research…

6/3/2026 · 3 min read · 37 views

How WeSearch handles this source

WeSearch's declared handling of arXiv cs.AI's content. Indexing, snippets, summaries, retrieval and training are separate questions — see the rights registry or read this source's machine-readable record.

Indexing: Allowed Snippet: Allowed AI summary: Limited Retrieval / RAG: Not asserted Model training: Not asserted Commercial reuse: Not permitted

More ai-research sources

Visit arXiv cs.AI directly →