Understanding Rollout Error in Graph World Models
Many planning environments, however, are not vectors or images; they are graphs of agents, tools, skills, routes, and…
Recent ai-research headlines from arXiv cs.AI.

Many planning environments, however, are not vectors or images; they are graphs of agents, tools, skills, routes, and…

Most current applications focus on static coding benchmarks. We extend this paradigm to algorithmic trading. This…

The paper discusses the concept of accelerating returns, which suggests that technological progress becomes…

Researchers have developed COrigami, an AI pipeline for co-designing flat-foldable visually recognizable origami. The…

The paper introduces OpenFinGym, a unified gym environment designed for evaluating quantitative‑finance agents across…

Computer Science > Artificial Intelligence arXiv:2606.26348 (cs) [Submitted on 24 Jun 2026] Title:What We are Missing…

Classical exact solvers suffer from combinatorial explosion for these types of problems, and standard reinforcement…

We introduce narration-of-thought (NoT), a system prompt that structures chain-of-thought into five sections:…

Electric bus fleets provide a relevant test case. Their operation requires continuous coordination between service…

Computer Science > Artificial Intelligence arXiv:2606.26346 (cs) [Submitted on 24 Jun 2026] Title:How Do…

For today's coding agents, this intuition is being inverted: as foundation models develop stronger reasoning…

We formalize this as compositional behavioral leakage (CBL): interference between modules sharing a context window.…

This paper observes that human institutions have governed powerful autonomous actors not by monitoring their reasoning…

Siegel, Arvind Narayanan View a PDF of the paper titled Life After Benchmark Saturation: A Case Study of CORE-Bench,…

Integrating them without conflating evidence and anecdote is especially consequential in psychiatry, where poorly…

These data pairs determine the degree to which interpretability frameworks can reliably detect model features…

However, they inherently suffer from response lag due to their exclusive reliance on match outcomes, neglecting the…

We introduce an LLM-powered comparative pipeline for large-scale governance discourse analysis, integrating automated…

Researchers have found that refusal and persona traits in chat models are interconnected, with a compliant persona…

The paper discusses the use of visual graph scaffolds to enhance structural reasoning in large language models (LLMs).…

The paper presents AURA-Mem, a novel memory architecture designed for robotic policies that operates with constant…

This study evaluates the effectiveness of Transformer and LSTM frameworks for predicting streamflow in ungauged…

The paper introduces BehaviorBench, a benchmark designed to evaluate personalized decision modeling using real-world…

ChatHealthAI is a proposed multimodal reasoning framework that aligns electronic health record (EHR) representations…

Traj-Evolve is a self-evolving multi-agent system designed for modeling patient trajectories in lung cancer early…

The paper explores novel methods for generating enemy morphologies in video games using player collision information.…

The paper evaluates the phenomenon of harmful overthinking in Large Reasoning Models (LRMs). It introduces a new…

The paper discusses a modular architecture for embedded AI agent systems designed to operate within the constraints of…

The paper titled 'Don't Gamble, GAMBLe' introduces a framework for analyzing AI-Driven Research Systems (ADRS). It…

The paper explores the impact of multi-agent debate on data cleaning processes. It finds that while debate can lead to…

The paper discusses the concept of 'handoff debt' in coding tasks where agents take over interrupted work. It…

A recent study explores the potential of large AI models in dental healthcare, highlighting the need for a unified…

The paper discusses the limitations of current benchmarks for evaluating autonomous agents, particularly their failure…

The paper presents WISE-HAR, an ensemble deep learning framework for recognizing human activities using WiFi signals.…

The paper introduces a method called Reasoning Primitive Induction, which aims to enhance the performance of…

The article discusses AuditFlow, a new framework designed for structured financial reporting verification. It utilizes…

TriEval is a new pipeline designed to assess bias, toxicity, and truthfulness in large language models (LLMs)…

The paper introduces RelGT-AC, a new model designed for autocomplete tasks in relational databases. It enhances the…

The paper introduces ToolGate, a system designed to improve the efficiency of tool-augmented vision-language agents.…

The paper introduces SkillDAG, a novel approach for selecting skills in large language models (LLMs) by modeling…

The article discusses a new framework called CORE, which stands for Conflict-Oriented Reasoning, designed to enhance…

The paper introduces DeltaMem, a framework designed to enhance memory management in Large Language Model (LLM) agents.…

The article discusses a new approach to budget allocation for Large Language Models (LLMs) based on economic…

The paper titled 'Decomposing how prompting steers behavior' explores how prompting influences the internal…

A new framework has been developed to enhance time series forecasting by integrating news articles. This approach…

The paper titled 'DeskCraft' introduces a new benchmark for evaluating desktop agents in professional workflows that…

EvoTrainer is a new autonomous training framework designed for co-evolving LLM policies and training harnesses. It…

The article discusses a new framework for Large Language Model (LLM) agents that aims to improve their performance in…

The paper introduces a new framework called Think-Before-Speak (TBS) for multi-agent social simulation. TBS separates…

The article introduces GTBench, a benchmark designed to evaluate large language models (LLMs) as mathematical research…
WeSearch's declared handling of arXiv cs.AI's content. Indexing, snippets, summaries, retrieval and training are separate questions — see the rights registry or read this source's machine-readable record.