WeSearch
Hub / ai-research / arXiv cs.AI
ai-research · source

arXiv cs.AI on WeSearch

Recent ai-research headlines from arXiv cs.AI.

GraphRAG on Consumer Hardware: Benchmarking Local LLMs for Healthcare EHR Schema Retrieval
arXiv cs.AI

GraphRAG on Consumer Hardware: Benchmarking Local LLMs for Healthcare EHR Schema Retrieval

The paper evaluates GraphRAG, a graph-based retrieval-augmented generation model, for healthcare EHR schema retrieval…

5/22/2026 · 3 min read · 35 views
ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models
arXiv cs.AI

ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models

The paper introduces ArchSIBench, a benchmark designed to evaluate the architectural spatial intelligence of…

5/22/2026 · 3 min read · 41 views
USV: Towards Understanding the User-generated Short-form Videos
arXiv cs.AI

USV: Towards Understanding the User-generated Short-form Videos

The paper titled 'USV: Towards Understanding the User-generated Short-form Videos' introduces a new dataset aimed at…

5/22/2026 · 2 min read · 29 views
DISC: Decoupling Instruction from State-Conditioned Control via Policy Generation
arXiv cs.AI

DISC: Decoupling Instruction from State-Conditioned Control via Policy Generation

The paper titled 'DISC: Decoupling Instruction from State-Conditioned Control via Policy Generation' introduces a new…

5/22/2026 · 3 min read · 28 views
Position: Let's Develop Data Probes to Fundamentally Understand How Data Affects LLM Performance
arXiv cs.AI

Position: Let's Develop Data Probes to Fundamentally Understand How Data Affects LLM Performance

A new position paper advocates for the development of data probes to better understand how data influences large…

5/20/2026 · 3 min read · 32 views
Operationalizing Document AI: A Microservice Architecture for OCR and LLM Pipelines in Production
arXiv cs.AI

Operationalizing Document AI: A Microservice Architecture for OCR and LLM Pipelines in Production

The article discusses a microservice architecture designed for operationalizing Document AI, focusing on OCR and large…

5/20/2026 · 3 min read · 32 views
Evaluating the Utility of Personal Health Records in Personalized Health AI
arXiv cs.AI

Evaluating the Utility of Personal Health Records in Personalized Health AI

The study evaluates the effectiveness of Personal Health Records (PHRs) in enhancing responses from large language…

5/20/2026 · 3 min read · 36 views
Learn-by-Wire Training Control Governance: Bounded Autonomous Training Under Stress for Stability and Efficiency
arXiv cs.AI

Learn-by-Wire Training Control Governance: Bounded Autonomous Training Under Stress for Stability and Efficiency

The paper introduces Learn-by-Wire Guard (LBW-Guard), a new training-control governance layer for language models.…

5/20/2026 · 3 min read · 33 views
AgentNLQ: A General-Purpose Agent for Natural Language to SQL
arXiv cs.AI

AgentNLQ: A General-Purpose Agent for Natural Language to SQL

The article discusses a new multi-agent method for converting natural language to SQL, known as AgentNLQ. This method…

5/20/2026 · 3 min read · 40 views
KAN-MLP-Mixer: A comprehensive investigation of the usage of Kolmogorov-Arnold Networks (KANs) for improving IMU-based Human Activity Recognition
arXiv cs.AI

KAN-MLP-Mixer: A comprehensive investigation of the usage of Kolmogorov-Arnold Networks (KANs) for improving IMU-based Human Activity Recognition

The article discusses a study on Kolmogorov-Arnold Networks (KANs) and their application in improving human activity…

5/20/2026 · 3 min read · 31 views
Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On
arXiv cs.AI

Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On

The paper discusses the importance of integrating trust into Agent-to-Agent (A2A) networks from the outset rather than…

5/20/2026 · 3 min read · 31 views
Interference-Aware Multi-Task Unlearning
arXiv cs.AI

Interference-Aware Multi-Task Unlearning

The paper introduces a framework for multi-task unlearning in machine learning, addressing the challenges of removing…

5/20/2026 · 2 min read · 34 views
Embedding by Elicitation: Dynamic Representations for Bayesian Optimization of System Prompts
arXiv cs.AI

Embedding by Elicitation: Dynamic Representations for Bayesian Optimization of System Prompts

The paper discusses a new framework called ReElicit for optimizing system prompts in AI using Bayesian methods. It…

5/20/2026 · 3 min read · 41 views
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows
arXiv cs.AI

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows

The paper introduces DecisionBench, a benchmark for emergent delegation in long-horizon agentic workflows. It…

5/20/2026 · 3 min read · 34 views
POLAR-Bench: A Diagnostic Benchmark for Privacy-Utility Trade-offs in LLM Agents
arXiv cs.AI

POLAR-Bench: A Diagnostic Benchmark for Privacy-Utility Trade-offs in LLM Agents

The article introduces POLAR-Bench, a diagnostic benchmark designed to evaluate privacy-utility trade-offs in large…

5/20/2026 · 3 min read · 30 views
Learning to Hand Off: Provably Convergent Workflow Learning under Interface Constraints
arXiv cs.AI

Learning to Hand Off: Provably Convergent Workflow Learning under Interface Constraints

The paper discusses a new approach to workflow learning in multi-agent systems where agents hand off control through a…

5/20/2026 · 3 min read · 33 views
Progressive Autonomy as Preference Learning: A Formalization of Trust Calibration for Agentic Tool Use
arXiv cs.AI

Progressive Autonomy as Preference Learning: A Formalization of Trust Calibration for Agentic Tool Use

The paper discusses a formalization of trust calibration for automated agents in tool use. It presents a method for…

5/20/2026 · 2 min read · 44 views
How Far Are We From True Auto-Research?
arXiv cs.AI

How Far Are We From True Auto-Research?

Recent advancements in auto-research systems have enabled the generation of complete research papers. However, a study…

5/20/2026 · 3 min read · 24 views
Discoverable Agent Knowledge -- A Formal Framework for Agentic KG Affordances (Extended Version)
arXiv cs.AI

Discoverable Agent Knowledge -- A Formal Framework for Agentic KG Affordances (Extended Version)

The article presents a formal framework for understanding agentic knowledge graph (KG) affordances. It critiques…

5/20/2026 · 3 min read · 34 views
Hallucination as Exploit: Evidence-Carrying Multimodal Agents
arXiv cs.AI

Hallucination as Exploit: Evidence-Carrying Multimodal Agents

The paper discusses the concept of hallucination in multimodal agents, where false visual claims can lead to…

5/20/2026 · 3 min read · 41 views
Not all uncertainty is alike: volatility, stochasticity, and exploration
arXiv cs.AI

Not all uncertainty is alike: volatility, stochasticity, and exploration

The paper discusses the differences between volatility and stochasticity in the context of adaptive decision-making in…

5/20/2026 · 2 min read · 35 views
SimGym: A Framework for A/B Test Simulation in E-Commerce with Traffic-Grounded VLM Agents
arXiv cs.AI

SimGym: A Framework for A/B Test Simulation in E-Commerce with Traffic-Grounded VLM Agents

SimGym is a new framework designed to simulate A/B tests in e-commerce using vision-language model agents. It aims to…

5/20/2026 · 3 min read · 31 views
Can Large Language Models Revolutionize Survey Research? Experiments with Disaster Preparedness Responses
arXiv cs.AI

Can Large Language Models Revolutionize Survey Research? Experiments with Disaster Preparedness Responses

The article discusses the potential of large language models (LLMs) to address challenges in survey research,…

5/20/2026 · 3 min read · 36 views
Causal Evidence for Attention Head Imbalance in Modality Conflict Hallucination
arXiv cs.AI

Causal Evidence for Attention Head Imbalance in Modality Conflict Hallucination

The paper presents findings on modality-conflict hallucination in multimodal large language models (MLLMs). It…

5/20/2026 · 3 min read · 43 views
AQuaUI: Visual Token Reduction for GUI Agents with Adaptive Quadtrees
arXiv cs.AI

AQuaUI: Visual Token Reduction for GUI Agents with Adaptive Quadtrees

The paper introduces AQuaUI, a novel method for visual token reduction in GUI agents using adaptive quadtrees. This…

5/20/2026 · 3 min read · 40 views
Swimming with Whales: Analysis of Power Imbalances in Stake-Weighted Governance
arXiv cs.AI

Swimming with Whales: Analysis of Power Imbalances in Stake-Weighted Governance

The paper titled 'Swimming with Whales: Analysis of Power Imbalances in Stake-Weighted Governance' explores the…

5/20/2026 · 2 min read · 32 views
MOCHA: Multi-Objective Chebyshev Annealing for Agent Skill Optimization
arXiv cs.AI

MOCHA: Multi-Objective Chebyshev Annealing for Agent Skill Optimization

The paper introduces MOCHA, a new method for optimizing agent skills in artificial intelligence. MOCHA employs…

5/20/2026 · 3 min read · 25 views
Agentic Trading: When LLM Agents Meet Financial Markets
arXiv cs.AI

Agentic Trading: When LLM Agents Meet Financial Markets

The paper titled 'Agentic Trading: When LLM Agents Meet Financial Markets' explores the integration of Large Language…

5/20/2026 · 3 min read · 30 views
Generative Recursive Reasoning
arXiv cs.AI

Generative Recursive Reasoning

The paper introduces Generative Recursive Reasoning Models (GRAM), a new framework for neural reasoning systems. GRAM…

5/20/2026 · 2 min read · 33 views
PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning
arXiv cs.AI

PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning

The article introduces PRISM, a new benchmark for evaluating programmatic spatial-temporal reasoning in video…

5/20/2026 · 3 min read · 35 views
Conflict-Resilient Multi-Agent Reasoning via Signed Graph Modeling
arXiv cs.AI

Conflict-Resilient Multi-Agent Reasoning via Signed Graph Modeling

The paper introduces SIGMA, a new framework for multi-agent reasoning that addresses the limitations of existing…

5/20/2026 · 3 min read · 28 views
What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents
arXiv cs.AI

What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents

The paper discusses a new framework called SERL for improving reinforcement learning in multi-turn agents. It focuses…

5/20/2026 · 2 min read · 30 views
Generative Auto-Bidding with Unified Modeling and Exploration
arXiv cs.AI

Generative Auto-Bidding with Unified Modeling and Exploration

The paper presents GUIDE, a new framework for automated bidding in digital advertising. It integrates directed…

5/20/2026 · 3 min read · 29 views
Beyond Mode Collapse: Distribution Matching for Diverse Reasoning
arXiv cs.AI

Beyond Mode Collapse: Distribution Matching for Diverse Reasoning

The paper discusses a new approach to on-policy reinforcement learning that addresses the issue of mode collapse. The…

5/20/2026 · 3 min read · 25 views
Attention-Guided Reward for Reinforcement Learning-based Jailbreak against Large Reasoning Models
arXiv cs.AI

Attention-Guided Reward for Reinforcement Learning-based Jailbreak against Large Reasoning Models

The paper discusses a novel approach to jailbreak attacks on Large Reasoning Models (LRMs) using reinforcement…

5/20/2026 · 3 min read · 28 views
Position: The Turing-Completeness of Real-World Autoregressive Transformers Relies Heavily on Context Management
arXiv cs.AI

Position: The Turing-Completeness of Real-World Autoregressive Transformers Relies Heavily on Context Management

The paper discusses the Turing-completeness of autoregressive Transformers, emphasizing the importance of context…

5/20/2026 · 3 min read · 33 views
BLINKG: A Benchmark for LLM-Integrated Knowledge Graph Generation
arXiv cs.AI

BLINKG: A Benchmark for LLM-Integrated Knowledge Graph Generation

The paper introduces BLINKG, a benchmark for evaluating the capabilities of Large Language Models (LLMs) in generating…

5/20/2026 · 3 min read · 32 views
Efficient Elicitation of Collective Disagreements
arXiv cs.AI

Efficient Elicitation of Collective Disagreements

The paper titled 'Efficient Elicitation of Collective Disagreements' explores how to analyze voter disagreements over…

5/20/2026 · 3 min read · 30 views
Generative-Evaluative Agreement: A Necessary Validity Criterion for LLM-Enabled Adaptive Assessment
arXiv cs.AI

Generative-Evaluative Agreement: A Necessary Validity Criterion for LLM-Enabled Adaptive Assessment

The paper introduces Generative-Evaluative Agreement (GEA) as a validity criterion for assessing LLM-enabled adaptive…

5/20/2026 · 2 min read · 30 views
Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries
arXiv cs.AI

Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries

The paper discusses a phenomenon called 'library drift' in self-evolving LLM skill libraries, which leads to…

5/20/2026 · 3 min read · 33 views
SceneCode: Executable World Programs for Editable Indoor Scenes with Articulated Objects
arXiv cs.AI

SceneCode: Executable World Programs for Editable Indoor Scenes with Articulated Objects

The paper introduces SceneCode, a framework designed for generating editable indoor scenes with articulated objects.…

5/20/2026 · 3 min read · 33 views
Towards Multi-Model LLM Schedulers: Empirical Insights into Offloading and Preemption
arXiv cs.AI

Towards Multi-Model LLM Schedulers: Empirical Insights into Offloading and Preemption

The paper discusses the challenges of scheduling multiple Large Language Models (LLMs) on shared hardware. It…

5/20/2026 · 3 min read · 24 views
Formal Skill: Programmable Runtime Skills for Efficient and Accurate LLM Agents
arXiv cs.AI

Formal Skill: Programmable Runtime Skills for Efficient and Accurate LLM Agents

The paper introduces Formal Skill, a new abstraction for enhancing the efficiency and accuracy of Large Language Model…

5/20/2026 · 3 min read · 25 views
EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection
arXiv cs.AI

EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection

The paper presents EMO-BOOST, a new framework aimed at improving deepfake detection by integrating emotion…

5/20/2026 · 2 min read · 27 views
When Tabular Foundation Models Meet Strategic Tabular Data: A Prior Alignment Approach
arXiv cs.AI

When Tabular Foundation Models Meet Strategic Tabular Data: A Prior Alignment Approach

The paper discusses the challenges faced by tabular foundation models in strategic data environments. It introduces a…

5/20/2026 · 3 min read · 36 views
Pseudocode-Guided Structured Reasoning for Automating Reliable Inference in Vision-Language Models
arXiv cs.AI

Pseudocode-Guided Structured Reasoning for Automating Reliable Inference in Vision-Language Models

The paper presents a new framework called Pseudocode-guided Structured Reasoning (PStar) aimed at improving the…

5/20/2026 · 3 min read · 25 views
Transforming Constraint Programs to Input for Local Search
arXiv cs.AI

Transforming Constraint Programs to Input for Local Search

The paper discusses the transformation of constraint programs into input for local search algorithms. It highlights…

5/20/2026 · 2 min read · 35 views
Beyond Rational Illusion: Behaviorally Realistic Strategic Classification
arXiv cs.AI

Beyond Rational Illusion: Behaviorally Realistic Strategic Classification

The paper titled 'Beyond Rational Illusion: Behaviorally Realistic Strategic Classification' introduces a new…

5/20/2026 · 3 min read · 36 views
Projecting Latent RL Actions: Towards Generalizable and Scalable Graph Combinatorial Optimization
arXiv cs.AI

Projecting Latent RL Actions: Towards Generalizable and Scalable Graph Combinatorial Optimization

The paper introduces a novel approach to graph combinatorial optimization using reinforcement learning. It addresses…

5/20/2026 · 3 min read · 25 views
EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design
arXiv cs.AI

EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design

EngiAI introduces a multi-agent framework and benchmark suite designed for LLM-driven engineering design tasks. The…

5/20/2026 · 3 min read · 28 views

How WeSearch handles this source

WeSearch's declared handling of arXiv cs.AI's content. Indexing, snippets, summaries, retrieval and training are separate questions — see the rights registry or read this source's machine-readable record.

Indexing: Allowed Snippet: Allowed AI summary: Limited Retrieval / RAG: Not asserted Model training: Not asserted Commercial reuse: Not permitted

More ai-research sources

Visit arXiv cs.AI directly →