WeSearch
Hub / ai-research / arXiv cs.AI
ai-research · source

arXiv cs.AI on WeSearch

Recent ai-research headlines from arXiv cs.AI.

NEXUS: Structured Runtime Safety for Tool-Using LLM Agents
Lead story
arXiv.org

NEXUS: Structured Runtime Safety for Tool-Using LLM Agents

Researchers have introduced NEXUS, a structured runtime safety monitor for tool-using LLM agents, which applies a formal intervention policy to ensure safe execution of high-impact actions. NEXUS combines deterministic…

5d · 2 min read · 46 views
The latest
Information Discernment in Large Language Models
arXiv.org

Information Discernment in Large Language Models

Do they weigh information appropriately -- updating more for reliable sources (source discernment) and more when…

5d · 3 min read · 18 views
Benchmarking Confidential GPU Inference on NVIDIA H100 under Intel TDX
arXiv.org

Benchmarking Confidential GPU Inference on NVIDIA H100 under Intel TDX

However, the performance cost of enabling confidential execution for GPU-accelerated large language model serving…

5d · 3 min read · 16 views
FormulaSPIN: Self-Play Fine-Tuning for Natural Language to Spreadsheet Formula Generation
arXiv.org

FormulaSPIN: Self-Play Fine-Tuning for Natural Language to Spreadsheet Formula Generation

Existing approaches rely on static supervised data, which quickly saturates on limited annotations. In this paper, we…

5d · 3 min read · 17 views
OpenEvoShield: Dual Non-Stationary Continual Defense for Open-World Multi-Agent System Attacks
arXiv.org

OpenEvoShield: Dual Non-Stationary Continual Defense for Open-World Multi-Agent System Attacks

Unlike static threats, these attacks are doubly dynamic: adversaries refine injection strategies against deployed…

5d · 3 min read · 37 views
Hybrid LSTM-Graph Neural Framework for Robust Financial Fraud Detection and Adversarial Resilience
arXiv.org

Hybrid LSTM-Graph Neural Framework for Robust Financial Fraud Detection and Adversarial Resilience

This paper proposes FraudShield AI, a hybrid framework that integrates Long Short-Term Memory (LSTM) networks with…

5d · 2 min read · 13 views
FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads
arXiv.org

FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads

Achieving low latency and high throughput under volatile demand requires deep understanding of real-world serving…

5d · 3 min read · 37 views
Spectral-LSH: Sub-Quadratic Prompt Compression via Krylov-Projected Locality-Sensitive Hashing
arXiv.org

Spectral-LSH: Sub-Quadratic Prompt Compression via Krylov-Projected Locality-Sensitive Hashing

We propose Spectral-LSH, a training-free prompt compression method that operates before the prompt enters the language…

5d · 3 min read · 22 views
Rethinking Uncertainty Evaluation in Large Language Models
arXiv.org

Rethinking Uncertainty Evaluation in Large Language Models

What we actually need is for LLM confidence estimates to satisfy the conditions required of coherent probabilistic…

5d · 2 min read · 21 views
Geometry-Guided Constraint Learning for LLM Safety Classification
arXiv.org

Geometry-Guided Constraint Learning for LLM Safety Classification

We show that sparse autoencoder (SAE) feature extraction resolves this: K=2 becomes optimal for 12/14 categories on…

5d · 2 min read · 22 views
Logic-Guided Data Extraction with Answer Set Programming and Large Language Models
arXiv.org

Logic-Guided Data Extraction with Answer Set Programming and Large Language Models

This paper proposes a logic-guided data extraction framework combining LLM-based extraction with Answer Set…

5d · 3 min read · 19 views
Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models
arXiv.org

Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models

Computer Science > Artificial Intelligence arXiv:2607.19364 (cs) [Submitted on 5 Jun 2026] Title:Statistically…

5d · 3 min read · 22 views
AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally
arXiv.org

AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally

Using simplified retrieval tasks and length generalization scenarios, we show -- both empirically and theoretically --…

5d · 3 min read · 19 views
GraphContainer: A Unified Platform for Comparing and Debugging Graph RAG Methods
arXiv.org

GraphContainer: A Unified Platform for Comparing and Debugging Graph RAG Methods

However, existing approaches remain highly fragmented and incompatible. The structural heterogeneity of graph formats…

5d · 3 min read · 17 views
Lifted Representation Hypothesis in Language Models
arXiv.org

Lifted Representation Hypothesis in Language Models

However, it remains unclear how these structures are stored, selected, and revised. To study this process, we propose…

5d · 2 min read · 15 views
Euclean: Automated Geometry Problem Formalization with Unified Verification in Lean
arXiv.org

Euclean: Automated Geometry Problem Formalization with Unified Verification in Lean

This split increases the trusted computing base and hinders unified model development. Existing geometry-in-Lean…

5d · 3 min read · 16 views
Mitigating Scaffolding Collapse in Socratic Tutors via Representation Alignment
arXiv.org

Mitigating Scaffolding Collapse in Socratic Tutors via Representation Alignment

Prior defenses primarily constrain observable responses through prompting, preference optimization, or filtering,…

5d · 3 min read · 15 views
Beyond Tracking or Shortcut: Composition-Bounded Predictive States in Poker Autoregressive Models
arXiv.org

Beyond Tracking or Shortcut: Composition-Bounded Predictive States in Poker Autoregressive Models

This paper studies this ambiguity in a no-range Limit Hold'em autoregressive model trained only on action and value…

5d · 3 min read · 13 views
Profile-Graph Memory for LLM Agents: Implicit Cross-Entity Traversal through Narrative Profiles
arXiv.org

Profile-Graph Memory for LLM Agents: Implicit Cross-Entity Traversal through Narrative Profiles

First, we introduce MemHop, a multi-hop memory benchmark of 1,000 questions at hop depths 1-5 across 10 social-network…

5d · 3 min read · 17 views
LISA: Linear-Indexed Sparse Attention for Efficient Long-Context Reasoning
arXiv.org

LISA: Linear-Indexed Sparse Attention for Efficient Long-Context Reasoning

However, the O(n^2) computational complexity of standard self-attention causes inference costs to grow sharply with…

5d · 3 min read · 17 views
Stochastic Primal-Dual Decoding for Multiobjective Generative Recommender Systems
arXiv.org

Stochastic Primal-Dual Decoding for Multiobjective Generative Recommender Systems

In practice, recommendation often involves constructing slates -- ordered lists of items -- that must satisfy multiple…

5d · 3 min read · 16 views
OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets
arXiv.org

OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets

Existing provenance systems operate at file or dataset level, forcing catastrophic over-deletion. We present ob, a…

7/16/2026 · 2 min read · 34 views
ProofCouncil: An LLM Agent for Solving Open Mathematical Problems
arXiv.org

ProofCouncil: An LLM Agent for Solving Open Mathematical Problems

However, their performance can be further improved through agentic workflows tailored to real-world mathematical…

7/13/2026 · 3 min read · 32 views
Communication-Efficient Digital-Twin Coordination for Heterogeneous LLM Embodied Agents over Computing Power Networks
arXiv.org

Communication-Efficient Digital-Twin Coordination for Heterogeneous LLM Embodied Agents over Computing Power Networks

Computer Science > Artificial Intelligence arXiv:2607.09330 (cs) [Submitted on 10 Jul 2026]…

7/13/2026 · 3 min read · 34 views
How Does Bayesian Causal Discovery Fail? Characterising Structural Consequences in Linear Gaussian Networks under Latent Confounding
arXiv.org

How Does Bayesian Causal Discovery Fail? Characterising Structural Consequences in Linear Gaussian Networks under Latent Confounding

The paper investigates how Bayesian causal discovery behaves when latent confounding is present in linear Gaussian…

7/13/2026 · 3 min read · 33 views
LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making
arXiv.org

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making

Prior evaluations of LLM-based medical agents have largely emphasized short-context knowledge QA and tool use.…

7/13/2026 · 3 min read · 37 views
Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review
arXiv.org

Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review

Two experiments across 20 diverse worldbuilding tasks, using GPT-OSS 120B and DeepSeek v3.2 as LLM backends,…

7/13/2026 · 3 min read · 30 views
OpenProver: Agentic and Interactive Theorem Proving with Lean 4
arXiv.org

OpenProver: Agentic and Interactive Theorem Proving with Lean 4

OpenProver integrates a Planner-Worker-Verifier architecture inspired by recent ATP agentic systems such as Aletheia.…

7/13/2026 · 2 min read · 36 views
Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents
arXiv.org

Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents

Equipped with broad knowledge, flexible reasoning, and tool use, they have the potential to autonomously explore and…

7/13/2026 · 2 min read · 76 views
Scoped Verification for Reliable Long-Horizon Agentic Context Evolution under Distribution Shift
arXiv.org

Scoped Verification for Reliable Long-Horizon Agentic Context Evolution under Distribution Shift

Computer Science > Artificial Intelligence arXiv:2607.09175 (cs) [Submitted on 10 Jul 2026] Title:Scoped Verification…

7/13/2026 · 3 min read · 34 views
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling
arXiv.org

KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling

However, existing PRMs are text-based: they re-encode the entire trajectory text from scratch. In long multi-agent…

7/13/2026 · 3 min read · 31 views
Neuro-Agentic Control: A Deep Learning-based LLM-Powered Agentic AI Framework for Controlling Security Controls
arXiv.org

Neuro-Agentic Control: A Deep Learning-based LLM-Powered Agentic AI Framework for Controlling Security Controls

While Large Language Models (LLMs) have strong semantic reasoning abilities to assist in decision support, their…

7/13/2026 · 3 min read · 35 views
ARCANA: A Reflective Multi-Agent Program Synthesis Framework for ARC-AGI-2 Reasoning
arXiv.org

ARCANA: A Reflective Multi-Agent Program Synthesis Framework for ARC-AGI-2 Reasoning

ARCANA decomposes each task into iterative perception, hypothesis generation, symbolic execution, and reflective…

7/13/2026 · 2 min read · 34 views
L-MAD: A Systematic Evaluation of Multi-Agent Debate Structures in Legal Reasoning
arXiv.org

L-MAD: A Systematic Evaluation of Multi-Agent Debate Structures in Legal Reasoning

The paper presents the Legal Multi-Agent Debate (L-MAD) framework for evaluating debate structures in legal textual…

7/13/2026 · 2 min read · 29 views
MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation
arXiv.org

MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation

Many rely on synthetic conversations or patient simulators, omit patient-uploaded medical images, or evaluate…

7/13/2026 · 3 min read · 30 views
A Formalization of the Mean-Field Derivation of the Vlasov Equation: AI-Assisted Lean Formalization as a Strategy Game
arXiv.org

A Formalization of the Mean-Field Derivation of the Vlasov Equation: AI-Assisted Lean Formalization as a Strategy Game

Computer Science > Artificial Intelligence arXiv:2607.08986 (cs) [Submitted on 9 Jul 2026] Title:A Formalization of…

7/13/2026 · 3 min read · 31 views
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
arXiv.org

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated…

7/13/2026 · 3 min read · 30 views
GATS: Graph-Augmented Tree Search with Layered World Models for Efficient Agent Planning
arXiv.org

GATS: Graph-Augmented Tree Search with Layered World Models for Efficient Agent Planning

We present \textbf{GATS} (Graph-Augmented Tree Search), a planning framework that combines systematic UCB1-based tree…

7/13/2026 · 3 min read · 27 views
CogniConsole: Externalizing Inference-Time Control as a Formal Abstraction for Reliable LLM Interactions
arXiv.org

CogniConsole: Externalizing Inference-Time Control as a Formal Abstraction for Reliable LLM Interactions

We challenge this by demonstrating that reliability is significantly influenced by \emph{inference-time control} --…

7/13/2026 · 3 min read · 34 views
Interval Certifications for Multilayered Perceptrons via Lattice Traversal
arXiv.org

Interval Certifications for Multilayered Perceptrons via Lattice Traversal

In particular, we show that the adversarial robustness problem can be reduced to a lattice traversal problem. Each…

7/13/2026 · 3 min read · 30 views
Ceci n'est pas une pipe: AI systems as semantic abstractions
arXiv cs.AI

Ceci n'est pas une pipe: AI systems as semantic abstractions

We propose a semantic framework to describe AI systems, to be able to examine the correctness of such representations.…

7/13/2026 · 2 min read · 25 views
Multimodal Reward Hacking in Reinforcement Learning
arXiv cs.AI

Multimodal Reward Hacking in Reinforcement Learning

This risk is amplified when visual evidence is evaluated by text-only or weakly grounded rewards. We study reward…

7/13/2026 · 3 min read · 24 views
Shared Selective Persistent Memory for Agentic LLM Systems
arXiv cs.AI

Shared Selective Persistent Memory for Agentic LLM Systems

Naively persisting entire conversation histories is token-inefficient and counterproductive: irrelevant context…

7/13/2026 · 3 min read · 27 views
SAGEAgent: A Self-Evolving Agent for Cost-Aware Modality Acquisition in Multimodal Survival Prediction
arXiv cs.AI

SAGEAgent: A Self-Evolving Agent for Cost-Aware Modality Acquisition in Multimodal Survival Prediction

In multimodal clinical oncology, diagnostic modalities follow a clinically mandated order of escalating burden -- from…

7/13/2026 · 3 min read · 21 views
Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI
arXiv cs.AI

Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI

These are powerful capabilities, but they share a structural limitation: the representational frame within which the…

7/13/2026 · 3 min read · 25 views
Knowledge Graphs and Explainable AI as Complementary Resources for Urban Mining
arXiv cs.AI

Knowledge Graphs and Explainable AI as Complementary Resources for Urban Mining

The relevant unit of value is not prediction accuracy alone, but the defensibility of the supported decisions: their…

7/13/2026 · 3 min read · 23 views
TrustX Agent Risk Classification Framework (ARC): Risk-Tiering Internally Created Agentic AI Systems
arXiv cs.AI

TrustX Agent Risk Classification Framework (ARC): Risk-Tiering Internally Created Agentic AI Systems

Computer Science > Artificial Intelligence arXiv:2607.09586 (cs) [Submitted on 10 Jul 2026] Title:TrustX Agent Risk…

7/13/2026 · 3 min read · 25 views
Agora: Enhancing LLM Agent Reasoning Via Auction-Based Task Allocation
arXiv cs.AI

Agora: Enhancing LLM Agent Reasoning Via Auction-Based Task Allocation

However, existing frameworks typically call APIs based on coarse-grained matching between tasks and the functions of…

7/13/2026 · 2 min read · 22 views
ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI
arXiv cs.AI

ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI

We introduce ConceptSMILE, a model-agnostic perturbation-based auditing framework for evaluating the reliability of…

7/13/2026 · 2 min read · 22 views
Minimal Decision Dynamics and Contextual Probability: A Quantum Tug-of-War Model
arXiv cs.AI

Minimal Decision Dynamics and Contextual Probability: A Quantum Tug-of-War Model

This paper develops a quantum-like extension of the Tug-of-War (QTOW) decision-making model to clarify when such…

7/13/2026 · 2 min read · 25 views

How WeSearch handles this source

WeSearch's declared handling of arXiv cs.AI's content. Indexing, snippets, summaries, retrieval and training are separate questions — see the rights registry or read this source's machine-readable record.

Indexing: Allowed Snippet: Allowed AI summary: Limited Retrieval / RAG: Not asserted Model training: Not asserted Commercial reuse: Not permitted

More ai-research sources

Visit arXiv cs.AI directly →