WeSearch
Hub / Ai Research
ai-research · WeSearch

Ai Research news.

The latest AI and machine-learning research — new papers, model architectures, transformers, reinforcement learning, benchmarks, and lab announcements.

NEXUS: Structured Runtime Safety for Tool-Using LLM Agents
Lead story
arXiv.org

NEXUS: Structured Runtime Safety for Tool-Using LLM Agents

Researchers have introduced NEXUS, a structured runtime safety monitor for tool-using LLM agents, which applies a formal intervention policy to ensure safe execution of high-impact actions. NEXUS combines deterministic…

4d · 2 min read · 45 views
Latest in Ai Research
Information Discernment in Large Language Models
arXiv.org

Information Discernment in Large Language Models

Do they weigh information appropriately -- updating more for reliable sources (source discernment) and more when…

4d · 3 min read · 17 views
Benchmarking Confidential GPU Inference on NVIDIA H100 under Intel TDX
arXiv.org

Benchmarking Confidential GPU Inference on NVIDIA H100 under Intel TDX

However, the performance cost of enabling confidential execution for GPU-accelerated large language model serving…

4d · 3 min read · 15 views
FormulaSPIN: Self-Play Fine-Tuning for Natural Language to Spreadsheet Formula Generation
arXiv.org

FormulaSPIN: Self-Play Fine-Tuning for Natural Language to Spreadsheet Formula Generation

Existing approaches rely on static supervised data, which quickly saturates on limited annotations. In this paper, we…

4d · 3 min read · 17 views
OpenEvoShield: Dual Non-Stationary Continual Defense for Open-World Multi-Agent System Attacks
arXiv.org

OpenEvoShield: Dual Non-Stationary Continual Defense for Open-World Multi-Agent System Attacks

Unlike static threats, these attacks are doubly dynamic: adversaries refine injection strategies against deployed…

4d · 3 min read · 37 views
Hybrid LSTM-Graph Neural Framework for Robust Financial Fraud Detection and Adversarial Resilience
arXiv.org

Hybrid LSTM-Graph Neural Framework for Robust Financial Fraud Detection and Adversarial Resilience

This paper proposes FraudShield AI, a hybrid framework that integrates Long Short-Term Memory (LSTM) networks with…

4d · 2 min read · 13 views
FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads
arXiv.org

FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads

Achieving low latency and high throughput under volatile demand requires deep understanding of real-world serving…

4d · 3 min read · 37 views
Google News

How AI Is Helping States Cut Through Decades of Red Tape - Stanford HAI

How AI Is Helping States Cut Through Decades of Red Tape Stanford HAI

4d · 17 views
Spectral-LSH: Sub-Quadratic Prompt Compression via Krylov-Projected Locality-Sensitive Hashing
arXiv.org

Spectral-LSH: Sub-Quadratic Prompt Compression via Krylov-Projected Locality-Sensitive Hashing

We propose Spectral-LSH, a training-free prompt compression method that operates before the prompt enters the language…

4d · 3 min read · 21 views
Rethinking Uncertainty Evaluation in Large Language Models
arXiv.org

Rethinking Uncertainty Evaluation in Large Language Models

What we actually need is for LLM confidence estimates to satisfy the conditions required of coherent probabilistic…

4d · 2 min read · 19 views
Geometry-Guided Constraint Learning for LLM Safety Classification
arXiv.org

Geometry-Guided Constraint Learning for LLM Safety Classification

We show that sparse autoencoder (SAE) feature extraction resolves this: K=2 becomes optimal for 12/14 categories on…

4d · 2 min read · 21 views
Logic-Guided Data Extraction with Answer Set Programming and Large Language Models
arXiv.org

Logic-Guided Data Extraction with Answer Set Programming and Large Language Models

This paper proposes a logic-guided data extraction framework combining LLM-based extraction with Answer Set…

4d · 3 min read · 19 views
Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models
arXiv.org

Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models

Computer Science > Artificial Intelligence arXiv:2607.19364 (cs) [Submitted on 5 Jun 2026] Title:Statistically…

4d · 3 min read · 19 views
AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally
arXiv.org

AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally

Using simplified retrieval tasks and length generalization scenarios, we show -- both empirically and theoretically --…

4d · 3 min read · 18 views
GraphContainer: A Unified Platform for Comparing and Debugging Graph RAG Methods
arXiv.org

GraphContainer: A Unified Platform for Comparing and Debugging Graph RAG Methods

However, existing approaches remain highly fragmented and incompatible. The structural heterogeneity of graph formats…

4d · 3 min read · 17 views
Lifted Representation Hypothesis in Language Models
arXiv.org

Lifted Representation Hypothesis in Language Models

However, it remains unclear how these structures are stored, selected, and revised. To study this process, we propose…

4d · 2 min read · 14 views
Euclean: Automated Geometry Problem Formalization with Unified Verification in Lean
arXiv.org

Euclean: Automated Geometry Problem Formalization with Unified Verification in Lean

This split increases the trusted computing base and hinders unified model development. Existing geometry-in-Lean…

4d · 3 min read · 16 views
Mitigating Scaffolding Collapse in Socratic Tutors via Representation Alignment
arXiv.org

Mitigating Scaffolding Collapse in Socratic Tutors via Representation Alignment

Prior defenses primarily constrain observable responses through prompting, preference optimization, or filtering,…

4d · 3 min read · 14 views
Beyond Tracking or Shortcut: Composition-Bounded Predictive States in Poker Autoregressive Models
arXiv.org

Beyond Tracking or Shortcut: Composition-Bounded Predictive States in Poker Autoregressive Models

This paper studies this ambiguity in a no-range Limit Hold'em autoregressive model trained only on action and value…

4d · 3 min read · 12 views
Profile-Graph Memory for LLM Agents: Implicit Cross-Entity Traversal through Narrative Profiles
arXiv.org

Profile-Graph Memory for LLM Agents: Implicit Cross-Entity Traversal through Narrative Profiles

First, we introduce MemHop, a multi-hop memory benchmark of 1,000 questions at hop depths 1-5 across 10 social-network…

4d · 3 min read · 16 views
LISA: Linear-Indexed Sparse Attention for Efficient Long-Context Reasoning
arXiv.org

LISA: Linear-Indexed Sparse Attention for Efficient Long-Context Reasoning

However, the O(n^2) computational complexity of standard self-attention causes inference costs to grow sharply with…

4d · 3 min read · 16 views
Stochastic Primal-Dual Decoding for Multiobjective Generative Recommender Systems
arXiv.org

Stochastic Primal-Dual Decoding for Multiobjective Generative Recommender Systems

In practice, recommendation often involves constructing slates -- ordered lists of items -- that must satisfy multiple…

4d · 3 min read · 14 views
OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets
arXiv.org

OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets

Existing provenance systems operate at file or dataset level, forcing catastrophic over-deletion. We present ob, a…

7/16/2026 · 2 min read · 33 views
ProofCouncil: An LLM Agent for Solving Open Mathematical Problems
arXiv.org

ProofCouncil: An LLM Agent for Solving Open Mathematical Problems

However, their performance can be further improved through agentic workflows tailored to real-world mathematical…

7/13/2026 · 3 min read · 32 views
Communication-Efficient Digital-Twin Coordination for Heterogeneous LLM Embodied Agents over Computing Power Networks
arXiv.org

Communication-Efficient Digital-Twin Coordination for Heterogeneous LLM Embodied Agents over Computing Power Networks

Computer Science > Artificial Intelligence arXiv:2607.09330 (cs) [Submitted on 10 Jul 2026]…

7/13/2026 · 3 min read · 33 views
How Does Bayesian Causal Discovery Fail? Characterising Structural Consequences in Linear Gaussian Networks under Latent Confounding
arXiv.org

How Does Bayesian Causal Discovery Fail? Characterising Structural Consequences in Linear Gaussian Networks under Latent Confounding

The paper investigates how Bayesian causal discovery behaves when latent confounding is present in linear Gaussian…

7/13/2026 · 3 min read · 32 views
LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making
arXiv.org

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making

Prior evaluations of LLM-based medical agents have largely emphasized short-context knowledge QA and tool use.…

7/13/2026 · 3 min read · 37 views
Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review
arXiv.org

Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review

Two experiments across 20 diverse worldbuilding tasks, using GPT-OSS 120B and DeepSeek v3.2 as LLM backends,…

7/13/2026 · 3 min read · 30 views
OpenProver: Agentic and Interactive Theorem Proving with Lean 4
arXiv.org

OpenProver: Agentic and Interactive Theorem Proving with Lean 4

OpenProver integrates a Planner-Worker-Verifier architecture inspired by recent ATP agentic systems such as Aletheia.…

7/13/2026 · 2 min read · 35 views
Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents
arXiv.org

Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents

Equipped with broad knowledge, flexible reasoning, and tool use, they have the potential to autonomously explore and…

7/13/2026 · 2 min read · 76 views
Scoped Verification for Reliable Long-Horizon Agentic Context Evolution under Distribution Shift
arXiv.org

Scoped Verification for Reliable Long-Horizon Agentic Context Evolution under Distribution Shift

Computer Science > Artificial Intelligence arXiv:2607.09175 (cs) [Submitted on 10 Jul 2026] Title:Scoped Verification…

7/13/2026 · 3 min read · 34 views
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling
arXiv.org

KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling

However, existing PRMs are text-based: they re-encode the entire trajectory text from scratch. In long multi-agent…

7/13/2026 · 3 min read · 31 views
Neuro-Agentic Control: A Deep Learning-based LLM-Powered Agentic AI Framework for Controlling Security Controls
arXiv.org

Neuro-Agentic Control: A Deep Learning-based LLM-Powered Agentic AI Framework for Controlling Security Controls

While Large Language Models (LLMs) have strong semantic reasoning abilities to assist in decision support, their…

7/13/2026 · 3 min read · 34 views
ARCANA: A Reflective Multi-Agent Program Synthesis Framework for ARC-AGI-2 Reasoning
arXiv.org

ARCANA: A Reflective Multi-Agent Program Synthesis Framework for ARC-AGI-2 Reasoning

ARCANA decomposes each task into iterative perception, hypothesis generation, symbolic execution, and reflective…

7/13/2026 · 2 min read · 33 views
L-MAD: A Systematic Evaluation of Multi-Agent Debate Structures in Legal Reasoning
arXiv.org

L-MAD: A Systematic Evaluation of Multi-Agent Debate Structures in Legal Reasoning

The paper presents the Legal Multi-Agent Debate (L-MAD) framework for evaluating debate structures in legal textual…

7/13/2026 · 2 min read · 29 views
MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation
arXiv.org

MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation

Many rely on synthetic conversations or patient simulators, omit patient-uploaded medical images, or evaluate…

7/13/2026 · 3 min read · 29 views
A Formalization of the Mean-Field Derivation of the Vlasov Equation: AI-Assisted Lean Formalization as a Strategy Game
arXiv.org

A Formalization of the Mean-Field Derivation of the Vlasov Equation: AI-Assisted Lean Formalization as a Strategy Game

Computer Science > Artificial Intelligence arXiv:2607.08986 (cs) [Submitted on 9 Jul 2026] Title:A Formalization of…

7/13/2026 · 3 min read · 31 views
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
arXiv.org

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated…

7/13/2026 · 3 min read · 29 views
GATS: Graph-Augmented Tree Search with Layered World Models for Efficient Agent Planning
arXiv.org

GATS: Graph-Augmented Tree Search with Layered World Models for Efficient Agent Planning

We present \textbf{GATS} (Graph-Augmented Tree Search), a planning framework that combines systematic UCB1-based tree…

7/13/2026 · 3 min read · 26 views
CogniConsole: Externalizing Inference-Time Control as a Formal Abstraction for Reliable LLM Interactions
arXiv.org

CogniConsole: Externalizing Inference-Time Control as a Formal Abstraction for Reliable LLM Interactions

We challenge this by demonstrating that reliability is significantly influenced by \emph{inference-time control} --…

7/13/2026 · 3 min read · 33 views
Interval Certifications for Multilayered Perceptrons via Lattice Traversal
arXiv.org

Interval Certifications for Multilayered Perceptrons via Lattice Traversal

In particular, we show that the adversarial robustness problem can be reduced to a lattice traversal problem. Each…

7/13/2026 · 3 min read · 28 views
Augmenting Fundamental Analysis with Large Language Models: A RAG-Based System for Generating Investor Briefs
arXiv cs.AI

Augmenting Fundamental Analysis with Large Language Models: A RAG-Based System for Generating Investor Briefs

Securities and Exchange Commission (SEC) which can be found in EDGAR. We were preprocessing those data and than…

7/13/2026 · 2 min read · 29 views
Event Stream based Multi-Modal Video Anomaly Detection: A Benchmark Dataset and Algorithms
arXiv cs.AI

Event Stream based Multi-Modal Video Anomaly Detection: A Benchmark Dataset and Algorithms

To address these limitations, we propose EVAD, an event enhanced VAD framework that jointly exploits conventional…

7/13/2026 · 3 min read · 27 views
Integrating Large Language Models and Graph Convolutional Networks for Semi-Supervised Image Classification
arXiv cs.AI

Integrating Large Language Models and Graph Convolutional Networks for Semi-Supervised Image Classification

Therefore, semi-supervised approaches such as Graph Convolutional Networks (GCNs), which learn from both labeled and…

7/13/2026 · 3 min read · 28 views
Beyond Metadata: CAPRA for Hidden Subgroup Analysis under Missing Metadata in Medical Imaging
arXiv cs.AI

Beyond Metadata: CAPRA for Hidden Subgroup Analysis under Missing Metadata in Medical Imaging

Once those metadata disappear, clinically critical failure modes can be masked by strong aggregate performance, and…

7/13/2026 · 3 min read · 27 views
A Coreset Selection Framework with Ensemble Aggregation for Image Classification
arXiv cs.AI

A Coreset Selection Framework with Ensemble Aggregation for Image Classification

Selecting representative training subsets, however, remains challenging: individual sample contributions are unclear,…

7/13/2026 · 3 min read · 29 views
PRecG: Legal Precedent Retrieval with Graph Neural Networks and Rhetorical Role Segmentation
arXiv cs.AI

PRecG: Legal Precedent Retrieval with Graph Neural Networks and Rhetorical Role Segmentation

Current approaches for automatic precedent retrieval map legal documents to a low-dimensional semantic space and…

7/13/2026 · 3 min read · 28 views
OmniMapBench: Benchmarking Visual-Centric Reasoning on Diverse Map Documents
arXiv cs.AI

OmniMapBench: Benchmarking Visual-Centric Reasoning on Diverse Map Documents

A critical limitation is identified in many document understanding benchmarks: visual content is often reducible to…

7/13/2026 · 3 min read · 27 views
Inside the Skill Market: From Software Engineering Activities to Reusable Agent Skills
arXiv cs.AI

Inside the Skill Market: From Software Engineering Activities to Reusable Agent Skills

SE) has continuously evolved through increasingly powerful forms of reuse, from source code and libraries to…

7/13/2026 · 3 min read · 23 views
On Locality and Length Generalization in Visual Reasoning
arXiv cs.AI

On Locality and Length Generalization in Visual Reasoning

This makes human vision distinctly different from most popular computer vision models in use today, which input images…

7/13/2026 · 3 min read · 22 views

Sources in Ai Research

Other categories