Information Discernment in Large Language Models
Do they weigh information appropriately -- updating more for reliable sources (source discernment) and more when…
The latest AI and machine-learning research — new papers, model architectures, transformers, reinforcement learning, benchmarks, and lab announcements.

Researchers have introduced NEXUS, a structured runtime safety monitor for tool-using LLM agents, which applies a formal intervention policy to ensure safe execution of high-impact actions. NEXUS combines deterministic…

Do they weigh information appropriately -- updating more for reliable sources (source discernment) and more when…

However, the performance cost of enabling confidential execution for GPU-accelerated large language model serving…

Existing approaches rely on static supervised data, which quickly saturates on limited annotations. In this paper, we…

Unlike static threats, these attacks are doubly dynamic: adversaries refine injection strategies against deployed…

This paper proposes FraudShield AI, a hybrid framework that integrates Long Short-Term Memory (LSTM) networks with…

Achieving low latency and high throughput under volatile demand requires deep understanding of real-world serving…
How AI Is Helping States Cut Through Decades of Red Tape Stanford HAI

We propose Spectral-LSH, a training-free prompt compression method that operates before the prompt enters the language…

What we actually need is for LLM confidence estimates to satisfy the conditions required of coherent probabilistic…

We show that sparse autoencoder (SAE) feature extraction resolves this: K=2 becomes optimal for 12/14 categories on…

This paper proposes a logic-guided data extraction framework combining LLM-based extraction with Answer Set…

Computer Science > Artificial Intelligence arXiv:2607.19364 (cs) [Submitted on 5 Jun 2026] Title:Statistically…

Using simplified retrieval tasks and length generalization scenarios, we show -- both empirically and theoretically --…

However, existing approaches remain highly fragmented and incompatible. The structural heterogeneity of graph formats…

However, it remains unclear how these structures are stored, selected, and revised. To study this process, we propose…

This split increases the trusted computing base and hinders unified model development. Existing geometry-in-Lean…

Prior defenses primarily constrain observable responses through prompting, preference optimization, or filtering,…

This paper studies this ambiguity in a no-range Limit Hold'em autoregressive model trained only on action and value…

First, we introduce MemHop, a multi-hop memory benchmark of 1,000 questions at hop depths 1-5 across 10 social-network…

However, the O(n^2) computational complexity of standard self-attention causes inference costs to grow sharply with…

In practice, recommendation often involves constructing slates -- ordered lists of items -- that must satisfy multiple…

Existing provenance systems operate at file or dataset level, forcing catastrophic over-deletion. We present ob, a…

However, their performance can be further improved through agentic workflows tailored to real-world mathematical…

Computer Science > Artificial Intelligence arXiv:2607.09330 (cs) [Submitted on 10 Jul 2026]…

The paper investigates how Bayesian causal discovery behaves when latent confounding is present in linear Gaussian…

Prior evaluations of LLM-based medical agents have largely emphasized short-context knowledge QA and tool use.…

Two experiments across 20 diverse worldbuilding tasks, using GPT-OSS 120B and DeepSeek v3.2 as LLM backends,…

OpenProver integrates a Planner-Worker-Verifier architecture inspired by recent ATP agentic systems such as Aletheia.…

Equipped with broad knowledge, flexible reasoning, and tool use, they have the potential to autonomously explore and…

Computer Science > Artificial Intelligence arXiv:2607.09175 (cs) [Submitted on 10 Jul 2026] Title:Scoped Verification…

However, existing PRMs are text-based: they re-encode the entire trajectory text from scratch. In long multi-agent…

While Large Language Models (LLMs) have strong semantic reasoning abilities to assist in decision support, their…

ARCANA decomposes each task into iterative perception, hypothesis generation, symbolic execution, and reflective…

The paper presents the Legal Multi-Agent Debate (L-MAD) framework for evaluating debate structures in legal textual…

Many rely on synthetic conversations or patient simulators, omit patient-uploaded medical images, or evaluate…

Computer Science > Artificial Intelligence arXiv:2607.08986 (cs) [Submitted on 9 Jul 2026] Title:A Formalization of…

However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated…

We present \textbf{GATS} (Graph-Augmented Tree Search), a planning framework that combines systematic UCB1-based tree…

We challenge this by demonstrating that reliability is significantly influenced by \emph{inference-time control} --…

In particular, we show that the adversarial robustness problem can be reduced to a lattice traversal problem. Each…

Securities and Exchange Commission (SEC) which can be found in EDGAR. We were preprocessing those data and than…

To address these limitations, we propose EVAD, an event enhanced VAD framework that jointly exploits conventional…

Therefore, semi-supervised approaches such as Graph Convolutional Networks (GCNs), which learn from both labeled and…

Once those metadata disappear, clinically critical failure modes can be masked by strong aggregate performance, and…

Selecting representative training subsets, however, remains challenging: individual sample contributions are unclear,…

Current approaches for automatic precedent retrieval map legal documents to a low-dimensional semantic space and…

A critical limitation is identified in many document understanding benchmarks: visual content is often reducible to…

SE) has continuously evolved through increasingly powerful forms of reuse, from source code and libraries to…

This makes human vision distinctly different from most popular computer vision models in use today, which input images…