WeSearch
Hub / ai-research / arXiv cs.AI
ai-research · source

arXiv cs.AI on WeSearch

Recent ai-research headlines from arXiv cs.AI.

Modeling Agentic Technical Debt and Stochastic Tax: A Standalone Framework for Measurement, Simulation, and Dashboarding
arXiv cs.AI

Modeling Agentic Technical Debt and Stochastic Tax: A Standalone Framework for Measurement, Simulation, and Dashboarding

The article presents a framework for measuring and simulating Agentic Technical Debt and Stochastic Tax in AI systems.…

5/27/2026 · 3 min read · 43 views
Maat: The Agentic Legal Research Assistant for Competition Protection
arXiv cs.AI

Maat: The Agentic Legal Research Assistant for Competition Protection

Maat is a new legal research assistant designed specifically for competition law analysis. It outperforms existing…

5/27/2026 · 3 min read · 26 views
2-ASP(Q) programs with weak constraints: Complexity and efficient implementation
arXiv cs.AI

2-ASP(Q) programs with weak constraints: Complexity and efficient implementation

The paper discusses 2-ASP(Q) programs with weak constraints, a significant area in Answer Set Programming. It provides…

5/27/2026 · 2 min read · 33 views
Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
arXiv cs.AI

Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases

The paper discusses a vulnerability known as alignment tampering in Reinforcement Learning from Human Feedback (RLHF).…

5/27/2026 · 3 min read · 40 views
Natural Language Query to Configuration for Retrieval Agents
arXiv cs.AI

Natural Language Query to Configuration for Retrieval Agents

The paper presents a new approach called BRANE for optimizing retrieval agent configurations based on natural language…

5/27/2026 · 3 min read · 32 views
MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
arXiv cs.AI

MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation

The MUSE-Autoskill framework introduces a new approach for self-evolving agents that enhances their ability to create…

5/27/2026 · 3 min read · 33 views
Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU
arXiv cs.AI

Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU

Xe-Forge is a new multi-stage pipeline designed to optimize kernel performance for Intel GPUs. It automates the…

5/27/2026 · 3 min read · 24 views
Edge AI Deployment Beyond Models: A BSP-Aware Systems Framework for Industrial Embedded Platforms
arXiv cs.AI

Edge AI Deployment Beyond Models: A BSP-Aware Systems Framework for Industrial Embedded Platforms

The paper presents a framework for deploying Edge AI in industrial embedded platforms, emphasizing the importance of a…

5/27/2026 · 3 min read · 29 views
GEM: Geometric Entropy Mixing for Optimal LLM Data Curation
arXiv cs.AI

GEM: Geometric Entropy Mixing for Optimal LLM Data Curation

The paper introduces GEM, a framework designed for optimal data curation in large language models (LLMs). It addresses…

5/27/2026 · 2 min read · 42 views
Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications
arXiv cs.AI

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications

The paper discusses Pretraining Data Exposure (PDE) in Large Language Models (LLMs), highlighting its implications for…

5/27/2026 · 2 min read · 34 views
Eroding Trust in Real Speech: A Large-Scale Study of Human Audio Deepfake Perception
arXiv cs.AI

Eroding Trust in Real Speech: A Large-Scale Study of Human Audio Deepfake Perception

A recent study investigates the impact of audio deepfakes on human trust in real speech. The research, which involved…

5/27/2026 · 3 min read · 35 views
AssetGen: Deployable 3D Asset Generation at Interactive Speed
arXiv cs.AI

AssetGen: Deployable 3D Asset Generation at Interactive Speed

AssetGen is a new 3D asset generation system that prioritizes user experience and deployability. It can produce…

5/27/2026 · 3 min read · 36 views
VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents
arXiv cs.AI

VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents

The article introduces VISTA, a benchmark designed to evaluate the capabilities of LLM-based agents in generating web…

5/27/2026 · 3 min read · 31 views
Augment Engineering: A Methodology for Multi-Tool AI Orchestration Across Professional Domains
arXiv cs.AI

Augment Engineering: A Methodology for Multi-Tool AI Orchestration Across Professional Domains

The paper titled 'Augment Engineering' introduces a methodology for orchestrating multiple AI tools across various…

5/27/2026 · 3 min read · 39 views
MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning
arXiv cs.AI

MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning

The paper introduces MemMorph, a novel attack method targeting long-term memory in LLM-driven agents. By injecting…

5/27/2026 · 3 min read · 30 views
When Does Adaptive Guidance Help? Belief-Aware Privileged Distillation for Autonomous Driving Under Partial Observability
arXiv cs.AI

When Does Adaptive Guidance Help? Belief-Aware Privileged Distillation for Autonomous Driving Under Partial Observability

The paper introduces Belief-Aware GSAC (BA-GSAC), which adapts the distillation coefficient in autonomous driving…

5/27/2026 · 3 min read · 38 views
Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges
arXiv cs.AI

Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges

The paper introduces BITE, a framework designed to exploit stylistic biases in LLM judges. It demonstrates that these…

5/27/2026 · 3 min read · 40 views
Furina: Fragmented Uncertainty-Driven Refusal Instability Attack
arXiv cs.AI

Furina: Fragmented Uncertainty-Driven Refusal Instability Attack

The paper titled 'Furina: Fragmented Uncertainty-Driven Refusal Instability Attack' explores safety alignment in large…

5/27/2026 · 2 min read · 37 views
TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models
arXiv cs.AI

TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models

The paper titled 'TSFMAudit' addresses the issue of data contamination in time series foundation models (TSFMs). It…

5/27/2026 · 3 min read · 35 views
On the Push-Based Asynchronous Federated Learning: A Bias-Correction Aggregation Approach
arXiv cs.AI

On the Push-Based Asynchronous Federated Learning: A Bias-Correction Aggregation Approach

The article discusses a new framework called PushCen-ADFL for asynchronous decentralized federated learning. This…

5/27/2026 · 3 min read · 32 views
Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets
arXiv cs.AI

Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets

The paper discusses the challenges faced by agentic RAG systems due to tool schemas consuming context windows needed…

5/27/2026 · 3 min read · 36 views
Enhancing Autonomous Online Intrusion Detection for IoT with Balanced Learning, Reliable Pseudo-Labels, and Lightweight Architectures
arXiv cs.AI

Enhancing Autonomous Online Intrusion Detection for IoT with Balanced Learning, Reliable Pseudo-Labels, and Lightweight Architectures

The paper discusses advancements in autonomous online intrusion detection systems (IDS) for IoT devices. It highlights…

5/27/2026 · 3 min read · 32 views
Planning Neural Dynamics with Lie Group Embedding through Supervised Projective Manifold Learning
arXiv cs.AI

Planning Neural Dynamics with Lie Group Embedding through Supervised Projective Manifold Learning

The article presents a novel approach to neural dynamics using Lie group embedding through supervised projective…

5/27/2026 · 3 min read · 37 views
A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orchestration
arXiv cs.AI

A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orchestration

The paper discusses the challenges of detecting cross-section defects in documents processed by language model…

5/27/2026 · 3 min read · 36 views
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
arXiv cs.AI

InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization

The paper titled 'InfoQuant' addresses the challenges of low-bit activation quantization in large language models. It…

5/27/2026 · 3 min read · 36 views
PitchBench: Measuring Pitch Hearing in Audio-Language Models
arXiv cs.AI

PitchBench: Measuring Pitch Hearing in Audio-Language Models

The article introduces PitchBench, a new evaluation suite designed to measure pitch hearing in audio-language models…

5/27/2026 · 3 min read · 35 views
RepoMirage: Probing Repository Context Reasoning in Code Agents with Perturbations
arXiv cs.AI

RepoMirage: Probing Repository Context Reasoning in Code Agents with Perturbations

RepoMirage is a new evaluation suite designed to assess repository context reasoning in code agents. The study reveals…

5/27/2026 · 3 min read · 35 views
AutoDFT: A Closed-Loop Multi-Agent Framework for Autonomous DFT Calculations
arXiv cs.AI

AutoDFT: A Closed-Loop Multi-Agent Framework for Autonomous DFT Calculations

AutoDFT is a new multi-agent framework designed to enhance autonomous DFT calculations in materials science. It…

5/27/2026 · 3 min read · 39 views
GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training
arXiv cs.AI

GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training

The paper presents GAC, a noise-aware adaptive mixing method for hybrid post-training in machine learning. This…

5/27/2026 · 2 min read · 31 views
SetupX: Can LLM Agents Learn from Past Failures in Functionality-Correct Code Repository Setup?
arXiv cs.AI

SetupX: Can LLM Agents Learn from Past Failures in Functionality-Correct Code Repository Setup?

The paper introduces SetupX, a framework designed to improve the setup of functionality-correct code repositories by…

5/27/2026 · 3 min read · 37 views
In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models
arXiv cs.AI

In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models

The paper explores the potential of large Vision-Language Models (VLMs) to replicate the open-ended creative processes…

5/26/2026 · 3 min read · 39 views
Confidence Calibration in Large Language Models
arXiv cs.AI

Confidence Calibration in Large Language Models

A recent study examines the confidence calibration of large language models (LLMs) across various tasks. The findings…

5/26/2026 · 2 min read · 40 views
How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning
arXiv cs.AI

How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning

The paper investigates the redundancy in reasoning processes of large language models (LLMs). It quantifies how much…

5/26/2026 · 3 min read · 33 views
Context: Proactive Goal-Directed Intelligence via Composable Sandboxed Programs, Declarative Wiring, and Structured Interaction
arXiv cs.AI

Context: Proactive Goal-Directed Intelligence via Composable Sandboxed Programs, Declarative Wiring, and Structured Interaction

The article introduces Context, a new intelligence layer designed to enhance proactive goal-directed interactions in…

5/26/2026 · 3 min read · 37 views
Toward Reliable Design of LLM-Enabled Agentic Workflows: Optimizing Latency-Reliability-Cost Tradeoffs
arXiv cs.AI

Toward Reliable Design of LLM-Enabled Agentic Workflows: Optimizing Latency-Reliability-Cost Tradeoffs

The paper discusses the design of workflows that utilize large language models (LLMs) alongside traditional…

5/26/2026 · 2 min read · 39 views
Quantum Frog: Emergent Cooperation and Difficulty Scaling in a Quantized-Time Cooperative Game
arXiv cs.AI

Quantum Frog: Emergent Cooperation and Difficulty Scaling in a Quantized-Time Cooperative Game

The paper introduces 'Quantum Frog', a two-player cooperative game that utilizes a quantized-time mechanic. It…

5/26/2026 · 3 min read · 46 views
BODHI: Precise OS Kernel Specification Inference
arXiv cs.AI

BODHI: Precise OS Kernel Specification Inference

The paper titled 'BODHI: Precise OS Kernel Specification Inference' introduces a method to enhance the formal…

5/26/2026 · 3 min read · 47 views
When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure
arXiv cs.AI

When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure

The paper discusses the challenges faced by large language models (LLMs) in maintaining correct medical diagnoses…

5/26/2026 · 2 min read · 43 views
Practical Quantum CIM Empowerment via All-Domestic-Core Agentic Large Model
arXiv cs.AI

Practical Quantum CIM Empowerment via All-Domestic-Core Agentic Large Model

The paper discusses the integration of a Coherent Ising Machine (CIM) with a large language model (LLM) to enhance…

5/26/2026 · 3 min read · 43 views
Operationalizing Reconstructive Authority: Runtime Construction, Dependency Resolution, and Execution Gating in Autonomous Agent Systems
arXiv cs.AI

Operationalizing Reconstructive Authority: Runtime Construction, Dependency Resolution, and Execution Gating in Autonomous Agent Systems

The paper discusses the concept of Reconstructive Authority in autonomous agent systems, focusing on how to enforce…

5/26/2026 · 3 min read · 44 views
Fuzzy, Neutrosophic, and Uncertain Graph Theory: Properties and Applications
arXiv cs.AI

Fuzzy, Neutrosophic, and Uncertain Graph Theory: Properties and Applications

The article discusses a new book on fuzzy, neutrosophic, and uncertain graph theory. It emphasizes the unifying role…

5/26/2026 · 2 min read · 46 views
BoxLitE: A Faithful Knowledge Base Embedding Based on Convex Optimization
arXiv cs.AI

BoxLitE: A Faithful Knowledge Base Embedding Based on Convex Optimization

The paper introduces BoxLitE, a knowledge base embedding model that utilizes convex optimization. This model aims to…

5/26/2026 · 3 min read · 42 views
Authority Inversion in LLM-Mediated Ubiquitous Systems: When Models Trust Users Over Sensors
arXiv cs.AI

Authority Inversion in LLM-Mediated Ubiquitous Systems: When Models Trust Users Over Sensors

The paper discusses the phenomenon of Authority Inversion in large language models (LLMs) used in ubiquitous systems.…

5/26/2026 · 3 min read · 43 views
DRIVE: Modeling Skills at the Reasoning and Interaction Levels for Web Agents under Continual Learning
arXiv cs.AI

DRIVE: Modeling Skills at the Reasoning and Interaction Levels for Web Agents under Continual Learning

The paper titled DRIVE proposes a dual-level skill modeling framework for web agents to enhance their reasoning and…

5/26/2026 · 3 min read · 39 views
Residual Drift Dominates Contradiction in Multi-Turn Constraint Reasoning
arXiv cs.AI

Residual Drift Dominates Contradiction in Multi-Turn Constraint Reasoning

The paper discusses how multi-turn reasoning systems often fail not due to logical contradictions, but rather due to a…

5/26/2026 · 2 min read · 36 views
MEMOR-E: In-Context and Fine-Tuned LLM Personalization for Alzheimer's Assistive Robotics
arXiv cs.AI

MEMOR-E: In-Context and Fine-Tuned LLM Personalization for Alzheimer's Assistive Robotics

The paper presents MEMOR-E, a mobile quadruped robot designed to assist Alzheimer's patients and their caregivers. It…

5/26/2026 · 3 min read · 33 views
A Dynamical Framework for Cognitive Processes Based on Transformations and Semantic Equivalence
arXiv cs.AI

A Dynamical Framework for Cognitive Processes Based on Transformations and Semantic Equivalence

The paper presents a framework for modeling cognitive processes through a cybernetic lens. It introduces a feedback…

5/26/2026 · 2 min read · 43 views
Spacetime Formation under Requirements: Contextual Realization and Form-Dependent Probability
arXiv cs.AI

Spacetime Formation under Requirements: Contextual Realization and Form-Dependent Probability

The paper titled 'Spacetime Formation under Requirements: Contextual Realization and Form-Dependent Probability' by…

5/26/2026 · 3 min read · 48 views
Right-Sizing Communication and Recommendation Set Size in AI-Assisted Search
arXiv cs.AI

Right-Sizing Communication and Recommendation Set Size in AI-Assisted Search

The paper discusses the interaction between users and AI-driven recommendation systems. It models how users convey…

5/26/2026 · 3 min read · 39 views
Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism
arXiv cs.AI

Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism

The paper presents a new method called PAT for improving the efficiency of Reinforcement Learning from Human Feedback…

5/26/2026 · 3 min read · 39 views

How WeSearch handles this source

WeSearch's declared handling of arXiv cs.AI's content. Indexing, snippets, summaries, retrieval and training are separate questions — see the rights registry or read this source's machine-readable record.

Indexing: Allowed Snippet: Allowed AI summary: Limited Retrieval / RAG: Not asserted Model training: Not asserted Commercial reuse: Not permitted

More ai-research sources

Visit arXiv cs.AI directly →