WeSearch
Hub / ai-research / arXiv cs.AI
ai-research · source

arXiv cs.AI on WeSearch

Recent ai-research headlines from arXiv cs.AI.

Mahjax: A GPU-Accelerated Mahjong Simulator for Reinforcement Learning in JAX
arXiv cs.AI

Mahjax: A GPU-Accelerated Mahjong Simulator for Reinforcement Learning in JAX

Mahjax is a new GPU-accelerated Mahjong simulator designed for reinforcement learning using JAX. It allows for…

5/22/2026 · 3 min read · 38 views
From Automated to Autonomous: Hierarchical Agent-native Network Architecture (HANA)
arXiv cs.AI

From Automated to Autonomous: Hierarchical Agent-native Network Architecture (HANA)

The paper presents a new architecture called Hierarchical Agent-native Network Architecture (HANA) aimed at achieving…

5/22/2026 · 3 min read · 27 views
COAgents: Multi-Agent Framework to Learn and Navigate Routing Problems Search Space
arXiv cs.AI

COAgents: Multi-Agent Framework to Learn and Navigate Routing Problems Search Space

The article introduces COAgents, a multi-agent framework designed to address the complexities of Vehicle Routing…

5/22/2026 · 3 min read · 28 views
Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines
arXiv cs.AI

Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines

The article discusses advancements in optimizing industrial asset operations through improved caching and workflow…

5/22/2026 · 3 min read · 30 views
Declarative Data Services: Structured Agentic Discovery for Composing Data Systems
arXiv cs.AI

Declarative Data Services: Structured Agentic Discovery for Composing Data Systems

The article discusses a new framework called Declarative Data Services (DDS) aimed at improving the composition of…

5/22/2026 · 3 min read · 28 views
VBFDD-Agent for Electric Vehicle Battery Fault Detection and Diagnosis: Descriptive Text Modeling of Battery Digital Signals
arXiv cs.AI

VBFDD-Agent for Electric Vehicle Battery Fault Detection and Diagnosis: Descriptive Text Modeling of Battery Digital Signals

The VBFDD-Agent is a new approach for detecting and diagnosing faults in electric vehicle batteries. It utilizes…

5/22/2026 · 3 min read · 31 views
Conflict-Aware Additive Guidance for Flow Models under Compositional Rewards
arXiv cs.AI

Conflict-Aware Additive Guidance for Flow Models under Compositional Rewards

The paper presents a novel method called Conflict-Aware Additive Guidance ($g^ ext{car}$) aimed at improving flow…

5/22/2026 · 3 min read · 30 views
Interaction Locality in Hierarchical Recursive Reasoning
arXiv cs.AI

Interaction Locality in Hierarchical Recursive Reasoning

The paper discusses a framework called interaction locality for measuring information flow in spatial reasoning tasks.…

5/22/2026 · 3 min read · 34 views
Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment
arXiv cs.AI

Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment

The paper discusses the conditional equivalence of Direct Preference Optimization (DPO) and Reinforcement Learning…

5/22/2026 · 3 min read · 26 views
PlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Models
arXiv cs.AI

PlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Models

PlanningBench is a new framework designed to generate scalable and verifiable planning data for evaluating and…

5/22/2026 · 3 min read · 39 views
Governance by Construction for Generalist Agents
arXiv cs.AI

Governance by Construction for Generalist Agents

The paper discusses the need for governance in autonomous enterprise agents. It introduces CUGA's policy system, which…

5/22/2026 · 3 min read · 20 views
For How Long Should We Be Punching? Learning Action Duration in Fighting Games
arXiv cs.AI

For How Long Should We Be Punching? Learning Action Duration in Fighting Games

A new study explores how reinforcement learning agents can improve their performance in fighting games by learning not…

5/22/2026 · 3 min read · 26 views
Playing Devil's Advocate: Off-the-Shelf Persona Vectors Rival Targeted Steering for Sycophancy
arXiv cs.AI

Playing Devil's Advocate: Off-the-Shelf Persona Vectors Rival Targeted Steering for Sycophancy

The study explores the impact of different persona vectors on sycophancy in AI models. It compares off-the-shelf…

5/22/2026 · 3 min read · 29 views
AutoRPA: Efficient GUI Automation through LLM-Driven Code Synthesis from Interactions
arXiv cs.AI

AutoRPA: Efficient GUI Automation through LLM-Driven Code Synthesis from Interactions

The article discusses AutoRPA, a framework designed to enhance GUI automation using large language models (LLMs). It…

5/22/2026 · 3 min read · 31 views
ScenePilot: Controllable Boundary-Driven Critical Scenario Generation for Autonomous Driving
arXiv cs.AI

ScenePilot: Controllable Boundary-Driven Critical Scenario Generation for Autonomous Driving

ScenePilot is a new framework designed for generating critical scenarios in autonomous driving. It focuses on creating…

5/22/2026 · 3 min read · 30 views
Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents
arXiv cs.AI

Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents

The Insights Generator (IG) is a new multi-agent system designed to diagnose failures in LLM agents by analyzing…

5/22/2026 · 3 min read · 33 views
Towards Resilient and Autonomous Networks: A BlueSky Vision on AI-Native 6G
arXiv cs.AI

Towards Resilient and Autonomous Networks: A BlueSky Vision on AI-Native 6G

The paper discusses the integration of Artificial Intelligence into 6G networks to enhance their resilience and…

5/22/2026 · 3 min read · 30 views
Teaching AI Through Benchmark Construction: QuestBench as a Course-Based Practice for Accountable Knowledge Work
arXiv cs.AI

Teaching AI Through Benchmark Construction: QuestBench as a Course-Based Practice for Accountable Knowledge Work

The article discusses a new educational approach to teaching AI through benchmark construction, specifically using a…

5/22/2026 · 3 min read · 37 views
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
arXiv cs.AI

PALS: Power-Aware LLM Serving for Mixture-of-Experts Models

The paper presents PALS, a power-aware runtime for serving large language models (LLMs) that optimizes GPU power…

5/22/2026 · 3 min read · 28 views
Mind the Sim-to-Real Gap & Think Like a Scientist
arXiv cs.AI

Mind the Sim-to-Real Gap & Think Like a Scientist

The paper discusses the challenges of bridging the gap between simulated and real-world decision-making in sequential…

5/22/2026 · 3 min read · 22 views
AiraXiv: An AI-Driven Open-Access Platform for Human and AI Scientists
arXiv cs.AI

AiraXiv: An AI-Driven Open-Access Platform for Human and AI Scientists

AiraXiv is a proposed AI-driven open-access platform designed for both human and AI scientists. It aims to address the…

5/22/2026 · 2 min read · 19 views
DeepWeb-Bench: A Deep Research Benchmark Demanding Massive Cross-Source Evidence and Long-Horizon Derivation
arXiv cs.AI

DeepWeb-Bench: A Deep Research Benchmark Demanding Massive Cross-Source Evidence and Long-Horizon Derivation

DeepWeb-Bench is a new benchmark designed to evaluate deep research capabilities of language models. It emphasizes the…

5/22/2026 · 3 min read · 31 views
Diverge to Induce Prompting: Multi-Rationale Induction for Zero-Shot Reasoning
arXiv cs.AI

Diverge to Induce Prompting: Multi-Rationale Induction for Zero-Shot Reasoning

The paper introduces a new framework called Diverge-to-Induce Prompting (DIP) aimed at improving zero-shot reasoning…

5/22/2026 · 2 min read · 32 views
Neural Estimation of Pairwise Mutual Information in Masked Discrete Sequence Models
arXiv cs.AI

Neural Estimation of Pairwise Mutual Information in Masked Discrete Sequence Models

A new paper proposes a neural framework for estimating pairwise conditional mutual information in masked discrete…

5/22/2026 · 3 min read · 34 views
GraphDiffMed: Knowledge-Constrained Differential Attention with Pharmacological Graph Priors for Medication Recommendation
arXiv cs.AI

GraphDiffMed: Knowledge-Constrained Differential Attention with Pharmacological Graph Priors for Medication Recommendation

GraphDiffMed is a new framework designed for medication recommendation that integrates pharmacological knowledge with…

5/22/2026 · 3 min read · 30 views
Improving Quantized Model Performance in Qualitative Analysis with Multi-Pass Prompt Verification
arXiv cs.AI

Improving Quantized Model Performance in Qualitative Analysis with Multi-Pass Prompt Verification

The study focuses on enhancing the performance of quantized large language models (LLMs) in qualitative analysis. It…

5/22/2026 · 3 min read · 29 views
Parallel LLM Reasoning for Bias-Resilient, Robust Conceptual Abstraction
arXiv cs.AI

Parallel LLM Reasoning for Bias-Resilient, Robust Conceptual Abstraction

The paper presents a framework for improving the performance of large language models (LLMs) in analyzing long…

5/22/2026 · 3 min read · 33 views
Pseudo-Siamese Network for Planning in Target-Oriented Proactive Dialogues
arXiv cs.AI

Pseudo-Siamese Network for Planning in Target-Oriented Proactive Dialogues

The article presents a novel approach to target-oriented proactive dialogue systems using a Forward-Focused…

5/22/2026 · 2 min read · 40 views
Data Scaling as Progressive Coverage of a Predictive Contribution Spectrum
arXiv cs.AI

Data Scaling as Progressive Coverage of a Predictive Contribution Spectrum

The paper explores the hypothesis that real-data scaling laws are influenced by a latent predictive contribution…

5/22/2026 · 3 min read · 34 views
FlowLM: Few-Step Language Modeling via Diffusion-to-Flow Adaptation
arXiv cs.AI

FlowLM: Few-Step Language Modeling via Diffusion-to-Flow Adaptation

FlowLM is a new language model that adapts pre-trained diffusion models for efficient few-step text generation. It…

5/22/2026 · 2 min read · 37 views
Evaluating multimodal emotion recognition in proactive conversational agents: A user study
arXiv cs.AI

Evaluating multimodal emotion recognition in proactive conversational agents: A user study

This article discusses a study on multimodal emotion recognition in proactive conversational agents. The research…

5/22/2026 · 3 min read · 35 views
Long-Context Reasoning Through Proxy-Based Chain-of-Thought Tuning
arXiv cs.AI

Long-Context Reasoning Through Proxy-Based Chain-of-Thought Tuning

The paper presents a novel training framework called ProxyCoT aimed at improving long-context reasoning in large…

5/22/2026 · 2 min read · 30 views
Under Pressure: Emotional Framing Induces Measurable Behavioral Shifts and Structured Internal Geometry in Small Language Models
arXiv cs.AI

Under Pressure: Emotional Framing Induces Measurable Behavioral Shifts and Structured Internal Geometry in Small Language Models

The study investigates how emotionally framed evaluations affect the behavior and internal representations of small…

5/22/2026 · 3 min read · 33 views
GrandGuard: Taxonomy, Benchmark, and Safeguards for Elderly-Chatbot Interaction Safety
arXiv cs.AI

GrandGuard: Taxonomy, Benchmark, and Safeguards for Elderly-Chatbot Interaction Safety

The article discusses the development of GrandGuard, a framework aimed at improving safety in interactions between…

5/22/2026 · 3 min read · 44 views
RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation
arXiv cs.AI

RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation

The paper introduces RealUserSim, a new user simulation framework designed to improve agent benchmarking by grounding…

5/22/2026 · 3 min read · 25 views
PrivacyAkinator: Articulating Key Privacy Design Decisions by Answering LLM-Generated Multiple-choice Questions
arXiv cs.AI

PrivacyAkinator: Articulating Key Privacy Design Decisions by Answering LLM-Generated Multiple-choice Questions

The paper introduces PrivacyAkinator, a tool designed to assist developers in making key privacy design decisions. It…

5/22/2026 · 3 min read · 27 views
Governance by Design: Architecting Agentic AI for Organizational Learning and Scalable Autonomy
arXiv cs.AI

Governance by Design: Architecting Agentic AI for Organizational Learning and Scalable Autonomy

The article discusses the transition of agentic AI systems from experimental prototypes to enterprise deployments. It…

5/22/2026 · 3 min read · 27 views
Leveraging Vision-Language Models to Detect Attention in Educational Videos
arXiv cs.AI

Leveraging Vision-Language Models to Detect Attention in Educational Videos

A recent study explores the use of Vision-Language Models (VLMs) to detect learner attention in educational videos.…

5/22/2026 · 3 min read · 35 views
Network-Based Interventions for HIV Prevention via Cascade-Aware Suppression of Transmission
arXiv cs.AI

Network-Based Interventions for HIV Prevention via Cascade-Aware Suppression of Transmission

A new approach to HIV prevention has been proposed through a method called Cascade-Aware Suppression of Transmission…

5/22/2026 · 3 min read · 34 views
AI-Assisted Competency Assessment from Egocentric Video in Simulation-Based Nursing Education
arXiv cs.AI

AI-Assisted Competency Assessment from Egocentric Video in Simulation-Based Nursing Education

A new study explores the use of AI-assisted competency assessment in nursing education through egocentric video…

5/22/2026 · 3 min read · 36 views
TabPFN-MT: A Natively Multitask In-Context Learner for Tabular Data
arXiv cs.AI

TabPFN-MT: A Natively Multitask In-Context Learner for Tabular Data

The paper introduces TabPFN-MT, a multitask in-context learner designed for tabular data. This model improves upon…

5/22/2026 · 3 min read · 30 views
Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine
arXiv cs.AI

Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine

A new paper presents a framework called Score-induced Latent Diffusion (SiLD) for learning diffusion models under the…

5/22/2026 · 3 min read · 33 views
Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry
arXiv cs.AI

Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry

The paper titled 'Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry' introduces a new method…

5/22/2026 · 3 min read · 30 views
LEAP: A closed-loop framework for perovskite precursor additive discovery
arXiv cs.AI

LEAP: A closed-loop framework for perovskite precursor additive discovery

The article discusses LEAP, a closed-loop framework designed for the discovery of perovskite precursor additives. This…

5/22/2026 · 3 min read · 25 views
Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search
arXiv cs.AI

Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search

The article discusses a new framework called Lean Refactor designed for optimizing Lean proofs. This framework…

5/22/2026 · 3 min read · 26 views
GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents
arXiv cs.AI

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents

The paper introduces GROW, a reinforcement learning framework designed for open-world vision-language model agents. It…

5/22/2026 · 3 min read · 42 views
CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning
arXiv cs.AI

CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning

The paper presents CP-MoE, a framework designed to tackle catastrophic forgetting in continual learning for large…

5/22/2026 · 3 min read · 34 views
Automated Kernel Discovery Towards Understanding High-dimensional Bayesian Optimization
arXiv cs.AI

Automated Kernel Discovery Towards Understanding High-dimensional Bayesian Optimization

The paper introduces a novel framework called Kernel Discovery for high-dimensional Bayesian optimization. This…

5/22/2026 · 3 min read · 38 views
ProcBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents
arXiv cs.AI

ProcBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents

The article introduces ProcBench, a new benchmark designed to evaluate process-level defects in LLM coding agents.…

5/22/2026 · 3 min read · 30 views
Efficient Table QA via TableGrid Navigation and Progressive Inference Prompting
arXiv cs.AI

Efficient Table QA via TableGrid Navigation and Progressive Inference Prompting

The paper presents a novel approach to Table Question-Answering (TQA) using two frameworks: TableGrid Navigation (TGN)…

5/22/2026 · 3 min read · 35 views

How WeSearch handles this source

WeSearch's declared handling of arXiv cs.AI's content. Indexing, snippets, summaries, retrieval and training are separate questions — see the rights registry or read this source's machine-readable record.

Indexing: Allowed Snippet: Allowed AI summary: Limited Retrieval / RAG: Not asserted Model training: Not asserted Commercial reuse: Not permitted

More ai-research sources

Visit arXiv cs.AI directly →