WeSearch
Hub / ai-research / arXiv cs.AI
ai-research · source

arXiv cs.AI on WeSearch

Recent ai-research headlines from arXiv cs.AI.

Boosting Inference with Guided Reasoning: Stochastic Exploration for Recursive Models
arXiv cs.AI

Boosting Inference with Guided Reasoning: Stochastic Exploration for Recursive Models

The paper discusses a novel approach to enhancing inference in recursive neural networks through guided reasoning and…

5/26/2026 · 3 min read · 38 views
Meta-Agent: From Task Descriptions to Verified Multi-Agent Systems
arXiv cs.AI

Meta-Agent: From Task Descriptions to Verified Multi-Agent Systems

The paper introduces Meta-Agent, a framework designed to improve the reliability of multi-agent systems. It automates…

5/26/2026 · 3 min read · 29 views
FrontierOR: Benchmarking LLMs' Capacity for Efficient Algorithm Design in Large-Scale Optimization
arXiv cs.AI

FrontierOR: Benchmarking LLMs' Capacity for Efficient Algorithm Design in Large-Scale Optimization

The paper introduces FrontierOR, a benchmark designed to evaluate the capacity of large language models (LLMs) in…

5/26/2026 · 3 min read · 36 views
LipoAgent: Coordinating Fine-Tuned LLM Agents for Safer Lipid Design
arXiv cs.AI

LipoAgent: Coordinating Fine-Tuned LLM Agents for Safer Lipid Design

LipoAgent is a new framework designed to enhance the safety and efficiency of lipid nanoparticles for nucleic acid…

5/26/2026 · 2 min read · 35 views
Whose Alignment? Comparing LLM Process Alignment Across Diverse Organizational Decision Contexts
arXiv cs.AI

Whose Alignment? Comparing LLM Process Alignment Across Diverse Organizational Decision Contexts

The paper discusses the alignment of AI systems with organizational decision-making, emphasizing the complexity of…

5/26/2026 · 3 min read · 32 views
AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
arXiv cs.AI

AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems

The paper titled 'AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems' introduces a framework for…

5/26/2026 · 3 min read · 35 views
Context-CoT: Enhancing Context Learning via High-Quality Reasoning Synthesis
arXiv cs.AI

Context-CoT: Enhancing Context Learning via High-Quality Reasoning Synthesis

The paper titled 'Context-CoT: Enhancing Context Learning via High-Quality Reasoning Synthesis' addresses the…

5/26/2026 · 2 min read · 33 views
Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models
arXiv cs.AI

Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models

The paper titled 'Second Guess' introduces a method for detecting uncertainty in small language models through a…

5/26/2026 · 2 min read · 29 views
Towards end-to-end LLM-based censoring-aware survival analysis
arXiv cs.AI

Towards end-to-end LLM-based censoring-aware survival analysis

The article discusses a new framework called LLMSurvival that enables censoring-aware survival analysis using large…

5/26/2026 · 3 min read · 37 views
CODESKILL: Learning Self-Evolving Skills for Coding Agents
arXiv cs.AI

CODESKILL: Learning Self-Evolving Skills for Coding Agents

CODESKILL is a proposed framework aimed at enhancing coding agents' abilities through self-evolving skills. It…

5/26/2026 · 2 min read · 34 views
Security of OpenClaw Agents: Fundamentals, Attacks, and Countermeasures
arXiv cs.AI

Security of OpenClaw Agents: Fundamentals, Attacks, and Countermeasures

The paper discusses the security challenges associated with OpenClaw agents, a new class of autonomous systems. It…

5/26/2026 · 3 min read · 33 views
A Signal-Language Foundation Model for Broad-Spectrum Cardiovascular Assessment from Routine Electrocardiography
arXiv cs.AI

A Signal-Language Foundation Model for Broad-Spectrum Cardiovascular Assessment from Routine Electrocardiography

A new AI model called ECGCLIP has been developed to enhance cardiovascular assessment from routine…

5/26/2026 · 3 min read · 34 views
ATWL: A Formal Language for Representing, Comparing, and Reusing Visual Analytics Workflows
arXiv cs.AI

ATWL: A Formal Language for Representing, Comparing, and Reusing Visual Analytics Workflows

The paper introduces the Artifact-Transform Workflow Language (ATWL), a formal language designed for visual analytics…

5/26/2026 · 3 min read · 33 views
Credit Assignment with Resets in Language Model Reasoning
arXiv cs.AI

Credit Assignment with Resets in Language Model Reasoning

The article discusses advancements in credit assignment methods for language model reasoning in reinforcement…

5/26/2026 · 3 min read · 27 views
What Gets Cited: Competitive GEO in AI Answer Engines
arXiv cs.AI

What Gets Cited: Competitive GEO in AI Answer Engines

The paper titled 'What Gets Cited: Competitive GEO in AI Answer Engines' explores how AI answer engines cite sources.…

5/26/2026 · 3 min read · 34 views
BOHM: Zero-Cost Hierarchical Attribution for Compound AI Systems
arXiv cs.AI

BOHM: Zero-Cost Hierarchical Attribution for Compound AI Systems

The article presents BOHM, a new method for hierarchical attribution in compound AI systems. Unlike traditional…

5/25/2026 · 3 min read · 26 views
NeuroNL2LTL: A Neurosymbolic Framework for Natural Language Translation of Linear Temporal Logic
arXiv cs.AI

NeuroNL2LTL: A Neurosymbolic Framework for Natural Language Translation of Linear Temporal Logic

NeuroNL2LTL is a neurosymbolic framework designed to translate natural language into Linear Temporal Logic (LTL). This…

5/25/2026 · 3 min read · 50 views
RMA: an Agentic System for Research-Level Mathematical Problems
arXiv cs.AI

RMA: an Agentic System for Research-Level Mathematical Problems

The article introduces Research Math Agents (RMA), a framework designed for automated reasoning on complex…

5/25/2026 · 3 min read · 22 views
SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research
arXiv cs.AI

SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research

SciAtlas is a large-scale knowledge graph aimed at enhancing automated scientific research. It integrates over 43…

5/25/2026 · 3 min read · 29 views
Energy per Successful Goal: Goal-Level Energy Accounting for Agentic AI Systems
arXiv cs.AI

Energy per Successful Goal: Goal-Level Energy Accounting for Agentic AI Systems

A new framework called A-LEMS has been introduced for measuring energy consumption in agentic AI systems. This…

5/25/2026 · 3 min read · 33 views
ImProver 2: Iteratively Self-Improving LMs for Neurosymbolic Proof Optimization
arXiv cs.AI

ImProver 2: Iteratively Self-Improving LMs for Neurosymbolic Proof Optimization

ImProver 2 is a neurosymbolic framework designed for automated proof optimization in formal mathematics. It addresses…

5/25/2026 · 3 min read · 29 views
Mediative Fuzzy Logic: From Type-1 Foundations to Type-2, Type-3 and Quantum Extensions
arXiv cs.AI

Mediative Fuzzy Logic: From Type-1 Foundations to Type-2, Type-3 and Quantum Extensions

The article discusses the development of Mediative Fuzzy Logic, which aims to reconcile conflicting assessments in…

5/25/2026 · 3 min read · 25 views
EVE-Agent: Evidence-Verifiable Self-Evolving Agents
arXiv cs.AI

EVE-Agent: Evidence-Verifiable Self-Evolving Agents

The paper introduces EVE-Agent, a self-evolving agent designed to enhance the reliability of AI-generated answers. It…

5/25/2026 · 3 min read · 25 views
The Deterministic Horizon: Impossibility Results as Design Specifications for Trustworthy AI Systems
arXiv cs.AI

The Deterministic Horizon: Impossibility Results as Design Specifications for Trustworthy AI Systems

The paper discusses the limitations of AI systems, particularly large language models, and proposes design…

5/25/2026 · 3 min read · 33 views
PathCal: State-Aware Reflection-Marker Calibration for Efficient Reasoning
arXiv cs.AI

PathCal: State-Aware Reflection-Marker Calibration for Efficient Reasoning

The paper introduces PathCal, a novel method for calibrating reasoning paths in Large Reasoning Language Models. It…

5/25/2026 · 3 min read · 36 views
Inductive Deductive Synthesis: Enabling AI to Generate Formally Verified Systems
arXiv cs.AI

Inductive Deductive Synthesis: Enabling AI to Generate Formally Verified Systems

The paper presents Inductive Deductive Synthesis (IDS), a novel approach for enabling AI to generate formally verified…

5/25/2026 · 3 min read · 32 views
Redrawing the AI Map: A Theory of Accountability Boundaries in Agentic Ecosystems
arXiv cs.AI

Redrawing the AI Map: A Theory of Accountability Boundaries in Agentic Ecosystems

The paper titled 'Redrawing the AI Map' explores accountability boundaries in agentic ecosystems. It introduces a…

5/25/2026 · 3 min read · 30 views
AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery
arXiv cs.AI

AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery

The article discusses the advancements in AI systems aimed at automating scientific research workflows. It highlights…

5/25/2026 · 3 min read · 29 views
Foundation Protocol: A Coordination Layer for Agentic Society
arXiv cs.AI

Foundation Protocol: A Coordination Layer for Agentic Society

The Foundation Protocol (FP) is introduced as a coordination layer for an emerging human-AI society. It aims to…

5/25/2026 · 3 min read · 27 views
GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models
arXiv cs.AI

GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models

The paper titled GENSTRAT introduces a new approach to evaluate strategic reasoning in large language models (LLMs).…

5/25/2026 · 3 min read · 35 views
Design and Report Benchmarks for Knowledge Work
arXiv cs.AI

Design and Report Benchmarks for Knowledge Work

The paper discusses the need for improved benchmarks in knowledge work AI, particularly in areas like coding and…

5/25/2026 · 3 min read · 31 views
Parallel Context Compaction for Long-Horizon LLM Agent Serving
arXiv cs.AI

Parallel Context Compaction for Long-Horizon LLM Agent Serving

The paper discusses a new method called parallel context compaction for managing long-horizon LLM agents. This…

5/25/2026 · 2 min read · 35 views
Ontological Knowledge Blocks: Executable Compliance and Profile-Based Validation for Trustworthy AI Systems
arXiv cs.AI

Ontological Knowledge Blocks: Executable Compliance and Profile-Based Validation for Trustworthy AI Systems

The paper introduces Ontological Knowledge Blocks (OKBs) as a solution for ensuring compliance in AI systems. OKBs…

5/25/2026 · 3 min read · 23 views
DART: Semantic Recoverability for Structured Tool Agents
arXiv cs.AI

DART: Semantic Recoverability for Structured Tool Agents

The paper titled 'DART: Semantic Recoverability for Structured Tool Agents' addresses the challenges faced by…

5/25/2026 · 3 min read · 28 views
Human-in-the-Loop Multi-Agent Ventilator Decision Support with Contextual Bandit Preference Learning
arXiv cs.AI

Human-in-the-Loop Multi-Agent Ventilator Decision Support with Contextual Bandit Preference Learning

A new paper presents a Human-in-the-Loop Multi-Agent Ventilator Decision Support System (VDSS) that utilizes…

5/25/2026 · 2 min read · 26 views
When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems
arXiv cs.AI

When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems

The paper discusses the phenomenon of epistemic miscalibration in planning within LLM-based multi-agent systems. It…

5/25/2026 · 2 min read · 28 views
EDGE-OPD: Internalizing Privileged Context with Evidence Guided On-Policy Distillation
arXiv cs.AI

EDGE-OPD: Internalizing Privileged Context with Evidence Guided On-Policy Distillation

The paper presents EDGE-OPD, a method for improving On-Policy Distillation (OPD) in machine learning. It addresses…

5/25/2026 · 3 min read · 29 views
CP or DP? Why Not Both: A Case Study in the Partial Shop Scheduling Problem
arXiv cs.AI

CP or DP? Why Not Both: A Case Study in the Partial Shop Scheduling Problem

The paper discusses the integration of Dynamic Programming (DP) and Constraint Programming (CP) in solving the Partial…

5/25/2026 · 3 min read · 32 views
Co-ReAct: Rubrics as Step-Level Collaborators for ReAct Agents
arXiv cs.AI

Co-ReAct: Rubrics as Step-Level Collaborators for ReAct Agents

The paper introduces Co-ReAct, a framework that enhances ReAct agents by using rubrics as step-level guidance during…

5/25/2026 · 3 min read · 21 views
Solving the Aircraft Disassembly Scheduling Problem
arXiv cs.AI

Solving the Aircraft Disassembly Scheduling Problem

The article discusses the complexities involved in dismantling aircrafts at the end of their life cycle. It emphasizes…

5/25/2026 · 2 min read · 28 views
One Policy, Infinite NPCs: Persona-Traceable Shared RL Policies for Scalable Game Agents
arXiv cs.AI

One Policy, Infinite NPCs: Persona-Traceable Shared RL Policies for Scalable Game Agents

The paper introduces a novel approach for controlling non-player characters (NPCs) in life simulation games using a…

5/25/2026 · 3 min read · 34 views
MemAudit: Post-hoc Auditing of Poisoned Agent Memory via Causal Attribution and Structural Anomaly Detection
arXiv cs.AI

MemAudit: Post-hoc Auditing of Poisoned Agent Memory via Causal Attribution and Structural Anomaly Detection

The article discusses a new framework called MemAudit designed for auditing the memory of language model agents. This…

5/25/2026 · 3 min read · 32 views
Agentic Proving for Program Verification
arXiv cs.AI

Agentic Proving for Program Verification

The paper discusses the application of agentic systems in program verification. It evaluates the performance of Claude…

5/25/2026 · 3 min read · 27 views
Beyond Binary Edits Robust Multimodal Knowledge Editing with Adversarial Subspace Alignment
arXiv cs.AI

Beyond Binary Edits Robust Multimodal Knowledge Editing with Adversarial Subspace Alignment

The paper discusses advancements in multimodal large language models (MLLMs) for knowledge editing. It addresses the…

5/25/2026 · 2 min read · 33 views
SPACENUM: Revisiting Spatial Numerical Understanding in VLMs
arXiv cs.AI

SPACENUM: Revisiting Spatial Numerical Understanding in VLMs

The paper titled 'SPACENUM: Revisiting Spatial Numerical Understanding in VLMs' explores the capabilities of…

5/25/2026 · 3 min read · 29 views
From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills
arXiv cs.AI

From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills

The study explores the lifecycle of model-generated agent skills, focusing on experience generation, skill extraction,…

5/25/2026 · 3 min read · 22 views
SkillOpt: Executive Strategy for Self-Evolving Agent Skills
arXiv cs.AI

SkillOpt: Executive Strategy for Self-Evolving Agent Skills

The paper introduces SkillOpt, a novel approach for optimizing agent skills in artificial intelligence. Unlike…

5/25/2026 · 3 min read · 25 views
An AI-Driven Framework for Energy-Efficient Environmental Monitoring in Smart Cities Using Edge Intelligence
arXiv cs.AI

An AI-Driven Framework for Energy-Efficient Environmental Monitoring in Smart Cities Using Edge Intelligence

A new AI-driven framework has been proposed for energy-efficient environmental monitoring in smart cities. This…

5/25/2026 · 3 min read · 43 views
KPI2KVI: A Multi Agent Workflow for Calculating Key Value Indicators from Service Descriptions
arXiv cs.AI

KPI2KVI: A Multi Agent Workflow for Calculating Key Value Indicators from Service Descriptions

The paper introduces KPI2KVI, a tool designed to compute Key Value Indicators (KVIs) from service descriptions. It…

5/25/2026 · 3 min read · 28 views
Evaluating Large Language Models in a Complex Hidden Role Game
arXiv cs.AI

Evaluating Large Language Models in a Complex Hidden Role Game

The study evaluates the deceptive capabilities of Large Language Models (LLMs) in the social deduction game Secret…

5/25/2026 · 3 min read · 31 views

How WeSearch handles this source

WeSearch's declared handling of arXiv cs.AI's content. Indexing, snippets, summaries, retrieval and training are separate questions — see the rights registry or read this source's machine-readable record.

Indexing: Allowed Snippet: Allowed AI summary: Limited Retrieval / RAG: Not asserted Model training: Not asserted Commercial reuse: Not permitted

More ai-research sources

Visit arXiv cs.AI directly →