WeSearch
Hub / ai-research / arXiv cs.AI
ai-research · source

arXiv cs.AI on WeSearch

Recent ai-research headlines from arXiv cs.AI.

LinAlg-Bench: A Forensic Benchmark Revealing Structural Failure Modes in LLM Mathematical Reasoning
arXiv cs.AI

LinAlg-Bench: A Forensic Benchmark Revealing Structural Failure Modes in LLM Mathematical Reasoning

LinAlg-Bench is a new diagnostic benchmark designed to evaluate large language models on linear algebra computations.…

5/19/2026 · 3 min read · 44 views
Enhancing Metacognitive AI: Knowledge-Graph Population with Graph-Theoretic LLM Enrichment
arXiv cs.AI

Enhancing Metacognitive AI: Knowledge-Graph Population with Graph-Theoretic LLM Enrichment

The article discusses a new system called MetaKGEnrich that enhances metacognitive abilities in AI. This system…

5/19/2026 · 2 min read · 36 views
Recall Isn't Enough: Bounding Commitments in Personalized Language Systems
arXiv cs.AI

Recall Isn't Enough: Bounding Commitments in Personalized Language Systems

The paper discusses the limitations of current personalized language systems, particularly in how they handle…

5/19/2026 · 3 min read · 20 views
GRID: Graph Representation of Intelligence Data for Security Text Knowledge Graph Construction
arXiv cs.AI

GRID: Graph Representation of Intelligence Data for Security Text Knowledge Graph Construction

The article discusses a new framework called GRID for constructing security text knowledge graphs from cyber threat…

5/19/2026 · 3 min read · 32 views
Baba in Wonderland: Online Self-Supervised Dynamics Discovery for Executable World Models
arXiv cs.AI

Baba in Wonderland: Online Self-Supervised Dynamics Discovery for Executable World Models

The paper titled 'Baba in Wonderland' explores online self-supervised dynamics discovery for executable world models.…

5/19/2026 · 3 min read · 29 views
A Global-Local Graph Attention Network for Traffic Forecasting
arXiv cs.AI

A Global-Local Graph Attention Network for Traffic Forecasting

A new paper introduces the Global-Local Graph Attention Network (GLGAT) aimed at improving traffic forecasting. This…

5/19/2026 · 2 min read · 22 views
PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play
arXiv cs.AI

PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play

The paper introduces PopuLoRA, a novel framework for reinforcement learning with large language models (LLMs). It…

5/19/2026 · 3 min read · 27 views
Body-Grounded Perspective Formation and Conative Attunement in Artificial Agents
arXiv cs.AI

Body-Grounded Perspective Formation and Conative Attunement in Artificial Agents

The paper discusses a new architecture for body-grounded perspective formation in artificial agents. It introduces…

5/19/2026 · 2 min read · 23 views
State Contamination in Memory-Augmented LLM Agents
arXiv cs.AI

State Contamination in Memory-Augmented LLM Agents

The paper discusses the issue of state contamination in memory-augmented LLM agents. It highlights a failure mode…

5/19/2026 · 3 min read · 26 views
NeuroMAS: Multi-Agent Systems as Neural Networks with Joint Reinforcement Learning
arXiv cs.AI

NeuroMAS: Multi-Agent Systems as Neural Networks with Joint Reinforcement Learning

NeuroMAS introduces a novel approach to multi-agent language systems by treating them as neural networks with joint…

5/19/2026 · 3 min read · 31 views
Multi-Paradigm Agent Interaction in Practice:A Systematic Analysis of Generator-Evaluator, ReAct Loop,and Adversarial Evaluation in the buddyMe Framework
arXiv cs.AI

Multi-Paradigm Agent Interaction in Practice:A Systematic Analysis of Generator-Evaluator, ReAct Loop,and Adversarial Evaluation in the buddyMe Framework

The paper presents a systematic analysis of multi-paradigm agent interaction within the buddyMe framework. It explores…

5/19/2026 · 3 min read · 38 views
Voices in the Loop: Mapping Participatory AI
arXiv cs.AI

Voices in the Loop: Mapping Participatory AI

The paper titled 'Voices in the Loop: Mapping Participatory AI' discusses the organization of participatory approaches…

5/19/2026 · 3 min read · 29 views
Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models
arXiv cs.AI

Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models

The paper presents a novel approach to optimizing Diffusion Multi-Modal Large Language Models (dMLLMs) using…

5/19/2026 · 3 min read · 36 views
Artificial Adaptive Intelligence: The Missing Stage Between Narrow and General Intelligence
arXiv cs.AI

Artificial Adaptive Intelligence: The Missing Stage Between Narrow and General Intelligence

A new paper introduces the concept of Artificial Adaptive Intelligence (AAI), a stage between narrow and general…

5/19/2026 · 3 min read · 22 views
Learning to Learn from Multimodal Experience
arXiv cs.AI

Learning to Learn from Multimodal Experience

The paper discusses a new approach to experience-driven learning in artificial intelligence. It emphasizes the need…

5/19/2026 · 2 min read · 25 views
Reasoning Can Be Restored by Correcting a Few Decision Tokens
arXiv cs.AI

Reasoning Can Be Restored by Correcting a Few Decision Tokens

A recent study explores the reasoning gap between large reasoning models and base models in artificial intelligence.…

5/19/2026 · 3 min read · 21 views
Virtual Nodes Guided Dynamic Graph Neural Network for Brain Tumor Segmentation with Missing Modalities
arXiv cs.AI

Virtual Nodes Guided Dynamic Graph Neural Network for Brain Tumor Segmentation with Missing Modalities

A new study presents a graph-based framework for brain tumor segmentation that addresses the common issue of missing…

5/19/2026 · 3 min read · 37 views
NGM: A Plug-and-Play Training-Free Memory Module for LLMs
arXiv cs.AI

NGM: A Plug-and-Play Training-Free Memory Module for LLMs

The paper introduces N-gram Memory (NGM), a training-free memory module designed for large language models (LLMs). NGM…

5/19/2026 · 3 min read · 23 views
TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents
arXiv cs.AI

TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents

The article introduces MM-ToolBench, a benchmark designed for evaluating task-oriented omni-modal tool-using agents.…

5/19/2026 · 3 min read · 21 views
From Static Risk to Dynamic Trajectories: Toward World-Model-Inspired Clinical Prediction
arXiv cs.AI

From Static Risk to Dynamic Trajectories: Toward World-Model-Inspired Clinical Prediction

The article discusses a new approach to clinical prediction that moves from static risk assessments to dynamic…

5/19/2026 · 3 min read · 30 views
How do Humans Process AI-generated Hallucination Contents: a Neuroimaging Study
arXiv cs.AI

How do Humans Process AI-generated Hallucination Contents: a Neuroimaging Study

A recent neuroimaging study investigates how humans process AI-generated hallucinations. The research reveals distinct…

5/19/2026 · 2 min read · 28 views
Harnessing AI for Inverse Partial Differential Equation Problems: Past, Present, and Prospects
arXiv cs.AI

Harnessing AI for Inverse Partial Differential Equation Problems: Past, Present, and Prospects

The paper discusses the application of artificial intelligence in solving inverse partial differential equation (PDE)…

5/19/2026 · 3 min read · 23 views
Brain Vascular Age Prediction Using Cerebral Blood Flow Velocity and Machine Learning Algorithms
arXiv cs.AI

Brain Vascular Age Prediction Using Cerebral Blood Flow Velocity and Machine Learning Algorithms

A recent study explores the prediction of brain vascular age using cerebral blood flow velocity and machine learning…

5/19/2026 · 3 min read · 28 views
A Conflict-aware Evidential Framework for Reliable Sleep Stage Classification
arXiv cs.AI

A Conflict-aware Evidential Framework for Reliable Sleep Stage Classification

The paper introduces ConfSleepNet, a framework designed for reliable sleep stage classification by addressing…

5/19/2026 · 3 min read · 31 views
Reliability and Effectiveness of Autonomous AI Agents in Supply Chain Management
arXiv cs.AI

Reliability and Effectiveness of Autonomous AI Agents in Supply Chain Management

The paper examines the performance of autonomous AI agents in supply chain management through the MIT Beer Game. It…

5/19/2026 · 3 min read · 30 views
Evidential Information Fusion on Possibilistic Structure
arXiv cs.AI

Evidential Information Fusion on Possibilistic Structure

The article discusses a new framework for evidential information fusion based on a possibilistic structure. This…

5/19/2026 · 2 min read · 23 views
PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models
arXiv cs.AI

PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models

The article discusses the development of PersonaArena, a dynamic simulation framework aimed at enhancing persona-level…

5/19/2026 · 2 min read · 32 views
Towards Human-Level Book-Writing Capability
arXiv cs.AI

Towards Human-Level Book-Writing Capability

A new paper discusses the challenges of aligning large language models with the requirements of high-quality creative…

5/19/2026 · 3 min read · 23 views
AnchorDiff: Topology-Aware Masked Diffusion with Confidence-based Rewriting for Radiology Report Generation
arXiv cs.AI

AnchorDiff: Topology-Aware Masked Diffusion with Confidence-based Rewriting for Radiology Report Generation

The paper introduces AnchorDiff, a novel framework for generating radiology reports using a topology-aware masked…

5/19/2026 · 3 min read · 27 views
RAGA: Reading-And-Graph-building-Agent for Autonomous Knowledge Graph Construction and Retrieval-Augmented Generation
arXiv cs.AI

RAGA: Reading-And-Graph-building-Agent for Autonomous Knowledge Graph Construction and Retrieval-Augmented Generation

The article introduces RAGA, a new framework for autonomous knowledge graph construction and retrieval-augmented…

5/19/2026 · 2 min read · 24 views
Scientific Logicality Enriched Methodology for LLM Reasoning: A Practice in Physics
arXiv cs.AI

Scientific Logicality Enriched Methodology for LLM Reasoning: A Practice in Physics

A new methodology for enhancing reasoning in Large Language Models (LLMs) has been proposed, focusing on the…

5/19/2026 · 3 min read · 29 views
Capturing LLM Capabilities via Evidence-Calibrated Query Clustering
arXiv cs.AI

Capturing LLM Capabilities via Evidence-Calibrated Query Clustering

The paper presents a new algorithm called ECC for clustering queries based on their latent capability demands. This…

5/19/2026 · 2 min read · 28 views
F2IND-IT! -- Multimodal Fuzzy Fake Indian News Detection using Images and Text
arXiv cs.AI

F2IND-IT! -- Multimodal Fuzzy Fake Indian News Detection using Images and Text

The paper presents a new framework for detecting fake news in Indian media by integrating visual and textual analysis.…

5/19/2026 · 2 min read · 31 views
Latent Heuristic Search: Continuous Optimization for Automated Algorithm Design
arXiv cs.AI

Latent Heuristic Search: Continuous Optimization for Automated Algorithm Design

The paper presents a new framework for automated algorithm design called Latent Heuristic Search, which utilizes…

5/19/2026 · 3 min read · 26 views
Dynamics of collective creativity in AI art competitions
arXiv cs.AI

Dynamics of collective creativity in AI art competitions

The study explores the dynamics of collective creativity in AI art competitions, specifically through the platform…

5/19/2026 · 3 min read · 34 views
MADP: A Multi-Agent Pipeline for Sustainable Document Processing with Human-in-the-Loop
arXiv cs.AI

MADP: A Multi-Agent Pipeline for Sustainable Document Processing with Human-in-the-Loop

The article introduces MADP, a multi-agent pipeline designed to automate document processing in enterprise…

5/19/2026 · 3 min read · 33 views
From Imitation to Interaction: Mastering Game of Schnapsen with Shallow Reinforcement Learning
arXiv cs.AI

From Imitation to Interaction: Mastering Game of Schnapsen with Shallow Reinforcement Learning

The paper explores the effectiveness of shallow neural network agents in mastering the card game Schnapsen. It…

5/19/2026 · 3 min read · 31 views
Responsible Agentic AI Requires Explicit Provenance
arXiv cs.AI

Responsible Agentic AI Requires Explicit Provenance

The paper discusses the need for explicit provenance in agentic AI to enhance public trust and accountability. It…

5/19/2026 · 3 min read · 30 views
CAREBench: Evaluating LLMs' Emotion Understanding by Assessing Cognitive Appraisal Reasoning
arXiv cs.AI

CAREBench: Evaluating LLMs' Emotion Understanding by Assessing Cognitive Appraisal Reasoning

The article introduces CAREBench, a new benchmark designed to evaluate the emotion understanding capabilities of large…

5/19/2026 · 3 min read · 28 views
ChemVA: Advancing Large Language Models on Chemical Reaction Diagrams Understanding
arXiv cs.AI

ChemVA: Advancing Large Language Models on Chemical Reaction Diagrams Understanding

The paper introduces ChemVA, a framework designed to enhance Large Language Models' understanding of chemical reaction…

5/19/2026 · 3 min read · 25 views
Towards Robust Argumentative Essay Understanding via TIDE: An Interactive Framework with Trial and Debate
arXiv cs.AI

Towards Robust Argumentative Essay Understanding via TIDE: An Interactive Framework with Trial and Debate

The article discusses a new framework called TIDE aimed at improving the understanding of argumentative essays. This…

5/19/2026 · 2 min read · 24 views
CatalyticMLLM: A Graph-Text Multimodal Large Language Model for Catalytic Materials
arXiv cs.AI

CatalyticMLLM: A Graph-Text Multimodal Large Language Model for Catalytic Materials

The paper introduces QE-Catalytic-V2, a multimodal large language model designed for catalytic materials. This model…

5/19/2026 · 3 min read · 30 views
CAM-Bench: A Benchmark for Computational and Applied Mathematics in Lean
arXiv cs.AI

CAM-Bench: A Benchmark for Computational and Applied Mathematics in Lean

CAM-Bench is a new benchmark designed for computational and applied mathematics within the Lean theorem-proving…

5/19/2026 · 3 min read · 33 views
Is VLA Reasoning Faithful? Probing Safety of Chain-of-Causation
arXiv cs.AI

Is VLA Reasoning Faithful? Probing Safety of Chain-of-Causation

A recent study investigates the faithfulness of Vision-Language-Action (VLA) driving models. The research reveals…

5/19/2026 · 2 min read · 33 views
A2RBench: An Automatic Paradigm for Formally Verifiable Abstract Reasoning Benchmark Generation
arXiv cs.AI

A2RBench: An Automatic Paradigm for Formally Verifiable Abstract Reasoning Benchmark Generation

The article introduces A2RBench, an automated system designed to generate benchmarks for evaluating abstract reasoning…

5/19/2026 · 3 min read · 30 views
MetaCogAgent: A Metacognitive Multi-Agent LLM Framework with Self-Aware Task Delegation
arXiv cs.AI

MetaCogAgent: A Metacognitive Multi-Agent LLM Framework with Self-Aware Task Delegation

The article introduces MetaCogAgent, a multi-agent large language model framework designed to enhance task delegation…

5/19/2026 · 3 min read · 36 views
CyberCorrect: A Cybernetic Framework for Closed-Loop Self-Correction in Large Language Models
arXiv cs.AI

CyberCorrect: A Cybernetic Framework for Closed-Loop Self-Correction in Large Language Models

The article introduces CyberCorrect, a framework designed for self-correction in large language models (LLMs). This…

5/19/2026 · 3 min read · 29 views
Reasoning Before Diagnosis: Physician-Inspired Structured Thinking for ECG Classification
arXiv cs.AI

Reasoning Before Diagnosis: Physician-Inspired Structured Thinking for ECG Classification

A new framework called CardioThink has been proposed to enhance ECG classification by incorporating structured…

5/19/2026 · 3 min read · 36 views
HyperPersona: A Multi-Level Hypergraph Framework for Text-Based Automatic Personality Prediction
arXiv cs.AI

HyperPersona: A Multi-Level Hypergraph Framework for Text-Based Automatic Personality Prediction

The article presents HyperPersona, a novel framework for text-based automatic personality prediction. It utilizes a…

5/19/2026 · 3 min read · 31 views
CBT-Audio: Evaluating Audio Language Models for Patient-Side Distress Intensity Estimation in CBT Session Recordings
arXiv cs.AI

CBT-Audio: Evaluating Audio Language Models for Patient-Side Distress Intensity Estimation in CBT Session Recordings

The article discusses the development of CBT-Audio, a dataset designed to evaluate patient distress estimation from…

5/19/2026 · 3 min read · 31 views

How WeSearch handles this source

WeSearch's declared handling of arXiv cs.AI's content. Indexing, snippets, summaries, retrieval and training are separate questions — see the rights registry or read this source's machine-readable record.

Indexing: Allowed Snippet: Allowed AI summary: Limited Retrieval / RAG: Not asserted Model training: Not asserted Commercial reuse: Not permitted

More ai-research sources

Visit arXiv cs.AI directly →