LinAlg-Bench: A Forensic Benchmark Revealing Structural Failure Modes in LLM Mathematical Reasoning
LinAlg-Bench is a new diagnostic benchmark designed to evaluate large language models on linear algebra computations.…
Recent ai-research headlines from arXiv cs.AI.

LinAlg-Bench is a new diagnostic benchmark designed to evaluate large language models on linear algebra computations.…

The article discusses a new system called MetaKGEnrich that enhances metacognitive abilities in AI. This system…

The paper discusses the limitations of current personalized language systems, particularly in how they handle…

The article discusses a new framework called GRID for constructing security text knowledge graphs from cyber threat…

The paper titled 'Baba in Wonderland' explores online self-supervised dynamics discovery for executable world models.…

A new paper introduces the Global-Local Graph Attention Network (GLGAT) aimed at improving traffic forecasting. This…

The paper introduces PopuLoRA, a novel framework for reinforcement learning with large language models (LLMs). It…

The paper discusses a new architecture for body-grounded perspective formation in artificial agents. It introduces…

The paper discusses the issue of state contamination in memory-augmented LLM agents. It highlights a failure mode…

NeuroMAS introduces a novel approach to multi-agent language systems by treating them as neural networks with joint…

The paper presents a systematic analysis of multi-paradigm agent interaction within the buddyMe framework. It explores…

The paper titled 'Voices in the Loop: Mapping Participatory AI' discusses the organization of participatory approaches…

The paper presents a novel approach to optimizing Diffusion Multi-Modal Large Language Models (dMLLMs) using…

A new paper introduces the concept of Artificial Adaptive Intelligence (AAI), a stage between narrow and general…

The paper discusses a new approach to experience-driven learning in artificial intelligence. It emphasizes the need…

A recent study explores the reasoning gap between large reasoning models and base models in artificial intelligence.…

A new study presents a graph-based framework for brain tumor segmentation that addresses the common issue of missing…

The paper introduces N-gram Memory (NGM), a training-free memory module designed for large language models (LLMs). NGM…

The article introduces MM-ToolBench, a benchmark designed for evaluating task-oriented omni-modal tool-using agents.…

The article discusses a new approach to clinical prediction that moves from static risk assessments to dynamic…

A recent neuroimaging study investigates how humans process AI-generated hallucinations. The research reveals distinct…

The paper discusses the application of artificial intelligence in solving inverse partial differential equation (PDE)…

A recent study explores the prediction of brain vascular age using cerebral blood flow velocity and machine learning…

The paper introduces ConfSleepNet, a framework designed for reliable sleep stage classification by addressing…

The paper examines the performance of autonomous AI agents in supply chain management through the MIT Beer Game. It…

The article discusses a new framework for evidential information fusion based on a possibilistic structure. This…

The article discusses the development of PersonaArena, a dynamic simulation framework aimed at enhancing persona-level…

A new paper discusses the challenges of aligning large language models with the requirements of high-quality creative…

The paper introduces AnchorDiff, a novel framework for generating radiology reports using a topology-aware masked…

The article introduces RAGA, a new framework for autonomous knowledge graph construction and retrieval-augmented…

A new methodology for enhancing reasoning in Large Language Models (LLMs) has been proposed, focusing on the…

The paper presents a new algorithm called ECC for clustering queries based on their latent capability demands. This…

The paper presents a new framework for detecting fake news in Indian media by integrating visual and textual analysis.…

The paper presents a new framework for automated algorithm design called Latent Heuristic Search, which utilizes…

The study explores the dynamics of collective creativity in AI art competitions, specifically through the platform…

The article introduces MADP, a multi-agent pipeline designed to automate document processing in enterprise…

The paper explores the effectiveness of shallow neural network agents in mastering the card game Schnapsen. It…

The paper discusses the need for explicit provenance in agentic AI to enhance public trust and accountability. It…

The article introduces CAREBench, a new benchmark designed to evaluate the emotion understanding capabilities of large…

The paper introduces ChemVA, a framework designed to enhance Large Language Models' understanding of chemical reaction…

The article discusses a new framework called TIDE aimed at improving the understanding of argumentative essays. This…

The paper introduces QE-Catalytic-V2, a multimodal large language model designed for catalytic materials. This model…

CAM-Bench is a new benchmark designed for computational and applied mathematics within the Lean theorem-proving…

A recent study investigates the faithfulness of Vision-Language-Action (VLA) driving models. The research reveals…

The article introduces A2RBench, an automated system designed to generate benchmarks for evaluating abstract reasoning…

The article introduces MetaCogAgent, a multi-agent large language model framework designed to enhance task delegation…

The article introduces CyberCorrect, a framework designed for self-correction in large language models (LLMs). This…

A new framework called CardioThink has been proposed to enhance ECG classification by incorporating structured…

The article presents HyperPersona, a novel framework for text-based automatic personality prediction. It utilizes a…

The article discusses the development of CBT-Audio, a dataset designed to evaluate patient distress estimation from…
WeSearch's declared handling of arXiv cs.AI's content. Indexing, snippets, summaries, retrieval and training are separate questions — see the rights registry or read this source's machine-readable record.