GraphRAG on Consumer Hardware: Benchmarking Local LLMs for Healthcare EHR Schema Retrieval
The paper evaluates GraphRAG, a graph-based retrieval-augmented generation model, for healthcare EHR schema retrieval…
Recent ai-research headlines from arXiv cs.AI.

The paper evaluates GraphRAG, a graph-based retrieval-augmented generation model, for healthcare EHR schema retrieval…

The paper introduces ArchSIBench, a benchmark designed to evaluate the architectural spatial intelligence of…

The paper titled 'USV: Towards Understanding the User-generated Short-form Videos' introduces a new dataset aimed at…

The paper titled 'DISC: Decoupling Instruction from State-Conditioned Control via Policy Generation' introduces a new…

A new position paper advocates for the development of data probes to better understand how data influences large…

The article discusses a microservice architecture designed for operationalizing Document AI, focusing on OCR and large…

The study evaluates the effectiveness of Personal Health Records (PHRs) in enhancing responses from large language…

The paper introduces Learn-by-Wire Guard (LBW-Guard), a new training-control governance layer for language models.…

The article discusses a new multi-agent method for converting natural language to SQL, known as AgentNLQ. This method…

The article discusses a study on Kolmogorov-Arnold Networks (KANs) and their application in improving human activity…

The paper discusses the importance of integrating trust into Agent-to-Agent (A2A) networks from the outset rather than…

The paper introduces a framework for multi-task unlearning in machine learning, addressing the challenges of removing…

The paper discusses a new framework called ReElicit for optimizing system prompts in AI using Bayesian methods. It…

The paper introduces DecisionBench, a benchmark for emergent delegation in long-horizon agentic workflows. It…

The article introduces POLAR-Bench, a diagnostic benchmark designed to evaluate privacy-utility trade-offs in large…

The paper discusses a new approach to workflow learning in multi-agent systems where agents hand off control through a…

The paper discusses a formalization of trust calibration for automated agents in tool use. It presents a method for…

Recent advancements in auto-research systems have enabled the generation of complete research papers. However, a study…

The article presents a formal framework for understanding agentic knowledge graph (KG) affordances. It critiques…

The paper discusses the concept of hallucination in multimodal agents, where false visual claims can lead to…

The paper discusses the differences between volatility and stochasticity in the context of adaptive decision-making in…

SimGym is a new framework designed to simulate A/B tests in e-commerce using vision-language model agents. It aims to…

The article discusses the potential of large language models (LLMs) to address challenges in survey research,…

The paper presents findings on modality-conflict hallucination in multimodal large language models (MLLMs). It…

The paper introduces AQuaUI, a novel method for visual token reduction in GUI agents using adaptive quadtrees. This…

The paper titled 'Swimming with Whales: Analysis of Power Imbalances in Stake-Weighted Governance' explores the…

The paper introduces MOCHA, a new method for optimizing agent skills in artificial intelligence. MOCHA employs…

The paper titled 'Agentic Trading: When LLM Agents Meet Financial Markets' explores the integration of Large Language…

The paper introduces Generative Recursive Reasoning Models (GRAM), a new framework for neural reasoning systems. GRAM…

The article introduces PRISM, a new benchmark for evaluating programmatic spatial-temporal reasoning in video…

The paper introduces SIGMA, a new framework for multi-agent reasoning that addresses the limitations of existing…

The paper discusses a new framework called SERL for improving reinforcement learning in multi-turn agents. It focuses…

The paper presents GUIDE, a new framework for automated bidding in digital advertising. It integrates directed…

The paper discusses a new approach to on-policy reinforcement learning that addresses the issue of mode collapse. The…

The paper discusses a novel approach to jailbreak attacks on Large Reasoning Models (LRMs) using reinforcement…

The paper discusses the Turing-completeness of autoregressive Transformers, emphasizing the importance of context…

The paper introduces BLINKG, a benchmark for evaluating the capabilities of Large Language Models (LLMs) in generating…

The paper titled 'Efficient Elicitation of Collective Disagreements' explores how to analyze voter disagreements over…

The paper introduces Generative-Evaluative Agreement (GEA) as a validity criterion for assessing LLM-enabled adaptive…

The paper discusses a phenomenon called 'library drift' in self-evolving LLM skill libraries, which leads to…

The paper introduces SceneCode, a framework designed for generating editable indoor scenes with articulated objects.…

The paper discusses the challenges of scheduling multiple Large Language Models (LLMs) on shared hardware. It…

The paper introduces Formal Skill, a new abstraction for enhancing the efficiency and accuracy of Large Language Model…

The paper presents EMO-BOOST, a new framework aimed at improving deepfake detection by integrating emotion…

The paper discusses the challenges faced by tabular foundation models in strategic data environments. It introduces a…

The paper presents a new framework called Pseudocode-guided Structured Reasoning (PStar) aimed at improving the…

The paper discusses the transformation of constraint programs into input for local search algorithms. It highlights…

The paper titled 'Beyond Rational Illusion: Behaviorally Realistic Strategic Classification' introduces a new…

The paper introduces a novel approach to graph combinatorial optimization using reinforcement learning. It addresses…

EngiAI introduces a multi-agent framework and benchmark suite designed for LLM-driven engineering design tasks. The…
WeSearch's declared handling of arXiv cs.AI's content. Indexing, snippets, summaries, retrieval and training are separate questions — see the rights registry or read this source's machine-readable record.