When to Re-Plan: Subgoal Persistence in Hierarchical Latent Reasoning
The paper titled 'When to Re-Plan: Subgoal Persistence in Hierarchical Latent Reasoning' explores the balance between…
Page 4 of Ai Research headlines on WeSearch — deduped and updated continuously from 10+ editorial sources.

The paper titled 'When to Re-Plan: Subgoal Persistence in Hierarchical Latent Reasoning' explores the balance between…

The paper introduces derivation graphs to enhance the understanding of do-calculus reasoning. These graphs help in…

The paper introduces Code-on-Graph (CoG), a new framework for integrating Large Language Models (LLMs) with Knowledge…

The paper introduces Dynamic Objective Selection with Safeguards (DOSS) for financial decision-making. DOSS aims to…

The article introduces SkillPyramid, a framework designed to enhance the skill consolidation of self-evolving AI…

The DeepSpeak-Agentic dataset consists of over 37 hours of semi-structured conversations between humans and AI agents.…

EvoDrive is a new framework designed for generating safety-critical scenarios in autonomous driving systems. It…

The paper introduces ChemCoTBench-V2, a benchmark designed for evaluating chemical reasoning in large language models.…

The paper introduces NovelAPIBench, a dynamic benchmark designed to evaluate large language models' ability to use…

The paper discusses advancements in propositional defeasible standpoint logic, focusing on non-monotonic entailment.…

A recent study investigates gender-dependent disparities in medical triage recommendations made by large language…

The paper introduces TSQAgent, a framework designed to improve the assessment of time series data quality using large…

The paper introduces a framework to improve instruction following in Large Reasoning Models (LRMs) by addressing the…

The paper discusses a new approach to optimize coding agents by reducing input-token costs. It introduces a middleware…

The paper introduces an SLM-based Agent Orchestration Gateway designed for AI-driven virtual worlds. This gateway…

The paper introduces SAGE, a framework for evaluating socialized evolution in agent ecosystems. It compares two…

The paper presents a new compositional authorization framework for managing delegation and scope in agentic AI…

The paper introduces ThoughtFold, a framework designed to improve the efficiency of Large Reasoning Models (LRMs) by…

The paper presents a formal definition and meta-model for the Machine Theory of Mind. It integrates insights from…

StepFinder is a new framework designed for failure attribution in multi-agent systems. It aims to improve the…

The Deterministic Memory Framework (DMF) aims to enhance memory systems for conversational AI agents. It replaces…

The paper investigates the effectiveness of interaction trajectories in training terminal agents. It reveals that…

The article introduces CP-Agent, a multimodal large language model designed for cellular morphological profiling under…

The paper introduces InfoMem, a new reward mechanism designed for training long-context memory agents in artificial…

The Violation Situation Pattern (VSP) is a new knowledge-graph pattern designed to improve compliance violation…

The article discusses the challenges of benchmark auditing in artificial intelligence, particularly regarding…

The article discusses LEAP, a new framework designed to enhance the capabilities of Large Language Models (LLMs) in…

The paper presents a negative result regarding cross-model activation transfer in a multi-hop reasoning setting using…

The article discusses a novel approach for enhancing Visual Question Answering (VQA) by distilling rules from Large…

The study investigates whether real-world datasets contain natural experiments, which are implicit interventions…

A recent paper argues that superintelligence developed from a solipsistic approach to AI design is unlikely to be…

The article discusses a new framework called the Pre-Reasoning Perception Framework (PRPF) designed to enhance…

The study investigates the impact of demographic bias on skin lesion classification using ResNet-based models. It…

MedCUA-Bench is a newly introduced benchmark designed specifically for clinical computer-use agents. It aims to…

The article introduces ClinicalMC, a benchmark designed for evaluating large language models in multi-course clinical…

The article introduces GTBench, a benchmark designed to evaluate large language models (LLMs) as mathematical research…

The paper introduces a new framework called Think-Before-Speak (TBS) for multi-agent social simulation. TBS separates…

The article discusses a new framework for Large Language Model (LLM) agents that aims to improve their performance in…

EvoTrainer is a new autonomous training framework designed for co-evolving LLM policies and training harnesses. It…

The paper titled 'DeskCraft' introduces a new benchmark for evaluating desktop agents in professional workflows that…

A new framework has been developed to enhance time series forecasting by integrating news articles. This approach…

The paper titled 'Decomposing how prompting steers behavior' explores how prompting influences the internal…

The article discusses a new approach to budget allocation for Large Language Models (LLMs) based on economic…

The paper introduces DeltaMem, a framework designed to enhance memory management in Large Language Model (LLM) agents.…

The article discusses a new framework called CORE, which stands for Conflict-Oriented Reasoning, designed to enhance…

The paper introduces SkillDAG, a novel approach for selecting skills in large language models (LLMs) by modeling…

The paper introduces ToolGate, a system designed to improve the efficiency of tool-augmented vision-language agents.…

The paper introduces RelGT-AC, a new model designed for autocomplete tasks in relational databases. It enhances the…

TriEval is a new pipeline designed to assess bias, toxicity, and truthfulness in large language models (LLMs)…

The article discusses AuditFlow, a new framework designed for structured financial reporting verification. It utilizes…