From Prompt to Service: An SLM-Based Agent Orchestration Gateway for AI-Driven Virtual Worlds
The paper introduces an SLM-based Agent Orchestration Gateway designed for AI-driven virtual worlds. This gateway…
Page 5 of Ai Research headlines on WeSearch — deduped and updated continuously from 10+ editorial sources.

The paper introduces an SLM-based Agent Orchestration Gateway designed for AI-driven virtual worlds. This gateway…

The paper discusses a new approach to optimize coding agents by reducing input-token costs. It introduces a middleware…

The paper introduces a framework to improve instruction following in Large Reasoning Models (LRMs) by addressing the…

The paper introduces TSQAgent, a framework designed to improve the assessment of time series data quality using large…

A recent study investigates gender-dependent disparities in medical triage recommendations made by large language…

The paper discusses advancements in propositional defeasible standpoint logic, focusing on non-monotonic entailment.…

The paper introduces NovelAPIBench, a dynamic benchmark designed to evaluate large language models' ability to use…

The paper introduces ChemCoTBench-V2, a benchmark designed for evaluating chemical reasoning in large language models.…

EvoDrive is a new framework designed for generating safety-critical scenarios in autonomous driving systems. It…

The DeepSpeak-Agentic dataset consists of over 37 hours of semi-structured conversations between humans and AI agents.…

The article introduces SkillPyramid, a framework designed to enhance the skill consolidation of self-evolving AI…

The paper introduces Dynamic Objective Selection with Safeguards (DOSS) for financial decision-making. DOSS aims to…

The paper introduces Code-on-Graph (CoG), a new framework for integrating Large Language Models (LLMs) with Knowledge…

The paper introduces derivation graphs to enhance the understanding of do-calculus reasoning. These graphs help in…

The paper titled 'When to Re-Plan: Subgoal Persistence in Hierarchical Latent Reasoning' explores the balance between…

The paper titled 'Proof-Refactor' addresses the challenges in generating formal proofs using Large Language Models…

The article introduces the Lab Agent Protocol (LAP), designed to enhance the interaction between autonomous agents and…
Comprehensive up-to-date news coverage, aggregated from sources all over the world by Google News.

The article presents BrickAnything, a new framework for generating buildable brick structures from 3D shapes. This…

The paper titled 'Can LLMs Introspect? A Reality Check' questions the ability of large language models (LLMs) to…

The paper discusses the need for persistent memory in long-running AI agents. It critiques current memory systems and…

The paper presents POLAR, a framework designed for personalizing embodied multimodal large language model agents…

The paper discusses the need for improved benchmarks in Constraint Acquisition (CA) research. Current benchmarks are…

The paper discusses the aging of AI agents deployed in operational systems and introduces a new benchmark called…

The paper discusses two innovative frameworks for creating autonomous AI systems to enhance scientific workflows.…

The paper introduces Anchor, a task-generation pipeline designed to address artifact drift in AI agent benchmark…

The paper introduces OmniToM, a benchmark designed to evaluate the Theory of Mind capabilities in large language…

The paper introduces JobBench, a new benchmark for evaluating AI agents based on human needs rather than economic…

The paper discusses a framework for managing uncertainty in procedural knowledge generated by large language models…

The paper titled 'ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence' presents a new…

A new study proposes an automated method for selecting layers in large language models to improve hallucination…

The paper discusses advancements in Hierarchical Reinforcement Learning (HRL) by focusing on the reuse of skills…

A new paper introduces MM-CreativityBench, a benchmark designed to evaluate creative problem-solving in large…

The article discusses a new approach to improving dialogue agents through a method called Calibrated Interactive RL.…

A recent study evaluates how Large Language Models (LLMs) perform on mathematical reasoning tasks when faced with…

The MiniMax-M2 series introduces a new family of Mixture-of-Experts language models. These models leverage mini…

The paper discusses the importance of distinguishing legally relevant changes in legal AI systems. It introduces a new…

PolyFusionAgent is a new multimodal foundation model designed to enhance polymer property prediction and inverse…

MobileExplorer is a new framework designed to enhance on-device inference for mobile GUI agents. It aims to reduce…

The article discusses MedGuideX, a new approach to integrating clinical practice guidelines into large language models…

The article discusses a new method called AGORA for improving prompt compression in large language model (LLM) agents.…

The paper presents FAST-GOAL, a method designed to improve the performance of vision-language models like CLIP when…

The article discusses a new method called Tail-Aware HiFloat4 for post-training quantization in low-bit text-to-video…

The article introduces UnityMAS-O, a general reinforcement learning optimization framework designed for large language…

The paper discusses long-horizon decision problems characterized by cumulative damage and the challenges faced by…

The paper introduces MemFail, a diagnostic benchmark designed to stress-test the failure modes of memory systems in…

The article discusses the challenges faced by medical AI agents when using external tools for diagnosis and treatment.…

The paper discusses advancements in self-evolving large language models (LLMs) for CUDA kernel generation. It…

A recent study challenges the assumption that higher-capability LLM models require less structural guidance. The…

A new dataset named MeDial-Speech has been introduced to enhance spoken language processing in medical consultations.…