WeSearch
Hub / Ai Research
ai-research · WeSearch

Ai Research news.

Page 5 of Ai Research headlines on WeSearch — deduped and updated continuously from 10+ editorial sources.

From Prompt to Service: An SLM-Based Agent Orchestration Gateway for AI-Driven Virtual Worlds
arXiv cs.AI

From Prompt to Service: An SLM-Based Agent Orchestration Gateway for AI-Driven Virtual Worlds

The paper introduces an SLM-based Agent Orchestration Gateway designed for AI-driven virtual worlds. This gateway…

6/2/2026 · 3 min read · 49 views
Cross-Lingual Token Arbitrage: Optimizing Code Agent Context Windows via Local LLM Preprocessing
arXiv cs.AI

Cross-Lingual Token Arbitrage: Optimizing Code Agent Context Windows via Local LLM Preprocessing

The paper discusses a new approach to optimize coding agents by reducing input-token costs. It introduces a middleware…

6/2/2026 · 3 min read · 51 views
Bridging Auxiliary Constraints to Resolve Instruction Following in Large Reasoning Models
arXiv cs.AI

Bridging Auxiliary Constraints to Resolve Instruction Following in Large Reasoning Models

The paper introduces a framework to improve instruction following in Large Reasoning Models (LRMs) by addressing the…

6/2/2026 · 3 min read · 59 views
TSQAgent: Rating Time Series Data Quality via Dedicated Agentic Reasoning
arXiv cs.AI

TSQAgent: Rating Time Series Data Quality via Dedicated Agentic Reasoning

The paper introduces TSQAgent, a framework designed to improve the assessment of time series data quality using large…

6/2/2026 · 3 min read · 58 views
Gender-Dependent Diagnostic Substitution in LLM Medical Triage: Same Symptoms, Unequal Urgency
arXiv cs.AI

Gender-Dependent Diagnostic Substitution in LLM Medical Triage: Same Symptoms, Unequal Urgency

A recent study investigates gender-dependent disparities in medical triage recommendations made by large language…

6/2/2026 · 3 min read · 49 views
Towards Non-Monotonic Entailment in Propositional Defeasible Standpoint Logic
arXiv cs.AI

Towards Non-Monotonic Entailment in Propositional Defeasible Standpoint Logic

The paper discusses advancements in propositional defeasible standpoint logic, focusing on non-monotonic entailment.…

6/2/2026 · 3 min read · 46 views
Diagnosing Knowledge Gaps in LLM Tool Use: An Agentic Benchmark for Novel API Acquisition
arXiv cs.AI

Diagnosing Knowledge Gaps in LLM Tool Use: An Agentic Benchmark for Novel API Acquisition

The paper introduces NovelAPIBench, a dynamic benchmark designed to evaluate large language models' ability to use…

6/2/2026 · 3 min read · 55 views
From Answers to States: Verifiable Process-Level Evaluation of Chemical Reasoning in Large Language Models
arXiv cs.AI

From Answers to States: Verifiable Process-Level Evaluation of Chemical Reasoning in Large Language Models

The paper introduces ChemCoTBench-V2, a benchmark designed for evaluating chemical reasoning in large language models.…

6/2/2026 · 3 min read · 54 views
EvoDrive: Pareto Evolution for Safety-Critical Autonomous Driving via Self-Improving LLM Agents
arXiv cs.AI

EvoDrive: Pareto Evolution for Safety-Critical Autonomous Driving via Self-Improving LLM Agents

EvoDrive is a new framework designed for generating safety-critical scenarios in autonomous driving systems. It…

6/2/2026 · 3 min read · 55 views
The DeepSpeak-Agentic Dataset
arXiv cs.AI

The DeepSpeak-Agentic Dataset

The DeepSpeak-Agentic dataset consists of over 37 hours of semi-structured conversations between humans and AI agents.…

6/2/2026 · 2 min read · 43 views
SkillPyramid: A Hierarchical Skill Consolidation Framework for Self-Evolving Agents
arXiv cs.AI

SkillPyramid: A Hierarchical Skill Consolidation Framework for Self-Evolving Agents

The article introduces SkillPyramid, a framework designed to enhance the skill consolidation of self-evolving AI…

6/2/2026 · 2 min read · 69 views
Dynamic Objective Selection with Safeguards and LLM Oversight for Financial Decision-Making
arXiv cs.AI

Dynamic Objective Selection with Safeguards and LLM Oversight for Financial Decision-Making

The paper introduces Dynamic Objective Selection with Safeguards (DOSS) for financial decision-making. DOSS aims to…

6/2/2026 · 3 min read · 66 views
Code-on-Graph: Iterative Programmatic Reasoning via Large Language Models on Knowledge Graphs
arXiv cs.AI

Code-on-Graph: Iterative Programmatic Reasoning via Large Language Models on Knowledge Graphs

The paper introduces Code-on-Graph (CoG), a new framework for integrating Large Language Models (LLMs) with Knowledge…

6/2/2026 · 3 min read · 71 views
Unveiling the Structure of Do-Calculus Reasoning via Derivation Graphs
arXiv cs.AI

Unveiling the Structure of Do-Calculus Reasoning via Derivation Graphs

The paper introduces derivation graphs to enhance the understanding of do-calculus reasoning. These graphs help in…

6/2/2026 · 2 min read · 78 views
When to Re-Plan: Subgoal Persistence in Hierarchical Latent Reasoning
arXiv cs.AI

When to Re-Plan: Subgoal Persistence in Hierarchical Latent Reasoning

The paper titled 'When to Re-Plan: Subgoal Persistence in Hierarchical Latent Reasoning' explores the balance between…

6/2/2026 · 3 min read · 61 views
Proof-Refactor: Refactoring Generated Formal Proofs into Modular Artifacts
arXiv cs.AI

Proof-Refactor: Refactoring Generated Formal Proofs into Modular Artifacts

The paper titled 'Proof-Refactor' addresses the challenges in generating formal proofs using Large Language Models…

6/2/2026 · 3 min read · 70 views
LAP: An Agent-to-Instrument Protocol for Autonomous Science
arXiv cs.AI

LAP: An Agent-to-Instrument Protocol for Autonomous Science

The article introduces the Lab Agent Protocol (LAP), designed to enhance the interaction between autonomous agents and…

6/2/2026 · 3 min read · 68 views
Google News

How AI is Transforming Scientific Discovery While Keeping Humans at the Center - Stanford HAI

Comprehensive up-to-date news coverage, aggregated from sources all over the world by Google News.

5/27/2026 · 52 views
BrickAnything: Geometry-Conditioned Buildable Brick Generation with Structure-Aware Tokenization
arXiv cs.AI

BrickAnything: Geometry-Conditioned Buildable Brick Generation with Structure-Aware Tokenization

The article presents BrickAnything, a new framework for generating buildable brick structures from 3D shapes. This…

5/26/2026 · 3 min read · 58 views
Can LLMs Introspect? A Reality Check
arXiv cs.AI

Can LLMs Introspect? A Reality Check

The paper titled 'Can LLMs Introspect? A Reality Check' questions the ability of large language models (LLMs) to…

5/26/2026 · 3 min read · 55 views
Is Agent Memory a Database? Rethinking Data Foundations for Long-Term AI Agent Memory
arXiv cs.AI

Is Agent Memory a Database? Rethinking Data Foundations for Long-Term AI Agent Memory

The paper discusses the need for persistent memory in long-running AI agents. It critiques current memory systems and…

5/26/2026 · 3 min read · 65 views
Personalizing Embodied Multimodal Large Language Model Agents over Long-term User Interactions
arXiv cs.AI

Personalizing Embodied Multimodal Large Language Model Agents over Long-term User Interactions

The paper presents POLAR, a framework designed for personalizing embodied multimodal large language model agents…

5/26/2026 · 3 min read · 52 views
Constraint acquisition needs better benchmarks
arXiv cs.AI

Constraint acquisition needs better benchmarks

The paper discusses the need for improved benchmarks in Constraint Acquisition (CA) research. Current benchmarks are…

5/26/2026 · 2 min read · 50 views
Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems
arXiv cs.AI

Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems

The paper discusses the aging of AI agents deployed in operational systems and introduces a new benchmark called…

5/26/2026 · 3 min read · 51 views
Experiments in Agentic AI for Science
arXiv cs.AI

Experiments in Agentic AI for Science

The paper discusses two innovative frameworks for creating autonomous AI systems to enhance scientific workflows.…

5/26/2026 · 2 min read · 51 views
Anchor: Mitigating Artifact Drift in Agent Benchmark Generation
arXiv cs.AI

Anchor: Mitigating Artifact Drift in Agent Benchmark Generation

The paper introduces Anchor, a task-generation pipeline designed to address artifact drift in AI agent benchmark…

5/26/2026 · 3 min read · 55 views
OmniToM: Benchmarking Theory of Mind in LLMs via Explicit Belief Modeling
arXiv cs.AI

OmniToM: Benchmarking Theory of Mind in LLMs via Explicit Belief Modeling

The paper introduces OmniToM, a benchmark designed to evaluate the Theory of Mind capabilities in large language…

5/26/2026 · 3 min read · 53 views
JobBench: Aligning Agent Work With Human Will
arXiv cs.AI

JobBench: Aligning Agent Work With Human Will

The paper introduces JobBench, a new benchmark for evaluating AI agents based on human needs rather than economic…

5/26/2026 · 3 min read · 64 views
Managing Uncertainty in LLM-Generated Procedural Knowledge for Virtual Laboratory Planning
arXiv cs.AI

Managing Uncertainty in LLM-Generated Procedural Knowledge for Virtual Laboratory Planning

The paper discusses a framework for managing uncertainty in procedural knowledge generated by large language models…

5/26/2026 · 3 min read · 46 views
ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence
arXiv cs.AI

ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence

The paper titled 'ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence' presents a new…

5/26/2026 · 3 min read · 45 views
Automatic Layer Selection for Hallucination Detection
arXiv cs.AI

Automatic Layer Selection for Hallucination Detection

A new study proposes an automated method for selecting layers in large language models to improve hallucination…

5/26/2026 · 3 min read · 41 views
Exploiting Local Dynamics Regularity for Reusable Skills in Offline Hierarchical RL
arXiv cs.AI

Exploiting Local Dynamics Regularity for Reusable Skills in Offline Hierarchical RL

The paper discusses advancements in Hierarchical Reinforcement Learning (HRL) by focusing on the reuse of skills…

5/26/2026 · 2 min read · 48 views
Advancing Creative Physical Intelligence in Large Multimodal Models
arXiv cs.AI

Advancing Creative Physical Intelligence in Large Multimodal Models

A new paper introduces MM-CreativityBench, a benchmark designed to evaluate creative problem-solving in large…

5/26/2026 · 3 min read · 56 views
From Static Context to Calibrated Interactive RL: Mitigating Distribution Shift in Multi-turn Dialogue with Aligned Simulator
arXiv cs.AI

From Static Context to Calibrated Interactive RL: Mitigating Distribution Shift in Multi-turn Dialogue with Aligned Simulator

The article discusses a new approach to improving dialogue agents through a method called Calibrated Interactive RL.…

5/26/2026 · 3 min read · 58 views
Reasoning, Code, or Both? How Large Language Models Handle Variations in Math Questions
arXiv cs.AI

Reasoning, Code, or Both? How Large Language Models Handle Variations in Math Questions

A recent study evaluates how Large Language Models (LLMs) perform on mathematical reasoning tasks when faced with…

5/26/2026 · 3 min read · 48 views
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
arXiv cs.AI

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

The MiniMax-M2 series introduces a new family of Mixture-of-Experts language models. These models leverage mini…

5/26/2026 · 5 min read · 50 views
Which Changes Matter? Towards Trustworthy Legal AI via Relevance-Sensitive Evaluation and Solver-Grounded Reasoning
arXiv cs.AI

Which Changes Matter? Towards Trustworthy Legal AI via Relevance-Sensitive Evaluation and Solver-Grounded Reasoning

The paper discusses the importance of distinguishing legally relevant changes in legal AI systems. It introduces a new…

5/26/2026 · 3 min read · 54 views
PolyFusionAgent: A Multimodal Foundation Model and Autonomous AI Assistant for Polymer Property Prediction and Inverse Design
arXiv cs.AI

PolyFusionAgent: A Multimodal Foundation Model and Autonomous AI Assistant for Polymer Property Prediction and Inverse Design

PolyFusionAgent is a new multimodal foundation model designed to enhance polymer property prediction and inverse…

5/26/2026 · 3 min read · 45 views
MobileExplorer: Accelerating On-Device Inference for Mobile GUI Agents via Online Exploration
arXiv cs.AI

MobileExplorer: Accelerating On-Device Inference for Mobile GUI Agents via Online Exploration

MobileExplorer is a new framework designed to enhance on-device inference for mobile GUI agents. It aims to reduce…

5/26/2026 · 3 min read · 51 views
MedGuideX: Internalizing Decision Logic from Executable Guidelines into Large Language Models for Clinical Reasoning
arXiv cs.AI

MedGuideX: Internalizing Decision Logic from Executable Guidelines into Large Language Models for Clinical Reasoning

The article discusses MedGuideX, a new approach to integrating clinical practice guidelines into large language models…

5/26/2026 · 3 min read · 50 views
AGORA: Adapter-Grounded Observation-Action Retention for Inference-Free Prompt Compression in LLM Agents
arXiv cs.AI

AGORA: Adapter-Grounded Observation-Action Retention for Inference-Free Prompt Compression in LLM Agents

The article discusses a new method called AGORA for improving prompt compression in large language model (LLM) agents.…

5/26/2026 · 3 min read · 49 views
FAST-GOAL: Fast and Efficient Global-local Object Alignment Learning
arXiv cs.AI

FAST-GOAL: Fast and Efficient Global-local Object Alignment Learning

The paper presents FAST-GOAL, a method designed to improve the performance of vision-language models like CLIP when…

5/26/2026 · 3 min read · 56 views
Tail-Aware HiFloat4: W4A4 Post-Training Quantization for Wan2.2
arXiv cs.AI

Tail-Aware HiFloat4: W4A4 Post-Training Quantization for Wan2.2

The article discusses a new method called Tail-Aware HiFloat4 for post-training quantization in low-bit text-to-video…

5/26/2026 · 2 min read · 41 views
UnityMAS-O: A General RL Optimization Framework for LLM-Based Multi-Agent Systems
arXiv cs.AI

UnityMAS-O: A General RL Optimization Framework for LLM-Based Multi-Agent Systems

The article introduces UnityMAS-O, a general reinforcement learning optimization framework designed for large language…

5/26/2026 · 3 min read · 55 views
Completion vs Optimality: Policy Gradient in Long-Horizon Cumulative-Damage Problems
arXiv cs.AI

Completion vs Optimality: Policy Gradient in Long-Horizon Cumulative-Damage Problems

The paper discusses long-horizon decision problems characterized by cumulative damage and the challenges faced by…

5/26/2026 · 3 min read · 49 views
MemFail: Stress-Testing Failure Modes of LLM Memory Systems
arXiv cs.AI

MemFail: Stress-Testing Failure Modes of LLM Memory Systems

The paper introduces MemFail, a diagnostic benchmark designed to stress-test the failure modes of memory systems in…

5/26/2026 · 2 min read · 49 views
Mind the Tool Failures: Achieving Synergistic Tool Gains for Medical Agents
arXiv cs.AI

Mind the Tool Failures: Achieving Synergistic Tool Gains for Medical Agents

The article discusses the challenges faced by medical AI agents when using external tools for diagnosis and treatment.…

5/26/2026 · 3 min read · 54 views
Towards Feedback-to-Plan Decisions for Self-Evolving LLM Agents in CUDA Kernel Generation
arXiv cs.AI

Towards Feedback-to-Plan Decisions for Self-Evolving LLM Agents in CUDA Kernel Generation

The paper discusses advancements in self-evolving large language models (LLMs) for CUDA kernel generation. It…

5/26/2026 · 3 min read · 62 views
It's Not the Capability: Harness Sensitivity Is Non-Monotone Across LLM Agent Tiers
arXiv cs.AI

It's Not the Capability: Harness Sensitivity Is Non-Monotone Across LLM Agent Tiers

A recent study challenges the assumption that higher-capability LLM models require less structural guidance. The…

5/26/2026 · 3 min read · 52 views
A Dataset of Robot-Patient and Doctor-Patient Medical Dialogues for Spoken Language Processing Tasks
arXiv cs.AI

A Dataset of Robot-Patient and Doctor-Patient Medical Dialogues for Spoken Language Processing Tasks

A new dataset named MeDial-Speech has been introduced to enhance spoken language processing in medical consultations.…

5/26/2026 · 3 min read · 51 views

Sources in Ai Research

Other categories