WeSearch
Hub / Ai Research
ai-research · WeSearch

Ai Research news.

Page 6 of Ai Research headlines on WeSearch — deduped and updated continuously from 10+ editorial sources.

The Compressive Knowledge Graph Hypothesis: Which Graph Facts Matter for Scientific Hypothesis Generation?
arXiv cs.AI

The Compressive Knowledge Graph Hypothesis: Which Graph Facts Matter for Scientific Hypothesis Generation?

The article discusses the Compressive Knowledge Graph Hypothesis, which explores the significance of various graph…

5/27/2026 · 2 min read · 31 views
Query Symbolically or Retrieve Semantically? A Dataset and Method for Semi-Structured Question Answering
arXiv cs.AI

Query Symbolically or Retrieve Semantically? A Dataset and Method for Semi-Structured Question Answering

The article discusses a new framework called DualGraph designed for semi-structured question answering. It combines…

5/27/2026 · 3 min read · 34 views
Detecting Is Not Resolving: The Monitoring Control Gap in Retrieval Augmented LLMs
arXiv cs.AI

Detecting Is Not Resolving: The Monitoring Control Gap in Retrieval Augmented LLMs

The paper discusses the limitations of retrieval-augmented language models (LLMs) in handling contradictory evidence.…

5/27/2026 · 3 min read · 32 views
VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions
arXiv cs.AI

VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions

VitaBench 2.0 is a new benchmark designed to evaluate personalized and proactive agents in long-term user…

5/27/2026 · 3 min read · 26 views
StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning
arXiv cs.AI

StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning

The paper presents StepOPSD, a new framework for improving reinforcement learning in multi-turn agents. This framework…

5/27/2026 · 3 min read · 35 views
ICCU: In-Context Continual Unlearning via Pattern-Induced Refusal Rules
arXiv cs.AI

ICCU: In-Context Continual Unlearning via Pattern-Induced Refusal Rules

The paper introduces ICCU, a framework for in-context continual unlearning in machine learning. It addresses…

5/27/2026 · 2 min read · 31 views
Scaling, Benchmarking, and Reasoning of Vision-Language Agents for Mobile GUI Navigation
arXiv cs.AI

Scaling, Benchmarking, and Reasoning of Vision-Language Agents for Mobile GUI Navigation

The paper discusses advancements in Vision-Language Models (VLMs) for mobile GUI navigation. It introduces HyperTrack,…

5/27/2026 · 2 min read · 33 views
Position: AI Safety Requires Effective Controllability
arXiv cs.AI

Position: AI Safety Requires Effective Controllability

The paper discusses the importance of controllability in AI safety, arguing that alignment alone is insufficient. It…

5/27/2026 · 3 min read · 29 views
Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation
arXiv cs.AI

Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation

The paper presents a novel approach called Counteraction-Aware Multi-Teacher On-Policy Distillation (CaMOPD) aimed at…

5/27/2026 · 3 min read · 45 views
Can Broad Biomedical Knowledge be Contextualized into Scenario-Grounded Propositions?
arXiv cs.AI

Can Broad Biomedical Knowledge be Contextualized into Scenario-Grounded Propositions?

The article discusses a new framework called SCENE that aims to contextualize broad biomedical knowledge into…

5/27/2026 · 3 min read · 39 views
Traceable Knowledge Graph Reasoning Enables LLM-Assisted Decision Support for Industrial VOCs in the Steel Industry
arXiv cs.AI

Traceable Knowledge Graph Reasoning Enables LLM-Assisted Decision Support for Industrial VOCs in the Steel Industry

A new system called Chat-ISV has been developed to assist in decision-making regarding volatile organic compounds…

5/27/2026 · 3 min read · 38 views
BatteryMFormer: Multi-level Learning for Battery Degradation Trajectory Forecasting
arXiv cs.AI

BatteryMFormer: Multi-level Learning for Battery Degradation Trajectory Forecasting

The article presents BatteryMFormer, a novel approach for forecasting battery degradation trajectories. This method…

5/27/2026 · 2 min read · 35 views
Boosting Knowledge Graph Foundation Models via Enhanced Negative Sampling
arXiv cs.AI

Boosting Knowledge Graph Foundation Models via Enhanced Negative Sampling

The paper discusses an innovative approach to improve Knowledge Graph Foundation Models (KGFMs) through enhanced…

5/27/2026 · 3 min read · 33 views
ORCA: An End-to-End Interactive Copilot for Optimized Root Cause Analysis
arXiv cs.AI

ORCA: An End-to-End Interactive Copilot for Optimized Root Cause Analysis

The article introduces ORCA, an interactive copilot designed for optimized root cause analysis. It aims to make causal…

5/27/2026 · 2 min read · 37 views
Generating Robust Portfolios of Optimization Models using Large Language Models
arXiv cs.AI

Generating Robust Portfolios of Optimization Models using Large Language Models

The paper discusses a novel algorithm for generating robust portfolios of optimization models using large language…

5/27/2026 · 3 min read · 33 views
LELA: An End-to-end LLM-based Entity Linking Framework with Zero-shot Domain Adaptation
arXiv cs.AI

LELA: An End-to-end LLM-based Entity Linking Framework with Zero-shot Domain Adaptation

The paper introduces LELA, an end-to-end framework for entity linking that utilizes large language models (LLMs) and…

5/27/2026 · 2 min read · 32 views
Neuro-Symbolic Verification of LLM Outputs for Data-Sensitive Domains (extended preprint)
arXiv cs.AI

Neuro-Symbolic Verification of LLM Outputs for Data-Sensitive Domains (extended preprint)

The paper discusses a new verification architecture for large language models (LLMs) used in sensitive domains. It…

5/27/2026 · 3 min read · 41 views
Developing a Totally Unimodular Linear Program for Optimal Conformance Checking: When and Why It Complements A*
arXiv cs.AI

Developing a Totally Unimodular Linear Program for Optimal Conformance Checking: When and Why It Complements A*

A new paper introduces a totally unimodular linear program for optimal conformance checking, enhancing the traditional…

5/27/2026 · 3 min read · 35 views
From Norms to Indicators (N2I-RAG): An Agentic Retrieval-Augmented Generation Framework for Legal Indicator Computation
arXiv cs.AI

From Norms to Indicators (N2I-RAG): An Agentic Retrieval-Augmented Generation Framework for Legal Indicator Computation

The paper introduces N2I-RAG, a framework aimed at automating the computation of legal indicators from normative…

5/27/2026 · 3 min read · 34 views
TADDLE: A Tool-Augmented Agent for Detecting Deficient LLM-Generated Peer Reviews
arXiv cs.AI

TADDLE: A Tool-Augmented Agent for Detecting Deficient LLM-Generated Peer Reviews

The article introduces TADDLE, a tool designed to detect deficiencies in peer reviews generated by large language…

5/27/2026 · 2 min read · 29 views
On the Detection of Commutative Factors in Factor Graphs: Necessary and Sufficient Conditions
arXiv cs.AI

On the Detection of Commutative Factors in Factor Graphs: Necessary and Sufficient Conditions

The paper discusses the detection of commutative factors in factor graphs, which are essential for efficient…

5/27/2026 · 3 min read · 38 views
Multi-Stakeholder LLM Alignment: Decomposing Estimation from Aggregation
arXiv cs.AI

Multi-Stakeholder LLM Alignment: Decomposing Estimation from Aggregation

The paper discusses the challenges of aligning large language models (LLMs) with multiple stakeholders who have…

5/27/2026 · 2 min read · 38 views
Helicase: Uncertainty-Guided Supply Chain Knowledge Graph Construction with Autonomous Multi-Agent LLMs
arXiv cs.AI

Helicase: Uncertainty-Guided Supply Chain Knowledge Graph Construction with Autonomous Multi-Agent LLMs

The article introduces Helicase, an autonomous multi-agent LLM system designed for constructing supply chain knowledge…

5/27/2026 · 3 min read · 36 views
What Makes Chain-of-Thought Work at Probe Time? Local Co-occurrence Rather Than Global Derivation
arXiv cs.AI

What Makes Chain-of-Thought Work at Probe Time? Local Co-occurrence Rather Than Global Derivation

The paper explores the effectiveness of chain-of-thought (CoT) prompting in language models, focusing on probe-time…

5/27/2026 · 3 min read · 32 views
Composition Collapse: Stable Factual Knowledge Does Not Imply Compositional Reasoning
arXiv cs.AI

Composition Collapse: Stable Factual Knowledge Does Not Imply Compositional Reasoning

The paper discusses the concept of composition collapse in artificial intelligence, where stable factual knowledge…

5/27/2026 · 3 min read · 38 views
LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations?
arXiv cs.AI

LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations?

The paper introduces LiveK12Bench, a benchmark designed to evaluate the reasoning abilities of large multimodal models…

5/27/2026 · 3 min read · 33 views
The Attribution Blind Spot: Detecting When Language Models Rely on Memory Rather Than Retrieved Context
arXiv cs.AI

The Attribution Blind Spot: Detecting When Language Models Rely on Memory Rather Than Retrieved Context

The paper discusses the challenges of verifying whether language models rely on retrieved context or their internal…

5/27/2026 · 3 min read · 32 views
Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal
arXiv cs.AI

Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal

The paper discusses how chain-of-thought (CoT) in large reasoning models (LRMs) complicates the control of refusal…

5/27/2026 · 3 min read · 33 views
A Dataset of Robot-Patient and Doctor-Patient Medical Dialogues for Spoken Language Processing Tasks
arXiv cs.AI

A Dataset of Robot-Patient and Doctor-Patient Medical Dialogues for Spoken Language Processing Tasks

A new dataset named MeDial-Speech has been introduced to enhance spoken language processing in medical consultations.…

5/27/2026 · 3 min read · 39 views
It's Not the Capability: Harness Sensitivity Is Non-Monotone Across LLM Agent Tiers
arXiv cs.AI

It's Not the Capability: Harness Sensitivity Is Non-Monotone Across LLM Agent Tiers

A recent study challenges the assumption that higher-capability LLM models require less structural guidance. The…

5/27/2026 · 3 min read · 31 views
Towards Feedback-to-Plan Decisions for Self-Evolving LLM Agents in CUDA Kernel Generation
arXiv cs.AI

Towards Feedback-to-Plan Decisions for Self-Evolving LLM Agents in CUDA Kernel Generation

The paper discusses advancements in self-evolving large language models (LLMs) for CUDA kernel generation. It…

5/27/2026 · 3 min read · 46 views
Mind the Tool Failures: Achieving Synergistic Tool Gains for Medical Agents
arXiv cs.AI

Mind the Tool Failures: Achieving Synergistic Tool Gains for Medical Agents

The article discusses the challenges faced by medical AI agents when using external tools for diagnosis and treatment.…

5/27/2026 · 3 min read · 44 views
MemFail: Stress-Testing Failure Modes of LLM Memory Systems
arXiv cs.AI

MemFail: Stress-Testing Failure Modes of LLM Memory Systems

The paper introduces MemFail, a diagnostic benchmark designed to stress-test the failure modes of memory systems in…

5/27/2026 · 2 min read · 39 views
Completion vs Optimality: Policy Gradient in Long-Horizon Cumulative-Damage Problems
arXiv cs.AI

Completion vs Optimality: Policy Gradient in Long-Horizon Cumulative-Damage Problems

The paper discusses long-horizon decision problems characterized by cumulative damage and the challenges faced by…

5/27/2026 · 3 min read · 38 views
UnityMAS-O: A General RL Optimization Framework for LLM-Based Multi-Agent Systems
arXiv cs.AI

UnityMAS-O: A General RL Optimization Framework for LLM-Based Multi-Agent Systems

The article introduces UnityMAS-O, a general reinforcement learning optimization framework designed for large language…

5/27/2026 · 3 min read · 42 views
Tail-Aware HiFloat4: W4A4 Post-Training Quantization for Wan2.2
arXiv cs.AI

Tail-Aware HiFloat4: W4A4 Post-Training Quantization for Wan2.2

The article discusses a new method called Tail-Aware HiFloat4 for post-training quantization in low-bit text-to-video…

5/27/2026 · 2 min read · 32 views
FAST-GOAL: Fast and Efficient Global-local Object Alignment Learning
arXiv cs.AI

FAST-GOAL: Fast and Efficient Global-local Object Alignment Learning

The paper presents FAST-GOAL, a method designed to improve the performance of vision-language models like CLIP when…

5/27/2026 · 3 min read · 38 views
AGORA: Adapter-Grounded Observation-Action Retention for Inference-Free Prompt Compression in LLM Agents
arXiv cs.AI

AGORA: Adapter-Grounded Observation-Action Retention for Inference-Free Prompt Compression in LLM Agents

The article discusses a new method called AGORA for improving prompt compression in large language model (LLM) agents.…

5/27/2026 · 3 min read · 38 views
MedGuideX: Internalizing Decision Logic from Executable Guidelines into Large Language Models for Clinical Reasoning
arXiv cs.AI

MedGuideX: Internalizing Decision Logic from Executable Guidelines into Large Language Models for Clinical Reasoning

The article discusses MedGuideX, a new approach to integrating clinical practice guidelines into large language models…

5/27/2026 · 3 min read · 35 views
MobileExplorer: Accelerating On-Device Inference for Mobile GUI Agents via Online Exploration
arXiv cs.AI

MobileExplorer: Accelerating On-Device Inference for Mobile GUI Agents via Online Exploration

MobileExplorer is a new framework designed to enhance on-device inference for mobile GUI agents. It aims to reduce…

5/27/2026 · 3 min read · 36 views
PolyFusionAgent: A Multimodal Foundation Model and Autonomous AI Assistant for Polymer Property Prediction and Inverse Design
arXiv cs.AI

PolyFusionAgent: A Multimodal Foundation Model and Autonomous AI Assistant for Polymer Property Prediction and Inverse Design

PolyFusionAgent is a new multimodal foundation model designed to enhance polymer property prediction and inverse…

5/27/2026 · 3 min read · 36 views
Which Changes Matter? Towards Trustworthy Legal AI via Relevance-Sensitive Evaluation and Solver-Grounded Reasoning
arXiv cs.AI

Which Changes Matter? Towards Trustworthy Legal AI via Relevance-Sensitive Evaluation and Solver-Grounded Reasoning

The paper discusses the importance of distinguishing legally relevant changes in legal AI systems. It introduces a new…

5/27/2026 · 3 min read · 38 views
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
arXiv cs.AI

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

The MiniMax-M2 series introduces a new family of Mixture-of-Experts language models. These models leverage mini…

5/27/2026 · 5 min read · 36 views
Reasoning, Code, or Both? How Large Language Models Handle Variations in Math Questions
arXiv cs.AI

Reasoning, Code, or Both? How Large Language Models Handle Variations in Math Questions

A recent study evaluates how Large Language Models (LLMs) perform on mathematical reasoning tasks when faced with…

5/27/2026 · 3 min read · 37 views
From Static Context to Calibrated Interactive RL: Mitigating Distribution Shift in Multi-turn Dialogue with Aligned Simulator
arXiv cs.AI

From Static Context to Calibrated Interactive RL: Mitigating Distribution Shift in Multi-turn Dialogue with Aligned Simulator

The article discusses a new approach to improving dialogue agents through a method called Calibrated Interactive RL.…

5/27/2026 · 3 min read · 46 views
Advancing Creative Physical Intelligence in Large Multimodal Models
arXiv cs.AI

Advancing Creative Physical Intelligence in Large Multimodal Models

A new paper introduces MM-CreativityBench, a benchmark designed to evaluate creative problem-solving in large…

5/27/2026 · 3 min read · 46 views
Exploiting Local Dynamics Regularity for Reusable Skills in Offline Hierarchical RL
arXiv cs.AI

Exploiting Local Dynamics Regularity for Reusable Skills in Offline Hierarchical RL

The paper discusses advancements in Hierarchical Reinforcement Learning (HRL) by focusing on the reuse of skills…

5/27/2026 · 2 min read · 38 views
Automatic Layer Selection for Hallucination Detection
arXiv cs.AI

Automatic Layer Selection for Hallucination Detection

A new study proposes an automated method for selecting layers in large language models to improve hallucination…

5/27/2026 · 3 min read · 35 views
ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence
arXiv cs.AI

ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence

The paper titled 'ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence' presents a new…

5/27/2026 · 3 min read · 31 views
Managing Uncertainty in LLM-Generated Procedural Knowledge for Virtual Laboratory Planning
arXiv cs.AI

Managing Uncertainty in LLM-Generated Procedural Knowledge for Virtual Laboratory Planning

The paper discusses a framework for managing uncertainty in procedural knowledge generated by large language models…

5/27/2026 · 3 min read · 36 views

Sources in Ai Research

Other categories