WeSearch
Hub / Ai Research
ai-research · WeSearch

Ai Research news.

Page 8 of Ai Research headlines on WeSearch — deduped and updated continuously from 10+ editorial sources.

GRAIL: AI translation for scientists application workflow on satellite data
arXiv cs.AI

GRAIL: AI translation for scientists application workflow on satellite data

The paper introduces GRAIL, an AI translation system designed to assist scientists in converting Python geospatial…

5/26/2026 · 2 min read · 44 views
PRIMA: Operational Patterns for Resilient Multi-Agent Research with Verifiable Identity and Convergent Feedback
arXiv cs.AI

PRIMA: Operational Patterns for Resilient Multi-Agent Research with Verifiable Identity and Convergent Feedback

The paper titled 'PRIMA: Operational Patterns for Resilient Multi-Agent Research with Verifiable Identity and…

5/26/2026 · 3 min read · 27 views
Uncertainty Decomposition via Cyclical SG-MCMC and Soft-label Learning for Subjective NLP
arXiv cs.AI

Uncertainty Decomposition via Cyclical SG-MCMC and Soft-label Learning for Subjective NLP

The paper discusses a novel approach to uncertainty decomposition in subjective natural language processing (NLP). It…

5/26/2026 · 2 min read · 34 views
Proper Scoring Rules for Agentic Uncertainty Quantification
arXiv cs.AI

Proper Scoring Rules for Agentic Uncertainty Quantification

The paper introduces the Trajectory Proper Score (TPS) for evaluating agentic uncertainty quantification in AI. It…

5/26/2026 · 3 min read · 23 views
Automated Detection and Classification of Delusion-related Content in Naturalistic Audio Diaries Using Multi-Agent Language Models
arXiv cs.AI

Automated Detection and Classification of Delusion-related Content in Naturalistic Audio Diaries Using Multi-Agent Language Models

A new study presents an automated pipeline using multi-agent language models to detect and classify delusion-related…

5/26/2026 · 3 min read · 31 views
Hylos: Operability Contracts for Model-Native Spatial Intelligence
arXiv cs.AI

Hylos: Operability Contracts for Model-Native Spatial Intelligence

The paper titled 'Hylos: Operability Contracts for Model-Native Spatial Intelligence' introduces a new systems…

5/26/2026 · 3 min read · 26 views
Fundamental Limitation in Explaining AI
arXiv cs.AI

Fundamental Limitation in Explaining AI

A recent paper discusses the inherent limitations in explaining AI systems, particularly large-scale models. The…

5/26/2026 · 3 min read · 29 views
MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional
arXiv cs.AI

MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional

The article discusses the MDIA, a Multi-Agent Diagnostic Intelligence Pipeline designed for clinical reasoning. It…

5/26/2026 · 3 min read · 33 views
Emotional intelligence in large language models is fragmented across perception, cognition, and interaction
arXiv cs.AI

Emotional intelligence in large language models is fragmented across perception, cognition, and interaction

A recent study highlights the fragmented nature of emotional intelligence in large language models (LLMs). The…

5/26/2026 · 3 min read · 40 views
Exploration of Perceptual Speech Features for Clinical Decision-Support in Mental Health Care
arXiv cs.AI

Exploration of Perceptual Speech Features for Clinical Decision-Support in Mental Health Care

A new study explores the use of perceptual speech features to support clinical decision-making in mental health care.…

5/26/2026 · 2 min read · 28 views
When Mean CE Fails: Median CE Can Better Track Language Model Quality
arXiv cs.AI

When Mean CE Fails: Median CE Can Better Track Language Model Quality

The paper discusses the limitations of mean cross-entropy (CE) as a metric for evaluating language model quality. It…

5/26/2026 · 3 min read · 26 views
Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework
arXiv cs.AI

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework

A new study proposes a multi-dimensional framework for evaluating reasoning quality in large language models (LLMs).…

5/26/2026 · 3 min read · 29 views
Beyond Inference-Only Deployment: Comparing Weight-Based Consolidation Against Cascading Compaction
arXiv cs.AI

Beyond Inference-Only Deployment: Comparing Weight-Based Consolidation Against Cascading Compaction

The article discusses a new approach to deploying large language models (LLMs) that goes beyond inference-only…

5/26/2026 · 3 min read · 35 views
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models
arXiv cs.AI

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models

AVBench is a newly introduced benchmark aimed at improving the evaluation of audio-video generative models,…

5/26/2026 · 3 min read · 37 views
GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration
arXiv cs.AI

GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration

GlobalDentBench is introduced as the first multinational benchmark for evaluating large language models (LLMs) in…

5/26/2026 · 4 min read · 40 views
Lattice theory and algebraic models for deep convolutional learning based on mathematical morphology
arXiv cs.AI

Lattice theory and algebraic models for deep convolutional learning based on mathematical morphology

A new paper presents an algebraic framework for deep convolutional learning based on lattice theory and mathematical…

5/26/2026 · 3 min read · 33 views
Agent-as-Peer-Debriefer: A Multi-Agent Framework with Perspective-Based Refinement for Qualitative Analysis
arXiv cs.AI

Agent-as-Peer-Debriefer: A Multi-Agent Framework with Perspective-Based Refinement for Qualitative Analysis

The article presents a new framework called Agent-as-Peer-Debriefer designed to enhance qualitative data analysis…

5/26/2026 · 3 min read · 34 views
Hera: Learning Long-Horizon Coordination for Device-Cloud Collaborative LLM Agents
arXiv cs.AI

Hera: Learning Long-Horizon Coordination for Device-Cloud Collaborative LLM Agents

The article discusses a new approach called Hera for coordinating device-cloud collaborative large language model…

5/26/2026 · 3 min read · 35 views
Learning to Reason Efficiently with A* Post-Training
arXiv cs.AI

Learning to Reason Efficiently with A* Post-Training

A recent study explores the use of A* search algorithms to improve reasoning in large language models (LLMs). The…

5/26/2026 · 3 min read · 32 views
HeartBeatAI: An Interpretable and Robust Deep Learning Framework for Multi-Label ECG Arrhythmia Detection
arXiv cs.AI

HeartBeatAI: An Interpretable and Robust Deep Learning Framework for Multi-Label ECG Arrhythmia Detection

HeartBeatAI is a new deep learning framework designed for multi-label ECG arrhythmia detection. It addresses…

5/26/2026 · 2 min read · 28 views
Associations between echocardiographic traits and AI-ECG predictions of heart failure
arXiv cs.AI

Associations between echocardiographic traits and AI-ECG predictions of heart failure

A recent study explored the relationship between echocardiographic traits and AI-ECG predictions of heart failure. The…

5/26/2026 · 3 min read · 25 views
Summoning the Oracle to Slay It: Mitigating Look-Ahead Bias in Financial Backtesting with Large Language Models
arXiv cs.AI

Summoning the Oracle to Slay It: Mitigating Look-Ahead Bias in Financial Backtesting with Large Language Models

A new paper proposes a method to mitigate look-ahead bias in financial backtesting using large language models. The…

5/26/2026 · 3 min read · 40 views
Jailbreak to Protect: Buffering and Reinforcing via Temporary Jailbreaking for Safe Fine-Tuning in Large Language Models
arXiv cs.AI

Jailbreak to Protect: Buffering and Reinforcing via Temporary Jailbreaking for Safe Fine-Tuning in Large Language Models

The paper discusses a new framework for safe fine-tuning of large language models (LLMs) called Buffer-and-Reinforce.…

5/26/2026 · 3 min read · 36 views
PALoRA: Projection-Adaptive LoRA for Preserving Reasoning in Large Language Models
arXiv cs.AI

PALoRA: Projection-Adaptive LoRA for Preserving Reasoning in Large Language Models

The paper introduces PALoRA, a framework designed to enhance the integration of new knowledge into Large Language…

5/26/2026 · 3 min read · 35 views
Beyond Control-Flow: Integrating the Resource Perspective into Multi-Collaborative Process Modeling from Text
arXiv cs.AI

Beyond Control-Flow: Integrating the Resource Perspective into Multi-Collaborative Process Modeling from Text

The paper discusses a novel approach to process modeling in Business Process Management (BPM) that integrates resource…

5/26/2026 · 3 min read · 33 views
Emission-Aware Reinforcement Learning for Sustainable Electric Vehicle Charging and Carbon Dioxide Reduction Under Varying Renewable Penetration
arXiv cs.AI

Emission-Aware Reinforcement Learning for Sustainable Electric Vehicle Charging and Carbon Dioxide Reduction Under Varying Renewable Penetration

A new study proposes an emission-aware reinforcement learning strategy for electric vehicle charging. This approach…

5/26/2026 · 3 min read · 44 views
DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations
arXiv cs.AI

DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations

The paper introduces DemoEvolve, a method for enhancing agent harness evolution using demonstrations. This approach…

5/26/2026 · 3 min read · 30 views
Hypothesis Generation and Inductive Inference in Children and Language Models
arXiv cs.AI

Hypothesis Generation and Inductive Inference in Children and Language Models

The study explores hypothesis generation and inductive inference in children and language models. It compares how both…

5/26/2026 · 3 min read · 32 views
Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs
arXiv cs.AI

Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs

The paper discusses the vulnerabilities of Large Reasoning Models (LRMs) to jailbreak attacks due to their…

5/26/2026 · 3 min read · 33 views
Market Regime Council for Dynamic Credit Assignment in Multi-Agent LLM Decision Systems
arXiv cs.AI

Market Regime Council for Dynamic Credit Assignment in Multi-Agent LLM Decision Systems

The article presents a new cooperative multi-agent decision system called Market Regime Council (MRC) for portfolio…

5/26/2026 · 3 min read · 35 views
TIGER: Text-Informed Generalized Enzyme-Reaction Retrieval
arXiv cs.AI

TIGER: Text-Informed Generalized Enzyme-Reaction Retrieval

The article presents TIGER, a framework designed for enzyme-reaction retrieval in computational biology. It addresses…

5/26/2026 · 2 min read · 34 views
AgentFugue: Agent Scaling for Long-Horizon Tasks through Collective Reasoning
arXiv cs.AI

AgentFugue: Agent Scaling for Long-Horizon Tasks through Collective Reasoning

The paper introduces AgentFugue, a framework designed for scaling agent capabilities in long-horizon tasks through…

5/26/2026 · 3 min read · 33 views
SPACE: Unifying Symmetric and Asymmetric Routing Problems for Generalist Neural Solver
arXiv cs.AI

SPACE: Unifying Symmetric and Asymmetric Routing Problems for Generalist Neural Solver

The paper titled 'SPACE: Unifying Symmetric and Asymmetric Routing Problems for Generalist Neural Solver' presents a…

5/26/2026 · 3 min read · 27 views
SAM: State-Adaptive Memory for Long-Horizon Reasoning Agent
arXiv cs.AI

SAM: State-Adaptive Memory for Long-Horizon Reasoning Agent

The article introduces State-Adaptive Memory (SAM), a framework designed for long-horizon reasoning in artificial…

5/26/2026 · 3 min read · 33 views
Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork
arXiv cs.AI

Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork

The paper explores the limitations of In-Context Reinforcement Learning (ICRL) in the context of Ad-Hoc Teamwork…

5/26/2026 · 3 min read · 34 views
JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data
arXiv cs.AI

JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data

The article introduces JT-Safe-V2, a large language model aimed at enhancing the safety and trustworthiness of…

5/26/2026 · 2 min read · 18 views
The Model Is Not the Product: A Dual-Pillar Architecture for Local-First Psychological Coaching
arXiv cs.AI

The Model Is Not the Product: A Dual-Pillar Architecture for Local-First Psychological Coaching

The paper introduces Psych LM, an iOS application designed for psychological coaching using a local-first…

5/26/2026 · 3 min read · 24 views
Advancing Graph Few-Shot Learning via In-Context Learning
arXiv cs.AI

Advancing Graph Few-Shot Learning via In-Context Learning

The article discusses advancements in graph few-shot learning through a novel model called VISION. This model…

5/26/2026 · 3 min read · 24 views
ConceptM$^3$oE: Concept-Guided Multimodal Mixture of Experts for Interpretable Computational Pathology
arXiv cs.AI

ConceptM$^3$oE: Concept-Guided Multimodal Mixture of Experts for Interpretable Computational Pathology

The article introduces ConceptM$^3$oE, a new framework for computational pathology that integrates multimodal…

5/26/2026 · 3 min read · 33 views
Understanding and Mitigating Premature Confidence for Better LLM Reasoning
arXiv cs.AI

Understanding and Mitigating Premature Confidence for Better LLM Reasoning

The paper discusses the issue of premature confidence in language models, which leads to flawed reasoning. It…

5/26/2026 · 3 min read · 34 views
A governance horizon for ethical-use constraints in open-weight AI models
arXiv cs.AI

A governance horizon for ethical-use constraints in open-weight AI models

The paper discusses ethical-use constraints in open-weight AI models and their implications for governance policy. It…

5/26/2026 · 3 min read · 28 views
Distilling Game Code World Model Generation into Lightweight Large Language Models
arXiv cs.AI

Distilling Game Code World Model Generation into Lightweight Large Language Models

The paper discusses the generation of Game Code World Models (GameCWMs) using Large Language Models (LLMs). It…

5/26/2026 · 3 min read · 34 views
Partner-Aware Hierarchical Skill Discovery for Robust Human-AI Collaboration
arXiv cs.AI

Partner-Aware Hierarchical Skill Discovery for Robust Human-AI Collaboration

The article discusses a new framework called Partner-Aware Skill Discovery (PASD) designed to enhance human-AI…

5/26/2026 · 3 min read · 33 views
Adaptive Human-AI Coordination via Hierarchical Action Disentanglement
arXiv cs.AI

Adaptive Human-AI Coordination via Hierarchical Action Disentanglement

The paper presents a new framework for adaptive human-AI coordination called Intrinsic Action Disentanglement (IAD).…

5/26/2026 · 2 min read · 26 views
When Does Synthetic Patent Data Help? Volume-Fidelity Trade-offs in Low-Resource Multi-Label Classification
arXiv cs.AI

When Does Synthetic Patent Data Help? Volume-Fidelity Trade-offs in Low-Resource Multi-Label Classification

The study investigates the effectiveness of LLM-generated synthetic data in low-resource multi-label patent…

5/26/2026 · 3 min read · 26 views
Safety-Oriented Routing Analysis of Mixtral MoE Under Benign and Harmful Prompts
arXiv cs.AI

Safety-Oriented Routing Analysis of Mixtral MoE Under Benign and Harmful Prompts

The paper analyzes the routing behavior of the Mixtral 8x7B-Instruct model under different prompt conditions. It finds…

5/26/2026 · 3 min read · 30 views
Toward Enactive Artificial Intelligence
arXiv cs.AI

Toward Enactive Artificial Intelligence

The paper titled 'Toward Enactive Artificial Intelligence' advocates for integrating enactive approaches to perception…

5/26/2026 · 3 min read · 34 views
How Well Do Models Follow Their Constitutions?
arXiv cs.AI

How Well Do Models Follow Their Constitutions?

The paper examines how well AI models adhere to their specified behavioral guidelines. It introduces a multi-method…

5/26/2026 · 3 min read · 30 views
Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows
arXiv cs.AI

Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows

The paper discusses the limitations of current hallucination benchmarks for Large Language Models (LLMs) in…

5/26/2026 · 2 min read · 38 views
Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks
arXiv cs.AI

Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks

The paper discusses the challenges of measuring performance in Large Language Models (LLMs) as they move into…

5/26/2026 · 3 min read · 44 views

Sources in Ai Research

Other categories