WeSearch
Hub / Ai Research
ai-research · WeSearch

Ai Research news.

Page 7 of Ai Research headlines on WeSearch — deduped and updated continuously from 10+ editorial sources.

JobBench: Aligning Agent Work With Human Will
arXiv cs.AI

JobBench: Aligning Agent Work With Human Will

The paper introduces JobBench, a new benchmark for evaluating AI agents based on human needs rather than economic…

5/27/2026 · 3 min read · 46 views
OmniToM: Benchmarking Theory of Mind in LLMs via Explicit Belief Modeling
arXiv cs.AI

OmniToM: Benchmarking Theory of Mind in LLMs via Explicit Belief Modeling

The paper introduces OmniToM, a benchmark designed to evaluate the Theory of Mind capabilities in large language…

5/27/2026 · 3 min read · 40 views
Anchor: Mitigating Artifact Drift in Agent Benchmark Generation
arXiv cs.AI

Anchor: Mitigating Artifact Drift in Agent Benchmark Generation

The paper introduces Anchor, a task-generation pipeline designed to address artifact drift in AI agent benchmark…

5/27/2026 · 3 min read · 45 views
Experiments in Agentic AI for Science
arXiv cs.AI

Experiments in Agentic AI for Science

The paper discusses two innovative frameworks for creating autonomous AI systems to enhance scientific workflows.…

5/27/2026 · 2 min read · 37 views
Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems
arXiv cs.AI

Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems

The paper discusses the aging of AI agents deployed in operational systems and introduces a new benchmark called…

5/27/2026 · 3 min read · 36 views
Constraint acquisition needs better benchmarks
arXiv cs.AI

Constraint acquisition needs better benchmarks

The paper discusses the need for improved benchmarks in Constraint Acquisition (CA) research. Current benchmarks are…

5/27/2026 · 2 min read · 36 views
Personalizing Embodied Multimodal Large Language Model Agents over Long-term User Interactions
arXiv cs.AI

Personalizing Embodied Multimodal Large Language Model Agents over Long-term User Interactions

The paper presents POLAR, a framework designed for personalizing embodied multimodal large language model agents…

5/27/2026 · 3 min read · 41 views
Is Agent Memory a Database? Rethinking Data Foundations for Long-Term AI Agent Memory
arXiv cs.AI

Is Agent Memory a Database? Rethinking Data Foundations for Long-Term AI Agent Memory

The paper discusses the need for persistent memory in long-running AI agents. It critiques current memory systems and…

5/27/2026 · 3 min read · 45 views
Can LLMs Introspect? A Reality Check
arXiv cs.AI

Can LLMs Introspect? A Reality Check

The paper titled 'Can LLMs Introspect? A Reality Check' questions the ability of large language models (LLMs) to…

5/27/2026 · 3 min read · 42 views
BrickAnything: Geometry-Conditioned Buildable Brick Generation with Structure-Aware Tokenization
arXiv cs.AI

BrickAnything: Geometry-Conditioned Buildable Brick Generation with Structure-Aware Tokenization

The article presents BrickAnything, a new framework for generating buildable brick structures from 3D shapes. This…

5/27/2026 · 3 min read · 41 views
What Gets Cited: Competitive GEO in AI Answer Engines
arXiv cs.AI

What Gets Cited: Competitive GEO in AI Answer Engines

The paper titled 'What Gets Cited: Competitive GEO in AI Answer Engines' explores how AI answer engines cite sources.…

5/26/2026 · 3 min read · 34 views
Credit Assignment with Resets in Language Model Reasoning
arXiv cs.AI

Credit Assignment with Resets in Language Model Reasoning

The article discusses advancements in credit assignment methods for language model reasoning in reinforcement…

5/26/2026 · 3 min read · 25 views
ATWL: A Formal Language for Representing, Comparing, and Reusing Visual Analytics Workflows
arXiv cs.AI

ATWL: A Formal Language for Representing, Comparing, and Reusing Visual Analytics Workflows

The paper introduces the Artifact-Transform Workflow Language (ATWL), a formal language designed for visual analytics…

5/26/2026 · 3 min read · 30 views
A Signal-Language Foundation Model for Broad-Spectrum Cardiovascular Assessment from Routine Electrocardiography
arXiv cs.AI

A Signal-Language Foundation Model for Broad-Spectrum Cardiovascular Assessment from Routine Electrocardiography

A new AI model called ECGCLIP has been developed to enhance cardiovascular assessment from routine…

5/26/2026 · 3 min read · 34 views
Security of OpenClaw Agents: Fundamentals, Attacks, and Countermeasures
arXiv cs.AI

Security of OpenClaw Agents: Fundamentals, Attacks, and Countermeasures

The paper discusses the security challenges associated with OpenClaw agents, a new class of autonomous systems. It…

5/26/2026 · 3 min read · 32 views
CODESKILL: Learning Self-Evolving Skills for Coding Agents
arXiv cs.AI

CODESKILL: Learning Self-Evolving Skills for Coding Agents

CODESKILL is a proposed framework aimed at enhancing coding agents' abilities through self-evolving skills. It…

5/26/2026 · 2 min read · 34 views
Towards end-to-end LLM-based censoring-aware survival analysis
arXiv cs.AI

Towards end-to-end LLM-based censoring-aware survival analysis

The article discusses a new framework called LLMSurvival that enables censoring-aware survival analysis using large…

5/26/2026 · 3 min read · 36 views
Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models
arXiv cs.AI

Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models

The paper titled 'Second Guess' introduces a method for detecting uncertainty in small language models through a…

5/26/2026 · 2 min read · 26 views
Context-CoT: Enhancing Context Learning via High-Quality Reasoning Synthesis
arXiv cs.AI

Context-CoT: Enhancing Context Learning via High-Quality Reasoning Synthesis

The paper titled 'Context-CoT: Enhancing Context Learning via High-Quality Reasoning Synthesis' addresses the…

5/26/2026 · 2 min read · 33 views
AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
arXiv cs.AI

AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems

The paper titled 'AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems' introduces a framework for…

5/26/2026 · 3 min read · 35 views
Whose Alignment? Comparing LLM Process Alignment Across Diverse Organizational Decision Contexts
arXiv cs.AI

Whose Alignment? Comparing LLM Process Alignment Across Diverse Organizational Decision Contexts

The paper discusses the alignment of AI systems with organizational decision-making, emphasizing the complexity of…

5/26/2026 · 3 min read · 30 views
LipoAgent: Coordinating Fine-Tuned LLM Agents for Safer Lipid Design
arXiv cs.AI

LipoAgent: Coordinating Fine-Tuned LLM Agents for Safer Lipid Design

LipoAgent is a new framework designed to enhance the safety and efficiency of lipid nanoparticles for nucleic acid…

5/26/2026 · 2 min read · 33 views
FrontierOR: Benchmarking LLMs' Capacity for Efficient Algorithm Design in Large-Scale Optimization
arXiv cs.AI

FrontierOR: Benchmarking LLMs' Capacity for Efficient Algorithm Design in Large-Scale Optimization

The paper introduces FrontierOR, a benchmark designed to evaluate the capacity of large language models (LLMs) in…

5/26/2026 · 3 min read · 36 views
Meta-Agent: From Task Descriptions to Verified Multi-Agent Systems
arXiv cs.AI

Meta-Agent: From Task Descriptions to Verified Multi-Agent Systems

The paper introduces Meta-Agent, a framework designed to improve the reliability of multi-agent systems. It automates…

5/26/2026 · 3 min read · 24 views
Boosting Inference with Guided Reasoning: Stochastic Exploration for Recursive Models
arXiv cs.AI

Boosting Inference with Guided Reasoning: Stochastic Exploration for Recursive Models

The paper discusses a novel approach to enhancing inference in recursive neural networks through guided reasoning and…

5/26/2026 · 3 min read · 37 views
DarkForest: Less Talk, Higher Accuracy for Multi-Agent LLMs
arXiv cs.AI

DarkForest: Less Talk, Higher Accuracy for Multi-Agent LLMs

The paper presents DarkForest, a new framework aimed at improving the accuracy of multi-agent large language models…

5/26/2026 · 3 min read · 23 views
SpecAlign: A Semantic Alignment Framework for SystemVerilog Assertion Generation
arXiv cs.AI

SpecAlign: A Semantic Alignment Framework for SystemVerilog Assertion Generation

The paper introduces SpecAlign, a framework designed to enhance the semantic alignment of SystemVerilog Assertions…

5/26/2026 · 2 min read · 27 views
SimuWoB: Simulating Real-World Mobile Apps for Fast and Faithful GUI Agent Benchmarking
arXiv cs.AI

SimuWoB: Simulating Real-World Mobile Apps for Fast and Faithful GUI Agent Benchmarking

The paper introduces SimuWoB, a synthetic benchmark designed for evaluating mobile GUI agents. It addresses the…

5/26/2026 · 3 min read · 32 views
Representation Without Control: Testing the Realization Effect in Language Models
arXiv cs.AI

Representation Without Control: Testing the Realization Effect in Language Models

The paper titled 'Representation Without Control: Testing the Realization Effect in Language Models' explores the…

5/26/2026 · 3 min read · 32 views
Beyond the Frontier: Stochastic Backtracking for Efficient Test-Time Scaling
arXiv cs.AI

Beyond the Frontier: Stochastic Backtracking for Efficient Test-Time Scaling

The paper introduces a method called stochastic backtracking for improving test-time scaling in language models. This…

5/26/2026 · 3 min read · 33 views
Trust but Verify: Prover-Verifier Deliberation for Selective LLM Prediction
arXiv cs.AI

Trust but Verify: Prover-Verifier Deliberation for Selective LLM Prediction

The paper introduces a new protocol called prover-verifier deliberation (PVD) for improving the reliability of…

5/26/2026 · 3 min read · 32 views
RECTOR: Priority-Aware Rule-Based Reranking for Compliance-Aware Autonomous Driving Trajectory Selection
arXiv cs.AI

RECTOR: Priority-Aware Rule-Based Reranking for Compliance-Aware Autonomous Driving Trajectory Selection

The paper presents RECTOR, a rule-based reranking system designed for autonomous driving trajectory selection. It…

5/26/2026 · 3 min read · 34 views
Evolutionary Enhanced Multi-Agent Reinforcement Learning for Cooperative Air Combat
arXiv cs.AI

Evolutionary Enhanced Multi-Agent Reinforcement Learning for Cooperative Air Combat

The paper presents a new framework for multi-agent reinforcement learning in cooperative air combat scenarios. It…

5/26/2026 · 2 min read · 32 views
AION: Next-Generation Tasks and Practical Harness for Time Series
arXiv cs.AI

AION: Next-Generation Tasks and Practical Harness for Time Series

The paper titled 'AION: Next-Generation Tasks and Practical Harness for Time Series' presents a new framework for time…

5/26/2026 · 3 min read · 27 views
Privacy-Preserving Local Language Models for Longitudinal Data Retrieval in Chronic Dermatologic Disease: Implementation in Pemphigus Patients
arXiv cs.AI

Privacy-Preserving Local Language Models for Longitudinal Data Retrieval in Chronic Dermatologic Disease: Implementation in Pemphigus Patients

A study evaluated the use of a privacy-preserving small language model (SLM) for retrieving clinical data in pemphigus…

5/26/2026 · 3 min read · 34 views
NeurIPS: Neuro-anatomical Inductive Priors for Sphere-based Brain Decoding
arXiv cs.AI

NeurIPS: Neuro-anatomical Inductive Priors for Sphere-based Brain Decoding

The paper presents NeurIPS, a framework designed to enhance surface-based brain decoding by utilizing neuro-anatomical…

5/26/2026 · 3 min read · 34 views
Mitigating Object Hallucinations in Vision-Language Models through Region-Aware Attention Recalibration
arXiv cs.AI

Mitigating Object Hallucinations in Vision-Language Models through Region-Aware Attention Recalibration

A new paper addresses the issue of object hallucination in Large Vision-Language Models (LVLMs). The authors propose a…

5/26/2026 · 3 min read · 32 views
Towards Multi-Turn Dialog Systems for Industrial Asset Operations and Maintenance
arXiv cs.AI

Towards Multi-Turn Dialog Systems for Industrial Asset Operations and Maintenance

The paper presents a multi-turn dialog system tailored for industrial asset operations and maintenance. This system…

5/26/2026 · 2 min read · 34 views
Energy Shields for Fairness
arXiv cs.AI

Energy Shields for Fairness

The paper titled 'Energy Shields for Fairness' introduces a novel approach to ensuring runtime fairness in…

5/26/2026 · 3 min read · 34 views
Noise-Robust Financial Numerical Entity Attribute Tagging
arXiv cs.AI

Noise-Robust Financial Numerical Entity Attribute Tagging

The paper introduces a new method called NORA for improving the understanding of financial numerical entities in…

5/26/2026 · 3 min read · 33 views
ProActor: Timing-Aware Reinforcement Learning for Proactive Task Scheduling Agents
arXiv cs.AI

ProActor: Timing-Aware Reinforcement Learning for Proactive Task Scheduling Agents

The paper introduces ProActor, a framework for proactive task scheduling using timing-aware reinforcement learning. It…

5/26/2026 · 3 min read · 32 views
TaBIIC2: Interactive Building of Ontological Taxonomies using Weighted Self-Organizing Maps
arXiv cs.AI

TaBIIC2: Interactive Building of Ontological Taxonomies using Weighted Self-Organizing Maps

The paper presents TaBIIC2, a tool designed for the interactive construction of ontological taxonomies using weighted…

5/26/2026 · 3 min read · 31 views
Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications
arXiv cs.AI

Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications

A new framework called POLARIS has been introduced to enhance safety testing for Large Language Models (LLMs). This…

5/26/2026 · 3 min read · 32 views
Clustering as Reasoning: A $k$-Means Interpretation of Chain-of-Thought Graph Learning
arXiv cs.AI

Clustering as Reasoning: A $k$-Means Interpretation of Chain-of-Thought Graph Learning

The paper presents a novel approach to Chain-of-Thought (CoT) graph learning by interpreting it through the lens of…

5/26/2026 · 3 min read · 34 views
Solving Combinatorial Counting Problems with Weighted First-Order Model Counting
arXiv cs.AI

Solving Combinatorial Counting Problems with Weighted First-Order Model Counting

The paper presents a new approach to solving combinatorial counting problems using a language called Cofola. This…

5/26/2026 · 3 min read · 33 views
Geo-Expert: Towards Expert-Level Geological Reasoning via Parameter-Efficient Fine-Tuning
arXiv cs.AI

Geo-Expert: Towards Expert-Level Geological Reasoning via Parameter-Efficient Fine-Tuning

The paper introduces Geo-Expert, a series of parameter-efficient geological language models designed to improve…

5/26/2026 · 2 min read · 27 views
Test-Time Deep Thinking to Explore Implicit Rules
arXiv cs.AI

Test-Time Deep Thinking to Explore Implicit Rules

A new framework called Test-Time Exploration (TTExplore) aims to improve the performance of intelligent agents in…

5/26/2026 · 3 min read · 24 views
Agent Manufacturing: Foundation-Model Agents as First-Class Industrial Entities
arXiv cs.AI

Agent Manufacturing: Foundation-Model Agents as First-Class Industrial Entities

The article discusses a new paradigm in manufacturing called Agent Manufacturing, which focuses on the role of…

5/26/2026 · 2 min read · 20 views
CoRe-Code: Collaborative Reinforcement Learning for Code Generation
arXiv cs.AI

CoRe-Code: Collaborative Reinforcement Learning for Code Generation

The paper introduces CoRe-Code, a framework for collaborative reinforcement learning aimed at improving code…

5/26/2026 · 3 min read · 27 views
PANDO: Efficient Multimodal AI Agents via Online Skill Distillation
arXiv cs.AI

PANDO: Efficient Multimodal AI Agents via Online Skill Distillation

The paper introduces PANDO, a framework designed to enhance the efficiency of multimodal AI agents through online…

5/26/2026 · 3 min read · 34 views

Sources in Ai Research

Other categories