JobBench: Aligning Agent Work With Human Will
The paper introduces JobBench, a new benchmark for evaluating AI agents based on human needs rather than economic…
Page 7 of Ai Research headlines on WeSearch — deduped and updated continuously from 10+ editorial sources.

The paper introduces JobBench, a new benchmark for evaluating AI agents based on human needs rather than economic…

The paper introduces OmniToM, a benchmark designed to evaluate the Theory of Mind capabilities in large language…

The paper introduces Anchor, a task-generation pipeline designed to address artifact drift in AI agent benchmark…

The paper discusses two innovative frameworks for creating autonomous AI systems to enhance scientific workflows.…

The paper discusses the aging of AI agents deployed in operational systems and introduces a new benchmark called…

The paper discusses the need for improved benchmarks in Constraint Acquisition (CA) research. Current benchmarks are…

The paper presents POLAR, a framework designed for personalizing embodied multimodal large language model agents…

The paper discusses the need for persistent memory in long-running AI agents. It critiques current memory systems and…

The paper titled 'Can LLMs Introspect? A Reality Check' questions the ability of large language models (LLMs) to…

The article presents BrickAnything, a new framework for generating buildable brick structures from 3D shapes. This…

The paper titled 'What Gets Cited: Competitive GEO in AI Answer Engines' explores how AI answer engines cite sources.…

The article discusses advancements in credit assignment methods for language model reasoning in reinforcement…

The paper introduces the Artifact-Transform Workflow Language (ATWL), a formal language designed for visual analytics…

A new AI model called ECGCLIP has been developed to enhance cardiovascular assessment from routine…

The paper discusses the security challenges associated with OpenClaw agents, a new class of autonomous systems. It…

CODESKILL is a proposed framework aimed at enhancing coding agents' abilities through self-evolving skills. It…

The article discusses a new framework called LLMSurvival that enables censoring-aware survival analysis using large…

The paper titled 'Second Guess' introduces a method for detecting uncertainty in small language models through a…

The paper titled 'Context-CoT: Enhancing Context Learning via High-Quality Reasoning Synthesis' addresses the…

The paper titled 'AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems' introduces a framework for…

The paper discusses the alignment of AI systems with organizational decision-making, emphasizing the complexity of…

LipoAgent is a new framework designed to enhance the safety and efficiency of lipid nanoparticles for nucleic acid…

The paper introduces FrontierOR, a benchmark designed to evaluate the capacity of large language models (LLMs) in…

The paper introduces Meta-Agent, a framework designed to improve the reliability of multi-agent systems. It automates…

The paper discusses a novel approach to enhancing inference in recursive neural networks through guided reasoning and…

The paper presents DarkForest, a new framework aimed at improving the accuracy of multi-agent large language models…

The paper introduces SpecAlign, a framework designed to enhance the semantic alignment of SystemVerilog Assertions…

The paper introduces SimuWoB, a synthetic benchmark designed for evaluating mobile GUI agents. It addresses the…

The paper titled 'Representation Without Control: Testing the Realization Effect in Language Models' explores the…

The paper introduces a method called stochastic backtracking for improving test-time scaling in language models. This…

The paper introduces a new protocol called prover-verifier deliberation (PVD) for improving the reliability of…

The paper presents RECTOR, a rule-based reranking system designed for autonomous driving trajectory selection. It…

The paper presents a new framework for multi-agent reinforcement learning in cooperative air combat scenarios. It…

The paper titled 'AION: Next-Generation Tasks and Practical Harness for Time Series' presents a new framework for time…

A study evaluated the use of a privacy-preserving small language model (SLM) for retrieving clinical data in pemphigus…

The paper presents NeurIPS, a framework designed to enhance surface-based brain decoding by utilizing neuro-anatomical…

A new paper addresses the issue of object hallucination in Large Vision-Language Models (LVLMs). The authors propose a…

The paper presents a multi-turn dialog system tailored for industrial asset operations and maintenance. This system…

The paper titled 'Energy Shields for Fairness' introduces a novel approach to ensuring runtime fairness in…

The paper introduces a new method called NORA for improving the understanding of financial numerical entities in…

The paper introduces ProActor, a framework for proactive task scheduling using timing-aware reinforcement learning. It…

The paper presents TaBIIC2, a tool designed for the interactive construction of ontological taxonomies using weighted…

A new framework called POLARIS has been introduced to enhance safety testing for Large Language Models (LLMs). This…

The paper presents a novel approach to Chain-of-Thought (CoT) graph learning by interpreting it through the lens of…

The paper presents a new approach to solving combinatorial counting problems using a language called Cofola. This…

The paper introduces Geo-Expert, a series of parameter-efficient geological language models designed to improve…

A new framework called Test-Time Exploration (TTExplore) aims to improve the performance of intelligent agents in…

The article discusses a new paradigm in manufacturing called Agent Manufacturing, which focuses on the role of…

The paper introduces CoRe-Code, a framework for collaborative reinforcement learning aimed at improving code…

The paper introduces PANDO, a framework designed to enhance the efficiency of multimodal AI agents through online…