Mahjax: A GPU-Accelerated Mahjong Simulator for Reinforcement Learning in JAX
Mahjax is a new GPU-accelerated Mahjong simulator designed for reinforcement learning using JAX. It allows for…
Recent ai-research headlines from arXiv cs.AI.

Mahjax is a new GPU-accelerated Mahjong simulator designed for reinforcement learning using JAX. It allows for…

The paper presents a new architecture called Hierarchical Agent-native Network Architecture (HANA) aimed at achieving…

The article introduces COAgents, a multi-agent framework designed to address the complexities of Vehicle Routing…

The article discusses advancements in optimizing industrial asset operations through improved caching and workflow…

The article discusses a new framework called Declarative Data Services (DDS) aimed at improving the composition of…

The VBFDD-Agent is a new approach for detecting and diagnosing faults in electric vehicle batteries. It utilizes…

The paper presents a novel method called Conflict-Aware Additive Guidance ($g^ ext{car}$) aimed at improving flow…

The paper discusses a framework called interaction locality for measuring information flow in spatial reasoning tasks.…

The paper discusses the conditional equivalence of Direct Preference Optimization (DPO) and Reinforcement Learning…

PlanningBench is a new framework designed to generate scalable and verifiable planning data for evaluating and…

The paper discusses the need for governance in autonomous enterprise agents. It introduces CUGA's policy system, which…

A new study explores how reinforcement learning agents can improve their performance in fighting games by learning not…

The study explores the impact of different persona vectors on sycophancy in AI models. It compares off-the-shelf…

The article discusses AutoRPA, a framework designed to enhance GUI automation using large language models (LLMs). It…

ScenePilot is a new framework designed for generating critical scenarios in autonomous driving. It focuses on creating…

The Insights Generator (IG) is a new multi-agent system designed to diagnose failures in LLM agents by analyzing…

The paper discusses the integration of Artificial Intelligence into 6G networks to enhance their resilience and…

The article discusses a new educational approach to teaching AI through benchmark construction, specifically using a…

The paper presents PALS, a power-aware runtime for serving large language models (LLMs) that optimizes GPU power…

The paper discusses the challenges of bridging the gap between simulated and real-world decision-making in sequential…

AiraXiv is a proposed AI-driven open-access platform designed for both human and AI scientists. It aims to address the…

DeepWeb-Bench is a new benchmark designed to evaluate deep research capabilities of language models. It emphasizes the…

The paper introduces a new framework called Diverge-to-Induce Prompting (DIP) aimed at improving zero-shot reasoning…

A new paper proposes a neural framework for estimating pairwise conditional mutual information in masked discrete…

GraphDiffMed is a new framework designed for medication recommendation that integrates pharmacological knowledge with…

The study focuses on enhancing the performance of quantized large language models (LLMs) in qualitative analysis. It…

The paper presents a framework for improving the performance of large language models (LLMs) in analyzing long…

The article presents a novel approach to target-oriented proactive dialogue systems using a Forward-Focused…

The paper explores the hypothesis that real-data scaling laws are influenced by a latent predictive contribution…

FlowLM is a new language model that adapts pre-trained diffusion models for efficient few-step text generation. It…

This article discusses a study on multimodal emotion recognition in proactive conversational agents. The research…

The paper presents a novel training framework called ProxyCoT aimed at improving long-context reasoning in large…

The study investigates how emotionally framed evaluations affect the behavior and internal representations of small…

The article discusses the development of GrandGuard, a framework aimed at improving safety in interactions between…

The paper introduces RealUserSim, a new user simulation framework designed to improve agent benchmarking by grounding…

The paper introduces PrivacyAkinator, a tool designed to assist developers in making key privacy design decisions. It…

The article discusses the transition of agentic AI systems from experimental prototypes to enterprise deployments. It…

A recent study explores the use of Vision-Language Models (VLMs) to detect learner attention in educational videos.…

A new approach to HIV prevention has been proposed through a method called Cascade-Aware Suppression of Transmission…

A new study explores the use of AI-assisted competency assessment in nursing education through egocentric video…

The paper introduces TabPFN-MT, a multitask in-context learner designed for tabular data. This model improves upon…

A new paper presents a framework called Score-induced Latent Diffusion (SiLD) for learning diffusion models under the…

The paper titled 'Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry' introduces a new method…

The article discusses LEAP, a closed-loop framework designed for the discovery of perovskite precursor additives. This…

The article discusses a new framework called Lean Refactor designed for optimizing Lean proofs. This framework…

The paper introduces GROW, a reinforcement learning framework designed for open-world vision-language model agents. It…

The paper presents CP-MoE, a framework designed to tackle catastrophic forgetting in continual learning for large…

The paper introduces a novel framework called Kernel Discovery for high-dimensional Bayesian optimization. This…

The article introduces ProcBench, a new benchmark designed to evaluate process-level defects in LLM coding agents.…

The paper presents a novel approach to Table Question-Answering (TQA) using two frameworks: TableGrid Navigation (TGN)…
WeSearch's declared handling of arXiv cs.AI's content. Indexing, snippets, summaries, retrieval and training are separate questions — see the rights registry or read this source's machine-readable record.