Modeling Agentic Technical Debt and Stochastic Tax: A Standalone Framework for Measurement, Simulation, and Dashboarding
The article presents a framework for measuring and simulating Agentic Technical Debt and Stochastic Tax in AI systems.…
Recent ai-research headlines from arXiv cs.AI.

The article presents a framework for measuring and simulating Agentic Technical Debt and Stochastic Tax in AI systems.…

Maat is a new legal research assistant designed specifically for competition law analysis. It outperforms existing…

The paper discusses 2-ASP(Q) programs with weak constraints, a significant area in Answer Set Programming. It provides…

The paper discusses a vulnerability known as alignment tampering in Reinforcement Learning from Human Feedback (RLHF).…

The paper presents a new approach called BRANE for optimizing retrieval agent configurations based on natural language…

The MUSE-Autoskill framework introduces a new approach for self-evolving agents that enhances their ability to create…

Xe-Forge is a new multi-stage pipeline designed to optimize kernel performance for Intel GPUs. It automates the…

The paper presents a framework for deploying Edge AI in industrial embedded platforms, emphasizing the importance of a…

The paper introduces GEM, a framework designed for optimal data curation in large language models (LLMs). It addresses…

The paper discusses Pretraining Data Exposure (PDE) in Large Language Models (LLMs), highlighting its implications for…

A recent study investigates the impact of audio deepfakes on human trust in real speech. The research, which involved…

AssetGen is a new 3D asset generation system that prioritizes user experience and deployability. It can produce…

The article introduces VISTA, a benchmark designed to evaluate the capabilities of LLM-based agents in generating web…

The paper titled 'Augment Engineering' introduces a methodology for orchestrating multiple AI tools across various…

The paper introduces MemMorph, a novel attack method targeting long-term memory in LLM-driven agents. By injecting…

The paper introduces Belief-Aware GSAC (BA-GSAC), which adapts the distillation coefficient in autonomous driving…

The paper introduces BITE, a framework designed to exploit stylistic biases in LLM judges. It demonstrates that these…

The paper titled 'Furina: Fragmented Uncertainty-Driven Refusal Instability Attack' explores safety alignment in large…

The paper titled 'TSFMAudit' addresses the issue of data contamination in time series foundation models (TSFMs). It…

The article discusses a new framework called PushCen-ADFL for asynchronous decentralized federated learning. This…

The paper discusses the challenges faced by agentic RAG systems due to tool schemas consuming context windows needed…

The paper discusses advancements in autonomous online intrusion detection systems (IDS) for IoT devices. It highlights…

The article presents a novel approach to neural dynamics using Lie group embedding through supervised projective…

The paper discusses the challenges of detecting cross-section defects in documents processed by language model…

The paper titled 'InfoQuant' addresses the challenges of low-bit activation quantization in large language models. It…

The article introduces PitchBench, a new evaluation suite designed to measure pitch hearing in audio-language models…

RepoMirage is a new evaluation suite designed to assess repository context reasoning in code agents. The study reveals…

AutoDFT is a new multi-agent framework designed to enhance autonomous DFT calculations in materials science. It…

The paper presents GAC, a noise-aware adaptive mixing method for hybrid post-training in machine learning. This…

The paper introduces SetupX, a framework designed to improve the setup of functionality-correct code repositories by…

The paper explores the potential of large Vision-Language Models (VLMs) to replicate the open-ended creative processes…

A recent study examines the confidence calibration of large language models (LLMs) across various tasks. The findings…

The paper investigates the redundancy in reasoning processes of large language models (LLMs). It quantifies how much…

The article introduces Context, a new intelligence layer designed to enhance proactive goal-directed interactions in…

The paper discusses the design of workflows that utilize large language models (LLMs) alongside traditional…

The paper introduces 'Quantum Frog', a two-player cooperative game that utilizes a quantized-time mechanic. It…

The paper titled 'BODHI: Precise OS Kernel Specification Inference' introduces a method to enhance the formal…

The paper discusses the challenges faced by large language models (LLMs) in maintaining correct medical diagnoses…

The paper discusses the integration of a Coherent Ising Machine (CIM) with a large language model (LLM) to enhance…

The paper discusses the concept of Reconstructive Authority in autonomous agent systems, focusing on how to enforce…

The article discusses a new book on fuzzy, neutrosophic, and uncertain graph theory. It emphasizes the unifying role…

The paper introduces BoxLitE, a knowledge base embedding model that utilizes convex optimization. This model aims to…

The paper discusses the phenomenon of Authority Inversion in large language models (LLMs) used in ubiquitous systems.…

The paper titled DRIVE proposes a dual-level skill modeling framework for web agents to enhance their reasoning and…

The paper discusses how multi-turn reasoning systems often fail not due to logical contradictions, but rather due to a…

The paper presents MEMOR-E, a mobile quadruped robot designed to assist Alzheimer's patients and their caregivers. It…

The paper presents a framework for modeling cognitive processes through a cybernetic lens. It introduces a feedback…

The paper titled 'Spacetime Formation under Requirements: Contextual Realization and Form-Dependent Probability' by…

The paper discusses the interaction between users and AI-driven recommendation systems. It models how users convey…

The paper presents a new method called PAT for improving the efficiency of Reinforcement Learning from Human Feedback…
WeSearch's declared handling of arXiv cs.AI's content. Indexing, snippets, summaries, retrieval and training are separate questions — see the rights registry or read this source's machine-readable record.