60 stories tagged with #reinforce, in publish-time order across the WeSearch catalog. Tag pages update as new stories ingest.
⌘ RSS feed for this tag → or search "Reinforce"
Alphabet: Q2 Strength Reinforces The Bullish Case, A Rare Pullback Worth Buying
Alphabet delivered robust Q2 FY26 results, with 24% top-line growth and strong momentum in AI-driven Cloud and Services segments. Click for more on GOOG stock.…
Ratnaveer Precision Engineering Reports 20% Revenue Growth and 21% PAT Growth in Q1 FY27; CCL Project, Credit Rating Upgrade and Rights Issue Approvals Reinforce Growth Strategy
Ratnaveer Precision Engineering Reports 20% Revenue Growth and 21% PAT Growth in Q1 FY27; CCL Project, Credit Rating Upgrade and Rights Issue Approvals Reinforce Growth Strategy…
Korea must help reinforce nuclear umbrella
A recent Foreign Affairs article by Jennifer Lind and Daryl G. Press, "The Broken Nuclear Umbrella," has reignited a debate on a decision South Kor...…
West Bank security situation rapidly snowballing, forces will be reinforced, military officials say
A total of 26 battalions are now operating across the sector, with a security official warning on Sunday that the situation was "a snowball rolling downhill."…
Real Madrid planning two further reinforcements after Yan Diomande deal
The manner in which Real Madrid are incessantly pushing for an agreement with RB Leipzig, Yan Diomande appears set to be Los Blancos‘ next recruit of the summer.After going trophy-…
Macron calls for EU reinforcements as France battles spreading wildfires
Wildfires force 80,000 evacuations in France and Spain as EU sends reinforcements
France and Spain sought EU help after wildfires forced about 80,000 people from homes, resorts and villages.…
Context Windows Forget What Matters — I Built a Usage-Reinforced Decay Engine for AI Agent Memory
Most AI memory systems keep the newest information—not the most important. Here's how I used the Ebbinghaus forgetting curve to build a better memory engine for LLMs. The post Cont…
Multimodal Reward Hacking in Reinforcement Learning
Reinforcement learning (RL) is increasingly used to align multimodal large language models (MLLMs), but higher rewards do not always imply better task performance. This risk is amp…
Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning
arXiv:2606.27483v1 Announce Type: new Abstract: Large language model (LLM) agents have demonstrated strong capability in sequential decision-making, yet they remains fundamentally …
DiDi Global: Strong Q1 Reinforces Our Long-Term Investment Thesis
DiDi Q1 2026 update: turnover +10%, Mobility EBITDA +26%, intl revenue +60%, $6.7B cash and buybacks. Read the full analysis here.…
Dead Sea archaea sport reinforced swimming tail for hypersalty waters
On This Day: With funds and reinforcements in hand, Washington returns to New York ready for battle
The Second Continental Congress is developing military measures in consultation with George Washington.…
InfoMem: Training Long-Context Memory Agents with Answer-Conditioned Information Gain
Long-context tasks require LLMs to identify and preserve answer-relevant information from large contexts. Chunk-wise memory agents address this issue by sequentially reading docume…
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning
Autonomous LLM training is often framed as recipe search, which leaves the training harness largely static. This limitation sharpens in agentic RL, where shifting bottlenecks and s…
Kevin O'Leary claims Chinese propaganda is to blame for anti-datacenter backlash, 'hundreds of millions of dollars' being spent to kill US dominance in AI — industry proponents and Trump administration reinforce claims of foreign interference
Just because you're paranoid doesn't mean they're not after you.…
Explainable Causal Reinforcement Learning for planetary geology survey missions with embodied agent feedback loops
It was 3 AM, and I was staring at a terminal window filled with telemetry data from a simulated Mars rover. The reinforcement learning (RL) agent I had trained overnight had just c…
Bean plants call for aerial reinforcements when caterpillars attack
NPR's Short Wave talks about a weakness in a well-known insect repellant, how plants call wasps to their defense and how bigger rewards speed up learning, in mice.…
Global Prosperity Summit 2026 reinforces Hong Kong’s strategic importance in advancing Apec cooperation and global governance
GPS 2026 highlights the city as a ‘bridge for exchange’ and regional growth driver ahead of this year’s Apec meeting in Shenzhen.…
Why I built the HuggingFace for RL agents — and why RL needs one
Showcase Video If you've ever tried MineRL or OpenAI Five, you know the feeling. The environment...…
Clark Vasey: Why winning a London council seat reinforces my belief in blue collar Conservatism
Ours was a relentlessly anti-Labour campaign. Even when our canvassing results showed us losing votes to Reform, we did not shift our focus from Labour. We didn’t attack Reform or …
ARTIST: RL-Powered Tool Use for LLM Agents Explained
How Microsoft's ARTIST framework uses outcome-based RL to train LLMs that interleave tool calls inside reasoning chains — no step supervision required.…
GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training
Hybrid post-training usually combines supervised fine-tuning and reinforcement learning, but fixed mixing schedules cannot adapt when the relative noise of the two signals changes …
Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
Reinforcement Learning from Human Feedback (RLHF) is the standard method to align Large Language Models (LLMs) with human preferences. In this work, we introduce alignment tamperin…
StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning
Reinforcement learning for multi-turn agents suffers from a credit-assignment mismatch: rewards are sparse and trajectory-level, while success often hinges on a few local decisions…
Mind the Tool Failures: Achieving Synergistic Tool Gains for Medical Agents
Medical AI agents increasingly use external tools for diagnosis, treatment recommendation, and evidence retrieval, yet most existing approaches assume that task-appropriate tools a…
UnityMAS-O: A General RL Optimization Framework for LLM-Based Multi-Agent Systems
LLM-based multi-agent systems decompose complex tasks into interacting roles, but most remain manually orchestrated by prompts, tools, and control rules, while agents are rarely op…
Exploiting Local Dynamics Regularity for Reusable Skills in Offline Hierarchical RL
Hierarchical Reinforcement Learning (HRL) promises to solve long-horizon Reinforcement Learning (RL) tasks more efficiently than non-hierarchical counterparts by discovering and re…
Polar: Agentic RL on Any Harness at Scale
Reinforcement learning for language agents increasingly depends on custom harnesses that manage long-running context, multi-turn tool use and multi-agent orchestration. However, po…
Credit Assignment with Resets in Language Model Reasoning
Contemporary reinforcement learning with verifiable reward methods post-train language models on multi-step reasoning by assigning a single outcome reward uniformly across all toke…
Evolutionary Enhanced Multi-Agent Reinforcement Learning for Cooperative Air Combat
As modern air combat evolves toward beyond-visual-range (BVR) multi-aircraft cooperative engagements, autonomous decision-making for unmanned combat aerial vehicles (UCAVs) faces s…
ProActor: Timing-Aware Reinforcement Learning for Proactive Task Scheduling Agents
Proactive task-oriented agents must autonomously anticipate user needs, identify actionable opportunities, and trigger software actions at appropriate moments - fundamentally shift…
CoRe-Code: Collaborative Reinforcement Learning for Code Generation
Large language models (LLMs) have achieved strong performance in code generation, but most methods rely on autoregressive decoding without global planning, often leading to locally…
Emission-Aware Reinforcement Learning for Sustainable Electric Vehicle Charging and Carbon Dioxide Reduction Under Varying Renewable Penetration
The rapid growth of Electric Vehicle (EV) adoption challenges power distribution networks through peak load spikes, voltage instability, and transformer overloads from uncoordinate…
Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork
In-Context Reinforcement Learning (ICRL) has enabled foundation agents to adapt instantaneously to novel tasks, yet its efficacy in Ad-Hoc Teamwork (AHT)-where coordination with un…
When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs
Multi-agent LLM workflows route inference through specialized roles to lift end-task accuracy, but jointly training those roles with reinforcement learning is unstable in ways that…
Quad Foreign Ministers meet LIVE: Amid West Asia turmoil, Quad Foreign Ministers gather to reinforce Indo-Pacific stability
Follow live updates from The Hindu as Quad Foreign Ministers from the United States, Australia and Japan hold a crucial meeting in New Delhi on May 26, 2026…
Understanding Reinforcement Learning with Human Feedback Part 5: Training the Reward Model with Loss Functions
In the previous article, we created a reward model. In this article, we will continue exploring how...…
Show HN: Reward Is Not Reinforcement Until Admitted
Contribute to nikitph/rewarder development by creating an account on GitHub.…
Spalletti laments lack of character in Juventus squad, expects summer reinforcement
Juventus manager Luciano Spalletti was displeased with his team after squandering a two-goal lead in the Derby della Mole.Due to the one-hour delay, the Bianconeri had already know…
If you use NVIDIA Isaac Sim for reinforcement learning, do you use Isaac Lab with it? Just want to get a sense of what the status quo is. [D]
Reflex: Reinforcement Learning with Reflection Symmetry Exploitation in State-Based Continuous Control
Reinforcement learning has long struggled with poor sample efficiency. One promising approach to mitigate this problem is leveraging group-invariant Markov Decision Processes ($G$-…
Curriculum reinforcement learning with measurable task representation learning
In curriculum reinforcement learning (CRL), an agent incrementally accumulates knowledge over a sequence of tasks (i.e., a curriculum), and the learning process is aimed at using t…
Score-Based One-step MeanFlow Policy Optimization
Diffusion and flow matching have emerged as expressive policy classes in reinforcement learning, but their reliance on multi-step denoising imposes substantial computational overhe…
Reinforcement Learning for Microcanonical Graph Ensemble with Assortativity Constraints
How network structure determines function is a fundamental question, and it can be investigated by graph ensembles with precisely controlled structural properties. Canonical approa…
Infra-Bayesian Reinforcement Learning Agents Outperform Classical RL For Worst-Case Robustness
Classical reinforcement learning assumes the agent interacts with a fixed environment whose behavior does not depend on the agent's policy. This assumption breaks down in non-reali…
Classical State Preparation for Variational Quantum Algorithms via Reinforcement Learning
Variational Quantum Algorithms (VQAs) potentially offer a pathway to practical quantum advantage, but their optimization is heavily hindered by barren plateaus and numerous local m…
Dreaming Smoothly and Sample Efficiently with Gradient Penalized Latent Dynamics
Model-based reinforcement learning improves sample efficiency by learning a world model. However, existing latent world models such as DreamerV3 do not explicitly enforce local smo…
One Policy, Infinite NPCs: Persona-Traceable Shared RL Policies for Scalable Game Agents
On a 300-persona life-simulation benchmark, pcsp achieves compositional zero-shot persona identification up to 17x above chance, Spearman rho approx 0.73 semantic-behavioral alignm…
How News Aggregators Reinforce Political Ignorance
How news aggregators reinforce political ignorance by filtering biased content and confirming beliefs.…
Mets Morning News: Reinforcements are coming, but will it matter?
Your Sunday morning dose of New York Mets and MLB news, notes, and links.…
Understanding Reinforcement Learning with Human Feedback Part 4: Teaching Models Human Preferences
In the previous article, we explored the part where we collect human preferences. In this article, we...…
News Aggregators’ Left-Wing Bias Reinforces Political Ignorance
The other day, a young lady got on the elevator and promptly whipped out her smartphone and began scrolling. “Excuse me,” I said, “can I ask you where you get your news?” She said,…
Two Loops: How China's Open AI Strategy Reinforces Its Industrial Dominance [pdf]
Nexpace Announces NXPC Buyback Program to Reinforce User-Centered Ecosystem Growth in MapleStory Universe
Distribution-Aware Reward: Reinforcement Learning over Predictive Distributions for LLM Regression
Large language models can predict real-valued quantities from heterogeneous inputs such as text, code, and molecular strings, but most training objectives score each decoded floati…
AGPO: Adaptive Group Policy Optimization with Dual Statistical Feedback
Reinforcement learning improves LLM reasoning, but PPO/GRPO typically use fixed clipping and decoding temperature, which makes training brittle and tuning-heavy. We propose Adaptiv…
Design for Manufacturing: A Manufacturability Knowledge-Integrated Reinforcement Learning Framework for Free-Form Pipe Routing in Aeroengines
Design for manufacturing plays a critical role in advanced aeroengine development, where complex components necessitate careful consideration of manufacturability. However, current…
Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs
We introduce a novel method that averages the logits of a frozen reference policy (e.g., SFT) and a trainable policy, and incorporate the method into Group Relative Policy Optimiza…
Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor
MXFP4 arithmetic can dramatically accelerate reinforcement learning (RL) post-training of large language models (LLMs), yet the quantization error introduces severe accuracy degrad…