WeSearch
Hub / Tags / Reinforce
TAG · #REINFORCE

Reinforce coverage.

Every story in the WeSearch catalog tagged with #reinforce, chronological, with view counts. Subscribe to the per-tag RSS feed to follow this topic in your reader of choice.

60 stories tagged with #reinforce, in publish-time order across the WeSearch catalog. Tag pages update as new stories ingest.

⌘ RSS feed for this tag →   or   search "Reinforce"

RELATED TAGS
#ai58#reinforcement-learning57#ml39#reinforcementlearning5#language-models4#machinelearning3#learning2#technology2#planning2#superintelligence1#tech-startups1#open-source1
SEEKING ALPHA

Alphabet: Q2 Strength Reinforces The Bullish Case, A Rare Pullback Worth Buying

Alphabet delivered robust Q2 FY26 results, with 24% top-line growth and strong momentum in AI-driven Cloud and Services segments. Click for more on GOOG stock.…

2 views ·
#alphabet#strength#reinforces
THE HINDU

Ratnaveer Precision Engineering Reports 20% Revenue Growth and 21% PAT Growth in Q1 FY27; CCL Project, Credit Rating Upgrade and Rights Issue Approvals Reinforce Growth Strategy

Ratnaveer Precision Engineering Reports 20% Revenue Growth and 21% PAT Growth in Q1 FY27; CCL Project, Credit Rating Upgrade and Rights Issue Approvals Reinforce Growth Strategy…

4 views ·
#ratnaveer#precision#engineering
KOREA TIMES NEWS

Korea must help reinforce nuclear umbrella

A recent Foreign Affairs article by Jennifer Lind and Daryl G. Press, "The Broken Nuclear Umbrella," has reignited a debate on a decision South Kor...…

4 views ·
#korea#must#help
THE JERUSALEM POST | JPOST.COM

West Bank security situation rapidly snowballing, forces will be reinforced, military officials say

A total of 26 battalions are now operating across the sector, with a security official warning on Sunday that the situation was "a snowball rolling downhill."…

5 views ·
#west#bank#security
YAHOO SPORTS

Real Madrid planning two further reinforcements after Yan Diomande deal

The manner in which Real Madrid are incessantly pushing for an agreement with RB Leipzig, Yan Diomande appears set to be Los Blancos‘ next recruit of the summer.After going trophy-…

8 views ·
#real#madrid#planning
RFI ENGLISH

Macron calls for EU reinforcements as France battles spreading wildfires

8 views ·
FORTUNE

Wildfires force 80,000 evacuations in France and Spain as EU sends reinforcements

France and Spain sought EU help after wildfires forced about 80,000 people from homes, resorts and villages.…

8 views ·
#wildfires#force#evacuations
TOWARDS DATA SCIENCE

Context Windows Forget What Matters — I Built a Usage-Reinforced Decay Engine for AI Agent Memory

Most AI memory systems keep the newest information—not the most important. Here's how I used the Ebbinghaus forgetting curve to build a better memory engine for LLMs. The post Cont…

12 views ·
#context#windows#forget
ARXIV CS.AI

Multimodal Reward Hacking in Reinforcement Learning

Reinforcement learning (RL) is increasingly used to align multimodal large language models (MLLMs), but higher rewards do not always imply better task performance. This risk is amp…

23 views ·
#multimodal#reward#hacking
ARXIV.ORG

Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning

arXiv:2606.27483v1 Announce Type: new Abstract: Large language model (LLM) agents have demonstrated strong capability in sequential decision-making, yet they remains fundamentally …

39 views ·
#artificial-intelligence#machine-learning#language-models
SEEKING ALPHA

DiDi Global: Strong Q1 Reinforces Our Long-Term Investment Thesis

DiDi Q1 2026 update: turnover +10%, Mobility EBITDA +26%, intl revenue +60%, $6.7B cash and buybacks. Read the full analysis here.…

32 views ·
#finance#earnings#investment
PHYS.ORG

Dead Sea archaea sport reinforced swimming tail for hypersalty waters

33 views ·
WASHINGTON EXAMINER

On This Day: With funds and reinforcements in hand, Washington returns to New York ready for battle

The Second Continental Congress is developing military measures in consultation with George Washington.…

30 views ·
#history#military#revolutionary war
ARXIV CS.AI

InfoMem: Training Long-Context Memory Agents with Answer-Conditioned Information Gain

Long-context tasks require LLMs to identify and preserve answer-relevant information from large contexts. Chunk-wise memory agents address this issue by sequentially reading docume…

39 views ·
#artificial intelligence#machine learning#reinforcement learning
ARXIV CS.AI

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning

Autonomous LLM training is often framed as recipe search, which leaves the training harness largely static. This limitation sharpens in agentic RL, where shifting bottlenecks and s…

43 views ·
#artificial intelligence#machine learning#reinforcement learning
TOM'S HARDWARE

Kevin O'Leary claims Chinese propaganda is to blame for anti-datacenter backlash, 'hundreds of millions of dollars' being spent to kill US dominance in AI — industry proponents and Trump administration reinforce claims of foreign interference

Just because you're paranoid doesn't mean they're not after you.…

39 views ·
#technology#ai#foreign interference
DEV.TO (TOP)

Explainable Causal Reinforcement Learning for planetary geology survey missions with embodied agent feedback loops

It was 3 AM, and I was staring at a terminal window filled with telemetry data from a simulated Mars rover. The reinforcement learning (RL) agent I had trained overnight had just c…

35 views ·
#ai#reinforcement learning#planetary science
NPR — SCIENCE

Bean plants call for aerial reinforcements when caterpillars attack

NPR's Short Wave talks about a weakness in a well-known insect repellant, how plants call wasps to their defense and how bigger rewards speed up learning, in mice.…

54 views ·
#science#plants#insects
SOUTH CHINA MORNING POST

Global Prosperity Summit 2026 reinforces Hong Kong’s strategic importance in advancing Apec cooperation and global governance

GPS 2026 highlights the city as a ‘bridge for exchange’ and regional growth driver ahead of this year’s Apec meeting in Shenzhen.…

39 views ·
#global governance#apec#hong kong
DEV.TO (TOP)

Why I built the HuggingFace for RL agents — and why RL needs one

Showcase Video If you've ever tried MineRL or OpenAI Five, you know the feeling. The environment...…

23 views ·
#reinforcement learning#technology#ai
CONSERVATIVEHOME

Clark Vasey: Why winning a London council seat reinforces my belief in blue collar Conservatism

Ours was a relentlessly anti-Labour campaign. Even when our canvassing results showed us losing votes to Reform, we did not shift our focus from Labour. We didn’t attack Reform or …

40 views ·
#politics#conservatism#local elections
DEV.TO (TOP)

ARTIST: RL-Powered Tool Use for LLM Agents Explained

How Microsoft's ARTIST framework uses outcome-based RL to train LLMs that interleave tool calls inside reasoning chains — no step supervision required.…

30 views ·
#reinforcementlearning#llmagents#tooluse
ARXIV CS.AI

GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training

Hybrid post-training usually combines supervised fine-tuning and reinforcement learning, but fixed mixing schedules cannot adapt when the relative noise of the two signals changes …

31 views ·
#machine learning#artificial intelligence#reinforcement learning
ARXIV CS.AI

Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases

Reinforcement Learning from Human Feedback (RLHF) is the standard method to align Large Language Models (LLMs) with human preferences. In this work, we introduce alignment tamperin…

39 views ·
#artificial intelligence#machine learning#bias
ARXIV CS.AI

StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning

Reinforcement learning for multi-turn agents suffers from a credit-assignment mismatch: rewards are sparse and trajectory-level, while success often hinges on a few local decisions…

35 views ·
#artificial intelligence#reinforcement learning#machine learning
ARXIV CS.AI

Mind the Tool Failures: Achieving Synergistic Tool Gains for Medical Agents

Medical AI agents increasingly use external tools for diagnosis, treatment recommendation, and evidence retrieval, yet most existing approaches assume that task-appropriate tools a…

42 views ·
#artificial intelligence#medical#reinforcement learning
ARXIV CS.AI

UnityMAS-O: A General RL Optimization Framework for LLM-Based Multi-Agent Systems

LLM-based multi-agent systems decompose complex tasks into interacting roles, but most remain manually orchestrated by prompts, tools, and control rules, while agents are rarely op…

39 views ·
#artificial intelligence#reinforcement learning#multi-agent systems
ARXIV CS.AI

Exploiting Local Dynamics Regularity for Reusable Skills in Offline Hierarchical RL

Hierarchical Reinforcement Learning (HRL) promises to solve long-horizon Reinforcement Learning (RL) tasks more efficiently than non-hierarchical counterparts by discovering and re…

38 views ·
#artificial intelligence#reinforcement learning#machine learning
ARXIV.ORG

Polar: Agentic RL on Any Harness at Scale

Reinforcement learning for language agents increasingly depends on custom harnesses that manage long-running context, multi-turn tool use and multi-agent orchestration. However, po…

36 views ·
#reinforcement learning#machine learning#software engineering
ARXIV CS.AI

Credit Assignment with Resets in Language Model Reasoning

Contemporary reinforcement learning with verifiable reward methods post-train language models on multi-step reasoning by assigning a single outcome reward uniformly across all toke…

25 views ·
#artificial intelligence#reinforcement learning#language models
ARXIV CS.AI

Evolutionary Enhanced Multi-Agent Reinforcement Learning for Cooperative Air Combat

As modern air combat evolves toward beyond-visual-range (BVR) multi-aircraft cooperative engagements, autonomous decision-making for unmanned combat aerial vehicles (UCAVs) faces s…

32 views ·
#artificial intelligence#reinforcement learning#air combat
ARXIV CS.AI

ProActor: Timing-Aware Reinforcement Learning for Proactive Task Scheduling Agents

Proactive task-oriented agents must autonomously anticipate user needs, identify actionable opportunities, and trigger software actions at appropriate moments - fundamentally shift…

32 views ·
#artificial intelligence#reinforcement learning#task scheduling
ARXIV CS.AI

CoRe-Code: Collaborative Reinforcement Learning for Code Generation

Large language models (LLMs) have achieved strong performance in code generation, but most methods rely on autoregressive decoding without global planning, often leading to locally…

27 views ·
#artificial intelligence#code generation#reinforcement learning
ARXIV CS.AI

Emission-Aware Reinforcement Learning for Sustainable Electric Vehicle Charging and Carbon Dioxide Reduction Under Varying Renewable Penetration

The rapid growth of Electric Vehicle (EV) adoption challenges power distribution networks through peak load spikes, voltage instability, and transformer overloads from uncoordinate…

43 views ·
#artificial intelligence#electric vehicles#renewable energy
ARXIV CS.AI

Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork

In-Context Reinforcement Learning (ICRL) has enabled foundation agents to adapt instantaneously to novel tasks, yet its efficacy in Ad-Hoc Teamwork (AHT)-where coordination with un…

34 views ·
#artificial intelligence#reinforcement learning#teamwork
ARXIV CS.AI

When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs

Multi-agent LLM workflows route inference through specialized roles to lift end-task accuracy, but jointly training those roles with reinforcement learning is unstable in ways that…

29 views ·
#artificial intelligence#machine learning#reinforcement learning
THE HINDU — TOP

Quad Foreign Ministers meet LIVE: Amid West Asia turmoil, Quad Foreign Ministers gather to reinforce Indo-Pacific stability

Follow live updates from The Hindu as Quad Foreign Ministers from the United States, Australia and Japan hold a crucial meeting in New Delhi on May 26, 2026…

45 views ·
#diplomacy#security#economy
DEV.TO (TOP)

Understanding Reinforcement Learning with Human Feedback Part 5: Training the Reward Model with Loss Functions

In the previous article, we created a reward model. In this article, we will continue exploring how...…

37 views ·
#ai#machinelearning#reinforcementlearning
GITHUB

Show HN: Reward Is Not Reinforcement Until Admitted

Contribute to nikitph/rewarder development by creating an account on GitHub.…

37 views ·
#research#technology#experiments
YAHOO SPORTS

Spalletti laments lack of character in Juventus squad, expects summer reinforcement

Juventus manager Luciano Spalletti was displeased with his team after squandering a two-goal lead in the Derby della Mole.Due to the one-hour delay, the Bianconeri had already know…

32 views ·
#football#juventus#spalletti
R/MACHINELEARNING

If you use NVIDIA Isaac Sim for reinforcement learning, do you use Isaac Lab with it? Just want to get a sense of what the status quo is. [D]

32 views ·
ARXIV CS.AI

Reflex: Reinforcement Learning with Reflection Symmetry Exploitation in State-Based Continuous Control

Reinforcement learning has long struggled with poor sample efficiency. One promising approach to mitigate this problem is leveraging group-invariant Markov Decision Processes ($G$-…

31 views ·
#machine learning#reinforcement learning#artificial intelligence
ARXIV CS.AI

Curriculum reinforcement learning with measurable task representation learning

In curriculum reinforcement learning (CRL), an agent incrementally accumulates knowledge over a sequence of tasks (i.e., a curriculum), and the learning process is aimed at using t…

34 views ·
#machine learning#artificial intelligence#reinforcement learning
ARXIV CS.AI

Score-Based One-step MeanFlow Policy Optimization

Diffusion and flow matching have emerged as expressive policy classes in reinforcement learning, but their reliance on multi-step denoising imposes substantial computational overhe…

37 views ·
#machine learning#reinforcement learning#artificial intelligence
ARXIV CS.AI

Reinforcement Learning for Microcanonical Graph Ensemble with Assortativity Constraints

How network structure determines function is a fundamental question, and it can be investigated by graph ensembles with precisely controlled structural properties. Canonical approa…

32 views ·
#machine learning#graph theory#reinforcement learning
ARXIV CS.AI

Infra-Bayesian Reinforcement Learning Agents Outperform Classical RL For Worst-Case Robustness

Classical reinforcement learning assumes the agent interacts with a fixed environment whose behavior does not depend on the agent's policy. This assumption breaks down in non-reali…

34 views ·
#machine learning#artificial intelligence#reinforcement learning
ARXIV CS.AI

Classical State Preparation for Variational Quantum Algorithms via Reinforcement Learning

Variational Quantum Algorithms (VQAs) potentially offer a pathway to practical quantum advantage, but their optimization is heavily hindered by barren plateaus and numerous local m…

38 views ·
#quantum physics#artificial intelligence#machine learning
ARXIV CS.AI

Dreaming Smoothly and Sample Efficiently with Gradient Penalized Latent Dynamics

Model-based reinforcement learning improves sample efficiency by learning a world model. However, existing latent world models such as DreamerV3 do not explicitly enforce local smo…

24 views ·
#machine learning#reinforcement learning#artificial intelligence
ARXIV CS.AI

One Policy, Infinite NPCs: Persona-Traceable Shared RL Policies for Scalable Game Agents

On a 300-persona life-simulation benchmark, pcsp achieves compositional zero-shot persona identification up to 17x above chance, Spearman rho approx 0.73 semantic-behavioral alignm…

32 views ·
#artificial intelligence#gaming#reinforcement learning
HOT AIR

How News Aggregators Reinforce Political Ignorance

How news aggregators reinforce political ignorance by filtering biased content and confirming beliefs.…

33 views ·
#media#politics#news
YAHOO SPORTS

Mets Morning News: Reinforcements are coming, but will it matter?

Your Sunday morning dose of New York Mets and MLB news, notes, and links.…

21 views ·
#mlb#new york mets#baseball
DEV.TO (TOP)

Understanding Reinforcement Learning with Human Feedback Part 4: Teaching Models Human Preferences

In the previous article, we explored the part where we collect human preferences. In this article, we...…

32 views ·
#ai#machinelearning#reinforcementlearning
DAILY SIGNAL

News Aggregators’ Left-Wing Bias Reinforces Political Ignorance

The other day, a young lady got on the elevator and promptly whipped out her smartphone and began scrolling. “Excuse me,” I said, “can I ask you where you get your news?” She said,…

35 views ·
#media#politics#bias
USCC

Two Loops: How China's Open AI Strategy Reinforces Its Industrial Dominance [pdf]

28 views ·
INVESTING.COM — NEWS

Nexpace Announces NXPC Buyback Program to Reinforce User-Centered Ecosystem Growth in MapleStory Universe

36 views ·
ARXIV CS.AI

Distribution-Aware Reward: Reinforcement Learning over Predictive Distributions for LLM Regression

Large language models can predict real-valued quantities from heterogeneous inputs such as text, code, and molecular strings, but most training objectives score each decoded floati…

27 views ·
#machine learning#reinforcement learning#language models
ARXIV CS.AI

AGPO: Adaptive Group Policy Optimization with Dual Statistical Feedback

Reinforcement learning improves LLM reasoning, but PPO/GRPO typically use fixed clipping and decoding temperature, which makes training brittle and tuning-heavy. We propose Adaptiv…

30 views ·
#machine learning#artificial intelligence#reinforcement learning
ARXIV CS.AI

Design for Manufacturing: A Manufacturability Knowledge-Integrated Reinforcement Learning Framework for Free-Form Pipe Routing in Aeroengines

Design for manufacturing plays a critical role in advanced aeroengine development, where complex components necessitate careful consideration of manufacturability. However, current…

33 views ·
#machine learning#manufacturing#aeroengines
ARXIV CS.AI

Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs

We introduce a novel method that averages the logits of a frozen reference policy (e.g., SFT) and a trainable policy, and incorporate the method into Group Relative Policy Optimiza…

29 views ·
#machine learning#artificial intelligence#reinforcement learning
ARXIV CS.AI

Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor

MXFP4 arithmetic can dramatically accelerate reinforcement learning (RL) post-training of large language models (LLMs), yet the quantization error introduces severe accuracy degrad…

24 views ·
#machine learning#artificial intelligence#quantization