The world model remembers, the actor forgets: dissecting AI forgetting on 1 GPU
We ask a question the continual-RL literature has assumed an answer to but never measured: which component forgets? Under never-clear replay, pre-registered component-level probes (n=3 seeds throughout) show that the world model retains essentially everything measurable about old tasks -- reward discrimination (retention ratio ~1.0), value estimates, and termination structure -- while the actor's behavior collapses. Forgetting in this regime is a channel problem, not a memory problem.
- ▪We ask a question the continual-RL literature has assumed an answer to but never measured: which component forgets?
- ▪Under never-clear replay, pre-registered component-level probes (n=3 seeds throughout) show that the world model retains essentially everything measurable about old tasks -- reward discrimination (retention ratio ~1.0), value estimates, and
- ▪Forgetting in this regime is a channel problem, not a memory problem.
Hacker News (AI / LLM) files mainly under ai. We currently carry 2,165 of its stories.
Opening excerpt (first ~120 words) tap to expand
Computer Science > Machine Learning arXiv:2607.19749 (cs) [Submitted on 22 Jul 2026] Title:The World Model Remembers, the Actor Forgets: Dream Rehearsal for Continual Model-Based RL Authors:Gurp Nijjer View a PDF of the paper titled The World Model Remembers, the Actor Forgets: Dream Rehearsal for Continual Model-Based RL, by Gurp Nijjer View PDF HTML (experimental) Abstract:Model-based reinforcement-learning agents of the DreamerV3 family forget catastrophically when trained on task sequences, even when an unbounded replay buffer preserves every earlier experience.
…
Excerpt limited to ~120 words for fair-use compliance. The full article is at arXiv.org.