Skip to content
AI News HubLIVE
Source content · Analysis pending3 min read

Self-discovering RL in the Era of Experience: Is Learning History an Asset or a Burden?

Summary

arXiv:2609.35897v1 Announce Type: new Abstract: The pursuit of recursive self-improvement (RSI) toward general intelligence is divided between macro-level language model scaling and the interaction-driven principles of "Era of Experience". Yet, any self-improving architecture ultimately rests upon its underlying optimization engine: if general intelligence requires learning from grounded interaction, the reinforcement learning (RL) update rule itself must be capable of cumulative adaptation. While algorithm self-discovery has produced Disco103 that surpassed PPO to achieve SOTA benchmark performance -- its internal update machinery remains an uninspected black box. We present the first causal mechanistic audit of a self-discovered RL rule, structured directly around the five pillars of th…

SourcearXiv AIAuthor: Haomin Luo (University of Cambridge, Models2 AI)
Self-discovering RL in the Era of Experience: Is Learning History an Asset or a Burden?
Report an error

The correction channel is not available yet. You can copy the article reference below for later.

Correction instructions
Read article

[Submitted on 27 Sep 2026]

Title:Self-discovering RL in the Era of Experience: Is Learning History an Asset or a Burden?

View a PDF of the paper titled Self-discovering RL in the Era of Experience: Is Learning History an Asset or a Burden?, by Haomin Luo (1 and 2) ((1) University of Cambridge and 1 other authors

View PDF HTML (experimental)

Abstract:The pursuit of recursive self-improvement (RSI) toward general intelligence is divided between macro-level language model scaling and the interaction-driven principles of "Era of Experience". Yet, any self-improving architecture ultimately rests upon its underlying optimization engine: if general intelligence requires learning from grounded interaction, the reinforcement learning (RL) update rule itself must be capable of cumulative adaptation. While algorithm self-discovery has produced Disco103 that surpassed PPO to achieve SOTA benchmark performance -- its internal update machinery remains an uninspected black box. We present the first causal mechanistic audit of a self-discovered RL rule, structured directly around the five pillars of the Era of Experience: extended horizon, grounded reward scales, continuing streams, within-lifetime change, and exploration depth. By surgically pinning, freezing, and transplanting recurrent states while holding meta-parameters fixed, we test when learning history acts as an asset or a burden. Three findings organize the audit: (1) Recurrent history actively expands usable reward scales, sustaining a six-decade window versus three under zero-pinning. (2) Decoupling historical content from its maintenance reveals that the penalty of mismatched history stems from perpetual clamping; allowing imported state to evolve naturally attenuates this burden. (3) Under environmental change, controlling replay retention reverses the apparent adaptation advantage over DQN, demonstrating that external data turnover can confound internal plasticity. Validated through capability thresholds and ported to a second rule (OPEN), this work grounds macro-RSI ambitions in micro-level learning dynamics, establishing a foundational audit standard for next-generation, self-evolving RL algorithms.

Comments: 52 pages, 17 figures

Subjects:

Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

Cite as: arXiv:2609.35897 [cs.AI]

(or arXiv:2609.35897v1 [cs.AI] for this version)

https://doi.org/10.48550/arXiv.2609.35897

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Haomin Luo [view email] [v1] Sun, 27 Sep 2026 20:56:26 UTC (2,527 KB)

Full-text links:

Access Paper:

View a PDF of the paper titled Self-discovering RL in the Era of Experience: Is Learning History an Asset or a Burden?, by Haomin Luo (1 and 2) ((1) University of Cambridge and 1 other authors

View PDF

HTML (experimental)

TeX Source

view license

Ancillary-file links:

Ancillary files (details):

README.txt

analysis_scripts/build_data_figures.py

analysis_scripts/figure_style.py

analysis_scripts/openscience.mplstyle

analysis_scripts/plot_sources/behavior_panorama.py

analysis_scripts/plot_sources/legacy_plots.py

analysis_scripts/plot_sources/recovery_panorama.py

analysis_scripts/plot_sources/replay_detail.py

analysis_scripts/plot_sources/scale_panorama.py

analysis_scripts/plot_sources/scale_transplant_detail.py

evidence_manifests/behavior_panorama.json

evidence_manifests/legacy_plots.json

evidence_manifests/recovery_panorama.json

evidence_manifests/replay_detail.json

evidence_manifests/scale_panorama.json

evidence_manifests/scale_transplant_detail.json

figure_data/app03_interface.data.json

figure_data/app04_streams.data.json

figure_data/app06_exploration.data.json

figure_data/app07_baselines.data.json

figure_data/integrated_behavior.data.json

figure_data/integrated_scale_interface.data.json

figure_data/replay_main.data.json

figure_data/scale_main.data.json

figure_data/transplant_main.data.json

figure_data/unified_behavior_summary.data.json

figure_data/unified_mapping.data.json

figure_data/unified_recovery_catch.data.json

figure_data/unified_recovery_overview.data.json

figure_data/unified_scale_full.data.json

figure_data/unified_state_scale.data.json

(26 additional files not shown)

Current browse context:

cs.AI

new | recent | 2026-09

Change to browse by:

cs cs.LG

References & Citations

NASA ADS

Google Scholar

Semantic Scholar

Loading...

Data provided by:

Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)

Related Papers

Recommenders and Search Tools

Link to Influence Flower

Influence Flower (What are Influence Flowers?)

Core recommender toggle

CORE Recommender (What is CORE?)

Author

Venue

Institution

Topic

About arXivLabs

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Key points and analysis

Article intelligence

EngineersAdvanced

Key points

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • arXiv:2609.35897v1 Announce Type: new Abstract: The pursuit of recursive self-improvement (RSI) toward general intelligence is divided between macro-level language model scaling a…

Highlights and analysis are generated automatically and may contain errors. Check the original source.