5 minute read

Recursive self-improvement (RSI) has shown great potential of pushing the boundaries of scientific discovery. While we see much progress being made with autonomous research in science [1, 2], we believe that there is still one big gap: a setup that connects scientific questions to a hill-climbable environment for agents.

RSI rarely works on a science problem out of the box, and this project aims to get a taste of what it takes to change that. In the spirit of NanoGPT, we made protein language model research hill-climbable, with an open 666M-protein corpus, a readable trainer, frozen evaluations and an explicit protocol, and we show that a plain sequential search with a human gate already produces a much better training recipe than the ESMC baseline.

A better PLM recipe by hill-climbing NanoProteinLM

If a research question can be turned into a loop that an agent runs without waiting on a person, the agent can try far more ideas than a lab could by hand, and each change it keeps becomes the starting point for the next. We wanted to see how far that goes on a real problem, so we tried it on protein language model training.

Training loss, contact precision and validation loss of the ESMC baseline and the scaled recipe over 100k training steps
The ESMC baseline (blue) and the recipe found with autoresearch (orange), trained on the same 48.4B tokens. Left, training loss. Right, contact precision (solid) and validation loss (dashed).

Starting from the ESMC recipe, a simple sequential search, with a human deciding which recipes to scale up, gave a recipe with about 30% higher long-range contact precision, from 28.2% to 36.6% P@L, at the same 48.4B-token budget.

Behind the scenes, the hard problem did not lie in the design of the autoresearch harness.1Side note 1We use a very simple Karpathy-style sequential search, which is certainly not the best harness one could build. Our aim is not a better autoresearch algorithm; we want to understand what makes a task hill-climbable. It lay in the environment around the scientific loop: the data, the evaluation, the budget and the rules that let the loop run unattended and still tell us something true about proteins.

Building hill-climbable environments for research tasks

Not every science question is hill-climbable.2Side note 2One way to see this is something like Amdahl’s law for autoresearch. Split a research round into the time it takes to get a useful scientific signal back, by running the experiment and measuring the result, and the time a researcher spends deciding what to try next. Even an agent as capable as the best human researcher can only shrink the second part, so the speed-up it buys is at most \(1 + T_\text{decide} / T_\text{signal}\). Hill-climbing pays off most when progress is capped by researcher effort, as with compute experiments that finish in an hour while a person needs a day to choose the next one. When the smallest loop that returns a useful signal is itself slow or expensive, such as a week of growing cells, automation buys little and the quality of each decision matters more. For tasks where hill-climbing does pay off, we apply the following general principles to design the environment.

Item What it fixes
Research objective The scientific outcome the campaign should improve, fixed before search.
Design space What the agent may change and what must stay fixed, with bounds that exclude gains the objective does not value.
Round budget The resources each candidate may spend, identical for the baseline and every candidate so rewards compare across rounds.
Round reward A scalar measured on held-out validation data after each round, which the loop compares with the incumbent’s to keep or discard the candidate.
Validation A larger-budget retraining of the selected recipe against the baseline, testing whether search-time gains carry over.

Failure mode: Perfect on reward, failure on science

Here we show an example of the most typical way reward hacking makes autoresearch go wrong. To see how agents behave in this environment, we let Codex Astra and Claude Opus 5.5 each run 72 rounds of search against validation loss. Both lowered the loss, and after 12 hours of training both recipes still beat our reference on it. Their contact precision (a proxy metric for downstream protein folding), though, shows a large gap to the best setting.

Retained validation loss of the two agents over 72 rounds, beside a table of search and 12-hour scale-up results
Left, retained validation loss over 72 rounds from a shared baseline. Right, search results and a fixed 12-hour scale-up, with each recipe's gap to our reference recipe in colour.

The agent optimizes the reward as written, and any gap between that reward and the scientific goal is open to it. That is also why the evaluation took most of our effort in the previous section. A reward that has not been tested carefully can be hacked in more ways than you can imagine in advance.

What survived the human gate

Our main campaign ran a Karpathy-style sequential hill-climber, deliberately simple and far from optimal (side note 1), with the following settings:

  • Every candidate trains with two seeds and is kept only if its mean gain beats the seed-to-seed spread.
  • P@L is too noisy to steer by at one hour, so we climbed validation loss while watching P@L, then ran a second search on P@L itself.
  • A final scaled-up ablation of the confirmed improvements.
Two sequential searches: validation loss over 38 candidates and contact precision over 35 candidates
The two searches, rewarded by validation loss (left) and contact precision (right). Points are two-seed means; blue steps track the retained recipe; numbers mark accepted changes.
Search-time and 24.2B-token results for each accepted change, with shaded rows for the changes kept
Each recipe adds one accepted change to the recipe named before the plus sign. Shaded rows are the changes kept in the final recipe; deltas are against that parent at 24.2B tokens.

Four of the seven changes survive the final ablation study. All four change the optimizer, the loss weighting or how batches are split across GPUs.

Takeaway

NanoProteinLM carries our vision of an open research environment that connects scientists with agentic researchers.

For protein researchers, we provide a minimal reproduction of ESMC-style model training: public data, readable PyTorch code, training recipes and evaluations in one place. We aim to contribute an open-source foundation that researchers can understand, reproduce and extend in support of open science.

For agentic researchers, it provides a controlled environment for iterative autoresearch on the same scientific objective. Deterministic data selection, fixed seeds, explicit compute budgets and frozen evaluation protocols make recipe changes measurable. An agent can modify the training recipe, train, evaluate and improve it; the choice of agent and search strategy remains yours.

For more details, please refer to our paper, accepted as a spotlight presentation at the AgenticLS workshop (Agentic AI for Biological Discovery) at NeurIPS 2026.

Muchen Li (University of British Columbia) and Chixiang Lu (The University of Hong Kong)

@inproceedings{nanoproteinlm2026,
  title     = {NanoProteinLM: Towards Hill-Climbable Protein Language Model Research},
  author    = {Li, Muchen and Lu, Chixiang},
  booktitle = {NeurIPS 2026 Workshop on Agentic AI for Biological Discovery},
  year      = {2026}
}

References

  1. Anthropic. How Claude is uplifting biomolecular modeling. September 2026.
  2. Recursive Superintelligence, Inc. First steps toward automated AI research. June 2026.

Updated: