PRIME Intellect
How Extropic Uses Prime Intellect to Train a Thermodynamic ML Research Agent
Case studies/Extropic

How Extropic Uses Prime Intellect to Train a Thermodynamic ML Research Agent

2.8x

held-out reward gain

100 GRPO steps

to achieve that improvement

~25 hours

training run wall-clock time

At a glance

Extropic post-trained an open model, Qwen3.6-35B-A3B, to reproduce classic thermodynamic machine learning experiments, nearly tripling its reward metric on held-out tasks in about 100 GRPO steps.

Using Prime Intellect’s Open Superintelligence Stack, they built a customized RL environment with verifiers and trained and evaluated it with Hosted Training, Prime Sandboxes, and Prime Inference.

“Prime Intellect has been immensely helpful in making our first experiments possible, providing optimized inference and training tools, reproducible code sandboxes, and GPU resources to train and benchmark open-source models. We're starting by teaching agents to reproduce classic thermodynamic machine learning experiments. These are the early innings of Thermo RSI. We look forward to further accelerating the co-evolution of the next generation of hardware and algorithms with these tools.”

Gill Verdon

Founder & CEO, Extropic

The problem

Extropic’s mission is to solve AI’s ever-increasing energy demands by building the thermodynamic computing stack, harnessing randomness in hardware rather than simulating it in software. On the hardware side, its thermodynamic sampling units (TSUs) are all-transistor circuits that leverage natural thermal noise to sample directly from probabilistic models. However, TSUs are built to run a family of models developed in the mid-2000s, so few modern algorithms exist for them.

To develop those algorithms, Extropic is betting on what it calls thermodynamic recursive self-improvement: research agents design and run experiments, developing new algorithms for TSUs, which enable better hardware and more capable agents. Their first step, described in Recursive intelligence for a new substrate, is teaching an agent to reproduce classic connectionist experiments in the form of coding tasks.

This meant running agentic RL on a 35B model, with untrusted model-written code executed on every rollout and a judge model scoring every attempt. It also required GPU clusters, RL training code, and sandboxing to all work together on day 1. For a small team, building out such extensive infrastructure from scratch would have been time-consuming and painful.

“For a team our size, this is the difference between doing the research and spinning our wheels on infrastructure.”

Gill Verdon

Founder & CEO, Extropic

Using Prime Intellect’s Open Superintelligence Stack

Extropic used Prime Intellect’s Open Superintelligence Stack to train and evaluate its research agent, combining Hosted Training, verifiers, Prime Sandboxes, and Prime Inference.

Extropic’s training workflow with Prime Intellect
Prime Intellect’s Open Superintelligence Stack. Components used by Extropic are highlighted in yellow.

This allowed them to nearly triple Qwen3.6’s baseline performance on held-out tasks in 100 steps with a wall clock of ~25 hours, all without Extropic managing the large scale multi-node GPU infrastructure required for this RL training run.

Designing the environment

Extropic built their RL environment with verifiers, Prime Intellect's open-source library for building environments. This environment contains about 50 coding tasks adapted from classic connectionist experiments. In each task, the model works in a Python REPL inside a Prime Sandbox. It writes code, runs it, reads any error, revises, and submits a solution.

Each submission is scored in two parts:

r=0.7⋅exec+0.3⋅rubricr = 0.7 \cdot \text{exec} + 0.3 \cdot \text{rubric}

Execution (70%) measures whether the code runs and whether it reproduces the reference metric from the original experiment. Rubric (30%) is graded by an LLM judge, Nemotron 3 Super 120B, which scores each solution against a per-task rubric.

In this experiment, execution carries most of the weight by design. Scoring on execution means running untrusted, model-written code on every rollout during training. Each change to the reward produced a new environment version, so Extropic always knew which scores could be compared.

RL loops

“On self-managed infrastructure, most of those runs would not have happened, and we would have shipped a smaller project”

Gill Verdon

Founder & CEO, Extropic

Extropic trained two Qwen models with reinforcement learning, using GRPO as the objective. They ran training on Hosted Training, Prime Intellect's managed RL training. For each run, Extropic's researchers chose the base model, built the environment with verifiers, and wrote the run config. Prime ran rollout generation and training on its managed GPU infrastructure, so Extropic's team never had to set up or manage any.

Each training step had two halves: rollouts and an update. Each rollout ran in its own Prime Sandbox, which executed the model’s code and returned the execution score. Prime Inference served the LLM judge, which graded the rubric score for every rollout.

In the update, GRPO compared the model's attempts at each task against one another. It reinforced the attempts that scored above the group's average and discouraged those below it.

During training, Extropic tracked performance on held-out tasks the model never trained on. After training, they deployed checkpoints on Prime Inference and ran their evals in parallel. The frontier baselines also ran through Prime Inference against the same environment.

Results

After 100 GRPO steps on Hosted Training, Qwen3.6-35B-A3B's reward on held-out problems rose from 0.127 to 0.361, a 2.8x gain.

Held-out task rewards before and after reinforcement learning
Held-out task rewards before and after reinforcement learning. Qwen3.6 improves from 0.127 to 0.361 after 100 GRPO steps.

The recipe transferred across model families: Qwen3.5, post-trained the same way, landed almost level with Qwen3.6. Both finished well ahead of other open models and closed much of the gap to Claude Opus 4.8, with a model that activates only 3B parameters per token.

What's next

Reproducing classic experiments was the first step. Extropic’s ultimate goal is thermo RSI: agents that discover new sampling-based learning rules for TSUs, expanding what the hardware can do, which then supports more capable agents.

Their next experiments will rely on the same infrastructure that made the first ones possible.

Read Extropic’s post: Recursive intelligence for a new substrate.

To post-train a model for your own task, get started with Prime Intellect.