RL systems · Architecture note · 14 min read
One sandbox can run RL. A pool is built for the loop.
When the reward comes from running code a model wrote, that code is untrusted and it runs millions of times. The trainer, a general sandbox, and a dedicated rollout plane can all run it. They draw different boundaries. The important question is not “can it run Python?” It is “where does the model's code execute, and how many round trips does the trainer pay?”
01 · Three models
The same policy loop, three execution boundaries.
Grading on the trainer runs model output on the same host as the policy. A general BoltzLabs sandbox moves that execution into an isolated machine. A dedicated RL pool keeps the trainer where the model runs and gives every attempt its own isolated environment behind a batched API.
| Model | Unit you manage | Step boundary | Best fit |
|---|---|---|---|
| On the trainer | Subprocesses on the trainer host | A local process per check | Trusted code only |
| General sandbox | One isolated machine | One run call per check | Isolating a small grading job |
| Dedicated RLPool | A pool of N environments | One request carrying N actions | Grading every rollout at scale |
02 · On the trainer
Grading locally works until the code is untrusted.
The simplest reward runs each sampled program against a few tests in a subprocess. There is no network boundary. There is also no isolation: the model's program shares the trainer's files, credentials, network, and GPU host, and early in training it writes whatever maximises reward, including code that deletes files or never exits.
import subprocess
TESTS = """
assert solve([3, 1, 2]) == [1, 2, 3]
assert solve([]) == []
"""
def grade(code):
# The model's program runs on the trainer host,
# beside your checkpoints, credentials and network.
check = subprocess.run(["python3", "-c", code + TESTS], capture_output=True, timeout=30)
return float(check.returncode == 0)
codes = sample(model, prompts) # your sampler: one program per prompt
rewards = [grade(code) for code in codes]03 · General sandbox
Run every check in one sandbox.
A BoltzLabs Sandbox is an isolated machine with
a writable workspace. Sending each program there instead of to a local subprocess is the direct
path from “runs on my trainer” to “runs somewhere it cannot hurt anything”.
import shlex
from boltzlabs import Sandbox
TESTS = """
assert solve([3, 1, 2]) == [1, 2, 3]
assert solve([]) == []
"""
codes = sample(model, prompts) # your sampler: one program per prompt
with Sandbox(machine="medium", environment="python") as sb:
rewards = []
for code in codes:
result = sb.exec(f"python3 -c {shlex.quote(code + TESTS)}", timeout=30)
rewards.append(float(result.exit_code == 0))The API creates the machine, runs each check, returns its output, and destroys the machine when the context manager exits. Two limits remain: every candidate shares that one workspace, so a program can leave files behind that change how the next one is graded, and the checks run one request at a time.
What a general sandbox does not provide
To grade a whole batch in parallel, with a clean workspace for every attempt, the generic
sandbox API has no environment-level step primitive. You would need to build and operate the bridge yourself.
# A general sandbox does not expose env.step() as a remote primitive.
# To keep the trainer local while environments run remotely, you would own:
for box in sandboxes:
upload_environment_code(box)
start_long_running_rpc_bridge(box)
while training:
actions = policy(observations)
results = fan_out_requests(sandboxes, actions)
observations = restore_index_order(results)
restart_dead_bridges()
reset_finished_environments()
# RLPool is the managed version of this control loop.04 · Dedicated RL
Make the batch the unit of work.
A trainer produces N actions before it can advance. RLPool therefore sends one request containing all N actions and receives one ordered batch containing
all N observations, rewards, done flags, and info objects. The network cost is paid once per
policy step, not once per environment.
The environment contract
Each environment is a small program started once in its own isolated worker. Here the
action is the model's program and the reward is whether your tests pass. The environment
receives newline-delimited JSON over stdin and writes one JSON reply per operation. The
SDK's serve() helper owns flushing, protects the protocol
from stray stdout, and converts user exceptions into terminal transitions.
# code_env/env.py — uploaded once, read-only at /env in every environment
from pathlib import Path
import subprocess
from boltzlabs.env import serve
def reset(seed):
Path("/workspace/candidate.py").unlink(missing_ok=True)
return {"task": Path("/env/TASK.md").read_text()}
def step(action):
Path("/workspace/candidate.py").write_text(action["code"])
check = subprocess.run(
["python3", "-m", "unittest", "discover", "-s", "/env/tests"],
cwd="/workspace", capture_output=True, text=True, timeout=30,
)
output = check.stdout + check.stderr
return {"test_output": output[-2000:]}, float(check.returncode == 0), True, {}
serve(reset=reset, step=step)The training loop
from boltzlabs import RLPool
with RLPool("./code_env", n=64) as pool:
obs = pool.reset(seed=0)
for _ in range(1_000):
codes = sample(model, [o["task"] for o in obs]) # your policy: one program per env
obs, rewards, dones, infos = pool.step([{"code": c} for c in codes]) # one HTTP request
update(model, codes, rewards) # your trainer
obs = pool.reset(where=dones) # a clean workspace for every finished attempt
print(pool.timing.worker_ms)
print(pool.timing.roundtrip_ms)
print(pool.timing.stragglers)05 · Inside the worker
Warm processes, not snapshots.
Package once
The SDK creates a deterministic compressed archive of the environment directory and vendors the tiny serve helper. The upload happens at pool creation, not on every step.
Mount code read-only
The worker extracts one code directory and mounts it read-only at /env in every environment. Each environment receives a private writable /workspace.
Start every environment
Creation is staggered to avoid an interpreter-startup thundering herd. The pool is not returned until every environment answers its initial reset handshake.
Keep processes resident
Python imports, simulator state, and heap objects remain alive between steps. There is no per-step process spawn and no process-memory snapshot store.
Fan out in index order
The worker sends each action to its environment and restores replies to their original batch positions.
Measure shared memory honestly
Status reports proportional memory use so shared interpreter and library pages are not counted in full for every environment.
06 · Resets and failure
Fast reset when possible. Respawn when necessary.
Soft reset
In-process episode reset
pool.reset(where=dones) sends reset only
to finished environments. Other observations remain in their original slots.
Hard reset
Kill, wipe, respawn
pool.reset(hard=True) recreates the private
workspace and process. It is a cold start, not snapshot restoration.
A process that misses its per-step deadline is killed so one wedged environment cannot
freeze the batch. Its result is explicit: null observation, zero reward, done=True, and info["boltzlabs_straggler"]. The
next reset respawns it.
07 · Choose deliberately
Use the smallest system that solves the bottleneck.
Stay local
The reward never runs model output, and it fits beside the trainer.
Use one sandbox
You need isolation for a small grading job and can run checks one at a time.
Use RLPool
Every rollout runs untrusted code and needs a clean, isolated environment, at the scale of a training batch.
Implementation status, without benchmark theatre
The pool engine, control-plane routes, SDK, partial resets, straggler path, and memory accounting are implemented and covered by worker and SDK tests. Production deployment and publishable amd64 benchmark results are not claimed here. The SDK deliberately reports worker time and end-to-end round-trip time separately so future measurements cannot hide network cost.