RL systems · Architecture note · 14 min read

One sandbox can run RL. A pool is built for the loop.

When the reward comes from running code a model wrote, that code is untrusted and it runs millions of times. The trainer, a general sandbox, and a dedicated rollout plane can all run it. They draw different boundaries. The important question is not “can it run Python?” It is “where does the model's code execute, and how many round trips does the trainer pay?”

01 · Three models

The same policy loop, three execution boundaries.

Grading on the trainer runs model output on the same host as the policy. A general BoltzLabs sandbox moves that execution into an isolated machine. A dedicated RL pool keeps the trainer where the model runs and gives every attempt its own isolated environment behind a batched API.

ModelUnit you manageStep boundaryBest fit
On the trainerSubprocesses on the trainer hostA local process per checkTrusted code only
General sandboxOne isolated machineOne run call per checkIsolating a small grading job
Dedicated RLPoolA pool of N environmentsOne request carrying N actionsGrading every rollout at scale

02 · On the trainer

Grading locally works until the code is untrusted.

The simplest reward runs each sampled program against a few tests in a subprocess. There is no network boundary. There is also no isolation: the model's program shares the trainer's files, credentials, network, and GPU host, and early in training it writes whatever maximises reward, including code that deletes files or never exits.

train.py
import subprocess

TESTS = """
assert solve([3, 1, 2]) == [1, 2, 3]
assert solve([]) == []
"""

def grade(code):
    # The model's program runs on the trainer host,
    # beside your checkpoints, credentials and network.
    check = subprocess.run(["python3", "-c", code + TESTS], capture_output=True, timeout=30)
    return float(check.returncode == 0)

codes = sample(model, prompts)  # your sampler: one program per prompt
rewards = [grade(code) for code in codes]
Do not distribute by reflex. A remote pool adds JSON serialization, a network hop, fleet placement, and another failure domain. If the reward never executes model output, stay local. Move execution out when isolation, CPU, memory, or reproducibility is the constraint.

03 · General sandbox

Run every check in one sandbox.

A BoltzLabs Sandbox is an isolated machine with a writable workspace. Sending each program there instead of to a local subprocess is the direct path from “runs on my trainer” to “runs somewhere it cannot hurt anything”.

sandbox_job.py
import shlex
from boltzlabs import Sandbox

TESTS = """
assert solve([3, 1, 2]) == [1, 2, 3]
assert solve([]) == []
"""

codes = sample(model, prompts)  # your sampler: one program per prompt

with Sandbox(machine="medium", environment="python") as sb:
    rewards = []
    for code in codes:
        result = sb.exec(f"python3 -c {shlex.quote(code + TESTS)}", timeout=30)
        rewards.append(float(result.exit_code == 0))

The API creates the machine, runs each check, returns its output, and destroys the machine when the context manager exits. Two limits remain: every candidate shares that one workspace, so a program can leave files behind that change how the next one is graded, and the checks run one request at a time.

What a general sandbox does not provide

To grade a whole batch in parallel, with a clean workspace for every attempt, the generic sandbox API has no environment-level step primitive. You would need to build and operate the bridge yourself.

orchestrator.py
# A general sandbox does not expose env.step() as a remote primitive.
# To keep the trainer local while environments run remotely, you would own:

for box in sandboxes:
    upload_environment_code(box)
    start_long_running_rpc_bridge(box)

while training:
    actions = policy(observations)
    results = fan_out_requests(sandboxes, actions)
    observations = restore_index_order(results)
    restart_dead_bridges()
    reset_finished_environments()

# RLPool is the managed version of this control loop.

04 · Dedicated RL

Make the batch the unit of work.

A trainer produces N actions before it can advance. RLPool therefore sends one request containing all N actions and receives one ordered batch containing all N observations, rewards, done flags, and info objects. The network cost is paid once per policy step, not once per environment.

policy → actions[N]
one HTTP step
worker fan-out × N

The environment contract

Each environment is a small program started once in its own isolated worker. Here the action is the model's program and the reward is whether your tests pass. The environment receives newline-delimited JSON over stdin and writes one JSON reply per operation. The SDK's serve() helper owns flushing, protects the protocol from stray stdout, and converts user exceptions into terminal transitions.

code_env/env.py
# code_env/env.py — uploaded once, read-only at /env in every environment
from pathlib import Path
import subprocess
from boltzlabs.env import serve

def reset(seed):
    Path("/workspace/candidate.py").unlink(missing_ok=True)
    return {"task": Path("/env/TASK.md").read_text()}

def step(action):
    Path("/workspace/candidate.py").write_text(action["code"])
    check = subprocess.run(
        ["python3", "-m", "unittest", "discover", "-s", "/env/tests"],
        cwd="/workspace", capture_output=True, text=True, timeout=30,
    )
    output = check.stdout + check.stderr
    return {"test_output": output[-2000:]}, float(check.returncode == 0), True, {}

serve(reset=reset, step=step)

The training loop

train_pool.py
from boltzlabs import RLPool

with RLPool("./code_env", n=64) as pool:
    obs = pool.reset(seed=0)

    for _ in range(1_000):
        codes = sample(model, [o["task"] for o in obs])  # your policy: one program per env
        obs, rewards, dones, infos = pool.step([{"code": c} for c in codes])  # one HTTP request
        update(model, codes, rewards)  # your trainer

        obs = pool.reset(where=dones)  # a clean workspace for every finished attempt

    print(pool.timing.worker_ms)
    print(pool.timing.roundtrip_ms)
    print(pool.timing.stragglers)

05 · Inside the worker

Warm processes, not snapshots.

01

Package once

The SDK creates a deterministic compressed archive of the environment directory and vendors the tiny serve helper. The upload happens at pool creation, not on every step.

02

Mount code read-only

The worker extracts one code directory and mounts it read-only at /env in every environment. Each environment receives a private writable /workspace.

03

Start every environment

Creation is staggered to avoid an interpreter-startup thundering herd. The pool is not returned until every environment answers its initial reset handshake.

04

Keep processes resident

Python imports, simulator state, and heap objects remain alive between steps. There is no per-step process spawn and no process-memory snapshot store.

05

Fan out in index order

The worker sends each action to its environment and restores replies to their original batch positions.

06

Measure shared memory honestly

Status reports proportional memory use so shared interpreter and library pages are not counted in full for every environment.

06 · Resets and failure

Fast reset when possible. Respawn when necessary.

Soft reset

In-process episode reset

pool.reset(where=dones) sends reset only to finished environments. Other observations remain in their original slots.

Hard reset

Kill, wipe, respawn

pool.reset(hard=True) recreates the private workspace and process. It is a cold start, not snapshot restoration.

A process that misses its per-step deadline is killed so one wedged environment cannot freeze the batch. Its result is explicit: null observation, zero reward, done=True, and info["boltzlabs_straggler"]. The next reset respawns it.

07 · Choose deliberately

Use the smallest system that solves the bottleneck.

Stay local

The reward never runs model output, and it fits beside the trainer.

Use one sandbox

You need isolation for a small grading job and can run checks one at a time.

Use RLPool

Every rollout runs untrusted code and needs a clean, isolated environment, at the scale of a training batch.

Implementation status, without benchmark theatre

The pool engine, control-plane routes, SDK, partial resets, straggler path, and memory accounting are implemented and covered by worker and SDK tests. Production deployment and publishable amd64 benchmark results are not claimed here. The SDK deliberately reports worker time and end-to-end round-trip time separately so future measurements cannot hide network cost.