Python SDK
RL pools
A pool runs your task in isolated workers. Your Python program sends each worker an action and gets an observation and reward back. The example below sends agent-written code to a worker, runs tests there, and returns whether they passed.
What happens when you call a pool?
RLPool(...)starts the requested number of workers.reset()asks each worker for its task.step(actions)sends one action to each worker and returns its result.close()stops the pool when you are done.
The pool does not train a model or call your agent for you. It runs the environment and returns the results to your Python code.
Before you run it
Install the Python SDK and set BOLTZLABS_API_KEY. Prepare a small my_env/ folder containing env.py, TASK.md, and your Python tests/. The folder is
uploaded once when the pool starts. A paid plan and available RL worker are required. The
environment file is shown below if you need to write one.
Run one worker from Python
Save the code produced by your agent as candidate.py next to this script. The script
reads that file but does not execute it on your machine; the pool runs it in the worker.
from pathlib import Path
from boltzlabs import RLPool
agent_code = Path("candidate.py").read_text()
pool = RLPool(env_dir="./my_env", n=1)
task = pool.reset()[0]
print("Task:", task["task"])
results, rewards, done, details = pool.step([{"code": agent_code}])
print("Passed:", bool(rewards[0]))
print("Worker output:", results[0] or details[0])
pool.close()Here n=1 means one worker. The brackets around the action matter: even with one
worker, step() takes a list with one action per worker. Replace the file read with your
agent's output when you connect a model.
Choose a mode
Omitting mode selects default. To request Boltz mode, change only the
pool creation line:
pool = RLPool(env_dir="./my_env", n=1, mode="boltz")| Mode | How it runs | Use it when |
|---|---|---|
default | One separate runtime per environment | You need snapshots or forks |
boltz | Warm environments on shared worker capacity | You want denser warm pools |
Boltz mode needs a configured Boltz worker. There is no automatic fallback to the other mode. Snapshot, restore, and fork work only in default mode. Neither mode is automatically faster for every task.
What is inside my_env/?
TASK.md describes the work. tests/ contains the tests you already use
for that task. env.py is the small adapter between those files and the pool API.
The uploaded files appear read-only under /env; each worker writes its candidate
under /workspace.
Show an env.py for Python code checks
This example runs Python's built-in unittest runner against the candidate. Replace the test command if your environment uses a different check.
from pathlib import Path
import subprocess
from boltzlabs.env import serve
def reset(seed):
Path("/workspace/candidate.py").unlink(missing_ok=True)
return {"task": Path("/env/TASK.md").read_text()}
def step(action):
Path("/workspace/candidate.py").write_text(action["code"])
check = subprocess.run(
["python3", "-m", "unittest", "discover", "-s", "/env/tests"],
cwd="/workspace", capture_output=True, text=True, timeout=30,
)
output = check.stdout + check.stderr
return {"test_output": output[-2000:]}, float(check.returncode == 0), True, {}
serve(reset=reset, step=step)These checks are useful for normal agent evaluation, but code running beside its tests is not a tamper-proof reward source. For adversarial grading, verify the result in a separate trusted step.
When you need more than one worker
Set n=2 for two workers, then send two actions in each step() call. The
returned results are in the same order as your actions. Start with one worker until the environment
works; pool size is limited by available worker capacity.
If your agent needs a shell, package installs, and ongoing edits to an entire repository, use a sandbox instead. RL pools are for batched reset-and-step environments.