RL environments
Give each agent attempt a place to work. It can change files, run code, and call tools without touching your machine. Check the result and send the outcome back to your evaluation or training code.
Two ways to use them
Choose based on what your agent needs to do. You can use a sandbox without using an RL pool.
01 / WORKSPACE
An agent working in a repository
Use a sandbox when the agent needs a shell, project files, package installs, and time to iterate. It can run the project's tests and leave its changes for inspection.
Example: give a coding agent a failing test, let it edit the repository, then check whether the test passes.
Sandbox docs02 / POOL
An environment in a training loop
Use an RL pool when your Python code already has a policy or agent and needs to run many independent environment steps. Send one action per environment; receive observations and rewards in the same order.
Example: submit agent-written code to separate workers, run the same checks, and compare the results.
RL pool docsA real coding-agent run
Suppose a repository has a failing test. You want to see whether an agent can fix it, and keep the evidence of what happened.
- 01Start Put the repository and task in a fresh sandbox for each attempt.
- 02Work The agent edits files and uses the shell inside that sandbox.
- 03Check Run your test command against the final files and keep the result.
- 04Review Inspect the session and diff before exporting or deleting the sandbox.
What BoltzLabs handles
BoltzLabs provides isolated execution, workspace lifecycle, and results from your checks. You provide the task, agent, and scoring rules. You can repeat a task across separate attempts and export the session, code changes, and check result.
BoltzLabs does not train the model for you, and current recordings are not a trainer-ready dataset. Available worker capacity still limits how many environments can run at once.