Skip to main content
Firedrill runs synthetic Tools and records what your agent does with them. Your agent and model stay in your own process. A test names an agent task and the changes you expect; a simulation runs that test in isolated Tool state. You can repeat cases or run several in parallel without making one case’s changes affect another.

1. Prepare the Tools

Sign in, select a project, and find the Tools your agent uses. Project creation also creates a default environment; it does not start Tools or run your agent.
If you already have a project, use firedrill project select instead of creating another. Choose a library ID from the list and create a setup:
Repeat --tool to include more Tools. Use --initial-state empty for no starting records. The CLI prints a setup ID when it is ready. You can also select Tools in the portal’s Tools page. A ready setup is not evidence that your agent ran.

2. Save a test

Open the ready setup in the portal and choose Add a test. Supply your agent’s task, a target ID matching the agent entry point you will run, and checks on observable behavior or Tool state. Keep the setup ID and the saved test and target IDs. You can also keep test definitions in your agent repository and create a derived setup:
That command creates a new setup and leaves the original unchanged. The JSON file contains typed targets, scenarios, drills, and optional suites. Start with the portal’s test form if you do not already have a source file; firedrill tools add --help shows the file-based command.

3. Run your agent

Create firedrill.config.json beside your agent. This example runs an existing Python entry point; use your own command, arguments, bindings, test ID, and target ID.
The command receives a task on stdin and a scoped Tool connection in its environment. It calls your existing agent. Firedrill does not need your model key or agent source. The entry point returns a target result; Firedrill evaluates the saved checks against recorded observations and Tool state.
The CLI prints progress and a Results link. An agent exiting successfully does not make failed checks pass. If a request is interrupted, use the CLI’s printed resume command instead of starting an uncertain case again.

Repeat or parallelize cases

In the portal’s Simulator, choose saved tests, repetitions, seeds, and parallelism. Starting a simulation prepares isolated cases; your runner must join the batch to execute your agent. The portal cannot run your agent process by itself.
Save that as firedrill.runner.json, then join the batch ID shown in the portal:
The same runner can execute tests from CI. Each independent case starts from its selected data; an explicitly continuous sequence keeps evolving state between its steps.

Read the result

Open Results in the portal or the link printed by the CLI. Start with the test verdict, then inspect failed checks, the Tool-call timeline, before/after state, logs, and any enabled screenshots or recordings. The portal shows recorded evidence; it does not invent a passing outcome from your agent’s final message.