Skip to main content
A simulation is a coordinated set of attempts against synthetic Tools. Your runner executes the agent; Firedrill prepares the cases, isolates their state, records consequences, and evaluates checks.

Start with a ready setup

Prepare Tools through init or the application. Save meaningful tasks and checks in Tests. Connect your customer-owned runner once rather than preparing a separate environment by hand for every attempt. In Simulator, select the setup, tests, seeds, repetitions, and concurrency. Preview the resolved plan before starting it. A batch waiting for a runner has not executed your agent. Connect the runner using the printed instructions or your SDK integration.

Independent attempts

Each independent case gets its own synthetic state. A mutation in one case does not alter another. Use different seeds or starting scenarios to vary data, permissions, faults, and scheduled events. Running more repetitions does not make a small or biased sample statistically conclusive. Several customer runners can participate in one batch. Each receives only its claimed case and scoped connection. Requested concurrency is capped by your plan and available capacity; queued work stays visible. Cleanup and retries belong to the coordinated plan, not to an implicit new request after a timeout.

Continuing steps

A continuing simulation intentionally keeps evolving state between ordered steps. Earlier agent actions, scheduled events, and time advancement affect what later tasks observe. This is different from independent attempts starting again from a baseline. Pause or continue through the supported controls. A pause preserves the exact position and ownership; it is not a reset. Advance Tool time with an explicit budget when you want scheduled work to become due. This does not speed up model inference or change your agent’s real clock.

Run from code or CI

runSimulation in TypeScript and run_simulation / run_simulation_async in Python coordinate the same cases with your callbacks. Target IDs must match the saved test definitions. Each callback gets the current task, exact binding, and cancellation signal, and returns its actual result—not a fabricated verdict. For command targets configured in your agent project:
See Run saved tests for the separate runner configuration. Use the same commands or SDK integration in CI on pull requests, pushes, schedules, or manual triggers. Customer execution may run on your workstation, CI, or server; the synthetic Tools remain managed by Firedrill.

Results, recovery, and replay

Open the batch’s Results link, then an attempt. Distinguish failed checks, runner failures, cancelled cases, and missing observations. Tool-state checks and runner-attributed browser checks remain separate evidence in the same result. Preserve recovery checkpoints privately. Resume the exact pending operation; uncertain or completed agent calls are never silently repeated. A rerun is a new attempt, while inspecting retained history does not execute the agent again. Pinned synthetic inputs are reproducible, but an external model may answer differently on another invocation.