Start with a ready setup
Prepare Tools through init or the application. Save meaningful tasks and checks in Tests. Connect your customer-owned runner once rather than preparing a separate environment by hand for every attempt. In Simulator, select the setup, tests, seeds, repetitions, and concurrency. Preview the resolved plan before starting it. A batch waiting for a runner has not executed your agent. Connect the runner using the printed instructions or your SDK integration.Independent attempts
Each independent case gets its own synthetic state. A mutation in one case does not alter another. Use different seeds or starting scenarios to vary data, permissions, faults, and scheduled events. Running more repetitions does not make a small or biased sample statistically conclusive. Several customer runners can participate in one batch. Each receives only its claimed case and scoped connection. Requested concurrency is capped by your plan and available capacity; queued work stays visible. Cleanup and retries belong to the coordinated plan, not to an implicit new request after a timeout.Continuing steps
A continuing simulation intentionally keeps evolving state between ordered steps. Earlier agent actions, scheduled events, and time advancement affect what later tasks observe. This is different from independent attempts starting again from a baseline. Pause or continue through the supported controls. A pause preserves the exact position and ownership; it is not a reset. Advance Tool time with an explicit budget when you want scheduled work to become due. This does not speed up model inference or change your agent’s real clock.Run from code or CI
runSimulation in TypeScript and run_simulation / run_simulation_async in
Python coordinate the same cases with your callbacks. Target IDs must match
the saved test definitions. Each callback gets the current task, exact binding,
and cancellation signal, and returns its actual result—not a fabricated verdict.
For command targets configured in your agent project: