Skip to main content
Compare configurations of your agent, not model names in isolation. Keep the synthetic data, tests, seeds and checks fixed, change the model settings in your runner, and inspect which configuration meets your requirements. Firedrill retains native behavioral outcomes and optional customer measurements. It does not proxy model calls, receive your provider key or decide which model is universally best.

Configure the same cases

Start with a reviewed runner config, optionally using a matrix. Add explicit, nonempty seeds to firedrill.config.json. Save firedrill.benchmark.json:
Replace agent with the actual saved target ID, the command with your adapter and both model values with settings your adapter reads. APP_MODEL is an example customer variable, not a Firedrill selector. You can use actual command arguments instead. Version 1 supports one to eight unique variants. All variants supply the same target IDs and exact canonical case selection. An optional configurationDigest identifies reviewed settings without uploading them. Provider credentials stay in your runner or CI secret manager; benchmark files are not a key store.

Preview and run

Preview does not invoke your agent, start batches or launch case compute. With editable test source it can save an immutable test/setup revision and plan metadata; it requires the corresponding write authority. It checks identical build/case selection and reports logical cases, potential maximum attempts, known service blockers and required drill allowance across variants. Attempts describe potential runner/model work, not a fictional billable-unit count. Actual provider spend depends on your agent’s calls. Preview is not a reservation. Run honors known blockers and normal service admission. Variants use ordinary durable simulations and isolated state; campaigns do not multiply your parallel allowance. A planned case is not an executed result. An interruption prints an exact recovery command:
Do not combine --resume with a new --config. Recovery retains the original campaign and inner simulation checkpoints. Changed source, commands or settings are rejected. Uncertain customer invocations are not silently executed twice.

Record measurements using existing instrumentation

JSON TargetResult output and SDK callbacks may include executionMetrics. TypeScript provides modelCallMetrics, executionMetrics and withExecutionMetrics; native Python provides model_call_metrics, execution_metrics and with_execution_metrics. Report actual durations and provider-response usage when available. Missing usage or cost is null, not zero. Costs require a currency and reported/estimated basis; estimates also require their pricing description. Customer measurements are not audited provider bills. The TypeScript importOpenTelemetryGenAi and native Python import_opentelemetry_genai helpers project supported GenAI measurement attributes from bounded OTLP JSON. They exclude prompts, messages, responses and raw trace bodies and report skipped/invalid spans. These versioned importers reuse your instrumentation; they are not another tracing backend. Export or filter spans for the current interaction, not the whole process on every callback. An imported GenAI span represents a logical model operation and may include automatic retries. Its modelCalls count is therefore not necessarily the number of physical provider requests, and its usage may not cover every billed attempt. Report actual attempt-level measurements separately when you need that accounting. The importer accepts string-encoded nanosecond timestamps; numeric timestamp encodings are not supported.
Review importer diagnostics: discarded spans are not zero-cost calls. No cost is inferred from a model name, and deprecated or unsupported usage fields remain unknown. You may instead use your provider response directly with the model-call measurement helpers; there is no mandatory tracing vendor. Optional customer evaluator scores identify their evaluator name, version and range. They stay separate from native state/operation checks: a high answer score cannot override a failed behavioral test.

Use SDK callbacks

TypeScript exposes previewModelBenchmark, runModelBenchmark and getSimulationComparison. Native Python exposes preview_model_benchmark, run_model_benchmark, async equivalents and get_simulation_comparison. Supply actual callback targets, the exact setup or environment/build, explicit seeds and the same target IDs across variants. Each callback chooses its model, consumes that case’s binding and returns the actual TargetResult. Python does not require Node.js. See SDK execution and authentication. Preserve private checkpoints through your application’s storage. These integration functions take two actual customer adapters as arguments. Configure their model settings in your application; Firedrill does not supply targetA / targetB or call a provider for you. Replace the project, setup, test and target IDs with your saved definitions. Both callbacks must use their assigned case binding rather than a shared production connection.
For durable application recovery, pass onCheckpoint / on_checkpoint to persist private checkpoints atomically, and resume to resume that exact saved campaign. Keep the callback configuration unchanged. Review preview output before execution in your workflow; run enforces known blockers again. The functions above demonstrate the invocation boundary, not a provider adapter or automatic approval of your case selection.

Interpret the report

Open the returned comparison URL. Reports show native outcomes and unfinished cases, wall-latency distributions, measurement coverage, known/unknown tokens, separate cost groups and customer evaluator scores. Cost groups preserve currency, basis and estimate description. Evaluator versions and ranges stay distinct. Unequal case selections, missing results and truncation are explicit. An incomplete variant is never labeled the winner. Choose the trade-off that matters for your agent: acceptable behavior first, then latency, cost and additional quality measures. Comparison reads remain in one authorized project. The portal displays reports; execution stays in CLI, SDK or CI.