Configure the same cases
Start with a reviewed runner config, optionally using a matrix. Add explicit, nonemptyseeds to firedrill.config.json. Save firedrill.benchmark.json:
agent with the actual saved target ID, the command with your adapter and both model values with settings your adapter reads. APP_MODEL is an example customer variable, not a Firedrill selector. You can use actual command arguments instead.
Version 1 supports one to eight unique variants. All variants supply the same target IDs and exact canonical case selection. An optional configurationDigest identifies reviewed settings without uploading them. Provider credentials stay in your runner or CI secret manager; benchmark files are not a key store.
Preview and run
--resume with a new --config. Recovery retains the original campaign and inner simulation checkpoints. Changed source, commands or settings are rejected. Uncertain customer invocations are not silently executed twice.
Record measurements using existing instrumentation
JSON TargetResult output and SDK callbacks may includeexecutionMetrics.
TypeScript provides modelCallMetrics, executionMetrics and
withExecutionMetrics; native Python provides model_call_metrics,
execution_metrics and with_execution_metrics.
Report actual durations and provider-response usage when available. Missing usage or cost is null, not zero. Costs require a currency and reported/estimated basis; estimates also require their pricing description. Customer measurements are not audited provider bills.
The TypeScript importOpenTelemetryGenAi and native Python
import_opentelemetry_genai helpers project supported GenAI measurement
attributes from bounded OTLP JSON. They exclude prompts, messages, responses and
raw trace bodies and report skipped/invalid spans. These versioned importers
reuse your instrumentation; they are not another tracing backend.
Export or filter spans for the current interaction, not the whole process on
every callback. An imported GenAI span represents a logical model operation and
may include automatic retries. Its modelCalls count is therefore not necessarily
the number of physical provider requests, and its usage may not cover every billed
attempt. Report actual attempt-level measurements separately when you need that
accounting. The importer accepts string-encoded nanosecond timestamps; numeric
timestamp encodings are not supported.
- TypeScript
- Python
Use SDK callbacks
TypeScript exposespreviewModelBenchmark, runModelBenchmark and getSimulationComparison. Native Python exposes preview_model_benchmark, run_model_benchmark, async equivalents and get_simulation_comparison.
Supply actual callback targets, the exact setup or environment/build, explicit seeds and the same target IDs across variants. Each callback chooses its model, consumes that case’s binding and returns the actual TargetResult. Python does not require Node.js. See SDK execution and authentication. Preserve private checkpoints through your application’s storage.
These integration functions take two actual customer adapters as arguments.
Configure their model settings in your application; Firedrill does not supply
targetA / targetB or call a provider for you. Replace the project, setup,
test and target IDs with your saved definitions. Both callbacks must use their
assigned case binding rather than a shared production connection.
- TypeScript
- Python
onCheckpoint / on_checkpoint to
persist private checkpoints atomically, and resume to resume that exact saved
campaign. Keep the callback configuration unchanged. Review preview output
before execution in your workflow; run enforces known blockers again. The
functions above demonstrate the invocation boundary, not a provider adapter or
automatic approval of your case selection.