Define your variations
Create a reviewed test. Add"matrixSource": "firedrill.matrix.json" to firedrill.config.json, then save this matrix. Replace first-test with your logical drill ID and both instructions with tasks your agent supports:
task, actorId, actorInstanceId, scenarioId or inlineScenario, assertions, or toolOverrides. Reference scenarios you actually authored in your test source. A saved dataset ID is not itself scenario state; use starting data and saved datasets to prepare the conditions.
Combine independent axes
Two tasks and three scenarios produce six combinations per template. Different axes cannot replace the same field: choices replace fields explicitly rather than applying ambiguous shallow merges.scenarioId and inlineScenario are alternative selectors, not values to combine.
actorId and actorInstanceId select one acting identity, so keep both in a
single axis when varying them. Omitted fields inherit the template: changing only
the scope preserves its actor, and changing only the actor preserves its explicit
scope. An explicit null or empty scope is invalid; it does not clear a named
scope. Choose an actor/scope pair actually declared in your scenario. A Tool-copy
name alone does not grant an actor access, and an ambiguous or foreign pair fails
validation rather than being guessed.
Optional exclude contains partial axis/choice mappings, for example [{"task":"detailed","data":"empty"}] when those axes and choices exist. Unknown choices fail validation instead of silently disappearing.
Version 1 allows up to 100 generated variants in total across the selected templates. The expanded source, including any tests outside the matrix, must also fit the existing bounds of 100 resources per kind and 512 KiB. Expansion beyond the reviewed bound fails; it is never silently truncated. The service then applies execution limits to the resolved plan.
Review, then execute
seeds, repetitions, concurrency and local capacity in your config. Concurrency requests service cases within your plan; capacity limits agent processes on this runner. Each parallel case must consume its own binding. Increasing either setting does not change your allowance or make model inference instantaneous.
Open the saved batch to inspect failing case IDs, seeds, checks and evidence. Preserve the exact source and selection for reproduction. A seed pins synthetic inputs, not an LLM’s answer. Use the same files in CI or model benchmarks.