> ## Documentation Index
> Fetch the complete documentation index at: https://docs.firedrill.run/llms.txt
> Use this file to discover all available pages before exploring further.

# Compare model configurations

> Run your agent against the same cases with different models, then compare behavior, latency and recorded cost.

Compare configurations of your agent, not model names in isolation. Keep the synthetic data, tests, seeds and checks fixed, change the model settings in your runner, and inspect which configuration meets your requirements.

Firedrill retains native behavioral outcomes and optional customer measurements. It does not proxy model calls, receive your provider key or decide which model is universally best.

## Configure the same cases

Start with a reviewed [runner config](/guides/run-first-drill), optionally using a [matrix](/guides/scenario-matrices). Add explicit, nonempty `seeds` to `firedrill.config.json`. Save `firedrill.benchmark.json`:

```json theme={null}
{
  "schemaVersion": 1,
  "config": "firedrill.config.json",
  "campaignId": "model-comparison",
  "variants": [
    {
      "id": "configuration-a",
      "label": "Configuration A",
      "targets": {
        "agent": {
          "command": "python",
          "arguments": ["tests/run_agent.py"],
          "bindings": ["mcp"],
          "output": "json"
        }
      },
      "environment": { "APP_MODEL": "YOUR_MODEL_A" }
    },
    {
      "id": "configuration-b",
      "label": "Configuration B",
      "targets": {
        "agent": {
          "command": "python",
          "arguments": ["tests/run_agent.py"],
          "bindings": ["mcp"],
          "output": "json"
        }
      },
      "environment": { "APP_MODEL": "YOUR_MODEL_B" }
    }
  ]
}
```

Replace `agent` with the actual saved target ID, the command with your adapter and both model values with settings your adapter reads. `APP_MODEL` is an example customer variable, not a Firedrill selector. You can use actual command arguments instead.

Version 1 supports one to eight unique variants. All variants supply the same target IDs and exact canonical case selection. An optional `configurationDigest` identifies reviewed settings without uploading them. Provider credentials stay in your runner or CI secret manager; benchmark files are not a key store.

## Preview and run

```sh theme={null}
firedrill benchmark preview --config firedrill.benchmark.json --project YOUR_PROJECT_ID --json
firedrill benchmark run --config firedrill.benchmark.json --project YOUR_PROJECT_ID --json
```

Preview does not invoke your agent, start batches or launch case compute. With
editable test source it can save an immutable test/setup revision and plan
metadata; it requires the corresponding write authority. It checks identical
build/case selection and reports logical cases, potential maximum attempts, known
service blockers and required drill allowance across variants. Attempts describe
potential runner/model work, not a fictional billable-unit count. Actual provider
spend depends on your agent's calls.

Preview is not a reservation. Run honors known blockers and normal service admission. Variants use ordinary durable simulations and isolated state; campaigns do not multiply your parallel allowance. A planned case is not an executed result.

An interruption prints an exact recovery command:

```sh theme={null}
firedrill benchmark run --resume YOUR_CHECKPOINT_PATH --project YOUR_PROJECT_ID --json
```

Do not combine `--resume` with a new `--config`. Recovery retains the original campaign and inner simulation checkpoints. Changed source, commands or settings are rejected. Uncertain customer invocations are not silently executed twice.

## Record measurements using existing instrumentation

JSON TargetResult output and SDK callbacks may include `executionMetrics`.
TypeScript provides `modelCallMetrics`, `executionMetrics` and
`withExecutionMetrics`; native Python provides `model_call_metrics`,
`execution_metrics` and `with_execution_metrics`.

Report actual durations and provider-response usage when available. Missing usage or cost is `null`, not zero. Costs require a currency and reported/estimated basis; estimates also require their pricing description. Customer measurements are not audited provider bills.

The TypeScript `importOpenTelemetryGenAi` and native Python
`import_opentelemetry_genai` helpers project supported GenAI measurement
attributes from bounded OTLP JSON. They exclude prompts, messages, responses and
raw trace bodies and report skipped/invalid spans. These versioned importers
reuse your instrumentation; they are not another tracing backend.

Export or filter spans for the **current interaction**, not the whole process on
every callback. An imported GenAI span represents a logical model operation and
may include automatic retries. Its `modelCalls` count is therefore not necessarily
the number of physical provider requests, and its usage may not cover every billed
attempt. Report actual attempt-level measurements separately when you need that
accounting. The importer accepts string-encoded nanosecond timestamps; numeric
timestamp encodings are not supported.

<Tabs>
  <Tab title="TypeScript">
    ```ts theme={null}
    import { importOpenTelemetryGenAi, executionMetrics } from "@firedrill-run/cloud";

    // otlpJson contains only spans for this interaction.
    const imported = importOpenTelemetryGenAi(otlpJson);
    const metrics = executionMetrics({ modelCalls: imported.modelCalls });
    console.log(imported.skippedSpans, imported.issues);
    // Return metrics as executionMetrics alongside your actual TargetResult.
    ```
  </Tab>

  <Tab title="Python">
    ```python theme={null}
    from firedrill_cloud import import_opentelemetry_genai, execution_metrics

    # otlp_json contains only spans for this interaction.
    imported = import_opentelemetry_genai(otlp_json)
    metrics = execution_metrics(model_calls=imported["modelCalls"])
    print(imported["skippedSpans"], imported["issues"])
    # Return metrics as executionMetrics alongside your actual TargetResult.
    ```
  </Tab>
</Tabs>

Review importer diagnostics: discarded spans are not zero-cost calls. No cost is
inferred from a model name, and deprecated or unsupported usage fields remain
unknown. You may instead use your provider response directly with the model-call
measurement helpers; there is no mandatory tracing vendor.

Optional customer evaluator scores identify their evaluator name, version and range. They stay separate from native state/operation checks: a high answer score cannot override a failed behavioral test.

## Use SDK callbacks

TypeScript exposes `previewModelBenchmark`, `runModelBenchmark` and `getSimulationComparison`. Native Python exposes `preview_model_benchmark`, `run_model_benchmark`, async equivalents and `get_simulation_comparison`.

Supply actual callback targets, the exact setup or environment/build, explicit seeds and the same target IDs across variants. Each callback chooses its model, consumes that case's binding and returns the actual TargetResult. Python does not require Node.js. See [SDK execution](/guides/simulations#run-from-the-sdk) and [authentication](/api-reference/authentication). Preserve private checkpoints through your application's storage.

These integration functions take two **actual customer adapters** as arguments.
Configure their model settings in your application; Firedrill does not supply
`targetA` / `targetB` or call a provider for you. Replace the project, setup,
test and target IDs with your saved definitions. Both callbacks must use their
assigned case binding rather than a shared production connection.

<Tabs>
  <Tab title="TypeScript">
    ```ts theme={null}
    import {
      FiredrillClient, previewModelBenchmark, runModelBenchmark,
      type ModelBenchmarkOptions, type SimulationTarget,
    } from "@firedrill-run/cloud";

    export async function compareModels(
      client: FiredrillClient,
      targetA: SimulationTarget,
      targetB: SimulationTarget,
    ) {
      const options: ModelBenchmarkOptions = {
        projectId: "prj_YOUR_PROJECT_ID",
        campaignId: "reviewed-model-comparison",
        selection: {
          toolSetupId: "setup_YOUR_READY_SETUP_ID",
          drillIds: ["YOUR_SAVED_TEST_ID"],
          seeds: ["42", "73"],
          concurrency: 1,
          retries: 0,
        },
        capacity: 1,
        bindingMode: "per-case",
        variants: [
          {
            id: "configuration-a", label: "Configuration A",
            targets: { "YOUR_SAVED_TARGET_ID": targetA },
            targetBindings: { "YOUR_SAVED_TARGET_ID": ["mcp"] },
          },
          {
            id: "configuration-b", label: "Configuration B",
            targets: { "YOUR_SAVED_TARGET_ID": targetB },
            targetBindings: { "YOUR_SAVED_TARGET_ID": ["mcp"] },
          },
        ],
      };
      const preview = await previewModelBenchmark(client, options);
      console.log(preview.totals, preview.allowance);
      const result = await runModelBenchmark(client, options);
      console.log(result.comparisonUrl);
      return result;
    }
    ```
  </Tab>

  <Tab title="Python">
    ```python theme={null}
    from firedrill_cloud import preview_model_benchmark, run_model_benchmark

    def compare_models(client, target_a, target_b):
        options = {
            "project_id": "prj_YOUR_PROJECT_ID",
            "campaign_id": "reviewed-model-comparison",
            "selection": {
                "tool_setup_id": "setup_YOUR_READY_SETUP_ID",
                "drill_ids": ["YOUR_SAVED_TEST_ID"],
                "seeds": ["42", "73"],
                "concurrency": 1,
                "retries": 0,
            },
            "capacity": 1,
            "binding_mode": "per-case",
            "variants": [
                {
                    "id": "configuration-a", "label": "Configuration A",
                    "targets": {"YOUR_SAVED_TARGET_ID": target_a},
                    "target_bindings": {"YOUR_SAVED_TARGET_ID": ["mcp"]},
                },
                {
                    "id": "configuration-b", "label": "Configuration B",
                    "targets": {"YOUR_SAVED_TARGET_ID": target_b},
                    "target_bindings": {"YOUR_SAVED_TARGET_ID": ["mcp"]},
                },
            ],
        }
        preview = preview_model_benchmark(client, **options)
        print(preview["totals"], preview["allowance"])
        result = run_model_benchmark(client, **options)
        print(result["comparisonUrl"])
        return result
    ```
  </Tab>
</Tabs>

For durable application recovery, pass `onCheckpoint` / `on_checkpoint` to
persist private checkpoints atomically, and `resume` to resume that exact saved
campaign. Keep the callback configuration unchanged. Review preview output
before execution in your workflow; `run` enforces known blockers again. The
functions above demonstrate the invocation boundary, not a provider adapter or
automatic approval of your case selection.

## Interpret the report

Open the returned comparison URL. Reports show native outcomes and unfinished cases, wall-latency distributions, measurement coverage, known/unknown tokens, separate cost groups and customer evaluator scores.

Cost groups preserve currency, basis and estimate description. Evaluator versions and ranges stay distinct. Unequal case selections, missing results and truncation are explicit. An incomplete variant is never labeled the winner.

Choose the trade-off that matters for your agent: acceptable behavior first, then latency, cost and additional quality measures. Comparison reads remain in one authorized project. The portal displays reports; execution stays in CLI, SDK or CI.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.