Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions docs/benchmarks.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,11 @@ Benchmarks are used to establish a baseline for agent performance on a given tas

Each benchmark is made up of a capabilities. These capabilities are used to determine if the agent is able to complete the task. Each capability has a name and a set of scenarios that are used to evaluate the agent's performance.

To compare additional treatments against an existing benchmark, reference it
from an [experiment](experiments.md#running-against-a-benchmark) with
`benchmark: '<filename-id>'`. The experiment reuses its capabilities and adds
Control and Benchmark reference treatments alongside your treatments.

As an example, you may be creating a design system benchmark. In it, one of the
capabilities is around the LLM's usage of icons from your system. You could have
different scenarios within this icon usage capability that test the different
Expand Down
51 changes: 46 additions & 5 deletions docs/experiments.md
Original file line number Diff line number Diff line change
Expand Up @@ -50,9 +50,7 @@ export default defineConfig({
})
```

## CLI

### Running against a benchmark
## Running against a benchmark

Use `benchmark: '<filename-id>'` instead of `scenarios` to evaluate every
capability and scenario in an existing benchmark. Specify exactly one of these
Expand Down Expand Up @@ -85,9 +83,46 @@ select a capability to see its scenario comparisons. Scenario links open the
corresponding capability's trial details. The same matrix is available for
historical runs and uses their saved benchmark metadata.

You can interact with experiments using the `experiments` subcommand of the `agent-eval` CLI. This sub-command gives you access to run experiments, create run plans to use for sharding, or merge the results of a plan.
For example, compare verification instructions to the standard configured in
`benchmarks/example.ts`:

```ts
// experiments/benchmark-comparison.ts
import {defineConfig} from '@primer/agent-eval/experiment'

export const experiment = defineConfig({
name: 'Benchmark comparison',
description: 'Compare verification instructions against the existing benchmark standard',
models: ['gpt-6-sol'],
benchmark: 'example',
treatments: [
{
name: 'Verification instructions',
async setup({sandbox}) {
await sandbox.addAgentInstruction('Before finishing, verify the implementation against the task requirements.')
},
},
],
})
```

Create a plan and inspect it before running:

```sh
npx agent-eval experiment plan create benchmark-comparison --output-path ./comparison-plan.json
npx agent-eval experiment plan run --plan-path ./comparison-plan.json --output-dir ./results/comparison-01
```

Use `agent-eval experiments --help` to see the available commands and options.
The plan includes Control, Benchmark, and Verification instructions for each
capability/scenario membership and model variant. An experiment's optional
`runners` array adds a runner dimension, or `--runner` selects one backend when
creating a new plan. To shard and merge, use the same experiment commands as for
scenario-based experiments.

No new overall benchmark score is calculated. Check metrics keep their existing
units, directions, and treatment-specific coverage; resource usage is averaged
per trial. A larger capability contributes more trials, not an implicitly
equal-weighted capability score.

## Loaded experiments

Expand All @@ -108,3 +143,9 @@ New experiment plan manifests include the same `type` discriminator.
Benchmark-backed plans require benchmark metadata and a `capabilityId` on every
trial. Scenario-backed plans contain neither. Replay accepts existing untagged
manifests and rejects changes to the experiment's source kind.

## CLI

You can interact with experiments using the `experiment` subcommand of the `agent-eval` CLI. This subcommand gives you access to run experiments, create run plans to use for sharding, or merge the results of a plan.

Use `agent-eval experiment --help` to see the available commands and options.
3 changes: 3 additions & 0 deletions skills/agent-eval/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,9 @@ not a prerequisite for every operation.
beyond the automatic control.
- Keep prompts and fixtures treatment-blind. Put the knowledge being tested in
treatment setup, not shared setup.
- For a benchmark-backed experiment, keep capability setup neutral. The automatic
Benchmark treatment supplies the existing standard without changing Control
or custom treatments.
- Withhold grader files using `files`. Verify the starter fails for the intended
reasons and a correct solution passes; restore the starter before execution.
- Inspect the plan's combinations before running. Use a fresh output directory.
Expand Down
5 changes: 5 additions & 0 deletions skills/agent-eval/references/benchmarks.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,11 @@
**Use when:** establishing a stable capability baseline across models or over
time. Use an [experiment](experiments.md) to compare multiple interventions.

An experiment can use `benchmark: '<filename-id>'` instead of a scenario list.
It reuses the capability suite and runs Control and Benchmark alongside its
custom treatments. Standalone benchmark commands and their two-treatment
behavior remain unchanged.

## Contract

| Item | Rule |
Expand Down
42 changes: 41 additions & 1 deletion skills/agent-eval/references/experiments.md
Original file line number Diff line number Diff line change
Expand Up @@ -88,6 +88,42 @@ your project.

## Run

### Compare against an existing benchmark

With `benchmarks/project.ts` from the [benchmark example](benchmarks.md), save
this as `experiments/benchmark-comparison.ts`:

```ts
import {defineConfig} from '@primer/agent-eval/experiment'

export const experiment = defineConfig({
name: 'Benchmark comparison',
description: 'Compare a task review instruction against the verification baseline',
models: [{name: 'gpt-5.4', reasoningEfforts: ['low']}],
benchmark: 'project',
treatments: [
{
name: 'Task review',
async setup({sandbox}) {
await sandbox.addAgentInstruction('Review the task requirements against the implementation before finishing.')
},
},
],
})
```

```sh
npx agent-eval experiment plan create benchmark-comparison --output-path ./benchmark-comparison-plan.json
npx agent-eval experiment plan run --plan-path ./benchmark-comparison-plan.json --output-dir ./results/benchmark-comparison-01
```

This one-model, one-scenario-membership example plans three trials: Control,
Benchmark, and Task review. Their prompt, fixture, graders, model, effort,
runner, and capability are held fixed. The Benchmark trial uses the benchmark's
verification instruction; Task review uses only its own instruction.

### Compare a scenario list

```sh
npx agent-eval experiment plan create verification --output-path ./verification-plan.json
npx agent-eval experiment plan run --plan-path ./verification-plan.json --output-dir ./results/verification-01
Expand All @@ -110,9 +146,13 @@ Start with one model, effort, scenario, runner, and treatment beyond control.
The matrix expands as follows:

```text
trials = model variants * scenarios * unique runners * (configured treatments + 1)
scenario experiment trials = model variants * scenarios * unique runners * (configured treatments + 1)
benchmark experiment trials = model variants * scenario memberships * unique runners * (configured treatments + 2)
```

Scenario memberships are the sum of the scenario counts in all capabilities,
including a separate membership when a scenario belongs to multiple capabilities.

`runners: ['copilot-cli', 'copilot-sdk']` adds a backend comparison. Changing both
runner and resource confounds their effects; see [models and runners](models-and-runners.md).

Expand Down
22 changes: 14 additions & 8 deletions skills/agent-eval/references/plans.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,14 +6,14 @@ evaluations.

## Contract

| Item | Rule |
| :--------------- | :----------------------------------------------------------------------------- |
| Input | An existing benchmark/experiment and resolvable scenarios |
| Manifest | Configuration ID/name and uniquely identified trial combinations |
| Trial dimensions | Scenario, treatment, model/effort, runner; also capability for benchmarks |
| Plan creation | Loads configs and validates references; no containers, setup hooks, or Copilot |
| Plan execution | Reloads local configs and resolves saved identifiers |
| Shards | One-based `order/total`; assignment by trial position |
| Item | Rule |
| :--------------- | :--------------------------------------------------------------------------------------------------------- |
| Input | An existing benchmark/experiment and resolvable scenarios |
| Manifest | Configuration ID/name and uniquely identified trial combinations |
| Trial dimensions | Scenario, treatment, model/effort, runner; also capability for benchmarks and benchmark-backed experiments |
| Plan creation | Loads configs and validates references; no containers, setup hooks, or Copilot |
| Plan execution | Reloads local configs and resolves saved identifiers |
| Shards | One-based `order/total`; assignment by trial position |

Load only trusted configurations: creation executes their top-level module code
even though it does not execute trials.
Expand All @@ -32,6 +32,12 @@ a `capabilityId` on every trial; scenario-backed plans contain neither. Replay
accepts older manifests without `type`, but rejects a change between scenario
and benchmark sources.

Benchmark-backed experiment plans also record the benchmark ID, name, and
capability/scenario grouping. Replay rejects a changed hierarchy; create a new
plan after editing it. This grouping snapshot does not freeze setup code or
scenario contents. Pass `--benchmarks <dir>` to both plan create and plan run
when the benchmark directory is not `./benchmarks`.

## Create and run

The following command templates require an existing experiment named
Expand Down
29 changes: 17 additions & 12 deletions skills/agent-eval/references/treatments.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,26 +5,31 @@ agent without changing its task or grader.

## Contract

| Item | Rule |
| :------------- | :------------------------------------------------------------- |
| Required field | Unique `name`; `Control` is reserved |
| Optional field | `async setup({sandbox})` |
| Identity | Name-derived ID; keep names stable when reusing plans |
| Setup API | Use the supplied [sandbox](sandbox.md), not the host workspace |
| Item | Rule |
| :------------- | :------------------------------------------------------------------------------------------------- |
| Required field | Unique `name`; `Control` is reserved, and `Benchmark` is reserved for benchmark-backed experiments |
| Optional field | `async setup({sandbox})` |
| Identity | Name-derived ID; keep names stable when reusing plans |
| Setup API | Use the supplied [sandbox](sandbox.md), not the host workspace |

The automatic `Control` has no treatment setup. It still receives the common
fixture, dependency installation, shared setup, and runtime tools.

| Setup location | Applies to |
| :--------------------------- | :------------------------------------------------ |
| Experiment top-level `setup` | All experiment trials, including control |
| Experiment treatment `setup` | That treatment only |
| Benchmark top-level `setup` | `Benchmark` treatment only |
| Benchmark capability `setup` | Both `Control` and `Benchmark` in that capability |
| Setup location | Applies to |
| :--------------------------- | :----------------------------------------------------------------- |
| Experiment top-level `setup` | All experiment trials, including control |
| Experiment treatment `setup` | That treatment only |
| Benchmark top-level `setup` | `Benchmark` treatment only |
| Benchmark capability `setup` | All treatments in that capability, including experiment treatments |

Shared setup runs before treatment setup. Do not accidentally give the tested
resource to control by installing it in shared setup or the fixture.

In a benchmark-backed experiment, setup runs in this order: scenario,
capability, experiment, selected treatment. Benchmark-level setup remains
exclusive to the automatic Benchmark treatment. Experiment models and runners
apply to Control, Benchmark, and every configured treatment.

## Setup fragments

Inside a treatment's `async setup({sandbox})`, use the appropriate method.
Expand Down
27 changes: 20 additions & 7 deletions skills/agent-eval/references/trials-and-results.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,13 +5,13 @@ treatments. Each trial has its own ID, sandbox, execution, and evidence.

## Contract

| Item | Rule |
| :-------------------------- | :------------------------------------------------------------------------- |
| Trial identity | Scenario, treatment, model variant, runner; also capability for benchmarks |
| Benchmark/experiment output | `output.json` manifest with `trials` mapping IDs to relative file paths |
| Trial result location | `artifacts/<trial-id>/<trial-id>.json` |
| Scenario output | `{id, results}`, with embedded results rather than a trial-file manifest |
| Quality evidence | Check outcomes/measurements and judge results, not command exit alone |
| Item | Rule |
| :-------------------------- | :---------------------------------------------------------------------------------------------------------- |
| Trial identity | Scenario, treatment, model variant, runner; also capability for benchmarks and benchmark-backed experiments |
| Benchmark/experiment output | `output.json` manifest with `trials` mapping IDs to relative file paths |
| Trial result location | `artifacts/<trial-id>/<trial-id>.json` |
| Scenario output | `{id, results}`, with embedded results rather than a trial-file manifest |
| Quality evidence | Check outcomes/measurements and judge results, not command exit alone |

## Lifecycle

Expand Down Expand Up @@ -43,6 +43,19 @@ artifacts/
paths, not embedded trial results. It includes scenario and treatment metadata;
benchmark output also includes capability metadata.

Benchmark-backed experiment output has a `benchmark` snapshot containing its
ID, name, and capabilities, while each trial records a `capabilityId`. Keep the
membership even when a scenario belongs to multiple capabilities; scenario ID
alone cannot distinguish those trials. Use `parseExperimentTrialOutput(json,
output.benchmark)` to parse and check membership against the saved snapshot.

The website compares treatments side by side within capabilities, with separate
model, effort, and runner selections. Select a metric and Control or Benchmark
reference, then select a capability to inspect its scenarios. Latest and
historical run pages use the saved hierarchy. Check metrics retain their units,
directions, errors, skips, and coverage; resource metrics use per-trial averages.
There is no new composite benchmark score.

Resolve each trial file path relative to the output directory. Trial files
include model, runner, treatment/scenario identifiers, checks, judges, agent
sessions, walkthrough information, and artifact locations. Preserve the whole
Expand Down
Loading