Skip to content

feat: compare benchmark experiment capabilities in the UI - #351

Closed
joshblack wants to merge 1 commit into
feat/benchmark-experiment-resultsfrom
feat/benchmark-experiment-ui
Closed

joshblack wants to merge 1 commit into
feat/benchmark-experiment-resultsfrom
feat/benchmark-experiment-ui

Conversation

@joshblack

@joshblack joshblack commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

Part of #329.

This PR adds capability-first results for benchmark-backed experiments. We compare Control, Benchmark, and custom treatments side by side, then select a capability to see its scenario comparisons.

Model, reasoning effort, runner, metric, and reference selections stay separate. Latest and historical run pages use saved capability metadata, with stable trial-section links for scenarios shared by capabilities. We keep the existing check and per-trial usage semantics instead of adding a composite benchmark score.

The screenshots use illustrative fixture results, not measured agent evaluations.

Before / After
Before After
Scenario-first experiment results Capability-first treatment matrix with scenario drilldown

Changelog

New

  • Add a capability comparison matrix with Control or Benchmark as the comparison reference.
  • Add scenario drilldown with the same metric and variant selections.
  • Add capability comparisons to historical experiment run pages.

Changed

  • Update experiment listings and overview links to identify their benchmark.
  • Update treatment summaries to keep runners separate.
  • Update artifact links to use stable capability/scenario targets.
  • Update missing-result presentation to distinguish unavailable comparisons from zero.
  • Update website data loading to narrow the loaded experiment variant and use getExperimentScenarios, while explicitly passing the benchmark directory.

Removed

  • Remove the pooled Checks column from benchmark-backed experiment run history, where it would combine different treatments.

**Validation (combined stack)

  • pnpm exec turbo run build: passed.
  • pnpm exec turbo run lint lint:npm type-check: passed.
  • pnpm exec turbo run test -- --run: passed (607 tests across 55 files).
  • pnpm run format:diff: passed.

Playwright preview: previously passed capability drilldown, metric/runner/reference selection preservation, historical run matrix, and capability-specific artifact targets; not repeated for this rebase.

No live agent evaluations were run.

@joshblack
joshblack added this pull request to stack #353 October 6, 2026 21:22
@joshblack
joshblack marked this pull request as ready for review October 7, 2026 15:54
Copilot AI balanced review requested due to automatic review settings October 7, 2026 15:54

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Saved benchmark metadata is misclassified in overview and mixed historical-run cases.

3 open findings
What changed in this PR

Adds capability-first comparison views for benchmark-backed experiments, including historical runs and scenario drilldowns.

Changes:

  • Adds capability/treatment matrices with variant, metric, and reference selection.
  • Preserves benchmark metadata and capability-specific artifact anchors.
  • Separates treatment summaries by runner and updates documentation.
File Description
website/​src/​scenario-anchor.ts Adds capability/scenario anchors.
website/​src/​run-details.ts Adds capability metadata to trial details.
website/​src/​experiments.ts Loads benchmark-backed experiments.
website/​src/​experiment-results.ts Adds benchmark and runner-aware summaries.
website/​src/​experiment-results.test.ts Tests runner-specific summaries.
website/​src/​experiment-page-data.ts Exposes benchmark overview metadata.
website/​src/​experiment-page-data.test.ts Tests saved benchmark names.
website/​src/​benchmark-experiment-results.ts Builds capability comparison data.
website/​src/​benchmark-experiment-results.test.ts Tests matrix aggregation semantics.
website/​src/​app/​experiments/​[id]/​runs/​[date]/​page.tsx Adds comparisons to historical runs.
website/​src/​app/​experiments/​[id]/​components/​Page.tsx Adjusts benchmark run history columns.
website/​src/​app/​components/​ScenarioAnchors.test.tsx Tests shared-scenario anchors.
website/​src/​app/​components/​RunDetailsView.tsx Renders matrices and stable anchors.
website/​src/​app/​components/​RunDetailsPage.tsx Forwards comparison data.
website/​src/​app/​components/​ResourceTables.tsx Identifies experiment benchmarks.
website/​src/​app/​components/​ExperimentResults.tsx Renders benchmark-aware latest results.
website/​src/​app/​components/​BenchmarkExperimentResults.tsx Implements the comparison UI.
website/​src/​app/​components/​BenchmarkExperimentResults.test.tsx Tests initial matrix rendering.
skills/​agent-eval/​references/​experiments.md Documents capability-first results.
docs/​experiments.md Documents website comparison behavior.

🧠 Review effort: Balanced


💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

field: 'checks',
align: 'end',
},
...(!results?.benchmark && !experiment.benchmark
experiments.map(async experiment => {
const run = await getLatestForExperiment(experiment.id)
const results = getExperimentResults(run ?? undefined)
const benchmark = results?.benchmark ?? experiment.benchmark
Comment on lines +81 to +83
<span className="text-caption text-muted">
{treatment.trials} trials / {treatment.scenarios} scenarios
</span>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: a259adf3-3caa-433f-a9d7-7ac4dbc5101c
Copilot AI balanced review requested due to automatic review settings October 7, 2026 16:43
@joshblack
joshblack force-pushed the feat/benchmark-experiment-ui branch from 3ce1fc3 to e12ad16 Compare October 7, 2026 16:43

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Benchmark reference notes can be mislabeled, filtered artifact links can target unmounted sections, and core interactions lack automated coverage.

6 open findings

🧠 Review effort: Balanced

Comment on lines +295 to +298
<Link
href={
`/experiments/${experimentId}/runs/${date}${getCapabilityScenarioAnchor(capability.id, row.id).fragment}` as Route
}
comparisons: Object.fromEntries(
references.map(reference => {
const baseline = summaries.get(reference.id)!
return [reference.id, String(formatCheckSummaries(summary, group, baseline).Checks ?? 'N/A')]
Comment on lines +48 to +52
function changeVariant(field: 'model' | 'reasoningEffort' | 'runner', value: string) {
const candidates = results.variants.filter(candidate => {
return (
candidate[field] === value &&
(field === 'model' || candidate.model === variant?.model) &&
@joshblack joshblack closed this Oct 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants