Repository navigation
Conversation
joshblack
added this pull request to stack #353
October 6, 2026 21:22
joshblack
marked this pull request as ready for review
October 7, 2026 15:54
Contributor
There was a problem hiding this comment.
🟡 Changes recommended
Saved benchmark metadata is misclassified in overview and mixed historical-run cases.
3 open findings
What changed in this PR
Adds capability-first comparison views for benchmark-backed experiments, including historical runs and scenario drilldowns.
Changes:
- Adds capability/treatment matrices with variant, metric, and reference selection.
- Preserves benchmark metadata and capability-specific artifact anchors.
- Separates treatment summaries by runner and updates documentation.
| File | Description |
|---|---|
website/src/scenario-anchor.ts |
Adds capability/scenario anchors. |
website/src/run-details.ts |
Adds capability metadata to trial details. |
website/src/experiments.ts |
Loads benchmark-backed experiments. |
website/src/experiment-results.ts |
Adds benchmark and runner-aware summaries. |
website/src/experiment-results.test.ts |
Tests runner-specific summaries. |
website/src/experiment-page-data.ts |
Exposes benchmark overview metadata. |
website/src/experiment-page-data.test.ts |
Tests saved benchmark names. |
website/src/benchmark-experiment-results.ts |
Builds capability comparison data. |
website/src/benchmark-experiment-results.test.ts |
Tests matrix aggregation semantics. |
website/src/app/experiments/[id]/runs/[date]/page.tsx |
Adds comparisons to historical runs. |
website/src/app/experiments/[id]/components/Page.tsx |
Adjusts benchmark run history columns. |
website/src/app/components/ScenarioAnchors.test.tsx |
Tests shared-scenario anchors. |
website/src/app/components/RunDetailsView.tsx |
Renders matrices and stable anchors. |
website/src/app/components/RunDetailsPage.tsx |
Forwards comparison data. |
website/src/app/components/ResourceTables.tsx |
Identifies experiment benchmarks. |
website/src/app/components/ExperimentResults.tsx |
Renders benchmark-aware latest results. |
website/src/app/components/BenchmarkExperimentResults.tsx |
Implements the comparison UI. |
website/src/app/components/BenchmarkExperimentResults.test.tsx |
Tests initial matrix rendering. |
skills/agent-eval/references/experiments.md |
Documents capability-first results. |
docs/experiments.md |
Documents website comparison behavior. |
🧠 Review effort: Balanced
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| field: 'checks', | ||
| align: 'end', | ||
| }, | ||
| ...(!results?.benchmark && !experiment.benchmark |
| experiments.map(async experiment => { | ||
| const run = await getLatestForExperiment(experiment.id) | ||
| const results = getExperimentResults(run ?? undefined) | ||
| const benchmark = results?.benchmark ?? experiment.benchmark |
Comment on lines
+81
to
+83
| <span className="text-caption text-muted"> | ||
| {treatment.trials} trials / {treatment.scenarios} scenarios | ||
| </span> |
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: a259adf3-3caa-433f-a9d7-7ac4dbc5101c
joshblack
force-pushed
the
feat/benchmark-experiment-ui
branch
from
October 7, 2026 16:43
3ce1fc3 to
e12ad16
Compare
Contributor
There was a problem hiding this comment.
🟡 Changes recommended
Benchmark reference notes can be mislabeled, filtered artifact links can target unmounted sections, and core interactions lack automated coverage.
6 open findings
Capability links cannot reveal filtered-out sections · New Benchmark notes are mislabeled as Control · New Use saved run metadata before falling back to current configuration Preserve per-run benchmark status in history column rendering Server-rendered test misses client interaction regressions · New Pluralize coverage labels based on their counts
🧠 Review effort: Balanced
Comment on lines
+295
to
+298
| <Link | ||
| href={ | ||
| `/experiments/${experimentId}/runs/${date}${getCapabilityScenarioAnchor(capability.id, row.id).fragment}` as Route | ||
| } |
| comparisons: Object.fromEntries( | ||
| references.map(reference => { | ||
| const baseline = summaries.get(reference.id)! | ||
| return [reference.id, String(formatCheckSummaries(summary, group, baseline).Checks ?? 'N/A')] |
Comment on lines
+48
to
+52
| function changeVariant(field: 'model' | 'reasoningEffort' | 'runner', value: string) { | ||
| const candidates = results.variants.filter(candidate => { | ||
| return ( | ||
| candidate[field] === value && | ||
| (field === 'model' || candidate.model === variant?.model) && |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.


Part of #329.
This PR adds capability-first results for benchmark-backed experiments. We compare Control, Benchmark, and custom treatments side by side, then select a capability to see its scenario comparisons.
Model, reasoning effort, runner, metric, and reference selections stay separate. Latest and historical run pages use saved capability metadata, with stable trial-section links for scenarios shared by capabilities. We keep the existing check and per-trial usage semantics instead of adding a composite benchmark score.
The screenshots use illustrative fixture results, not measured agent evaluations.
Before / After
Changelog
New
Changed
getExperimentScenarios, while explicitly passing the benchmark directory.Removed
**Validation (combined stack)
pnpm exec turbo run build: passed.pnpm exec turbo run lint lint:npm type-check: passed.pnpm exec turbo run test -- --run: passed (607 tests across 55 files).pnpm run format:diff: passed.Playwright preview: previously passed capability drilldown, metric/runner/reference selection preservation, historical run matrix, and capability-specific artifact targets; not repeated for this rebase.
No live agent evaluations were run.