Skip to content

feat: run experiments across benchmark scenarios - #348

Closed
joshblack with Copilot wants to merge 2 commits into
mainfrom
copilot/create-experiment-benchmark-option
Closed

joshblack with Copilot wants to merge 2 commits into
mainfrom
copilot/create-experiment-benchmark-option

Conversation

Copilot AI commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Experiments currently require callers to list scenarios manually, even when those scenarios already belong to a benchmark. This adds a benchmark reference as an alternative.

Changelog

New

  • Set benchmark to a benchmark filename ID instead of scenarios:

    export const experiment = defineConfig({
      name: 'Treatment comparison',
      description: 'Compare treatments across the design system benchmark',
      models: ['gpt-5.5'],
      benchmark: 'design-system',
      treatments: [{name: 'Skill'}],
    })
  • Use --benchmarks <directory> with experiment run and plan commands when benchmarks are outside ./benchmarks.

Changed

  • Expand the benchmark’s capabilities into distinct experiment scenarios. Experiment models and setup remain independent of benchmark models and setup.
  • Require exactly one of scenarios or benchmark in experiment configuration.

Co-authored-by: joshblack <3901764+joshblack@users.noreply.github.com>
Copilot AI changed the title [WIP] Add option to run experiment against a benchmark feat: run experiments across benchmark scenarios Oct 6, 2026
Copilot AI requested a review from joshblack October 6, 2026 16:53

@joshblack joshblack left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@copilot I think we need to find a good way to show result for experiments based on benchmarks that is distinct from how we currently compare scenarios. I think we'll want the taxonomy to include if the capability improved or not alongside the scenario diff. It'd be great if this could be its own path separate from the current experiment stuff so that we don't have to branch each time (instead its just a different call chain) but don't hold close to that if it makes things worse.

Copilot AI commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

@copilot I think we need to find a good way to show result for experiments based on benchmarks that is distinct from how we currently compare scenarios. I think we'll want the taxonomy to include if t...

The current experiment report has no capability membership or control-vs-treatment deltas, so this needs a benchmark-specific results path. Before I change code, could you confirm the intended scope and meaning of “improved”? My proposal is a separate CLI report showing capability → scenario comparisons against Control, with improvement based on directional check results (and N/A when checks are missing or incomparable). Should this also be represented in saved output/the website, and how should a capability with mixed improving and regressing checks be classified?

Copilot AI requested a review from joshblack October 6, 2026 17:38
@joshblack joshblack closed this Oct 6, 2026
@joshblack
joshblack deleted the copilot/create-experiment-benchmark-option branch October 6, 2026 20:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Create a way to run an experiment on a benchmark

2 participants