Repository navigation
Conversation
Co-authored-by: joshblack <3901764+joshblack@users.noreply.github.com>
joshblack
left a comment
There was a problem hiding this comment.
@copilot I think we need to find a good way to show result for experiments based on benchmarks that is distinct from how we currently compare scenarios. I think we'll want the taxonomy to include if the capability improved or not alongside the scenario diff. It'd be great if this could be its own path separate from the current experiment stuff so that we don't have to branch each time (instead its just a different call chain) but don't hold close to that if it makes things worse.
The current experiment report has no capability membership or control-vs-treatment deltas, so this needs a benchmark-specific results path. Before I change code, could you confirm the intended scope and meaning of “improved”? My proposal is a separate CLI report showing capability → scenario comparisons against Control, with improvement based on directional check results (and N/A when checks are missing or incomparable). Should this also be represented in saved output/the website, and how should a capability with mixed improving and regressing checks be classified? |
Experiments currently require callers to list scenarios manually, even when those scenarios already belong to a benchmark. This adds a benchmark reference as an alternative.
Changelog
New
Set
benchmarkto a benchmark filename ID instead ofscenarios:Use
--benchmarks <directory>with experiment run and plan commands when benchmarks are outside./benchmarks.Changed
scenariosorbenchmarkin experiment configuration.