Skip to content

NanoBench: public review closed, request for a final version #60

Description

@will-lamerton

@RONAK-AI647 the public review window closed on 4 August, so NanoBench now moves from a proposal under review to a settled document. This issue is the last step before that happens.

Credit first. Every substantive correction raised during the window was addressed, and addressed fast: #30, #31, #32 and #33 all landed in #45 within days. The dual mode design that came out of #31, and the deterministic rule engine with confidence levels that came out of #33, both made the document meaningfully stronger than the version that arrived. That was a genuinely good review cycle, and the collective's first whitepaper from outside the core team setting that standard is not a small thing.

The four remaining threads were open questions rather than defects, and three of them were collective calls rather than yours. I have now closed all four with decisions, so nothing is left hanging on you to arbitrate:

What I would like in the final version

  1. §14 rewritten. The four open questions are now answered, so the section should record the decisions rather than pose the questions. If anything genuinely remains open, keep only that.
  2. §1 and §6 made consistent on orchestrator language, per Feedback for “NanoBench”: open question — orchestration language (Python vs shared toolchain) #34. Right now §1 says Python while the §6 diagram labels the ingestion layer "(TS)".
  3. §11 and §7 corrected on contamination, per Feedback for “NanoBench”: open question — community contribution validation gates #36. Claim the post cutoff merge dates and pinned SHAs, which are defensible today, and move similarity screening to v2.
  4. §7 gains the three consecutive run determinism gate. It is currently missing, and it matters more than it looks: a task that is flaky before the agent touches it can never produce a trustworthy score.
  5. §9.2 and §14 record the v1 sole reviewer limitation and the blinding protocol, per Feedback for “NanoBench”: open question — automated vs human-in-the-loop scoring #35.

Two loose ends

The upstream usage block has not been requested. You offered in #30 to raise a feature request on Nanocoder for token counts in the --json run report, and to write the PR. It has not been opened. Until it lands, §9.1's context_window_exceeded rule and §8's token telemetry have no data source, which means one of the six taxonomy categories is currently unimplementable and §8.5's sampling policy references a field that does not exist. Worth opening now. There are adjacent issues already on Nanocoder, Nano-Collective/nanocoder#756 and Nano-Collective/nanocoder#796, so check whether one can be extended rather than duplicated.

Nano-Collective/nanobench does not exist. Every v1 deliverable in §15, the CI gates in #36, and the dispute path in #37 all assume that repository. Say the word and I will create it so v1 has somewhere to live.

Ping me when the final version is up. I will do a last pass and move the status to accepted.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions