You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
@RONAK-AI647 the public review window closed on 4 August, so NanoBench now moves from a proposal under review to a settled document. This issue is the last step before that happens.
Credit first. Every substantive correction raised during the window was addressed, and addressed fast: #30, #31, #32 and #33 all landed in #45 within days. The dual mode design that came out of #31, and the deterministic rule engine with confidence levels that came out of #33, both made the document meaningfully stronger than the version that arrived. That was a genuinely good review cycle, and the collective's first whitepaper from outside the core team setting that standard is not a small thing.
The four remaining threads were open questions rather than defects, and three of them were collective calls rather than yours. I have now closed all four with decisions, so nothing is left hanging on you to arbitrate:
Feedback for “NanoBench”: open question — automated vs human-in-the-loop scoring #35 Automated versus human evaluation. Scoring stays fully automated with no human override. A sampled human audit checks Level 4 classification only, publishes an agreement rate, and collapses the two Level 4 categories if agreement falls below 90%. Blinding is mandatory from the second reviewer onward.
§14 rewritten. The four open questions are now answered, so the section should record the decisions rather than pose the questions. If anything genuinely remains open, keep only that.
§7 gains the three consecutive run determinism gate. It is currently missing, and it matters more than it looks: a task that is flaky before the agent touches it can never produce a trustworthy score.
The upstream usage block has not been requested. You offered in #30 to raise a feature request on Nanocoder for token counts in the --json run report, and to write the PR. It has not been opened. Until it lands, §9.1's context_window_exceeded rule and §8's token telemetry have no data source, which means one of the six taxonomy categories is currently unimplementable and §8.5's sampling policy references a field that does not exist. Worth opening now. There are adjacent issues already on Nanocoder, Nano-Collective/nanocoder#756 and Nano-Collective/nanocoder#796, so check whether one can be extended rather than duplicated.
Nano-Collective/nanobench does not exist. Every v1 deliverable in §15, the CI gates in #36, and the dispute path in #37 all assume that repository. Say the word and I will create it so v1 has somewhere to live.
Ping me when the final version is up. I will do a last pass and move the status to accepted.
@RONAK-AI647 the public review window closed on 4 August, so NanoBench now moves from a proposal under review to a settled document. This issue is the last step before that happens.
Credit first. Every substantive correction raised during the window was addressed, and addressed fast: #30, #31, #32 and #33 all landed in #45 within days. The dual mode design that came out of #31, and the deterministic rule engine with confidence levels that came out of #33, both made the document meaningfully stronger than the version that arrived. That was a genuinely good review cycle, and the collective's first whitepaper from outside the core team setting that standard is not a small thing.
The four remaining threads were open questions rather than defects, and three of them were collective calls rather than yours. I have now closed all four with decisions, so nothing is left hanging on you to arbitrate:
--jsonreport shape is consumed through a checked in schema, with CI validating a real payload against it.What I would like in the final version
Two loose ends
The upstream
usageblock has not been requested. You offered in #30 to raise a feature request on Nanocoder for token counts in the--jsonrun report, and to write the PR. It has not been opened. Until it lands, §9.1'scontext_window_exceededrule and §8's token telemetry have no data source, which means one of the six taxonomy categories is currently unimplementable and §8.5's sampling policy references a field that does not exist. Worth opening now. There are adjacent issues already on Nanocoder, Nano-Collective/nanocoder#756 and Nano-Collective/nanocoder#796, so check whether one can be extended rather than duplicated.Nano-Collective/nanobenchdoes not exist. Every v1 deliverable in §15, the CI gates in #36, and the dispute path in #37 all assume that repository. Say the word and I will create it so v1 has somewhere to live.Ping me when the final version is up. I will do a last pass and move the status to accepted.