Skip to content

docs(autotrain): continuous-openui-scheduled-bcz5ho c1-c8 closeout (non-positive) - #1501

Merged
Tyler-R-Kendrick merged 8 commits into
mainfrom
claude/great-dirac-bcz5ho
Aug 8, 2026
Merged

docs(autotrain): continuous-openui-scheduled-bcz5ho c1-c8 closeout (non-positive)#1501
Tyler-R-Kendrick merged 8 commits into
mainfrom
claude/great-dirac-bcz5ho

Conversation

@Tyler-R-Kendrick

@Tyler-R-Kendrick Tyler-R-Kendrick commented Aug 8, 2026

Copy link
Copy Markdown
Owner

Summary

Tracks incremental commits from a scheduled /autotrain continuous-loop session (loop-id openui-scheduled-bcz5ho), run fixture-scale and local (no paid/remote compute — no session authority was granted for that; environment has no GPU). This was a fresh container (no venv, no node_modules), so this firing also bootstrapped the Python venv (.venv, pip install -e ".[dev,hf]") and the OpenUI/DESIGN.md JS bridges before the first cycle.

Eight cycles ran this firing, all non-positive per the SDLC Phase A gate (no primary-metric win, no ship-quality win, no unblock) — so this PR is documentation/provenance only, not a model or harness change:

  • c1 (bounds/control, TwoTower, fixture wf_smoke_v2): honest ship-gate reject on fixture n (insufficient_n actual=3 need>=20); control/bounds landed on an identical primary metric (structural_similarity=0.0575).
  • c2 (control/component-plan): evaluation measurement incomplete — decode_timeout_count=3 on control (CPU-only, no-GPU, 3-minute wall cap), missing scoreboard on the component-plan candidate.
  • c3 (frozen replay of c2): retried the identical incomplete arm; control hit the same decode-timeout signature again (2nd consecutive occurrence on that arm — one more identical occurrence would have crossed the loop's repeated-hard-block threshold; it did not recur on that specific arm again this firing).
  • c4 (control/bounds): honest ship-gate reject on fixture n; null primary-metric delta (0.4167 both arms).
  • c5 (control/component-plan): measurement incomplete, missing scoreboard on the candidate (soft failure, different arm than c2/c3).
  • c6 (frozen replay of c5, retry_measurement): component-plan hit a decode-timeout this time (decode_timeout_count=3); still measurement-incomplete.
  • c7 (control/canvas): honest ship-gate reject on fixture n; null primary-metric delta (0.0575 both arms).
  • c8 (control/component-plan): honest ship-gate reject on fixture n; a small mpr_per_ms efficiency gain (7.89e-58.18e-5, +3.6%) was measured but rejected as sub-threshold (needs ≥5%).

No cycle crossed the 3-consecutive-identical-hard-block threshold, so the loop kept going in-process rather than reporting blocked. Two distinct arms (c2/c3's control, c6's component-plan) each hit CPU-only decode-timeout signatures — worth watching on a future firing since a 3rd identical occurrence on the same arm would be a real infra signal to route through improve-openui-harnesses.

Per sdlc / autotrain-iteration-delivery, non-positive cycles only get local commits + iron-law docs, never a new stack layer. Each cycle's doc note to README/MODEL_CARD also required a no-bump: version-stamp history entry for harness.experiments.slm228_spectral_disposition (behavior unchanged) per scripts.verify_version_stamps.

Note: this environment does not have the gh CLI available, so this PR is opened directly against main (no gh stack) — matching how prior non-positive screening batches in this repo have landed when a stack tool wasn't available. Commit authorship on the driver's self-heal commits was rebased/amended to the session's identity before each push (GitHub verified-commit requirement).

This firing is now wrapping up (no positive-result cycle landed, so there is no stack layer to close out per Phase B — only this docs-only PR). The next scheduled firing will continue the loop from a fresh container.

Test plan

  • python -m scripts.verify_version_stamps --check --stagedok (each cycle)
  • Local pytest subset triggered by the pre-commit changed-check hook passed (tests/test_scripts/test_verify_checkpoint_references.py, tests/test_versioning/*)
  • Full ship-gate suite (not run — fixture-scale screening cycles only, honest-ship-eval still pending a positive/ship-scale run)

@vercel

vercel Bot commented Aug 8, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
slm-training Ready Ready Preview Aug 8, 2026 3:12pm

Request Review

@coderabbitai

coderabbitai Bot commented Aug 8, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@Tyler-R-Kendrick, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 46 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 2bf34093-7ffa-43a7-ac19-db191aa642b8

📥 Commits

Reviewing files that changed from the base of the PR and between 0f5a1f0 and a390a5e.

📒 Files selected for processing (17)
  • README.md
  • docs/MODEL_CARD.md
  • docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c2-results.json
  • docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c2-results.md
  • docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c3-results.json
  • docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c3-results.md
  • docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c4-results.json
  • docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c4-results.md
  • docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c5-results.json
  • docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c5-results.md
  • docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c6-results.json
  • docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c6-results.md
  • docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c7-results.json
  • docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c7-results.md
  • docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c8-results.json
  • docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c8-results.md
  • src/slm_training/resources/versions.json
📝 Walkthrough

Walkthrough

The pull request adds a fixture-only continuous autotrain screening record for the 2026-08-08 OpenUI campaign. It records candidate and control metrics, rejection reasons, checkpoint provenance, non-promotion status, and unchanged SLM-228 behavior.

Changes

Continuous autotrain screening

Layer / File(s) Summary
Screening results and outcome record
docs/design/continuous-loop-...-results.json, docs/design/continuous-loop-...-results.md
The results artifacts record campaign metadata, candidate and control metrics, completed measurements, a non-positive outcome, insufficiency reasons, and fixture-only status.
Campaign documentation and provenance
README.md, docs/MODEL_CARD.md, src/slm_training/resources/versions.json
The README and model card list the campaign and checkpoint paths. The version history records fixture provenance and unchanged SLM-228 behavior.

Estimated code review effort: 1 (Trivial) | ~5 minutes

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly identifies the autotrain closeout, campaign, cycle scope, and non-positive outcome documented by the pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch claude/great-dirac-bcz5ho

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In
`@docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c1-results.json`:
- Around line 37-41: Preserve the numeric sample-size gate inputs in both result
artifacts: in
docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c1-results.json
lines 37-41, add canonical observed and required sample-count fields for both
arms, including the rejected values actual=3 and need>=20; in
docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c1-results.md
lines 9-11, render those same counts in the historical narrative.

In `@README.md`:
- Around line 796-800: Update README.md lines 796-800 to add campaign
continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c1 to the model-card
summary with roster, evaluations, recipe metadata, honesty/no-sync status,
non-promotion status, and both exact checkpoint paths. Update docs/MODEL_CARD.md
lines 1452-1456 with newest-first checkpoint-history coverage and the
corresponding roster and evaluation sections, preserving the same metadata and
resolvable paths at both canonical model-card surfaces.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 3e7c1782-295b-49a0-985a-764da6a6c446

📥 Commits

Reviewing files that changed from the base of the PR and between b12eb2c and 0f5a1f0.

📒 Files selected for processing (5)
  • README.md
  • docs/MODEL_CARD.md
  • docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c1-results.json
  • docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c1-results.md
  • src/slm_training/resources/versions.json

Comment on lines +37 to +41
"reasons": [
"fixture_insufficient_n:c20260808-openui-scheduled-bcz5ho-5720c848-c1-bounds",
"fixture_insufficient_n:c20260808-openui-scheduled-bcz5ho-5720c848-c1-control",
"primary_metric_null_or_worse:smoke.structural_similarity:control=0.057499999999999996 candidate=0.057499999999999996 improvement=0.0",
"fixture_insufficient_n_alone"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Preserve the numeric sample-size gate inputs in both result artifacts.

The rejection currently depends on actual=3 and need>=20, but neither artifact stores those values. Add canonical sample-count fields to the JSON and render the same values in the Markdown.

  • docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c1-results.json#L37-L41: add observed and required sample counts for both arms.
  • docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c1-results.md#L9-L11: include the same counts in the historical narrative.
📍 Affects 2 files
  • docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c1-results.json#L37-L41 (this comment)
  • docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c1-results.md#L9-L11
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In
`@docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c1-results.json`
around lines 37 - 41, Preserve the numeric sample-size gate inputs in both
result artifacts: in
docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c1-results.json
lines 37-41, add canonical observed and required sample-count fields for both
arms, including the rejected values actual=3 and need>=20; in
docs/design/continuous-loop-20260808-openui-scheduled-bcz5ho-5720c848-c1-results.md
lines 9-11, render those same counts in the historical narrative.

Comment thread README.md
@Tyler-R-Kendrick Tyler-R-Kendrick changed the title docs(autotrain): continuous-openui-scheduled-bcz5ho c1 closeout (non-positive) docs(autotrain): continuous-openui-scheduled-bcz5ho c1-c4 closeout (non-positive) Aug 8, 2026
@Tyler-R-Kendrick Tyler-R-Kendrick changed the title docs(autotrain): continuous-openui-scheduled-bcz5ho c1-c4 closeout (non-positive) docs(autotrain): continuous-openui-scheduled-bcz5ho c1-c8 closeout (non-positive) Aug 8, 2026
@Tyler-R-Kendrick
Tyler-R-Kendrick force-pushed the claude/great-dirac-bcz5ho branch from 482995c to a390a5e Compare August 8, 2026 15:11
@Tyler-R-Kendrick
Tyler-R-Kendrick merged commit 01a5d62 into main Aug 8, 2026
1 of 2 checks passed
@Tyler-R-Kendrick
Tyler-R-Kendrick deleted the claude/great-dirac-bcz5ho branch August 8, 2026 15:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants