Skip to content

docs(autotrain): continuous-openui-scheduled-216646 c1-c4 closeout (non-positive) - #1518

Merged
Tyler-R-Kendrick merged 4 commits into
mainfrom
claude/great-dirac-n45ffs
Aug 9, 2026
Merged

docs(autotrain): continuous-openui-scheduled-216646 c1-c4 closeout (non-positive)#1518
Tyler-R-Kendrick merged 4 commits into
mainfrom
claude/great-dirac-n45ffs

Conversation

@Tyler-R-Kendrick

@Tyler-R-Kendrick Tyler-R-Kendrick commented Aug 8, 2026

Copy link
Copy Markdown
Owner

Summary

Scheduled /autotrain continuous-loop iteration (loop_id=continuous-openui-scheduled-216646), 4 supervised screening cycles at fixture scale (train_version=wf_smoke_v2, steps=20), each wall-capped by MAX_RUN_MINUTES=3.

Cycle Campaign Result
c1 ...eaa78363-c1 (bounds vs control) non-positive — smoke fixture insufficient_n, ship gates reject (expected at n=3)
c2 ...eaa78363-c2 (component-plan vs control) primary-metric win on smoke.structural_similarity (0.327→0.383) but fixture-scale only, no tracked code/docs delta beyond the auto-written results docs → stack_action=positive_no_tracked_delta_skip_stack
c3 ...eaa78363-c3 (fresh-seed confirmation of c2) non-positive — confirmation rejected the c2 candidate fingerprint (fixture-seed noise, not a real win)
c4 ...eaa78363-c4 (batch1 vs control) non-positive — ship gates reject

Per autotrain/sdlc iteration-delivery rules, a stacked PR layer is opened only for a positive run with a tracked code/docs delta. None of these four cycles qualify (c2's win did not survive fresh-seed confirmation in c3), so this PR carries only the iron-law documentation + auto-refreshed model-card/version-stamp bookkeeping for each cycle — no harness or model code changes.

Changes

  • docs/design/continuous-loop-20260808-continuous-openui-schedu-eaa78363-c{1,2,3,4}-results.{md,json} — per-cycle campaign results (iron law)
  • docs/MODEL_CARD.md, README.md — model-card summary refresh for the touched cycles
  • src/slm_training/resources/versions.json — version-stamp bookkeeping for the docs/version-stamp component bump

Test plan

  • scripts.verify_version_stamps --check --staged passed on each cycle's commit
  • Changed-file-scoped pytest targets (tests/test_scripts/test_verify_checkpoint_references.py, tests/test_versioning/*) passed on each cycle's commit
  • git fetch origin main / rebase-free merge check confirmed clean before push (branch was level with origin/main)
  • No ship-quality claim is made — smoke suite is fixture n=3, expected to fail volume/quality gates per honest-ship-eval

Generated by Claude Code

Summary by CodeRabbit

  • Documentation
    • Added continuous training campaign notes covering four screening cycles, checkpoint locations, evaluation results, and non-promotion status.
    • Added detailed reports and machine-readable records for candidate-versus-control comparisons, quality metrics, measurement status, and rejection reasons.
    • Documented fixture limitations, including insufficient sample sizes and non-positive outcomes.
  • Chores
    • Updated training history records with scheduled-loop checkpoints.
    • Confirmed no behavior changes or production-ready model promotion.

@vercel

vercel Bot commented Aug 8, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
slm-training Ready Ready Preview Aug 9, 2026 4:05am

Request Review

@coderabbitai

coderabbitai Bot commented Aug 8, 2026

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@Tyler-R-Kendrick, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 44 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 89ddc476-428e-428c-ad56-3e4a8fa5487d

📥 Commits

Reviewing files that changed from the base of the PR and between 62d896b and 0d4f5de.

📒 Files selected for processing (3)
  • README.md
  • docs/MODEL_CARD.md
  • src/slm_training/resources/versions.json
📝 Walkthrough

Walkthrough

The PR adds fixture-only screening records for scheduled campaign eaa78363, covering cycles c1–c4. It also adds matching checkpoint notes to the README, model card, and version history. No behavior or public entities change.

Changes

Continuous autotrain campaign

Layer / File(s) Summary
Cycle screening evidence
docs/design/continuous-loop-20260808-continuous-openui-schedu-eaa78363-c*-results.{json,md}
Added cycle c1–c4 metrics, metadata, fixture evidence, evaluation outcomes, rejection reasons, and non-ship qualifications.
Campaign checkpoint and provenance notes
README.md, docs/MODEL_CARD.md, src/slm_training/resources/versions.json
Added c1–c4 checkpoint paths and fixture-only, non-promotion history entries.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly describes the four-cycle Autotrain closeout and its non-positive outcome, which matches the main changes.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch claude/great-dirac-n45ffs

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

autotrain added 4 commits August 8, 2026 23:05
…uous-loop-20260808-continuous-openui-schedu-eaa78363-c1 closeout
…uous-loop-20260808-continuous-openui-schedu-eaa78363-c2 closeout
…uous-loop-20260808-continuous-openui-schedu-eaa78363-c3 closeout
…uous-loop-20260808-continuous-openui-schedu-eaa78363-c4 closeout
@Tyler-R-Kendrick
Tyler-R-Kendrick force-pushed the claude/great-dirac-n45ffs branch from 62d896b to 0d4f5de Compare August 9, 2026 04:05
@Tyler-R-Kendrick
Tyler-R-Kendrick merged commit 11e8860 into main Aug 9, 2026
1 of 3 checks passed
@Tyler-R-Kendrick
Tyler-R-Kendrick deleted the claude/great-dirac-n45ffs branch August 9, 2026 04:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant