docs(autotrain): scheduled continuous loop 0ln96zd — 4 screening cycles (non-positive, docs only) - #1482
docs(autotrain): scheduled continuous loop 0ln96zd — 4 screening cycles (non-positive, docs only)#1482Tyler-R-Kendrick wants to merge 4 commits into
Conversation
…-positive, reduced-steps lever) Cycle c1 of scheduled continuous-loop continuous-openui-scheduled-0ln96zd tried the untried "reduce screening training steps" lever from continuous-openui-scheduled-gmyilq c3's rank-1 priority (steps=10 instead of 20, wf_smoke_v2, size-matched bounds/control). Both arms still hit the same decode-capacity wall on this CPU-only sandbox (structural_similarity=0.0575, insufficient_n at fixture scale); non-positive per SDLC Phase A (fixture_insufficient_n, primary_metric_null_or_worse). No stack layer per autotrain-iteration-delivery. Adds a no-bump version-stamp history entry for harness.experiments.slm228_spectral_disposition (README/MODEL_CARD checkpoint-provenance note only; SLM-228 behavior unchanged), matching the established pattern from prior scheduled loops.
…-positive, component-plan lever) Cycle c2 tried the size-matched "component-plan" quality lever (steps=10, wf_smoke_v2) queued by c1's rank-1 priority. structural_similarity improves 0.0575 -> 0.0964 vs c1, and decode again completes without wall-timeout, but component-plan == control on this cycle's own matched pair (both 0.0964, no within-cycle delta) and remains far under the 0.35 ship gate at n=3. Non-positive per SDLC Phase A (fixture_insufficient_n, primary_metric_null_or_worse within-cycle). No stack layer. no-bump version-stamp entry for harness.experiments.slm228_spectral_disposition (README/MODEL_CARD checkpoint-provenance note only).
…-positive, component-edge lever) Cycle c3 tried the size-matched "component-edge" quality lever (steps=10, wf_smoke_v2). structural_similarity keeps climbing (0.0964 -> 0.174167), and an efficiency signal (mpr_per_ms) is close to but under the 5% min-effect threshold (gain_fraction=0.045). component-edge == control within this cycle's matched pair, so no within-cycle delta; still smoke n=3, still under the 0.35 ship gate. Non-positive per SDLC Phase A. No stack layer. no-bump version-stamp entry for harness.experiments.slm228_spectral_disposition.
…-positive, component-plan regression signal) Cycle c4 re-ran the "component-plan" lever (steps=10, wf_smoke_v2) and hit a degenerate long-decode seed: latency_ms_p50 jumped ~5-6x to ~28s (compiler_ms_mean=16226ms, tokens_emitted_mean=765 vs ~27-30 in prior cycles), and structural_similarity regressed to 0.0575, back to c1's level. This is a new diagnostic signal (seed variance under the reduced-steps lever can occasionally approach the decode wall again) rather than a metric win. Non-positive per SDLC Phase A. No stack layer. no-bump version-stamp entry for harness.experiments.slm228_spectral_disposition.
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
|
Warning Review limit reached
Next review available in: 1 minute You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (11)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Summary
Scheduled
/autotraincontinuous-loop firing. Ran 4 supervised screeningcycles of loop
continuous-openui-scheduled-0ln96zd(fixture-scale,CPU-only sandbox,
train_version=wf_smoke_v2,--ship-gateshonest, stepsreduced from the prior loop's default 20 → 10 per
continuous-openui-scheduled-gmyilqc3's queued rank-1 priority: "reducescreening training steps to leave more of the shared 180s
MAX_RUN_MINUTESwall budget for eval").
Per
autotrain/sdlcautotrain-iteration-delivery: no stack layer isopened, because none of the 4 cycles met the positive-result gate (primary
metric win / ship-quality win / executable unblock). This PR is docs-only
iron-law closeout + the version-stamp no-bump entries the driver's self-heal
required — landed per Phase B guidance ("open a single intentional PR ...
independently valuable", not a stack of failed-experiment noise).
smoke.structural_similaritympr_per_msefficiency signal close to but under the 5% min-effect gate (0.045)All cycles still fail ship gates on fixture evidence volume
(
smoke:insufficient_n actual=3 need>=20, missingheld_out/adversarial/ood/rico_heldsuites) — expected at smoke scale, not a ship claim.Changes
docs/design/continuous-loop-20260808-continuous-openui-schedu-33d4c6ef-c{1,2,3,4}-results.{md,json}— iron-law run docs for each cycledocs/MODEL_CARD.md,README.md— checkpoint-provenance notes for each cycle's fixture checkpoints (honesty stub, not a ship promotion)src/slm_training/resources/versions.json— 4no-bump:history entries forharness.experiments.slm228_spectral_disposition(README/MODEL_CARD touched, SLM-228 behavior unchanged)Next priorities (queued for the next scheduled firing)
Test plan
python -m scripts.verify_version_stamps --check— ok, 0 pending componentspython -m scripts.refresh_test_cases --check --changed— clean, no driftGenerated by Claude Code