Skip to content

SLM-258: publish LotusOpenUIDispositionV1 (inconclusive_missing_evidence) - #943

Merged
Tyler-R-Kendrick merged 10 commits into
mainfrom
slm-258-lot4-02-disposition
Jul 26, 2026
Merged

SLM-258: publish LotusOpenUIDispositionV1 (inconclusive_missing_evidence)#943
Tyler-R-Kendrick merged 10 commits into
mainfrom
slm-258-lot4-02-disposition

Conversation

@Tyler-R-Kendrick

Copy link
Copy Markdown
Owner

Summary

Publishes SLM-258 (LOT4-02) LotusOpenUIDispositionV1 — the final fail-closed disposition for the LOTUS-to-OpenUI transfer, generated from the real committed upstream gate artifacts (no retyped numbers):

  • SLM-248 needs_target_trace_contract, SLM-249 inconclusive, SLM-250…257 all not_authorized → adoption verdict inconclusive_missing_evidence (SLM-248 explicitly ruled out reject_transfer; this is a scoped prerequisite, not a rejection)
  • Every minimum-evidence leg blocked; no interpretable-reasoning/planning language (causal-use gate not positive); no equal-quality cost/latency/FLOPs/energy claims
  • Mechanism dispositions: all model mechanisms not_identifiable; CompilerReasoningTraceV1 + gate evaluators retained as keep_diagnostic assets (trace/eval assets survive the negative model disposition, as the issue allows)
  • Production defaults unchanged — V1 has no adopt_default; a default switch would need a separate rollout issue
  • Negative-result registry + required follow-ups recorded (run the SLM-249 oracle-ceiling campaign; if positive, the live gate evaluators flip and LOT1 work gets re-filed)

Stacked on PR #942 (chain: #927#930#934#936#937#939#941#942 → this).

Verification

  • pytest tests/test_harnesses/experiments/test_lotus_openui_disposition.py tests/test_scripts/test_publish_lotus_openui_disposition.py — 8 passed (incl. missing-artifact fail-closed, interpretability-language block, no-default-change, trace-asset retention)
  • python -m scripts.publish_lotus_openui_disposition — exit 0; version-stamped docs/design/iter-slm258-lot4-02-disposition-20260725.{json,md}
  • python -m scripts.verify_version_stamps --check — ok (harness.experiments v117, new harness.experiments.lotus_openui_disposition v1)
  • python -m scripts.repo_policy — ok; git diff --check — clean

Honest scope

  • Docs/disposition only: no training, deployment, model-card/README checkpoint claims (no checkpoint exists), or production change.

claude and others added 9 commits July 25, 2026 10:51
LOT1-01's own hard activation gates require SLM-248's fidelity-contract
verdict to be authorize_bounded_implementation and SLM-249's trace gate
verdict to be oracle_ceiling_positive. The real committed upstream
artifacts report needs_target_trace_contract and inconclusive
respectively (SLM-249's own allowed_lot1_implementation field says
"none: ... not authorized by this issue"). Per the issue's own text,
this closes not_authorized in plan-only mode without any Kxc
model/training code.

Adds a small, reusable, tested LotusOpenUIModelContractV1 evaluator
that reads the two upstream contracts and derives the verdict from
their real published fields (not hardcoded to always fail — a
synthetic both-gates-met case is tested to flip the result to
authorized_wiring_only). Re-running the same CLI after SLM-249 gets a
real oracle-ceiling campaign will honestly reflect the new
disposition.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KShriKrGosZr67yVPFgi78
LOT1-02's launch prerequisites are unmet against the real committed
upstream contracts: SLM-248 reports needs_target_trace_contract (not
authorize_bounded_implementation), SLM-249's trace gate reports
inconclusive (not oracle_ceiling_positive), its own
allowed_lot1_implementation authorizes no LOT1 implementation, and
LOT1-01 (SLM-250) itself closed not_authorized -- so no faithful K x c
model path, curriculum hooks, or Stage 0 parent exist to train, and
there is no treatment arm to attribute against a continued-explicit
control. Per LOT1-01's own gate law this closes LOT1-02 not_authorized
in plan-only mode with no training/curriculum/model code.

Adds a small tested Lot102LaunchGateV1 evaluator that reuses the
SLM-250 lot1_01_activation_gate module (no parallel path) and adds the
SLM-249 explicit-allowance check; synthetic positive fixtures prove it
flips to authorized_wiring_only when prerequisites are honestly met.
LOT2-01's activation requires SLM-251 (LOT1-02) to authorize LOT2 with
a selected parent, faithful treatment checkpoint/recipe, continued
explicit control, curriculum, and fairness manifests. The LOT1-02
launch gate reports not_authorized against the real committed upstream
contracts (SLM-248 needs_target_trace_contract, SLM-249 inconclusive
with allowed_lot1_implementation=none, LOT1-01 itself not_authorized),
so none of those artifacts exist and the supervision routing x timing x
grounding factorial cannot launch honestly.

Adds a generic LotDownstreamGateV1 evaluator covering all LOT2+ issues
through one registry (no per-issue parallel paths), reusing the
lot1_02 gate. Even a synthetic positive LOT1 chain yields only
upstream_authorized_pending_campaign_gate, never a launch or quality
claim. No training/factorial/readout code is added.
LOT2-02 requires SLM-251's qualified faithful checkpoint/recipe plus
continued explicit control and SLM-252's selected supervision routing;
the LOT1-02 launch gate is not_authorized against the real committed
contracts, so neither exists and the PCL vs structured/set-valued/
block-autoregressive readout matrix cannot launch honestly. Registers
SLM-253 in the shared LOT downstream gate (structured-latent-readout-gate-v1)
with not_authorized iter docs; no readout/objective code added.
LOT3-01's activation gates require a qualified faithful latent
checkpoint (SLM-251), a selected routing contract (SLM-252), a selected
readout contract (SLM-253), and a nonzero oracle substitution ceiling;
per the issue text, when these fail the issue closes with
causal_study_not_authorized rather than manufacturing interpretability
evidence. The LOT1-02 launch gate is not_authorized against the real
committed contracts, so there is no latent checkpoint to capture,
probe, or intervene on. Registers SLM-254 in the shared LOT downstream
gate (causal-latent-use-gate-v1) with not_authorized iter docs; no
capture/intervention/probe code added.
LOT3-02 requires SLM-253's selected readout contract and SLM-254's
valid causal factor/intervention encoding; the LOT1-02 launch gate is
not_authorized against the real committed contracts, so neither exists
and there is no latent neighborhood to learn or intervene on. Registers
SLM-255 in the shared LOT downstream gate
(alternative-valid-latent-gate-v1) with not_authorized iter docs; no
set-valued objective, generator, or intervention code added.
LOT2-03 requires SLM-251's qualified faithful treatment/control and
SLM-252's selected routing contract; the LOT1-02 launch gate is
not_authorized against the real committed contracts, so there is no
looped-latent workspace to sweep and no routing contract to hold fixed.
Registers SLM-256 in the shared LOT downstream gate
(latent-workspace-capacity-gate-v1) with not_authorized iter docs; no
capacity/order matrix code added.
LOT4-01 requires SLM-251's faithful treatment and continued explicit
checkpoints with fairness manifests plus the LOT2/LOT3
objective/readout/workspace selections; the LOT1-02 launch gate is
not_authorized against the real committed contracts, so there are no
checkpoints to benchmark and no equal-quality frontier to measure.
Registers SLM-257 in the shared LOT downstream gate
(lotus-openui-compute-frontier-v1) with not_authorized iter docs; no
benchmark/telemetry code added.
…nce)

Aggregates the real committed upstream gate artifacts into the final
fail-closed LOT4-02 disposition: SLM-248 needs_target_trace_contract,
SLM-249 inconclusive, SLM-250..257 all not_authorized. Adoption is
blocked on every minimum-evidence leg, no interpretable-reasoning or
cost language is allowed, production defaults are unchanged (V1 has no
adopt_default), the trace contract and gate evaluators are retained as
diagnostic assets, and the negative-result registry + required
follow-ups (oracle-ceiling campaign, then re-run the live gate
evaluators) are recorded. All fields are generated from the artifacts,
not retyped.
@vercel

vercel Bot commented Jul 25, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
slm-training Ready Ready Preview, Comment Jul 26, 2026 11:13pm

Request Review

@coderabbitai

coderabbitai Bot commented Jul 25, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@Tyler-R-Kendrick, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 46 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 5bda7838-3b36-4270-a1d1-d56873c696a2

📥 Commits

Reviewing files that changed from the base of the PR and between 8efc34c and 5566949.

📒 Files selected for processing (7)
  • docs/design/iter-slm258-lot4-02-disposition-20260725.json
  • docs/design/iter-slm258-lot4-02-disposition-20260725.md
  • scripts/publish_lotus_openui_disposition.py
  • src/slm_training/harnesses/experiments/lotus_openui_disposition.py
  • src/slm_training/resources/versions.json
  • tests/test_harnesses/experiments/test_lotus_openui_disposition.py
  • tests/test_scripts/test_publish_lotus_openui_disposition.py
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch slm-258-lot4-02-disposition

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Base automatically changed from slm-257-lot4-01-not-authorized to main July 26, 2026 23:11
…sition

# Conflicts:
#	src/slm_training/resources/versions.json
@Tyler-R-Kendrick
Tyler-R-Kendrick merged commit 9d33549 into main Jul 26, 2026
4 of 6 checks passed
@Tyler-R-Kendrick
Tyler-R-Kendrick deleted the slm-258-lot4-02-disposition branch July 26, 2026 23:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants