Skip to content

release(sdk): 0.20.0 — migration note first, ordering documented - #114

Merged
maltsev-dev merged 20 commits into
masterfrom
release/0.20.0
Oct 1, 2026
Merged

maltsev-dev merged 20 commits into
masterfrom
release/0.20.0

Conversation

@maltsev-dev

Copy link
Copy Markdown
Member

Summary

Cuts v0.20.0 — the second half of DEF-MP-TS12-ENF-01. Where 0.19.0
closed the paths found by reading the code, 0.20 closes the ones that
only showed up when the properties were asserted end-to-end: a forged
allow authorised by decision_source absence, a production-usable
sensitive-tool fail-open opt-out, and an @protect-outside-@tool
collapse that made enforcement silently disappear under LangChain's
type-resolution errors.

This is a minor, not a patch: an unclassifiable refusal now
raises where 0.19.0 let the call proceed. That is a working loop
becoming a throwing one — a behavioural break, not a bug fix. The
release notes lead with the migration rather than burying it under the
security items; the load-bearing entry is #3.

Migration (read first if upgrading from 0.19.x)

Six things differ from 0.19.0. Only the first three can surprise you
at runtime, and the third is the one worth reading twice.

  1. A /gate test double must carry decision_source. Real
    backend answers always do — it is a non-Option String on
    GateResponse. A hand-written fixture that omits it now raises
    NullRunMalformedGateResponseError. Fix: add
    "decision_source": "gateway" next to "decision": "allow".

  2. NULLRUN_SENSITIVE_FAIL_OPEN against production now needs a
    second variable.
    NULLRUN_ALLOW_SENSITIVE_FAIL_OPEN=1 must be
    set too, otherwise the opt-out is refused, enforcement falls back
    to its own fail-CLOSED default, and the attempt is logged at ERROR
    with a metric. Non-production behaviour is unchanged.

  3. An unclassifiable refusal now raises where 0.19.0 let the call
    proceed.
    NullRunUnclassifiedRefusalError is importable from
    nullrun.breaker.categories (it is not re-exported at the top
    level) and carries error_code="NR-P003" and retryable=True.
    It is a NullRunInfrastructureError and a sibling of
    NullRunTransportError
    — deliberately not a subclass of it.
    So an existing except NullRunTransportError: arm, which in
    0.19.0 caught everything the gate could not classify and failed
    open, will not catch this one. That is the point: it is what
    stops a failed-open arm from swallowing a refusal. If you have
    such an arm, you have three options and they are not equivalent:

    from nullrun.breaker.categories import NullRunUnclassifiedRefusalError
    from nullrun.breaker.exceptions import NullRunError, NullRunTransportError
    
    try:
        ...
    except NullRunUnclassifiedRefusalError:
        raise                      # recommended: the SDK was right
    except NullRunTransportError:
        ...                        # still fails open, as before

    Catching it alongside NullRunTransportError restores 0.19.0's
    behaviour exactly — which means it restores the bypass. Do not
    widen the arm to NullRunError: that also swallows real policy
    refusals, which is a larger hole than the one this closes.

  4. on_denied is new. Default "raise", which is 0.19.0's
    behaviour. Set "message" to get a NullRunDeniedError carrying
    the server-authored agent_message for category="denied" only.

  5. NULLRUN_SKIP_BUDGET_CHECK outside production now logs and
    increments a metric
    when it skips the check. Enforcement
    behaviour is unchanged.

  6. /execute refusal bodies carry category and agent_message.
    Additive on the wire.

Carried over from 0.19.0 and still true on 0.20.0, because it is the
same class of break and the same shape of fix:

  • NullRunRuntime.execute(..., mode="inline") is gone and has no
    replacement — every call goes through /execute.
  • nullrun.runtime.register_strict_mode_forced /
    is_strict_mode_forced
    are gone, along with @guarded,
    nullrun.handle (renamed nullrun.guard in 0.18.5),
    nullrun.status(), and nullrun.auto_instrument.

Security

  • DEFS-MP-TS12-ENF-01.4 (c459379) — A /gate body must state
    who decided before it counts as a verdict.
    _require_gate_decision rejects a body whose decision_source is
    absent or unrecognised, raising
    NullRunMalformedGateResponseError. The runtime's rule was
    decision_source != fallback → honour the wire decision, which
    read a missing decision_source as more trustworthy than
    fallback. So {"decision": "allow"} — a captive portal's login
    JSON, an intercepting proxy's stub, anything on-path — was enough
    to authorise a call no policy engine evaluated. The field is a
    non-Option String on the backend's GateResponse
    (gate/internal.rs:638) and every producer sets "gateway", so
    a real answer always carries one; there is no legitimate body
    this rejects.
  • DEFS-MP-TS12-ENF-01.5 (cdf94c3) — NULLRUN_SENSITIVE_FAIL_OPEN
    is refused against production. It was read straight into the
    enforcement path, letting a sensitive tool's body run with no
    policy evaluation at all. Its sibling NULLRUN_SKIP_BUDGET_CHECK
    has been production-guarded since it was caught doing the same;
    the asymmetry was an oversight, and this half is the more
    dangerous one — that one skips a pre-flight, this one skips the
    gate. Requires both NULLRUN_SENSITIVE_FAIL_OPEN=1 and
    NULLRUN_ALLOW_SENSITIVE_FAIL_OPEN=1; the documented dev/test
    use is unchanged.
  • DEFS-MP-TS12-ENF-01.6 (700fea7) — The non-prod budget bypass
    is now visible. When NULLRUN_SKIP_BUDGET_CHECK=1 skips the check,
    it logs and increments a metric instead of being
    indistinguishable from a normal call.
  • DEFS-MP-TS12-ENF-01.7 (83960b4 + 4cb009c + d191a7f) —
    Every gate refusal is classified. on_denied="message" is the
    operator-facing "explain, don't crash" path for category="denied"
    only (budget and halt deliberately excluded — an agent told "that
    tool is not allowed" when the truth is "you are out of money" will
    look for another way to spend). /execute refusal bodies now
    carry category and agent_message; tests pin §4.7 option 2
    (a 503 refusal blocks, an outage fails open).

Fixed

  • DEFS-MP-TS12-ENF-01.8 (b38fcf5) — @protect no longer
    loses a LangChain tool when it is the outer decorator. Applied to
    a @tool result, @protect wrapped the object with
    functools.wraps and returned a plain function — so .invoke,
    .name, .args_schema and .description were gone, an agent
    loop could not bind the tool, and a tool the loop cannot see
    produces no refusal. protect now duck-types the argument
    (.invoke + .name, then .coroutine before .func) and
    returns the same object. Duck-typed rather than isinstance
    because langchain_core is an optional dependency. This is a
    silent enforcement loss users may have shipped against
    :
    enforcement was absent whenever @protect was the outer
    decorator on a LangChain tool, and the user-visible error
    pointed at LangChain's type-resolution layer rather than at the
    real cause.

Documentation

  • 939fd7d + 660954a + c2e52d5 — README states the three
    enforcement limitations precisely (TLS-interception residual,
    server-side trust, old-SDK version skew), and the provenanced
    boundary of server-authoritative enforcement.
  • 642ebef + 467cb1a — CHANGELOG migration note for the
    refusal-category SDK; records the second half of
    DEF-MP-TS12-ENF-01 (the half that only showed up when the
    properties were asserted end-to-end).

Verification

Check Result
ruff check src tests All checks passed
mypy src/nullrun Success: no issues found in 37 source files
pytest -q 1597 passed, 1 skipped in ~116s (vs baseline 1482 / 1 — +115 new tests)
Scratch diff clean
nullrun.__version__ 0.20.0
Wire-format additive only on /execute refusal bodies (category, agent_message); non-additive on /gate (decision_source is now REQUIRED where it was previously optional on the SDK side — but a +it was always present on a real backend answer)

Commits included

836b9eb fix(gate): a /gate body with no verdict must never read as "allowed"
cdf94c3 fix(gate): refuse NULLRUN_SENSITIVE_FAIL_OPEN against production
700fea7 fix(gate): make the non-prod budget bypass visible in logs and metrics
3c45138 test(gate): pin that a 503 is an allow and a 403 is a stop (ADR-063 1.3f)
838a905 test: correct the ADR-063 fixtures to the real wire nesting
83960b4 fix(breaker): classify every gate refusal, and raise when it cannot be
4cb009c feat(breaker): add on_denied, a message path for `denied` and nothing else
408668a test(adr063): pin §4.7 option 2 — a 503 refusal blocks, an outage fails open
d191a7f fix(breaker): carry the category and the agent message onto /execute refusals
939fd7d docs(readme): state the three enforcement limitations, verified
a33f863 fix(docs): name ADR-064's discriminator, not a marker that is not on the wire
c459379 fix(gate): a /gate body must state WHO decided before it is a verdict
c2e52d5 docs(readme): state the provenance boundary of server-authoritative enforcement
a3ea0ce style: clear the three lint errors this branch's own test files carry
660954a docs(readme): name the trust boundary precisely, and the second version skew
6d0285b test(protect): prove refusals and provenance at the decorator level
467cb1a docs(changelog): record the second half of DEF-MP-TS12-ENF-01
642ebef docs(changelog): migration note for the refusal-category SDK
b38fcf5 fix(protect): make @tool/@protect ordering irrelevant
f8bfbdc release(sdk): 0.20.0 — migration note first, ordering documented

`check_workflow_budget` read the verdict as

    decision = response.get("decision", "allow")

The default is not merely redundant, it is a hole. `decision` is a
non-Option field with no `skip_serializing_if` on the backend's
GateResponse (gate/internal.rs:637), so every real NULLRUN backend
serialises it on every answer. A body without one did not come from
NULLRUN -- a proxy error page, a captive portal, a TLS interception
box -- so the default authorised calls that no policy engine had
evaluated. ADR-008 grants fail-OPEN to *transport* failures, where the
gate never got to rule. A body that arrived but carried no verdict is
the opposite case, and the SDK read it as the former.

It also hid a reserved refusal. `GateDecision::Deny`
(gate/internal.rs:579, reserved by ADR-046, no producer yet) had no arm
in the method, so it fell off the end -- which the caller reads as "no
block raised, proceed". The one refusal on the wire contract that
executed. Deny now raises, so the day ADR-046 ships its producer the
SDK already fails CLOSED.

The non-JSON body was a second path to the same place.
`Transport.check` ends in `response.json()`, and the resulting
JSONDecodeError (a ValueError) matched neither the auth arm nor the
fail-OPEN arm in the cached branch, while the uncached branch's
`except Exception` swallowed it as "gate unavailable" -- a non-NULLRUN
responder read as an outage. Both branches now convert it to a typed
error, placed before the broad arm it must precede.

New: `NullRunMalformedGateResponseError(NullRunProtocolError)`. A
subclass rather than a new root so existing `except
NullRunProtocolError` handlers keep catching it, and it shares the
NR-P* family without overloading NR-P001, whose user_action is
specifically "upgrade the SDK". Its own code is NR-P002.

Verified by mutation, not by assertion. Reverting the extraction to
the old default fails 5 of 11; disabling the deny arm fails
test_deny_decision_raises; removing the ValueError arm fails
test_non_json_body_raises_typed_error with the raw JSONDecodeError.
The first mutation pass was incomplete -- it reverted only the cached
branch, and the test passed, because the uncached branch still had its
own arm. Reverted both, and the test bit.

Suite: 1493 passed, 1 skipped, 0 failed.
The opt-out was read straight into the enforcement path with no
environment check:

    fail_open = os.environ.get("NULLRUN_SENSITIVE_FAIL_OPEN", "") == "1"

Its sibling NULLRUN_SKIP_BUDGET_CHECK has been production-guarded
since it was caught doing the same thing. The asymmetry was an
oversight, and this is the more dangerous of the two: the budget
opt-out skips a pre-flight, while this one lets the body of a
sensitive tool run with no policy evaluation at all -- an unblocked
charge_card during an outage, which ADR-008 calls a security
regression rather than an availability trade-off.

Resolution now lives in `NullRunRuntime.sensitive_fail_open_enabled`
and the decorator asks the runtime, so the environment policy stays in
one module with `_is_production_environment` and the two opt-outs
cannot drift apart the way their predicates did under
DEF-MP-TS12-ENF-01.

In production the flag alone is refused: ignored, logged at ERROR,
counted as `sensitive_fail_open_blocked_in_prod`. Enforcement falls
back to its own fail-CLOSED default, so an operator who set the var
carelessly keeps a working agent rather than a crash loop, and the
attempt is visible instead of silent. An incident-response runbook can
still force it with NULLRUN_ALLOW_SENSITIVE_FAIL_OPEN=1, mirroring
NULLRUN_ALLOW_SKIP_BUDGET_CHECK -- logged at WARNING and counted as
`sensitive_fail_open_allowed_in_prod`.

The guard is verified two ways, and the second is the one that
matters. Seven predicate tests cover the env matrix. The e2e test
covers the caller, because a correct helper with a caller that still
reads the raw env var looks identical from the predicate alone: under
mutation (decorator reading os.environ again) the seven still pass and
only `test_sensitive_body_does_not_run_in_prod` fails, logging
"body will run" against the production host.

`_StubRuntime` in test_langchain_enforcement.py gained the method.
Its "every gate passes" contract means reporting the flag OFF, not
mirroring the raw read -- mirroring would have hidden the very
regression the new tests exist to catch.

Suite: 1502 passed, 1 skipped, 0 failed.
NULLRUN_SKIP_BUDGET_CHECK=1 outside production was a `logger.debug`
and a bare `return`. DEBUG is below every default handler, and
nothing was counted, so a test suite running the whole budget path
with the bypass on left no trace: tests went green, the dashboard
showed the org spending nothing, and the only evidence the gate was
never consulted was the absence of a block.

CLAUDE.md's rule is that a test passing only with this flag set is
evidence of a broken gate, not of a working one. That rule is not
actionable unless setting the flag is loud enough to find. Now: a
WARNING naming the flag and stating that no budget, rate-limit or
tool-block check ran (so a reader can tell a bypass from an allow),
plus a `skip_budget_used_non_prod` counter a CI dashboard can alert
on. The production refusal already emitted
`skip_budget_blocked_in_prod`; this is its dev/test counterpart.

Verified by mutation: reverting to the `logger.debug` line fails all
three new tests. The fourth asserts the production refusal is
untouched by making the dev path louder.

Two test-infrastructure corrections this surfaced:

`mock_prod_api` moved to conftest. Both guard test files need a
runtime that genuinely looks like production, and the first version
duplicated the prod auth response in each module -- the same
duplication that let the AUTH_ERROR predicate drift between
`check_workflow_budget` and `_run_tool_policy_gate` under
DEF-MP-TS12-ENF-01.

The new tests set the env vars with `os.environ[...] = ...` instead of
`monkeypatch.setenv`, which leaked NULLRUN_SKIP_BUDGET_CHECK into
every later test in the session. Symptom was confusing: the files
passed alone and the full suite failed 9 tests in
test_v3_wire_contract.py, because those tests' budget pre-flight was
being skipped by a flag set three modules earlier. Now
monkeypatch-scoped, and the full suite is green.

Suite: 1506 passed, 1 skipped, 0 failed.
….3f)

The backend now answers a failed workflow-state read with 503 and a
tripped breaker with 403. Both are correct on the backend -- the gate
blocks either way -- but the SDK does something different with each,
and the difference is invisible from the gate's own tests.

  403 CIRCUIT_BREAKER_TRIPPED     -> real gateway decision, honoured,
                                     raises, ONE attempt
  503 *_LOOKUP_FAILED             -> retried 3x, then a synthetic
                                     FALLBACK decision, which
                                     check_workflow_budget treats as a
                                     transport error and returns
                                     WITHOUT raising. An allow.

So the current SDK behaves identically to 0.18.5 here. That is not a
stale-SDK problem an SDK release fixes -- it is ADR-008's documented
fail-OPEN on transport error, unchanged. Pinned so the asymmetry is a
visible fact rather than something discovered during an incident, and
so a future release that tightens it has to come with an ADR amendment.

The 503 test asserts the CURRENT fail-OPEN on purpose. Asserting the
desired behaviour would ship a red test; asserting "it raises" would be
a lie. The comment says so at the assertion.

Also pins that the "do not retry" instruction survives the round trip.
The backend ships no agent_message for a halt-category refusal and no
SDK version reads that field, so the explanation is the only channel a
model sees -- worth a test, since nothing else would notice its loss.

4 passed; 29 with the neighbouring fallback/gate-path suites.
The two response bodies in this file put `error_code` at the top level.
It is not there. Captured from a live gate run:

  top_level_keys=["category", "decision", "decision_source", "details",
    "explanation", "policy_id", "policy_version",
    "projected_cost_cents", "remaining_budget_cents", "reservation_id",
    "staleness_ms", "user_message"]

`error_code` lives at `details.error_code` (`internal.rs:716`), which is
where the backend's own status mapper reads it (`gate.rs:88-90`).
`category` / `user_message` are the top-level fields, set by
`attach_refusal_surface` (`gate.rs:163-172`).

These fixtures were hand-assembled when ADR-064 was written and the key
landed where a reader would expect it. Harmless as a behaviour pin --
the SDK reads `explanation`, so all four assertions are unaffected and
still pass. Not harmless as a reference: ADR-064's SDK-side plan reads
exactly this body, and a client written against the wrong nesting
classifies every refusal as unparseable, which is fail-OPEN.

ADR-064 records the correction. 4 passed.
ADR-062 §2.2. The gate does not send a bare "no" — it sends a
`category`, one of denied | budget | halt | infra, and that is what
tells the SDK whether the answer is a policy decision, a money wall,
a stop, or a backend fault.

The SDK had no vocabulary for it, so every refusal became the same
thing. This adds `nullrun.breaker.categories` mirroring the backend
enum, resolves the category at the wire boundary in
`Transport.check`, and puts it on the returned decision.

The load-bearing half is what happens when the category is ABSENT or
UNRECOGNISED. It raises `NullRunUnclassifiedRefusalError` rather
than guessing. Guessing is the dangerous direction specifically: a
refusal rounded to something permissive can be rendered as a
friendly, model-readable "that's not allowed", and an agent handed
that adapts and retries against a stop an operator deliberately
placed. That is the bypass ADR-061 closed server-side, re-opened
here.

Two runtime arms had to change for the raise to mean anything.
`check_workflow_budget`'s ADR-008 fail-OPEN arms catch
`NullRunError`, and `NullRunUnclassifiedRefusalError` is one, so
without an explicit re-raise the unclassifiable refusal was caught by
the fail-OPEN arm and returned as `None` — the agent proceeding on a
call the backend refused. Verified by mutation: dropping the two
re-raise arms turns both branches red with
"check_workflow_budget: /gate unavailable, failing open: Gate refused
the call (decision="block") but sent no 'category' field". That is
DEF-MP-TS12-ENF-01's exact shape with a different trigger.

ADR-008 is not narrowed. A real outage still fails open; the
counter-test pins it, and the arm order is what makes the two
distinguishable — the re-raise is listed before the `NullRunError`
and bare `except Exception` arms precisely because both are supersets
of it.

Also re-files the `DEF-NR-TOOLBLOCKED-PARSER` source-pin comment
into `_parse_v3_error_envelope`, where the parser actually is. It
was filed against a `check()` branch that does not exist, so the pin
pointed readers at the wrong function; the pin now asserts the
function it means.
… else

A `denied` refusal is the one category where telling the model is
both useful and safe: a rule judged *this call* unacceptable, where
another call or a different tool might pass. The other three
describe a wall the model cannot climb, and a model told about a
wall walks into it — it tool-shops after "your budget is exhausted",
routes around a pause an operator deliberately placed, and tries to
solve a backend fault that is not the model's problem.

`on_denied="message"` therefore sits behind a single guard,
`category is DecisionCategory.DENIED`, at the one block site in
`check_workflow_budget`. No arrangement of the flag turns a budget,
halt, or infra refusal into model-readable text — verified by
mutation: widening the guard to consult `on_denied` alone turns the
three negative cases and the budget-exception case red.

The resulting `NullRunDeniedError` (NR-D001, non-retryable) carries
`agent_message` — server-authored text the backend guarantees is
model-safe (`attach_refusal_surface` sends it only for `Denied`) —
behind one accessor, so a host cannot reach for `str(exc)` and put
operator internals in the model's context by accident. The default is
`"raise"`, and an unrecognised value is rejected at construction
rather than treated as `"raise"`, so a typo cannot quietly disable
the path a host asked for.

`halt` and `infra` refusals still surface as `NullRunBudgetError`.
Recorded, not fixed: the exception family does not yet match the
category. It is a diagnostics defect, not a bypass — no model text is
produced on either path and `on_denied` cannot reach them.
…ls open

Product decision, 2026-09-30. The previous version of this file
pinned an asymmetry that was not a property of the categories but an
accident of the status code: a 403 stopped the agent and a 503 did
not, purely because `400 <= status < 500` happened to be the branch
that produced a gateway decision.

The rule now follows the failure's MEANING rather than its number.
Ordinary unavailability stays fail-open; a failure OF THE CHECK
ITSELF gets a marker the SDK treats as a block. SDKs predating the
category work keep failing open, by definition — they never had a way
to read the marker.

The marker already existed (`category: "infra"`, and the backend
already splits the two 503 groups via `GateErrorCode::is_fail_closed()`),
so this was a client-side mapping change, not a re-architecture.

Verified by mutation: reverting `Transport.check` to its pre-fix
4xx-only condition turns the 503 cases red with
"synthetic decision_source='fallback', treating as transport error" —
the gate's own fail-CLOSED refusal answered to the agent as
"allowed". That was DEF-MP-TS12-ENF-01's exact shape with a
different trigger, and it was wider than a single code path: the
whole 5xx band was being converted into a synthetic transport failure.

The 502 counter-test is what keeps ADR-008's promise intact. A 5xx
with no refusal body is a proxy that never reached the gate, not a
verdict, and must still fail open — otherwise every deploy would
freeze every agent.
…refusals

`on_denied` and the ADR-062 §2.2 category rule were only ever wired
into `check_workflow_budget`, which calls `/gate`. Every MCP tool
call goes through `runtime.execute` → `/api/v1/execute` instead, so on
that path the flag was inert: a host that set `on_denied="message"` got
a promise it did not keep for the whole class of tool calls that MCP
adapters mediate, and the server-authored `agent_message` never
reached it.

Two independent pieces were missing on `/execute`.

`_extract_error_envelope` had no arm for a gate refusal envelope. A real
refusal puts `error_code` at `details.error_code` (backend
`internal.rs:716`), not at the top level, so a `/execute` 403
TOOL_BLOCKED fell through to the legacy slug arm and arrived at the
host as `NullRunAuthenticationError: Auth failed on execute (status
403, error_code='')`. Wrong class, wrong `format_user_message`, and an
operator action — "check your credentials" — that cannot fix a policy
decision. Added the missing envelope shape.

Exceptions built by the parser carried neither the category nor the
server's text. Split `_parse_v3_error_envelope` into a thin wrapper
over `_parse_v3_error_envelope_uncategorised` so all ~15 return points
stamp `wire_category` (left ABSENT when the body has no `category` —
never defaulted, or a missing field would become a confident answer),
`wire_error_code`, and the server-authored `agent_message` /
`user_message`. Then `runtime.execute` converts a
`NullRunBlockedException` to `NullRunDeniedError` under the same
single-guard discipline as the `/gate` block site: `on_denied` is
consulted only for `DecisionCategory.DENIED`, read from the wire. No
arrangement of the flag turns a budget, halt, or infra refusal into
model-readable text, and an absent category still raises.

Test notes, since two of them are not what they first look like.

`infra` is absent from the negative parameter set on purpose. ADR-063
§4.7 sends infra refusals as 503, and `_retry_with_backoff` maps the
whole 5xx band on `/execute` to NullRunTransportError/GATEWAY_ERROR
before the envelope parser is reached; the runtime then raises a
STRICT fallback block at `runtime.py:3670`, downstream of the
`on_denied` arm. The property is real and is asserted, but it is
enforced by the 5xx band and not by the category check — so the test
pins the concrete class to say which mechanism is in play.

The negatives originally used a synthetic `SOME_CODE`, which the
parser maps to a non-block class, so the exceptions never entered the
arm and all three parameters passed regardless of the guard. Widening
the guard to consult `on_denied` alone left them green. They now use
real codes verified to map into `NullRunBlockedException`. All three
mutations are recorded in the test docstrings.

The `DEF-NR-TOOLBLOCKED-PARSER` source pins are repointed in the same
commit because they are not separable from the parser split — the
split alone leaves the suite red, and splitting them across commits
would violate the one-green-commit-per-step rule. They resolved the
parser by name, which after the split points at the wrapper and
asserts nothing. They now resolve the dispatching function by content
(the branch text, not the exception name: the name also appears in the
import block, so resolving on it left `test_branch_in_parser` green
after the branch was deleted — verified). Deleting the branch now
turns 13 of 14 red.

1543 passed, 1 skipped.
The README carried no fail-open claim and no limitations section, so
the two most consequential behaviours of the gate were only
discoverable by reading the source. Adds both.

Each limitation was checked against the code before being written
down, and two of them are not what the shorthand description of them
suggests.

The 503 case fails OPEN on old SDKs, not closed. A failed security
check is served as a 503 carrying a `category`; the current SDK reads
it and refuses. Pre-category there was nothing to read, the 503 became
a synthetic FALLBACK decision, and `check_workflow_budget` fails open
on a FALLBACK source — so the call proceeds. Verified by probing both
cases through a real runtime and by reading the pre-category source
(`git show 83960b4~1:src/nullrun/transport.py`): the 5xx band had no
`is_refusal` condition and fell straight to the FALLBACK return.

The breaker setting is server-side, not an SDK default. `LogOnly` is
`NULLRUN_GATE_CB_TRIP_ENFORCEMENT_MODE` in the backend, not a fallback
mode in this SDK (there is no `LogOnly` anywhere in `src/`). The
production boot check refuses to start unless it is explicitly
`Enforce` or `LogOnly` (backend `main.rs:290-293`); `detect_mode()`
still defaults to `LogOnly` on unset outside production
(`main.rs:294-295`). In `LogOnly` the trip is recorded and alerted but
tripped workflows still pass `/check` — the backend's own words at
`main.rs:286-288`.

Pause and kill really are 403-only, and that is now stated as the
limitation it is: `WORKFLOW_PAUSED` and `WORKFLOW_INACTIVE` are served
as the same 403 from the same key, distinguishable only by
operator-facing text, which the backend deliberately keeps distinct
(`error_codes.rs:3400-3412`). Support tooling should key off the error
code, not the status.

Also adds the fail-open summary, matching the ADR-008 table in
`runtime.py` — which the module docstring requires to be kept in
lockstep with any README claim. There was no README claim to keep in
step before this; there is one now.
…the wire

The transport comment and two test docstrings repeated ADR-063 §4.7's
claim that the backend "already distinguishes the two 503 groups
internally via GateErrorCode::is_fail_closed()". That is wrong in the
way ADR-063 §4.7 was wrong: is_fail_closed is an ordinary Rust method
that is never serialised, so no client could have read it. The SDK
never did — the code has always keyed on decision == "block"
(categories.py:168), which is ADR-064's discriminator.

The comment also claimed BUDGET_DATA_UNAVAILABLE "arrives as 503 with
decision=block". It does not, and cannot: grepping the backend, that
string appears only in error_codes.rs and budget.rs, never in
internal.rs or orchestrator.rs, so no gate producer emits it. The one
response that does carry it (budget.rs:246,
BudgetUnavailableResponse) has no decision field at all — which is
correctly treated as a non-refusal and fails open per ADR-008.

Logic unchanged; this is a claim correction, and every ADR-063 §4.7
citation for the 503 rule retargeted to ADR-064, which now owns it.

44 passed.
The refusal classifier splits a 5xx into "the gate ruled" and
"nothing ruled" by reading the body. Every other category change on
this branch rests on the same unexamined premise: that a body on the
SDK's HTTPS connection was written by NULLRUN.

A test for that premise found the premise is false in one direction.
`{"decision": "allow"}` from an arbitrary responder passes every check
`_require_gate_decision` applied -- it is a dict, and `decision` is a
known string -- and was then honoured as a gateway decision, because
the runtime's rule reads a MISSING decision_source as "not fallback"
and therefore as authoritative. So an on-path responder that gets
JSON in front of the SDK: a captive portal on a hotel WLAN, a
corporate TLS-interception proxy. It needs neither the API key nor a
HMAC bypass, only to answer before the real backend. The call is
authorised, /track books its cost against a policy nobody consulted,
the audit trail records an allow, and there is no later point at
which it is caught.

Fix: require `decision_source` to be a known provenance value BEFORE
looking at `decision`. It is the right field because it cannot be
absent from a real answer -- `GateResponse.decision_source` is a
non-Option String with no skip_serializing_if (`gate/internal.rs:638`).
`GateResponseBody` in schemas.rs declares it optional, but that struct
is referenced only from openapi.rs: it is the documentation schema,
not the wire.

Mutation-verified: with the guard absent, the new test reports DID
NOT RAISE on `{"decision": "allow"}`. 18 passed with it.

The same guard caught a drifted fixture, not a second defect:
test_v3_wire_contract.py's require_approval body omitted all three
non-Option fields on GateResponse. Corrected to the real shape.

Full suite 1561 passed / 1 skipped.
…nforcement

Limitation 4. The three existing entries describe what the SDK does
not enforce; this one describes what it cannot verify. Every
enforcement claim in the README rests on a premise it never said out
loud: that the JSON on the HTTPS connection was written by NullRun.

The SDK now checks for it -- a /gate body with no usable
decision_source is rejected rather than acted on -- but that is a
field in the body, not a signature. An on-path responder that can
produce a well-formed body including decision_source defeats it, and
needs neither the API key nor an HMAC bypass. NullRun's answers are
not signed, so there is no after-the-fact detection either: the audit
trail would record the allow.

Also extends the fail-open summary with the not-a-verdict case, which
is newly user-visible as NullRunMalformedGateResponseError.

This is the honest boundary of "server-authoritative": authoritative
against a client, not against an attacker who owns the path.
ruff runs over src/ only in CI, so none of these would have failed a
build -- but all three files are new on this branch (neither exists on
origin/master), so they are this branch's to leave clean.

Two F841: the `rt = make_runtime(...)` binding was never used, because
@Protect resolves the runtime from the context, not the local. The call
is load-bearing; the assignment was not.

One I001: nullrun.runtime sorted after nullrun.toolbox.
…on skew

Limitation 4 said "use https://", which understates the problem. It
implied https was sufficient, when the residual is specifically an
on-path responder holding a certificate the OS already trusts for
api.nullrun.io -- i.e. corporate TLS interception, which is ordinary
on a managed network and which https alone does nothing about.

Corrected to state what each layer does and does not give:

- certificate verification is on and CANNOT be disabled by
  configuration. `verify_cert = True` (transport.py:572) and the only
  override is NULLRUN_TLS_CA_CERT (:568), which swaps in a trust
  anchor you chose and is still verification. Grepped exhaustively:
  the only TLS env vars in the SDK are NULLRUN_TLS_CA_CERT /
  _CLIENT_CERT / _CLIENT_KEY, none of which turns verification off.
- plain http is refused (InsecureTransportError, :526).
- responses are not signed, so no after-the-fact detection: the audit
  trail would faithfully record an allow the gate never gave.

The consequence is stated as the operational control it is -- exclude
api.nullrun.io from interception and verify it stays excluded -- and
not as something the SDK does for you.

Also adds the old-SDK note beside limitation 1, since it is the same
version skew pointing the same way and a reader who pins for the 503
guarantee should know they also need it for the forged-allow case:
both a real refusal and a fabricated permission are read permissively
on an older SDK, and neither shows up in the SDK's output.
Every other file in this cluster calls runtime.check_workflow_budget()
directly. That is the wrong level for the question a user has, because
@Protect is what they call -- and @Protect is a separate code path with
its own try, its own except BaseException, and its own two wrappers
(sync and async).

Reading decorators.py says the refusals propagate. That is exactly the
reasoning that let DEF-MP-TS12-ENF-01 ship: check_workflow_budget's
except Exception looked harmless in isolation too, until it was reached
by a real 401.

14 tests, four properties:

  - a denied / budget / unclassifiable refusal, and a body with no
    decision_source, do NOT run the decorated body;
  - the counter-tests: a real allow, a connection failure under
    on_denied="message", and a 502 with no refusal body all DO run it
    (a suite of only "did not run" assertions passes against a
    decorator that refuses everything);
  - on_denied reaches @Protect for category=denied and, proven by the
    negative parametrisation, does not reach budget/halt;
  - the async wrapper is not a separate property, so the refusal, the
    forged allow and the real allow are each proven on it too.

Mutation-verified: deleting the decision_source guard in
_require_gate_decision turns exactly the sync and async forged-allow
tests red (DID NOT RAISE), and the other 12 stay green.

Verified in a clean python:3.11 container on the unmutated tree:
  ruff check src/ tests/  -- All checks passed
  mypy src/              -- no issues in 37 source files
  pytest -q              -- 1575 passed, 1 skipped
0.19.0 is already on origin/master and its entry closed the paths
found by reading the code. These commits closed the paths that only
appeared once the properties were asserted end-to-end, and none of
them were in the release notes -- so the two most serious items
(a forged allow authorised by decision_source absence, and the
production-usable sensitive fail-open opt-out) were documented only in
README.

Filed as [Unreleased] rather than inventing a version number. That is
deliberate: the behaviour change is real and needs a migration note,
not a patch number chosen by accident. An unclassifiable refusal now
raises where 0.19.0 let the call proceed, and a hand-rolled /gate
test double must now carry decision_source. Whoever cuts the release
picks 0.19.1 or 0.20.0 with that in hand.

Verified citation: GateResponse.decision_source is a non-Option String
at backend/src/proxy/http/gate/internal.rs:638.
Six differences from 0.19.0, written as things you can act on rather
than as a list of changed internals. The third is the one worth
reading twice: an unclassifiable refusal now raises where 0.19.0 let
the call proceed, and NullRunUnclassifiedRefusalError is deliberately
a SIBLING of NullRunTransportError, not a subclass -- so an existing
`except NullRunTransportError:` fail-open arm will not catch it.
That is the fix working, and it is also the thing most likely to turn
a working loop into a throwing one.

The note gives the exact import path (nullrun.breaker.categories --
it is not re-exported at the top level, which was verified, not
assumed), the error_code, and the three options as three, with the
consequence of each stated. Widening the arm to NullRunError is
called out as worse than the hole being closed, because it also
swallows real policy refusals.

Every code shape in the note was executed against the installed
package rather than reasoned about:

    recommended two-arm  -> propagates (recommended)
    one-arm permissive   -> swallowed == 0.19.0 behaviour, bypass restored
    widened to NullRunError -> a real policy refusal is swallowed too
    issubclass(NRError, NullRunError) = True
Priority 1 for DEF-MP-TS12-ENF-01 was "one @Protect, rest under the
hood", and the part nobody had tested is whether that survives the
framework layer. It did not, in one direction.

`@protect` above `@tool` collapsed the tool to a plain function.
`functools.wraps`-based wrapping cannot preserve `.invoke`, `.name`,
`.args_schema` or `.description`, so the agent loop could not bind it --
and a tool the loop cannot see produces no refusal. Measured on
langchain-core 0.3.86:

    @Protect          # outer
    @tool             # inner -> StructuredTool
    def f(query: str) -> str: ...

    # f is now <function>, not StructuredTool
    # convert_to_openai_tool(f) -> NameError: name 'Annotated' is not defined

That last line is the failure mode worth naming: the error a user sees
is a langchain type-resolution error, not "your tool is not a tool any
more". Enforcement was gone and the symptom pointed somewhere else.

`protect` now duck-types the argument (`.invoke` + `.name`, then
`.coroutine` before `.func` -- an async tool sets both, with `func` as a
sync fallback that raises if called), wraps whichever callable is live,
and returns the SAME object. Every other attribute is untouched, so the
loop sees the tool it saw before. Duck-typed rather than isinstance
because `langchain_core` is an optional dependency; the three attributes
matched are the ones `BaseTool.run()`/`arun()` dispatch through.

Three properties were established by probe on the real library before
being written down, and all three are now pinned:

- Ordering is irrelevant in BOTH directions. `@tool` outside already
  worked; the reverse now does too, and each has an allow counter-test
  so "the body did not run" cannot pass against a wrapper that refuses
  everything.
- `handle_tool_error=True` cannot turn a refusal into model-visible
  text. It only catches `ToolException`; a `NullRunBudgetError` reaches
  the bare `except (Exception, KeyboardInterrupt)` arm and is re-raised.
  This is the DEF-MP-TS12-ENF-01 shape in a different costume, so it is
  pinned rather than assumed.
- `@protect` with no `init()` fails LOUD. The lazy resolver propagates
  `NullRunAuthenticationError` (NR-A001) and the body does not run.
  FIX-4 removed the fallback that used to swallow this; the AST guard
  keeps it removed, and an AST check rather than a substring search
  because the docstring describes the very `except` being asserted gone.

`on_denied="message"` is now exercised through a real `@tool` too, in
both orders and both sync/async, with the negative set (budget, halt)
kept at the framework layer: an agent told "that tool is not allowed"
when the truth is "you are out of money" will look for another way to
spend.

Mutation-verified; each mutant ran the suite and the restore was
byte-identical:

  sync arm converts refusal to text   ->  2 red
  async arm converts refusal to text  ->  1 red
  FIX-4 fallback in lazy resolver     ->  2 red
  BaseTool in-place wrap removed      ->  5 red
  on_denied category guard dropped    ->  8 red across 3 files

Three earlier mutation attempts were invalid and are not counted: an
inner `try` cannot observe a `_protect_body.__enter__` raise, and one
anchor matched `get_protected_runtime` instead of
`_get_or_create_runtime`.

Suite: 1597 passed, 1 skipped. ruff and mypy clean.
Minor, not patch: an unclassifiable refusal now RAISES where 0.19.0 let
the call proceed. That turns a working loop into a throwing one, which
is a behavioural break, so the release note leads with the migration
rather than burying it under the security items.

`### Migration` is the first section under 0.20.0 and enumerates all six
behavioural differences from 0.19.0, then repeats the three inherited
from 0.19.0 that are the same class of break and the same shape of fix
(`mode="inline"` removed, `decision_source` now required in a `/gate`
test double, `handle` -> `guard`). A reader upgrading from 0.18.x sees
the whole break list in one place instead of reconstructing it across
two releases.

The load-bearing item is #3: `NullRunUnclassifiedRefusalError` is a
sibling of `NullRunTransportError`, not a subclass, so an existing
`except NullRunTransportError:` arm will not catch it. Catching it
alongside restores 0.19.0's behaviour exactly -- which restores the
bypass. The note says so and says not to widen to `NullRunError`, which
is a larger hole than the one this closes.

Also records the `@tool`/`@protect` fix under `### Fixed`, since it is a
silent enforcement loss users may have shipped against: enforcement was
absent whenever `@protect` was the outer decorator on a LangChain tool.

README gains the two things a user would otherwise get wrong:

- decorator order, both directions, with the reason (an unbound tool
  cannot refuse) and the `NameError` that was the real symptom;
- `on_denied="message"` as the operator-facing "explain, don't crash"
  mode -- `category="denied"` only, budget and halt deliberately
  excluded -- and why LangChain's `handle_tool_error=True` is not an
  equivalent (it catches `ToolException` and does not know which
  exceptions are refusals).

Every claim in the new README section was executed before being written
down, including the pre-fix `NameError`, which is reproduced rather than
asserted from memory.

Not published to PyPI. Suite: 1597 passed, 1 skipped; ruff and mypy
clean.
@codecov

codecov Bot commented Oct 1, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 93.10345% with 12 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
src/nullrun/runtime.py 88.88% 8 Missing ⚠️
src/nullrun/breaker/categories.py 95.23% 1 Missing and 1 partial ⚠️
src/nullrun/decorators.py 85.71% 1 Missing and 1 partial ⚠️

📢 Thoughts on this report? Let us know!

@maltsev-dev
maltsev-dev merged commit 46d6ee0 into master Oct 1, 2026
5 checks passed
@maltsev-dev
maltsev-dev deleted the release/0.20.0 branch October 1, 2026 15:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant