release(sdk): 0.20.0 — migration note first, ordering documented - #114
Merged
Merged
Conversation
`check_workflow_budget` read the verdict as
decision = response.get("decision", "allow")
The default is not merely redundant, it is a hole. `decision` is a
non-Option field with no `skip_serializing_if` on the backend's
GateResponse (gate/internal.rs:637), so every real NULLRUN backend
serialises it on every answer. A body without one did not come from
NULLRUN -- a proxy error page, a captive portal, a TLS interception
box -- so the default authorised calls that no policy engine had
evaluated. ADR-008 grants fail-OPEN to *transport* failures, where the
gate never got to rule. A body that arrived but carried no verdict is
the opposite case, and the SDK read it as the former.
It also hid a reserved refusal. `GateDecision::Deny`
(gate/internal.rs:579, reserved by ADR-046, no producer yet) had no arm
in the method, so it fell off the end -- which the caller reads as "no
block raised, proceed". The one refusal on the wire contract that
executed. Deny now raises, so the day ADR-046 ships its producer the
SDK already fails CLOSED.
The non-JSON body was a second path to the same place.
`Transport.check` ends in `response.json()`, and the resulting
JSONDecodeError (a ValueError) matched neither the auth arm nor the
fail-OPEN arm in the cached branch, while the uncached branch's
`except Exception` swallowed it as "gate unavailable" -- a non-NULLRUN
responder read as an outage. Both branches now convert it to a typed
error, placed before the broad arm it must precede.
New: `NullRunMalformedGateResponseError(NullRunProtocolError)`. A
subclass rather than a new root so existing `except
NullRunProtocolError` handlers keep catching it, and it shares the
NR-P* family without overloading NR-P001, whose user_action is
specifically "upgrade the SDK". Its own code is NR-P002.
Verified by mutation, not by assertion. Reverting the extraction to
the old default fails 5 of 11; disabling the deny arm fails
test_deny_decision_raises; removing the ValueError arm fails
test_non_json_body_raises_typed_error with the raw JSONDecodeError.
The first mutation pass was incomplete -- it reverted only the cached
branch, and the test passed, because the uncached branch still had its
own arm. Reverted both, and the test bit.
Suite: 1493 passed, 1 skipped, 0 failed.
The opt-out was read straight into the enforcement path with no
environment check:
fail_open = os.environ.get("NULLRUN_SENSITIVE_FAIL_OPEN", "") == "1"
Its sibling NULLRUN_SKIP_BUDGET_CHECK has been production-guarded
since it was caught doing the same thing. The asymmetry was an
oversight, and this is the more dangerous of the two: the budget
opt-out skips a pre-flight, while this one lets the body of a
sensitive tool run with no policy evaluation at all -- an unblocked
charge_card during an outage, which ADR-008 calls a security
regression rather than an availability trade-off.
Resolution now lives in `NullRunRuntime.sensitive_fail_open_enabled`
and the decorator asks the runtime, so the environment policy stays in
one module with `_is_production_environment` and the two opt-outs
cannot drift apart the way their predicates did under
DEF-MP-TS12-ENF-01.
In production the flag alone is refused: ignored, logged at ERROR,
counted as `sensitive_fail_open_blocked_in_prod`. Enforcement falls
back to its own fail-CLOSED default, so an operator who set the var
carelessly keeps a working agent rather than a crash loop, and the
attempt is visible instead of silent. An incident-response runbook can
still force it with NULLRUN_ALLOW_SENSITIVE_FAIL_OPEN=1, mirroring
NULLRUN_ALLOW_SKIP_BUDGET_CHECK -- logged at WARNING and counted as
`sensitive_fail_open_allowed_in_prod`.
The guard is verified two ways, and the second is the one that
matters. Seven predicate tests cover the env matrix. The e2e test
covers the caller, because a correct helper with a caller that still
reads the raw env var looks identical from the predicate alone: under
mutation (decorator reading os.environ again) the seven still pass and
only `test_sensitive_body_does_not_run_in_prod` fails, logging
"body will run" against the production host.
`_StubRuntime` in test_langchain_enforcement.py gained the method.
Its "every gate passes" contract means reporting the flag OFF, not
mirroring the raw read -- mirroring would have hidden the very
regression the new tests exist to catch.
Suite: 1502 passed, 1 skipped, 0 failed.
NULLRUN_SKIP_BUDGET_CHECK=1 outside production was a `logger.debug` and a bare `return`. DEBUG is below every default handler, and nothing was counted, so a test suite running the whole budget path with the bypass on left no trace: tests went green, the dashboard showed the org spending nothing, and the only evidence the gate was never consulted was the absence of a block. CLAUDE.md's rule is that a test passing only with this flag set is evidence of a broken gate, not of a working one. That rule is not actionable unless setting the flag is loud enough to find. Now: a WARNING naming the flag and stating that no budget, rate-limit or tool-block check ran (so a reader can tell a bypass from an allow), plus a `skip_budget_used_non_prod` counter a CI dashboard can alert on. The production refusal already emitted `skip_budget_blocked_in_prod`; this is its dev/test counterpart. Verified by mutation: reverting to the `logger.debug` line fails all three new tests. The fourth asserts the production refusal is untouched by making the dev path louder. Two test-infrastructure corrections this surfaced: `mock_prod_api` moved to conftest. Both guard test files need a runtime that genuinely looks like production, and the first version duplicated the prod auth response in each module -- the same duplication that let the AUTH_ERROR predicate drift between `check_workflow_budget` and `_run_tool_policy_gate` under DEF-MP-TS12-ENF-01. The new tests set the env vars with `os.environ[...] = ...` instead of `monkeypatch.setenv`, which leaked NULLRUN_SKIP_BUDGET_CHECK into every later test in the session. Symptom was confusing: the files passed alone and the full suite failed 9 tests in test_v3_wire_contract.py, because those tests' budget pre-flight was being skipped by a flag set three modules earlier. Now monkeypatch-scoped, and the full suite is green. Suite: 1506 passed, 1 skipped, 0 failed.
….3f)
The backend now answers a failed workflow-state read with 503 and a
tripped breaker with 403. Both are correct on the backend -- the gate
blocks either way -- but the SDK does something different with each,
and the difference is invisible from the gate's own tests.
403 CIRCUIT_BREAKER_TRIPPED -> real gateway decision, honoured,
raises, ONE attempt
503 *_LOOKUP_FAILED -> retried 3x, then a synthetic
FALLBACK decision, which
check_workflow_budget treats as a
transport error and returns
WITHOUT raising. An allow.
So the current SDK behaves identically to 0.18.5 here. That is not a
stale-SDK problem an SDK release fixes -- it is ADR-008's documented
fail-OPEN on transport error, unchanged. Pinned so the asymmetry is a
visible fact rather than something discovered during an incident, and
so a future release that tightens it has to come with an ADR amendment.
The 503 test asserts the CURRENT fail-OPEN on purpose. Asserting the
desired behaviour would ship a red test; asserting "it raises" would be
a lie. The comment says so at the assertion.
Also pins that the "do not retry" instruction survives the round trip.
The backend ships no agent_message for a halt-category refusal and no
SDK version reads that field, so the explanation is the only channel a
model sees -- worth a test, since nothing else would notice its loss.
4 passed; 29 with the neighbouring fallback/gate-path suites.
The two response bodies in this file put `error_code` at the top level.
It is not there. Captured from a live gate run:
top_level_keys=["category", "decision", "decision_source", "details",
"explanation", "policy_id", "policy_version",
"projected_cost_cents", "remaining_budget_cents", "reservation_id",
"staleness_ms", "user_message"]
`error_code` lives at `details.error_code` (`internal.rs:716`), which is
where the backend's own status mapper reads it (`gate.rs:88-90`).
`category` / `user_message` are the top-level fields, set by
`attach_refusal_surface` (`gate.rs:163-172`).
These fixtures were hand-assembled when ADR-064 was written and the key
landed where a reader would expect it. Harmless as a behaviour pin --
the SDK reads `explanation`, so all four assertions are unaffected and
still pass. Not harmless as a reference: ADR-064's SDK-side plan reads
exactly this body, and a client written against the wrong nesting
classifies every refusal as unparseable, which is fail-OPEN.
ADR-064 records the correction. 4 passed.
ADR-062 §2.2. The gate does not send a bare "no" — it sends a `category`, one of denied | budget | halt | infra, and that is what tells the SDK whether the answer is a policy decision, a money wall, a stop, or a backend fault. The SDK had no vocabulary for it, so every refusal became the same thing. This adds `nullrun.breaker.categories` mirroring the backend enum, resolves the category at the wire boundary in `Transport.check`, and puts it on the returned decision. The load-bearing half is what happens when the category is ABSENT or UNRECOGNISED. It raises `NullRunUnclassifiedRefusalError` rather than guessing. Guessing is the dangerous direction specifically: a refusal rounded to something permissive can be rendered as a friendly, model-readable "that's not allowed", and an agent handed that adapts and retries against a stop an operator deliberately placed. That is the bypass ADR-061 closed server-side, re-opened here. Two runtime arms had to change for the raise to mean anything. `check_workflow_budget`'s ADR-008 fail-OPEN arms catch `NullRunError`, and `NullRunUnclassifiedRefusalError` is one, so without an explicit re-raise the unclassifiable refusal was caught by the fail-OPEN arm and returned as `None` — the agent proceeding on a call the backend refused. Verified by mutation: dropping the two re-raise arms turns both branches red with "check_workflow_budget: /gate unavailable, failing open: Gate refused the call (decision="block") but sent no 'category' field". That is DEF-MP-TS12-ENF-01's exact shape with a different trigger. ADR-008 is not narrowed. A real outage still fails open; the counter-test pins it, and the arm order is what makes the two distinguishable — the re-raise is listed before the `NullRunError` and bare `except Exception` arms precisely because both are supersets of it. Also re-files the `DEF-NR-TOOLBLOCKED-PARSER` source-pin comment into `_parse_v3_error_envelope`, where the parser actually is. It was filed against a `check()` branch that does not exist, so the pin pointed readers at the wrong function; the pin now asserts the function it means.
… else A `denied` refusal is the one category where telling the model is both useful and safe: a rule judged *this call* unacceptable, where another call or a different tool might pass. The other three describe a wall the model cannot climb, and a model told about a wall walks into it — it tool-shops after "your budget is exhausted", routes around a pause an operator deliberately placed, and tries to solve a backend fault that is not the model's problem. `on_denied="message"` therefore sits behind a single guard, `category is DecisionCategory.DENIED`, at the one block site in `check_workflow_budget`. No arrangement of the flag turns a budget, halt, or infra refusal into model-readable text — verified by mutation: widening the guard to consult `on_denied` alone turns the three negative cases and the budget-exception case red. The resulting `NullRunDeniedError` (NR-D001, non-retryable) carries `agent_message` — server-authored text the backend guarantees is model-safe (`attach_refusal_surface` sends it only for `Denied`) — behind one accessor, so a host cannot reach for `str(exc)` and put operator internals in the model's context by accident. The default is `"raise"`, and an unrecognised value is rejected at construction rather than treated as `"raise"`, so a typo cannot quietly disable the path a host asked for. `halt` and `infra` refusals still surface as `NullRunBudgetError`. Recorded, not fixed: the exception family does not yet match the category. It is a diagnostics defect, not a bypass — no model text is produced on either path and `on_denied` cannot reach them.
…ls open Product decision, 2026-09-30. The previous version of this file pinned an asymmetry that was not a property of the categories but an accident of the status code: a 403 stopped the agent and a 503 did not, purely because `400 <= status < 500` happened to be the branch that produced a gateway decision. The rule now follows the failure's MEANING rather than its number. Ordinary unavailability stays fail-open; a failure OF THE CHECK ITSELF gets a marker the SDK treats as a block. SDKs predating the category work keep failing open, by definition — they never had a way to read the marker. The marker already existed (`category: "infra"`, and the backend already splits the two 503 groups via `GateErrorCode::is_fail_closed()`), so this was a client-side mapping change, not a re-architecture. Verified by mutation: reverting `Transport.check` to its pre-fix 4xx-only condition turns the 503 cases red with "synthetic decision_source='fallback', treating as transport error" — the gate's own fail-CLOSED refusal answered to the agent as "allowed". That was DEF-MP-TS12-ENF-01's exact shape with a different trigger, and it was wider than a single code path: the whole 5xx band was being converted into a synthetic transport failure. The 502 counter-test is what keeps ADR-008's promise intact. A 5xx with no refusal body is a proxy that never reached the gate, not a verdict, and must still fail open — otherwise every deploy would freeze every agent.
…refusals `on_denied` and the ADR-062 §2.2 category rule were only ever wired into `check_workflow_budget`, which calls `/gate`. Every MCP tool call goes through `runtime.execute` → `/api/v1/execute` instead, so on that path the flag was inert: a host that set `on_denied="message"` got a promise it did not keep for the whole class of tool calls that MCP adapters mediate, and the server-authored `agent_message` never reached it. Two independent pieces were missing on `/execute`. `_extract_error_envelope` had no arm for a gate refusal envelope. A real refusal puts `error_code` at `details.error_code` (backend `internal.rs:716`), not at the top level, so a `/execute` 403 TOOL_BLOCKED fell through to the legacy slug arm and arrived at the host as `NullRunAuthenticationError: Auth failed on execute (status 403, error_code='')`. Wrong class, wrong `format_user_message`, and an operator action — "check your credentials" — that cannot fix a policy decision. Added the missing envelope shape. Exceptions built by the parser carried neither the category nor the server's text. Split `_parse_v3_error_envelope` into a thin wrapper over `_parse_v3_error_envelope_uncategorised` so all ~15 return points stamp `wire_category` (left ABSENT when the body has no `category` — never defaulted, or a missing field would become a confident answer), `wire_error_code`, and the server-authored `agent_message` / `user_message`. Then `runtime.execute` converts a `NullRunBlockedException` to `NullRunDeniedError` under the same single-guard discipline as the `/gate` block site: `on_denied` is consulted only for `DecisionCategory.DENIED`, read from the wire. No arrangement of the flag turns a budget, halt, or infra refusal into model-readable text, and an absent category still raises. Test notes, since two of them are not what they first look like. `infra` is absent from the negative parameter set on purpose. ADR-063 §4.7 sends infra refusals as 503, and `_retry_with_backoff` maps the whole 5xx band on `/execute` to NullRunTransportError/GATEWAY_ERROR before the envelope parser is reached; the runtime then raises a STRICT fallback block at `runtime.py:3670`, downstream of the `on_denied` arm. The property is real and is asserted, but it is enforced by the 5xx band and not by the category check — so the test pins the concrete class to say which mechanism is in play. The negatives originally used a synthetic `SOME_CODE`, which the parser maps to a non-block class, so the exceptions never entered the arm and all three parameters passed regardless of the guard. Widening the guard to consult `on_denied` alone left them green. They now use real codes verified to map into `NullRunBlockedException`. All three mutations are recorded in the test docstrings. The `DEF-NR-TOOLBLOCKED-PARSER` source pins are repointed in the same commit because they are not separable from the parser split — the split alone leaves the suite red, and splitting them across commits would violate the one-green-commit-per-step rule. They resolved the parser by name, which after the split points at the wrapper and asserts nothing. They now resolve the dispatching function by content (the branch text, not the exception name: the name also appears in the import block, so resolving on it left `test_branch_in_parser` green after the branch was deleted — verified). Deleting the branch now turns 13 of 14 red. 1543 passed, 1 skipped.
The README carried no fail-open claim and no limitations section, so the two most consequential behaviours of the gate were only discoverable by reading the source. Adds both. Each limitation was checked against the code before being written down, and two of them are not what the shorthand description of them suggests. The 503 case fails OPEN on old SDKs, not closed. A failed security check is served as a 503 carrying a `category`; the current SDK reads it and refuses. Pre-category there was nothing to read, the 503 became a synthetic FALLBACK decision, and `check_workflow_budget` fails open on a FALLBACK source — so the call proceeds. Verified by probing both cases through a real runtime and by reading the pre-category source (`git show 83960b4~1:src/nullrun/transport.py`): the 5xx band had no `is_refusal` condition and fell straight to the FALLBACK return. The breaker setting is server-side, not an SDK default. `LogOnly` is `NULLRUN_GATE_CB_TRIP_ENFORCEMENT_MODE` in the backend, not a fallback mode in this SDK (there is no `LogOnly` anywhere in `src/`). The production boot check refuses to start unless it is explicitly `Enforce` or `LogOnly` (backend `main.rs:290-293`); `detect_mode()` still defaults to `LogOnly` on unset outside production (`main.rs:294-295`). In `LogOnly` the trip is recorded and alerted but tripped workflows still pass `/check` — the backend's own words at `main.rs:286-288`. Pause and kill really are 403-only, and that is now stated as the limitation it is: `WORKFLOW_PAUSED` and `WORKFLOW_INACTIVE` are served as the same 403 from the same key, distinguishable only by operator-facing text, which the backend deliberately keeps distinct (`error_codes.rs:3400-3412`). Support tooling should key off the error code, not the status. Also adds the fail-open summary, matching the ADR-008 table in `runtime.py` — which the module docstring requires to be kept in lockstep with any README claim. There was no README claim to keep in step before this; there is one now.
…the wire The transport comment and two test docstrings repeated ADR-063 §4.7's claim that the backend "already distinguishes the two 503 groups internally via GateErrorCode::is_fail_closed()". That is wrong in the way ADR-063 §4.7 was wrong: is_fail_closed is an ordinary Rust method that is never serialised, so no client could have read it. The SDK never did — the code has always keyed on decision == "block" (categories.py:168), which is ADR-064's discriminator. The comment also claimed BUDGET_DATA_UNAVAILABLE "arrives as 503 with decision=block". It does not, and cannot: grepping the backend, that string appears only in error_codes.rs and budget.rs, never in internal.rs or orchestrator.rs, so no gate producer emits it. The one response that does carry it (budget.rs:246, BudgetUnavailableResponse) has no decision field at all — which is correctly treated as a non-refusal and fails open per ADR-008. Logic unchanged; this is a claim correction, and every ADR-063 §4.7 citation for the 503 rule retargeted to ADR-064, which now owns it. 44 passed.
The refusal classifier splits a 5xx into "the gate ruled" and
"nothing ruled" by reading the body. Every other category change on
this branch rests on the same unexamined premise: that a body on the
SDK's HTTPS connection was written by NULLRUN.
A test for that premise found the premise is false in one direction.
`{"decision": "allow"}` from an arbitrary responder passes every check
`_require_gate_decision` applied -- it is a dict, and `decision` is a
known string -- and was then honoured as a gateway decision, because
the runtime's rule reads a MISSING decision_source as "not fallback"
and therefore as authoritative. So an on-path responder that gets
JSON in front of the SDK: a captive portal on a hotel WLAN, a
corporate TLS-interception proxy. It needs neither the API key nor a
HMAC bypass, only to answer before the real backend. The call is
authorised, /track books its cost against a policy nobody consulted,
the audit trail records an allow, and there is no later point at
which it is caught.
Fix: require `decision_source` to be a known provenance value BEFORE
looking at `decision`. It is the right field because it cannot be
absent from a real answer -- `GateResponse.decision_source` is a
non-Option String with no skip_serializing_if (`gate/internal.rs:638`).
`GateResponseBody` in schemas.rs declares it optional, but that struct
is referenced only from openapi.rs: it is the documentation schema,
not the wire.
Mutation-verified: with the guard absent, the new test reports DID
NOT RAISE on `{"decision": "allow"}`. 18 passed with it.
The same guard caught a drifted fixture, not a second defect:
test_v3_wire_contract.py's require_approval body omitted all three
non-Option fields on GateResponse. Corrected to the real shape.
Full suite 1561 passed / 1 skipped.
…nforcement Limitation 4. The three existing entries describe what the SDK does not enforce; this one describes what it cannot verify. Every enforcement claim in the README rests on a premise it never said out loud: that the JSON on the HTTPS connection was written by NullRun. The SDK now checks for it -- a /gate body with no usable decision_source is rejected rather than acted on -- but that is a field in the body, not a signature. An on-path responder that can produce a well-formed body including decision_source defeats it, and needs neither the API key nor an HMAC bypass. NullRun's answers are not signed, so there is no after-the-fact detection either: the audit trail would record the allow. Also extends the fail-open summary with the not-a-verdict case, which is newly user-visible as NullRunMalformedGateResponseError. This is the honest boundary of "server-authoritative": authoritative against a client, not against an attacker who owns the path.
ruff runs over src/ only in CI, so none of these would have failed a build -- but all three files are new on this branch (neither exists on origin/master), so they are this branch's to leave clean. Two F841: the `rt = make_runtime(...)` binding was never used, because @Protect resolves the runtime from the context, not the local. The call is load-bearing; the assignment was not. One I001: nullrun.runtime sorted after nullrun.toolbox.
…on skew Limitation 4 said "use https://", which understates the problem. It implied https was sufficient, when the residual is specifically an on-path responder holding a certificate the OS already trusts for api.nullrun.io -- i.e. corporate TLS interception, which is ordinary on a managed network and which https alone does nothing about. Corrected to state what each layer does and does not give: - certificate verification is on and CANNOT be disabled by configuration. `verify_cert = True` (transport.py:572) and the only override is NULLRUN_TLS_CA_CERT (:568), which swaps in a trust anchor you chose and is still verification. Grepped exhaustively: the only TLS env vars in the SDK are NULLRUN_TLS_CA_CERT / _CLIENT_CERT / _CLIENT_KEY, none of which turns verification off. - plain http is refused (InsecureTransportError, :526). - responses are not signed, so no after-the-fact detection: the audit trail would faithfully record an allow the gate never gave. The consequence is stated as the operational control it is -- exclude api.nullrun.io from interception and verify it stays excluded -- and not as something the SDK does for you. Also adds the old-SDK note beside limitation 1, since it is the same version skew pointing the same way and a reader who pins for the 503 guarantee should know they also need it for the forged-allow case: both a real refusal and a fabricated permission are read permissively on an older SDK, and neither shows up in the SDK's output.
Every other file in this cluster calls runtime.check_workflow_budget() directly. That is the wrong level for the question a user has, because @Protect is what they call -- and @Protect is a separate code path with its own try, its own except BaseException, and its own two wrappers (sync and async). Reading decorators.py says the refusals propagate. That is exactly the reasoning that let DEF-MP-TS12-ENF-01 ship: check_workflow_budget's except Exception looked harmless in isolation too, until it was reached by a real 401. 14 tests, four properties: - a denied / budget / unclassifiable refusal, and a body with no decision_source, do NOT run the decorated body; - the counter-tests: a real allow, a connection failure under on_denied="message", and a 502 with no refusal body all DO run it (a suite of only "did not run" assertions passes against a decorator that refuses everything); - on_denied reaches @Protect for category=denied and, proven by the negative parametrisation, does not reach budget/halt; - the async wrapper is not a separate property, so the refusal, the forged allow and the real allow are each proven on it too. Mutation-verified: deleting the decision_source guard in _require_gate_decision turns exactly the sync and async forged-allow tests red (DID NOT RAISE), and the other 12 stay green. Verified in a clean python:3.11 container on the unmutated tree: ruff check src/ tests/ -- All checks passed mypy src/ -- no issues in 37 source files pytest -q -- 1575 passed, 1 skipped
0.19.0 is already on origin/master and its entry closed the paths found by reading the code. These commits closed the paths that only appeared once the properties were asserted end-to-end, and none of them were in the release notes -- so the two most serious items (a forged allow authorised by decision_source absence, and the production-usable sensitive fail-open opt-out) were documented only in README. Filed as [Unreleased] rather than inventing a version number. That is deliberate: the behaviour change is real and needs a migration note, not a patch number chosen by accident. An unclassifiable refusal now raises where 0.19.0 let the call proceed, and a hand-rolled /gate test double must now carry decision_source. Whoever cuts the release picks 0.19.1 or 0.20.0 with that in hand. Verified citation: GateResponse.decision_source is a non-Option String at backend/src/proxy/http/gate/internal.rs:638.
Six differences from 0.19.0, written as things you can act on rather
than as a list of changed internals. The third is the one worth
reading twice: an unclassifiable refusal now raises where 0.19.0 let
the call proceed, and NullRunUnclassifiedRefusalError is deliberately
a SIBLING of NullRunTransportError, not a subclass -- so an existing
`except NullRunTransportError:` fail-open arm will not catch it.
That is the fix working, and it is also the thing most likely to turn
a working loop into a throwing one.
The note gives the exact import path (nullrun.breaker.categories --
it is not re-exported at the top level, which was verified, not
assumed), the error_code, and the three options as three, with the
consequence of each stated. Widening the arm to NullRunError is
called out as worse than the hole being closed, because it also
swallows real policy refusals.
Every code shape in the note was executed against the installed
package rather than reasoned about:
recommended two-arm -> propagates (recommended)
one-arm permissive -> swallowed == 0.19.0 behaviour, bypass restored
widened to NullRunError -> a real policy refusal is swallowed too
issubclass(NRError, NullRunError) = True
Priority 1 for DEF-MP-TS12-ENF-01 was "one @Protect, rest under the hood", and the part nobody had tested is whether that survives the framework layer. It did not, in one direction. `@protect` above `@tool` collapsed the tool to a plain function. `functools.wraps`-based wrapping cannot preserve `.invoke`, `.name`, `.args_schema` or `.description`, so the agent loop could not bind it -- and a tool the loop cannot see produces no refusal. Measured on langchain-core 0.3.86: @Protect # outer @tool # inner -> StructuredTool def f(query: str) -> str: ... # f is now <function>, not StructuredTool # convert_to_openai_tool(f) -> NameError: name 'Annotated' is not defined That last line is the failure mode worth naming: the error a user sees is a langchain type-resolution error, not "your tool is not a tool any more". Enforcement was gone and the symptom pointed somewhere else. `protect` now duck-types the argument (`.invoke` + `.name`, then `.coroutine` before `.func` -- an async tool sets both, with `func` as a sync fallback that raises if called), wraps whichever callable is live, and returns the SAME object. Every other attribute is untouched, so the loop sees the tool it saw before. Duck-typed rather than isinstance because `langchain_core` is an optional dependency; the three attributes matched are the ones `BaseTool.run()`/`arun()` dispatch through. Three properties were established by probe on the real library before being written down, and all three are now pinned: - Ordering is irrelevant in BOTH directions. `@tool` outside already worked; the reverse now does too, and each has an allow counter-test so "the body did not run" cannot pass against a wrapper that refuses everything. - `handle_tool_error=True` cannot turn a refusal into model-visible text. It only catches `ToolException`; a `NullRunBudgetError` reaches the bare `except (Exception, KeyboardInterrupt)` arm and is re-raised. This is the DEF-MP-TS12-ENF-01 shape in a different costume, so it is pinned rather than assumed. - `@protect` with no `init()` fails LOUD. The lazy resolver propagates `NullRunAuthenticationError` (NR-A001) and the body does not run. FIX-4 removed the fallback that used to swallow this; the AST guard keeps it removed, and an AST check rather than a substring search because the docstring describes the very `except` being asserted gone. `on_denied="message"` is now exercised through a real `@tool` too, in both orders and both sync/async, with the negative set (budget, halt) kept at the framework layer: an agent told "that tool is not allowed" when the truth is "you are out of money" will look for another way to spend. Mutation-verified; each mutant ran the suite and the restore was byte-identical: sync arm converts refusal to text -> 2 red async arm converts refusal to text -> 1 red FIX-4 fallback in lazy resolver -> 2 red BaseTool in-place wrap removed -> 5 red on_denied category guard dropped -> 8 red across 3 files Three earlier mutation attempts were invalid and are not counted: an inner `try` cannot observe a `_protect_body.__enter__` raise, and one anchor matched `get_protected_runtime` instead of `_get_or_create_runtime`. Suite: 1597 passed, 1 skipped. ruff and mypy clean.
Minor, not patch: an unclassifiable refusal now RAISES where 0.19.0 let the call proceed. That turns a working loop into a throwing one, which is a behavioural break, so the release note leads with the migration rather than burying it under the security items. `### Migration` is the first section under 0.20.0 and enumerates all six behavioural differences from 0.19.0, then repeats the three inherited from 0.19.0 that are the same class of break and the same shape of fix (`mode="inline"` removed, `decision_source` now required in a `/gate` test double, `handle` -> `guard`). A reader upgrading from 0.18.x sees the whole break list in one place instead of reconstructing it across two releases. The load-bearing item is #3: `NullRunUnclassifiedRefusalError` is a sibling of `NullRunTransportError`, not a subclass, so an existing `except NullRunTransportError:` arm will not catch it. Catching it alongside restores 0.19.0's behaviour exactly -- which restores the bypass. The note says so and says not to widen to `NullRunError`, which is a larger hole than the one this closes. Also records the `@tool`/`@protect` fix under `### Fixed`, since it is a silent enforcement loss users may have shipped against: enforcement was absent whenever `@protect` was the outer decorator on a LangChain tool. README gains the two things a user would otherwise get wrong: - decorator order, both directions, with the reason (an unbound tool cannot refuse) and the `NameError` that was the real symptom; - `on_denied="message"` as the operator-facing "explain, don't crash" mode -- `category="denied"` only, budget and halt deliberately excluded -- and why LangChain's `handle_tool_error=True` is not an equivalent (it catches `ToolException` and does not know which exceptions are refusals). Every claim in the new README section was executed before being written down, including the pre-fix `NameError`, which is reproduced rather than asserted from memory. Not published to PyPI. Suite: 1597 passed, 1 skipped; ruff and mypy clean.
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Cuts
v0.20.0— the second half ofDEF-MP-TS12-ENF-01. Where 0.19.0closed the paths found by reading the code, 0.20 closes the ones that
only showed up when the properties were asserted end-to-end: a forged
allow authorised by
decision_sourceabsence, a production-usablesensitive-tool fail-open opt-out, and an
@protect-outside-@toolcollapse that made enforcement silently disappear under LangChain's
type-resolution errors.
This is a minor, not a patch: an unclassifiable refusal now
raises where 0.19.0 let the call proceed. That is a working loop
becoming a throwing one — a behavioural break, not a bug fix. The
release notes lead with the migration rather than burying it under the
security items; the load-bearing entry is #3.
Migration (read first if upgrading from 0.19.x)
Six things differ from 0.19.0. Only the first three can surprise you
at runtime, and the third is the one worth reading twice.
A
/gatetest double must carrydecision_source. Realbackend answers always do — it is a non-
OptionStringonGateResponse. A hand-written fixture that omits it now raisesNullRunMalformedGateResponseError. Fix: add"decision_source": "gateway"next to"decision": "allow".NULLRUN_SENSITIVE_FAIL_OPENagainst production now needs asecond variable.
NULLRUN_ALLOW_SENSITIVE_FAIL_OPEN=1must beset too, otherwise the opt-out is refused, enforcement falls back
to its own fail-CLOSED default, and the attempt is logged at ERROR
with a metric. Non-production behaviour is unchanged.
An unclassifiable refusal now raises where 0.19.0 let the call
proceed.
NullRunUnclassifiedRefusalErroris importable fromnullrun.breaker.categories(it is not re-exported at the toplevel) and carries
error_code="NR-P003"andretryable=True.It is a
NullRunInfrastructureErrorand a sibling ofNullRunTransportError— deliberately not a subclass of it.So an existing
except NullRunTransportError:arm, which in0.19.0 caught everything the gate could not classify and failed
open, will not catch this one. That is the point: it is what
stops a failed-open arm from swallowing a refusal. If you have
such an arm, you have three options and they are not equivalent:
Catching it alongside
NullRunTransportErrorrestores 0.19.0'sbehaviour exactly — which means it restores the bypass. Do not
widen the arm to
NullRunError: that also swallows real policyrefusals, which is a larger hole than the one this closes.
on_deniedis new. Default"raise", which is 0.19.0'sbehaviour. Set
"message"to get aNullRunDeniedErrorcarryingthe server-authored
agent_messageforcategory="denied"only.NULLRUN_SKIP_BUDGET_CHECKoutside production now logs andincrements a metric when it skips the check. Enforcement
behaviour is unchanged.
/executerefusal bodies carrycategoryandagent_message.Additive on the wire.
Carried over from 0.19.0 and still true on 0.20.0, because it is the
same class of break and the same shape of fix:
NullRunRuntime.execute(..., mode="inline")is gone and has noreplacement — every call goes through
/execute.nullrun.runtime.register_strict_mode_forced/is_strict_mode_forcedare gone, along with@guarded,nullrun.handle(renamednullrun.guardin 0.18.5),nullrun.status(), andnullrun.auto_instrument.Security
c459379) — A/gatebody must statewho decided before it counts as a verdict.
_require_gate_decisionrejects a body whosedecision_sourceisabsent or unrecognised, raising
NullRunMalformedGateResponseError. The runtime's rule wasdecision_source != fallback → honour the wire decision, whichread a missing
decision_sourceas more trustworthy thanfallback. So{"decision": "allow"}— a captive portal's loginJSON, an intercepting proxy's stub, anything on-path — was enough
to authorise a call no policy engine evaluated. The field is a
non-
OptionStringon the backend'sGateResponse(
gate/internal.rs:638) and every producer sets"gateway", soa real answer always carries one; there is no legitimate body
this rejects.
cdf94c3) —NULLRUN_SENSITIVE_FAIL_OPENis refused against production. It was read straight into the
enforcement path, letting a sensitive tool's body run with no
policy evaluation at all. Its sibling
NULLRUN_SKIP_BUDGET_CHECKhas been production-guarded since it was caught doing the same;
the asymmetry was an oversight, and this half is the more
dangerous one — that one skips a pre-flight, this one skips the
gate. Requires both
NULLRUN_SENSITIVE_FAIL_OPEN=1andNULLRUN_ALLOW_SENSITIVE_FAIL_OPEN=1; the documented dev/testuse is unchanged.
700fea7) — The non-prod budget bypassis now visible. When
NULLRUN_SKIP_BUDGET_CHECK=1skips the check,it logs and increments a metric instead of being
indistinguishable from a normal call.
83960b4+4cb009c+d191a7f) —Every gate refusal is classified.
on_denied="message"is theoperator-facing "explain, don't crash" path for
category="denied"only (budget and halt deliberately excluded — an agent told "that
tool is not allowed" when the truth is "you are out of money" will
look for another way to spend).
/executerefusal bodies nowcarry
categoryandagent_message; tests pin §4.7 option 2(a 503 refusal blocks, an outage fails open).
Fixed
b38fcf5) —@protectno longerloses a LangChain tool when it is the outer decorator. Applied to
a
@toolresult,@protectwrapped the object withfunctools.wrapsand returned a plain function — so.invoke,.name,.args_schemaand.descriptionwere gone, an agentloop could not bind the tool, and a tool the loop cannot see
produces no refusal.
protectnow duck-types the argument(
.invoke+.name, then.coroutinebefore.func) andreturns the same object. Duck-typed rather than
isinstancebecause
langchain_coreis an optional dependency. This is asilent enforcement loss users may have shipped against:
enforcement was absent whenever
@protectwas the outerdecorator on a LangChain tool, and the user-visible error
pointed at LangChain's type-resolution layer rather than at the
real cause.
Documentation
939fd7d+660954a+c2e52d5— README states the threeenforcement limitations precisely (TLS-interception residual,
server-side trust, old-SDK version skew), and the provenanced
boundary of server-authoritative enforcement.
642ebef+467cb1a— CHANGELOG migration note for therefusal-category SDK; records the second half of
DEF-MP-TS12-ENF-01(the half that only showed up when theproperties were asserted end-to-end).
Verification
ruff check src testsmypy src/nullrunpytest -qnullrun.__version__0.20.0/executerefusal bodies (category,agent_message); non-additive on/gate(decision_sourceis now REQUIRED where it was previously optional on the SDK side — but a+itwas always present on a real backend answer)Commits included