Ship AI agents with real-time budget, policy, and human-approval gates.
Zero-refactor cost control, tool policy enforcement, and audit trail for any
LLM-powered agent — works with any LLM SDK that uses httpx, plus your own stack.
Quickstart · Docs · Examples
⚠️ Status: alpha (v0.17.1). The public API may shift between minor versions. Pin your dependency and read the CHANGELOG before upgrading.
AI agents can overspend, call dangerous tools, and act without audit trails. Existing observability tools tell you after the fact. NullRun enforces before the action.
| Without NullRun | With NullRun |
|---|---|
Agent calls gpt-4o 10,000 times → surprise $5,000 invoice |
Hard budget cap → SDK blocks at 402 before invocation |
Agent runs bash rm -rf / |
Tool policy → SDK blocks at 403 before execution |
| Sensitive action with no human in the loop | Approval flow → SDK pauses and waits for WS approval_resolved push |
| Cost & calls scattered across 4 libraries | Single source of truth: per-org, per-workflow, per-execution |
Runaway SDK loop calling /gate without /track |
Per-reservation rate cap → 402 budget error (see docs/errors/NR-R001.md) |
| Hard & soft budget gates — atomic Redis-enforced | Tool policy enforcement — block dangerous tools before execution |
Human-in-the-loop approvals — pause agent and await approval_resolved via WS push |
Immutable audit trail — every decision, every tool call, every cent |
Zero-code instrumentation — nullrun.init() patches httpx once for any vendor |
No vendor lock-in — works with any LLM SDK that uses httpx |
| Memory-safe streaming — 16 MiB response body; full body for usage extraction | Lightweight — no LLM-key storage, no proxy required |
| Server-authoritative cost — server-minted execution IDs | MCP support — expose tools to agents via Model Context Protocol |
%%{init: {
'flowchart': {
'curve': 'basis',
'htmlLabels': true,
'nodeSpacing': 80,
'rankSpacing': 90
}
}}%%
flowchart LR
%% =========================
%% AI RUNTIME
%% =========================
subgraph USER ["👤 AI Runtime"]
direction TB
A["🤖 Agent"]
end
%% =========================
%% NULLRUN LAYER
%% =========================
subgraph LIB ["📦 NullRun Enforcement Layer"]
direction TB
B["NullRun SDK<br/>Interceptor"]
C["🚦 Runtime Gate"]
P["📜 Policy Engine"]
H["👤 Human Approval"]
end
%% =========================
%% PRODUCTION
%% =========================
subgraph PROD ["⚙️ Production Actions"]
direction TB
T["🛠 Tools"]
API["🌐 External APIs"]
DB["🗄 Databases"]
end
STATE["🗂 Audit + Runtime State"]
%% =========================
%% FLOW
%% =========================
A -->|"protected action"| B
B -->|"authorize"| C
C --> P
P -->|"allow"| T
P -->|"allow"| API
P -->|"allow"| DB
C -->|"require approval"| H
H -->|"approved"| T
C --> STATE
%% =========================
%% COLORS
%% =========================
classDef user fill:#dbeafe,stroke:#2563eb,color:#0f172a
classDef sdk fill:#dcfce7,stroke:#16a34a,color:#0f172a
classDef srv fill:#fed7aa,stroke:#ea580c,color:#0f172a
classDef store fill:#f5d0fe,stroke:#a21caf,color:#0f172a
classDef ok fill:#bbf7d0,stroke:#16a34a,color:#0f172a
classDef wait fill:#fef08a,stroke:#ca8a04,color:#0f172a
class A user
class B sdk
class C,P,H srv
class STATE store
class T,API,DB ok
class H wait
style USER fill:#f8fafc,stroke:#64748b,stroke-width:1px
style LIB fill:#f8fafc,stroke:#64748b,stroke-width:1px
style PROD fill:#f8fafc,stroke:#64748b,stroke-width:1px
The gate is server-authoritative — the SDK never trusts client-supplied cost. Redis is the source of truth for budget and tool-policy state; Postgres holds the immutable audit log.
sequenceDiagram
participant Agent
participant SDK
participant Gate
participant Policy
participant Human
participant Tool
Agent->>SDK: execute(tool)
SDK->>Gate: authorize(action)
Gate->>Policy: evaluate rules
alt Allowed
Policy-->>Gate: allow
Gate-->>SDK: continue
SDK->>Tool: execute
else Approval required
Policy-->>Gate: approval_required
Gate-->>SDK: wait
Gate->>Human: request approval
Human-->>Gate: approved
Gate-->>SDK: resume
SDK->>Tool: execute
else Blocked
Policy-->>Gate: deny
Gate-->>SDK: exception
end
Install:
pip install nullrun
export NULLRUN_API_KEY="nr_..." # get one at https://nullrun.io/control-center/api-keysfrom nullrun import protect
@protect
def my_agent(prompt: str) -> str:
return call_llm(prompt)If you call @protect before init(), the SDK lazy-initializes
the runtime from NULLRUN_API_KEY on the first decorated call. You can
write your agent code with the decorator first and the init second — or
skip init entirely if your environment is already configured.
For CLI scripts that want fail-fast on missing config, pass
fail_on_exit=True — the SDK prints a four-line developer report and
exits with code 1 instead of raising. nullrun.shutdown() is
auto-registered via atexit inside init(), so a clean WS close on
process exit happens without any explicit call.
Both orders gate. @protect recognises a LangChain tool, wraps the
tool's func/coroutine in place, and returns the same object, so this
is not a rule you have to remember:
from langchain_core.tools import tool
from nullrun import protect
@tool # fine
@protect
def charge(amount: int) -> str: ...
@protect # also fine
@tool
def charge(amount: int) -> str: ...Before 0.20.0 the second form silently produced a plain function. The
agent loop could not bind it, and a tool the loop cannot bind cannot
refuse — so the gate was not running. If you saw
NameError: name 'Annotated' is not defined from
convert_to_openai_tool, that was this.
By default a refusal raises, which is right for most code: the caller
decides what happens next. on_denied="message" is the operator-facing
alternative for a policy denial — the agent gets the
server-authored explanation and the run continues:
rt = nullrun.init(on_denied="message")The text is authored by the backend, never assembled by the SDK, and it
applies to category="denied" only. Budget and halt refusals keep their
own exceptions under the same flag: an agent told "that tool is not
allowed" when the truth is "you are out of money" will go looking for
another way to spend.
LangChain's own handle_tool_error=True is not an equivalent. It
catches ToolException and stringifies it, and it does not know which
exceptions are refusals — use on_denied="message".
| NullRun | LangChain callbacks | Helicone | Portkey | OpenLLMetry | |
|---|---|---|---|---|---|
| Enforce before execution | ✅ | ❌ | ❌ | ||
| Server-authoritative budget | ✅ | ❌ | ❌ | ❌ | ❌ |
| Tool-call policy | ✅ | ❌ | ❌ | ❌ | |
| Human-in-the-loop approvals | ✅ | ❌ | ❌ | ❌ | ❌ |
| Zero-code instrumentation | ✅ | ✅ | ✅ | ✅ | ✅ |
| Immutable audit trail | ✅ | ✅ | ✅ | ✅ | |
| Streaming memory cap (anti-OOM) | ✅ | ❌ | ❌ | ||
| MCP support | ✅ | ❌ | ❌ |
NullRun is the only option that blocks expensive or dangerous calls before they happen, not just observes them.
Every gate decision, approval resolution, and execution lifecycle event
is written to the org's hash-chained audit_events table on the backend.
The SDK surfaces a typed read API at runtime.audit.* so backends on
ADR-009 (schema_version = 3) return typed dataclasses — not raw dicts.
from nullrun import NullRunRuntime, AuditQuery
from datetime import datetime, timezone, timedelta
runtime = NullRunRuntime(api_key="nr_...")
# 1) Last 50 governance decisions in the last 24h.
since = (datetime.now(timezone.utc) - timedelta(hours=24)).isoformat()
page = runtime.audit.list(
AuditQuery(event_type="authorization_decision", since=since, limit=50)
)
for entry in page.entries:
print(entry.timestamp, entry.decision, entry.tool_name, entry.reason_code)Available surfaces:
| Method | Returns | Endpoint |
|---|---|---|
runtime.audit.list(query=...) |
AuditLogPage (entries + meta) |
GET /api/v1/orgs/{org}/audit-log |
runtime.audit.verify(since=...) |
AuditVerifyResult (chain head/tail/reason) |
GET /api/v1/orgs/{org}/audit-log/verify |
runtime.audit.list_exports() |
list[AuditExportJob] |
GET /api/v1/orgs/{org}/audit-log/export |
runtime.audit.create_export() |
dict (job_id, status) |
POST /api/v1/orgs/{org}/audit-log/export |
runtime.audit.export_status(job_id) |
AuditExportStatus |
GET /api/v1/orgs/{org}/audit-log/export/{job_id}/status |
AuditQuery filters on the canonical ADR-009 columns: event_type
(authorization_decision / approval_decision / execution_lifecycle),
decision, policy_id, execution_id, actor, since, until, limit.
Pre-ADR-009 backends return legacy fields only — AuditEntry.is_governance
is False for those rows, and the 13 governance columns default to None.
If you call runtime.audit.* before nullrun.init() (no org binding),
the proxy raises NullRunAuthenticationError — not a silent 404 — so a
misconfigured CI step fails loudly at the audit call site rather than
silently dropping the query.
An approval row that lands at status='APPROVED' but never flips to
CONSUMED is an "orphan grant" — the operator sees it on the dashboard
forever (or until the sweeper runs). Two paths close the orphan:
- Success path — when the WebSocket approval push resolves
outcome=approved, the SDK auto-callsPOST /api/v1/approvals/{approval_id}/consumeso the row flips toCONSUMEDbefore the function body runs. Best-effort: a network blip is logged atDEBUGand the success path is not blocked. - Exception path —
@protect's_safe_cancel_active_executioncallscancel_executionandconsume_approval(in that order) when an exception fires after/gatesucceeded. The reverse-index lookupexecution_id → approval_idis populated by the WS push handler, so if the SDK never reached the WS-approval branch the lookup returnsNoneandconsume_approvalis a no-op.
The new endpoint is structurally distinct from the orchestrator's
consume_approved SQL (no execution_id binding per ADR-046, so it
does not participate in the cached-replay arm race window) and carries
organization_id for C2 closure. Idempotent: replay returns
already_consumed; PENDING/DENIED/EXPIRED rows return not_approved,
both with HTTP 200. See src/nullrun/runtime.py::consume_approval
and src/nullrun/transport.py::consume_approval.
Runnable, copy-pastable examples live in a separate repo so you can adapt without cloning the SDK source:
- Custom tools — register your own tools for policy
- Multi-agent — shared budget across sub-agents
| Version | Status | Highlights |
|---|---|---|
| v0.14.x | ✅ alpha | Wire protocol v3.31, server-minted execution IDs, MCP, anti-OOM streaming cap |
| v0.15.x | ✅ alpha | ADR-009 governance audit surface, typed runtime.audit.*, capability probes for /audit-log/verify, fail-OPEN observability closure |
| v0.16.x | ✅ alpha | Phase-1+ action_digest on /gate, /execute tools propagation, transient-5xx retry on gate (NR-006), error-code parity (NR-007, 41→56 entries) |
| v0.17.x | ✅ alpha | Chain-setter Token discipline, _GATE_CACHE staleness closure, lazy-export repair, circuit-breaker lock unification (sync+async), op_id mint-fresh (DEF-OPID-REUSE-HASH-MISMATCH), error-code map closure (DEF-SDKT-004) |
| v0.18.x (current) | ✅ alpha | Close-orphan fix (ADR-047): SDK auto-calls POST /api/v1/approvals/{id}/consume after WS approval resolves to outcome=approved and on the @protect exception path. Closes the structural orphan where mode="inline" tools left approval rows at status=APPROVED past expires_at. |
| v1.0 | 🎯 beta target | Stable wire contract, full async support, type-safe decisions |
git clone https://github.com/nullrunio/nullrun-sdk-python
cd nullrun-sdk-python
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pytest -qWe follow Conventional Commits,
require tests for new public API, and run ruff + mypy in CI.
Four things this SDK does not do. All four are enforcement-relevant, so they are stated here rather than left to be discovered during an incident. Each was verified against the code before being written down.
1. A failed security check arrives as a 503, and only this version of the SDK stops on it. When the backend cannot evaluate the security check itself, it refuses with a 503 carrying a category field. This SDK reads that field and fails closed — the call is refused. An SDK older than the category work has nothing to read: the 503 is turned into a synthetic FALLBACK decision, and check_workflow_budget fails open on a FALLBACK source — the call proceeds. So during a partial backend outage, enforcement differs by SDK version. A genuine outage (not a failed check) still fails open on every version, which is deliberate: a dead backend must not freeze your agent loop.
If you need this guarantee today, pin the SDK version. Do not assume a refused call implies the backend rejected the call.
Older SDKs fail open on a forged allow, too — and this is the second version-skewed behaviour, so pin for it too. The decision_source check described in limitation 4 arrived with the same release. An SDK older than it accepts {"decision": "allow"} from any responder, because it reads a missing decision_source as "not fallback" and therefore as authoritative. The two skews point the same way and are worth knowing together: on an older SDK, both a real backend refusal and a fabricated permission are read the permissive way, and neither is visible in the SDK's own output — the call simply proceeds. Pin the version if either matters to you.
2. The gate circuit-breaker's trip mode is a server-side setting, and LogOnly does not block. When the gate's circuit breaker trips, what happens is decided by NULLRUN_GATE_CB_TRIP_ENFORCEMENT_MODE on the server, not by anything in this SDK. In LogOnly the trip is recorded and alerted on, but tripped workflows still pass /check. The production boot check refuses to start unless the variable is explicitly set to Enforce or LogOnly, so a deploy cannot inherit the dev default (detect_mode() still falls back to LogOnly when unset outside production) — but an operator who chooses LogOnly is choosing non-enforcement, knowingly. If your compliance story depends on breaker trips being enforced, confirm that value with whoever operates the deployment.
3. Pause and kill both reach the agent as a 403. There is no separate status to branch on. WORKFLOW_PAUSED and WORKFLOW_INACTIVE are served as the same 403 from the same key; the only thing distinguishing them is the operator-facing text, which the backend deliberately keeps distinct because they mean opposite things about whether the run will resume. If you write support tooling, key off the error code, not the status. Separately, the SDK can observe pause/kill ahead of the next gate call via check_control_plane (WebSocket push, or a /status poll), which raises WorkflowPausedException / NullRunWorkflowKilledError locally.
4. The SDK trusts the channel, and says so rather than pretending otherwise. Everything above rests on one premise: that the JSON arriving on the SDK's HTTPS connection was written by NullRun. The SDK checks for it — a /gate body with no usable decision_source is rejected as NullRunMalformedGateResponseError rather than acted on — but that is a field in the body, not a signature, and a field is only as trustworthy as the channel carrying it.
What the channel does give you, verified in transport.py: certificate verification is on and cannot be switched off by configuration — verify_cert is True (:572) and there is no env var that sets it to False; the only override, NULLRUN_TLS_CA_CERT (:568), replaces the trust anchor with one you chose explicitly and is still verification. Plain http:// is refused outright (InsecureTransportError, :526). So a passive network observer cannot pose as NullRun, and the ordinary captive portal — which returns a login page, or JSON without a decision — is rejected.
What remains is specific, and it is not "use https". It is: an on-path responder that can present a certificate the operating system already trusts for api.nullrun.io. That is what corporate TLS interception installs, and it is not exotic — it is a normal thing to find on a managed network. Such a responder can return a perfectly well-formed body including decision_source: "gateway"; it needs neither your API key nor an HMAC bypass, only to answer before the real backend. NullRun's responses are not signed, so there is no after-the-fact detection either — the audit trail would faithfully record an allow that the gate never gave.
So the honest boundary is: server-authoritative enforcement is authoritative against a client and against a network observer, not against an attacker who terminates TLS inside your trust store. If you operate on such a network, exclude api.nullrun.io from interception and check that it stays excluded — that is an operational control, not something the SDK can do for you.
Fail-open here is narrow and deliberate, and the authoritative table lives in runtime.py (ADR-008). In short: a transport failure on the check path is open, so an unreachable backend cannot freeze your agent; a wire response that names an enforcement failure is closed, because the backend made a decision and the SDK will not overrule it; a body that is not a verdict at all — no decision, no decision_source, or an unrecognised value in either — is closed, because something answered and what it said was not a decision, and reading that as permission would let a non-NullRun responder authorise a call no policy engine evaluated; a 401 is closed, because no retry fixes a revoked key; and the /execute path is closed by default (FallbackMode.STRICT).
NullRun does not store or proxy your LLM provider keys — it sits beside your existing clients and observes the calls. The gate is server-authoritative for cost: even a malicious SDK cannot inflate spend by sending a fake cost_cents to /track.
See the security policy for the threat model and disclosure policy.
Made with care by NullRun and contributors.