Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
67 commits
Select commit Hold shift + click to select a range
edd0f82
Bench(docs[orchestration]): Specify active scale
tony Aug 18, 2026
caeaede
Bench(docs[orchestration]): Repeat wait gates
tony Aug 18, 2026
ccff71a
Bench(feat[fuzzer]): Generate active streams
tony Aug 18, 2026
9558e37
Bench(fix[fuzzer]): Record sentinel timing
tony Aug 18, 2026
8b64f2e
Bench(docs[fuzzer]): Exercise service examples
tony Aug 18, 2026
b20b304
Bench(feat[model]): Guard active topology
tony Aug 18, 2026
9e41c45
Bench(fix[model]): Harden run evidence
tony Aug 18, 2026
e56a3f2
Bench(fix[model]): Validate ramp state
tony Aug 18, 2026
57b03ae
Bench(fix[model]): Enforce ramp transitions
tony Aug 18, 2026
cc146b2
Bench(fix[model]): Complete model audit
tony Aug 18, 2026
cfd2e2e
Bench(fix[model]): Enforce ramp checkpoints
tony Aug 18, 2026
995649a
Bench(feat[topology]): Build active servers
tony Aug 18, 2026
f672ee3
Bench(fix[topology]): Harden live ownership
tony Aug 18, 2026
e47e16d
Bench(fix[cleanup]): Roll back private dirs
tony Aug 18, 2026
c40c1c5
Bench(feat[phases]): Measure active queries
tony Aug 18, 2026
5b7c464
Bench(fix[phases]): Harden live acceptance
tony Aug 18, 2026
9df37eb
Bench(feat[wait]): Time delayed output
tony Aug 18, 2026
ec491b3
Bench(fix[wait]): Bound delayed capture
tony Aug 18, 2026
13d1b1d
Bench(feat[runner]): Supervise active runs
tony Aug 18, 2026
e12ad7b
Bench(fix[runner]): Harden supervised recovery
tony Aug 18, 2026
0f6a332
Bench(fix[runner]): Prove safe handoff
tony Aug 18, 2026
019d450
Bench(fix[runner]): Finalize exactly once
tony Aug 18, 2026
438f31e
Bench(fix[runner]): Drain repeated interrupts
tony Aug 18, 2026
a84ea32
Bench(fix[runner]): Bind socket ownership
tony Aug 18, 2026
127ede6
Bench(fix[runner]): Publish ownership atomically
tony Aug 18, 2026
de4a81b
Bench(feat[runner]): Add evidence CLI
tony Aug 19, 2026
3ccf0d1
Bench(fix[runner]): Gate pane freshness
tony Aug 19, 2026
47d6d16
Bench(fix[fuzzer]): Consume requests
tony Aug 19, 2026
b81e138
Bench(fix[runner]): Bound pane readiness
tony Aug 19, 2026
30e4aa1
Bench(fix[content]): Retain fresh sentinels
tony Aug 19, 2026
2899561
Bench(fix[content]): Route control during setup
tony Aug 19, 2026
e58f308
Bench(fix[content]): Retain monotonic pulses
tony Aug 19, 2026
424909e
Engine(fix[control]): Await reconnect readiness
tony Aug 19, 2026
37dcf98
Bench(fix[control]): Quiesce pane output
tony Aug 19, 2026
86fd0de
Bench(docs[fuzzer]): Describe workload options
tony Aug 19, 2026
51ef6cb
Bench(docs[orchestration]): Publish scale guide
tony Aug 19, 2026
56a8349
Bench(feat[matrix]): Supervise lane comparisons
tony Aug 19, 2026
9a5e811
Bench(fix[cleanup]): Ignore socket mode churn
tony Aug 19, 2026
38a5427
Bench(feat[matrix]): Render per-phase lane comparison
tony Aug 19, 2026
aadb767
Bench(feat[orm]): Measure the classic reference
tony Aug 19, 2026
66e035f
Bench(feat[orm]): Admit the reference cells
tony Aug 19, 2026
bee35f8
Bench(docs[orchestration]): Document the reference cells
tony Aug 19, 2026
b5dbeda
Bench(docs[plans]): Correct the cleanup root cause
tony Aug 19, 2026
7251eb7
Bench(fix[matrix]): Refuse ratios inside the noise
tony Aug 19, 2026
ef15036
Bench(feat[stress]): Find where each axis buckles
tony Aug 19, 2026
d4e176b
Bench(fix[stress]): Align shapes in ladder output
tony Aug 19, 2026
9ff6ae2
Bench(docs[plans]): Record the measured comparisons
tony Aug 19, 2026
1898b4c
Bench(docs[orchestration]): Document the pressure ladder
tony Aug 19, 2026
cae3a4b
Bench(fix[matrix]): Default to a shape that fits the budget
tony Aug 19, 2026
191d9d5
Bench(fix[matrix]): Name the interpreter a child needs
tony Aug 20, 2026
651c6f5
Bench(fix[runner]): Teach the pidfd refusal its remedy
tony Aug 20, 2026
7cf8698
Bench(docs[orchestration]): Correct the standalone claim
tony Aug 20, 2026
42fc0d0
Bench(docs[plans]): Record why the interpreter check won
tony Aug 20, 2026
eb4661f
Engines(feat[instrumentation]): Add the async observation wrapper
tony Aug 20, 2026
57d3050
Engines(test[instrumentation]): Pin per-scope attribution under overlap
tony Aug 20, 2026
a1aa0b0
Docs(experimental[instrumentation]): Document the observation seam
tony Aug 20, 2026
28df7c7
Bench(fix[scripts]): Mark the runnable benchmark scripts executable
tony Aug 23, 2026
ee4fc02
Bench(fix[matrix]): Give the phase-table casts their type arguments
tony Aug 23, 2026
9738ab2
Bench(fix[docs]): Exclude both instrumentation modules from the engin…
tony Aug 23, 2026
b962b81
Bench(fix[tests]): Keep the capability socket inside the AF_UNIX budget
tony Aug 23, 2026
32075f0
Bench(fix[tests]): Record tmux 3.3a as the live harness floor
tony Aug 23, 2026
19038c1
Bench(fix[tests]): Stop the reaping stub racing its own timeout
tony Aug 20, 2026
685d3ae
Bench(fix[cleanup]): Stop reading tmux's absent server as a failure
tony Aug 21, 2026
8d74006
Bench(fix[stall]): Stop an orphaned test worker outliving its test
tony Aug 21, 2026
1153816
Docs(fix[bench]): Follow the engine grid to its new path
tony Aug 23, 2026
80dbe81
Bench(refactor): Gather the orchestration tooling into one directory
tony Aug 23, 2026
45cfb97
Tests(refactor): Mirror the orchestration tests under tests/scripts
tony Aug 23, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions docs/experimental/engines.md
Original file line number Diff line number Diff line change
Expand Up @@ -51,6 +51,11 @@ A tmux command rejection is normally result data. Missing executables, dead
persistent connections, protocol mismatches, and similar transport failures
raise at the engine boundary.

## Observe an engine

Any engine can be wrapped to count or trace the tmux commands it dispatches. See
{ref}`instrumentation`.

## Tutorials

Each engine row links to its tested workflow. Start with
Expand Down
8 changes: 8 additions & 0 deletions docs/experimental/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,12 @@ result.
- [Engines](engines.md) explains how to choose a transport without changing the
operation contract.
- [Plans](plans.md) covers deferred execution, planners, and the fluent builder.
- [Instrumentation](instrumentation.md) explains how to count and trace the tmux
commands an engine dispatches, without changing a program that does not.
- [Active orchestration benchmark](orchestration-benchmark.md) documents the
persistent, high-cardinality workload used to measure construction, mutation,
waiting, enumeration, capture, and search under pane activity, how to run it,
and how to read the evidence it retains.

## Run one operation

Expand Down Expand Up @@ -50,4 +56,6 @@ operations/index
results
engines
plans
instrumentation
orchestration-benchmark
```
293 changes: 293 additions & 0 deletions docs/experimental/instrumentation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,293 @@
(instrumentation)=

# Instrumentation

Wrap an engine to count or trace every tmux command that passes through it. A
program that does not wrap one pays nothing, because there is nothing to pay:
{class}`~libtmux.experimental.engines.instrumentation.InstrumentedEngine`
implements the same protocol as the engine it wraps, so observation is a
substitution rather than a feature the engine carries.

## Count what a run costs

{class}`~libtmux.experimental.engines.instrumentation.CountingSink` accumulates
the traffic. {func}`~libtmux.experimental.engines.instrumentation.instrument`
wraps an engine with it and returns something usable anywhere the engine was:

```python
>>> from libtmux.experimental.engines import SubprocessEngine, instrument
>>> from libtmux.experimental.engines.instrumentation import CountingSink
>>> from libtmux.experimental.ops import ListSessions, ListWindows, run
>>> counts = CountingSink()
>>> engine = instrument(SubprocessEngine.for_server(server), counts)
>>> run(ListSessions(), engine).status
'complete'
>>> run(ListWindows(), engine).status
'complete'
>>> counts.requests, counts.tmux_commands, counts.inlined
(2, 2, 0)
>>> counts.elapsed_ns > 0
True
```

## What the three counts mean

`requests` is what the caller asked for, `tmux_commands` is what tmux was told
to do, and
{attr}`~libtmux.experimental.engines.instrumentation.CountingSink.inlined` is
the difference: commands that rode inside another request's argv rather than
costing a dispatch of their own.

The three separate because a request may carry a command group. Sending two
tmux commands in one argv is one dispatch and two commands:

```python
>>> from libtmux.experimental.engines import MockEngine, instrument
>>> from libtmux.experimental.engines.base import CommandRequest, CommandSeparator
>>> from libtmux.experimental.engines.instrumentation import CountingSink
>>> counts = CountingSink()
>>> engine = instrument(MockEngine(), counts)
>>> _ = engine.run(
... CommandRequest.from_args(
... "set-option", "-g", "@x", "1", CommandSeparator(";"), "show-options", "-g"
... )
... )
>>> counts.requests, counts.tmux_commands, counts.inlined
(1, 2, 1)
```

Read the pair against the transport to see what a lane actually costs. A
subprocess engine starts one tmux process per request, so `requests` is also its
process count and every inlined command is a process it did not start. A control
mode engine holds one client for its whole life, so the same `requests` figure
costs no process starts at all — the counts stay comparable across lanes while
the resource they imply does not.

## Write a sink

A sink implements three methods:
{meth}`~libtmux.experimental.engines.instrumentation.Sink.before_command`,
{meth}`~libtmux.experimental.engines.instrumentation.Sink.after_command`, and
{meth}`~libtmux.experimental.engines.instrumentation.Sink.handle_error`. Those
are the hook names OpenTelemetry and Sentry already attach to on SQLAlchemy, so
an exporter written for one transposes here without monkeypatching.

Whatever `before_command` returns comes back as `state`, which is how a sink
carries a span or a timestamp for one command without keeping a map keyed by
request:

```python
>>> import time
>>> from libtmux.experimental.engines import MockEngine, instrument
>>> from libtmux.experimental.engines.base import CommandRequest
>>> class SlowCommandSink:
... """Record commands that took longer than a threshold."""
...
... def __init__(self, threshold_ns: int) -> None:
... self.threshold_ns = threshold_ns
... self.slow: list[str] = []
...
... def before_command(self, request):
... return time.perf_counter_ns()
...
... def after_command(self, request, result, state):
... if time.perf_counter_ns() - state > self.threshold_ns:
... self.slow.append(str(request.args[0]))
...
... def handle_error(self, request, error, state):
... self.slow.append(f"{request.args[0]} (failed)")
>>> slow = SlowCommandSink(threshold_ns=0)
>>> engine = instrument(MockEngine(), slow)
>>> _ = engine.run(CommandRequest.from_args("list-panes", "-a"))
>>> slow.slow
['list-panes']

Sinks stack, and each sees every command independently, so counting and tracing
compose without either knowing about the other:

>>> from libtmux.experimental.engines.instrumentation import CountingSink
>>> counts, slow = CountingSink(), SlowCommandSink(threshold_ns=0)
>>> engine = instrument(MockEngine(), counts, slow)
>>> _ = engine.run(CommandRequest.from_args("list-windows"))
>>> counts.requests, len(slow.slow)
(1, 1)
```

An error reaches `handle_error` and then keeps propagating. A sink observes
failures; it does not handle them.

## Export to OpenTelemetry

A tracing sink opens a span in `before_command`, returns it as state, and ends
it in `after_command` or `handle_error`. Mirror the database client
conventions, which map cleanly onto tmux: the span name is the tmux command,
`tmux.command` is the subcommand, `tmux.statement` is the joined argv, and
`tmux.commands` plus `tmux.inlined` carry the counts above so a query can find
the requests that batched work.

```python
>>> from libtmux.experimental.engines import MockEngine, instrument
>>> from libtmux.experimental.engines.base import CommandRequest, CommandSeparator
>>> from libtmux.experimental.engines.control_mode import command_count
>>> class SpanSink:
... """An exporter's shape, driven here by a recording tracer."""
...
... def __init__(self, tracer) -> None:
... self.tracer = tracer
...
... def before_command(self, request):
... argv = tuple(str(arg) for arg in request.args)
... span = self.tracer.start_span(f"tmux {argv[0]}")
... span.set_attribute("tmux.command", argv[0])
... span.set_attribute("tmux.statement", " ".join(argv)[:512])
... commands = command_count(tuple(request.args))
... span.set_attribute("tmux.commands", commands)
... span.set_attribute("tmux.inlined", commands - 1)
... return span
...
... def after_command(self, request, result, state):
... state.set_attribute("tmux.returncode", result.returncode)
... state.end()
...
... def handle_error(self, request, error, state):
... state.record_exception(error)
... state.end()

Driving it proves the lifecycle: one span per command, ended exactly once,
carrying the attributes a query will filter on.

>>> class RecordingSpan:
... def __init__(self, name):
... self.name, self.attributes, self.ended = name, {}, False
... def set_attribute(self, key, value):
... self.attributes[key] = value
... def end(self):
... self.ended = True
>>> class RecordingTracer:
... def __init__(self):
... self.spans = []
... def start_span(self, name):
... span = RecordingSpan(name)
... self.spans.append(span)
... return span
>>> tracer = RecordingTracer()
>>> engine = instrument(MockEngine(), SpanSink(tracer))
>>> _ = engine.run(
... CommandRequest.from_args(
... "set-option", "-g", "@x", "1", CommandSeparator(";"), "show-options", "-g"
... )
... )
>>> span = tracer.spans[0]
>>> span.name, span.ended
('tmux set-option', True)
>>> span.attributes["tmux.commands"], span.attributes["tmux.inlined"]
(2, 1)
>>> span.attributes["tmux.returncode"]
0
```

Swap the recording tracer for `opentelemetry.trace.get_tracer(...)` and the same
sink exports over OTLP, where `{ span.tmux.commands > 1 }` finds every request
that batched work. Nothing else changes, because the sink never learns which
engine it is observing.

## Async

{func}`~libtmux.experimental.engines.instrumentation.instrument` returns
{class}`~libtmux.experimental.engines.instrumentation.AsyncInstrumentedEngine`
for an async engine, chosen from the engine's own `run`, so a caller does not
pick the wrapper:

```python
>>> import asyncio
>>> from libtmux.experimental.engines import AsyncMockEngine, instrument
>>> from libtmux.experimental.engines.base import CommandRequest
>>> from libtmux.experimental.engines.instrumentation import CountingSink
>>> counts = CountingSink()
>>> engine = instrument(AsyncMockEngine(), counts)
>>> type(engine).__name__
'AsyncInstrumentedEngine'
>>> async def probe():
... await asyncio.gather(
... engine.run(CommandRequest.from_args("list-panes")),
... engine.run(CommandRequest.from_args("list-windows")),
... )
>>> asyncio.run(probe())
>>> counts.requests
2
```

The wrapper awaits the inner engine, so commands that overlapped before still
overlap. Sink callbacks are synchronous and run on the event loop between the
await points, which makes a blocking sink an event-loop stall: keep them to
arithmetic and span bookkeeping, and hand anything slower to a background
exporter.

## Why observation is not ambient

An observer could have been ambient instead: a scope that engines consult, so
no call site changes at all. That design was built and measured against this
one, and lost twice.

It costs everyone, always. Consulting a scope means a lookup on every engine
call whether or not anyone is observing, which measured about eight percent per
call and never goes away. Wrapping costs only the programs that wrap.

It also reports numbers that are wrong rather than absent. An ambient scope has
to be consulted by each engine, so covering the engines is a list somebody
maintains. Control mode dispatches some commands over its persistent connection
and others through a subprocess fallback, so an engine missing from that list
does not report zero -- it reports the fallback commands and hides the rest. A
wrapper sits at the boundary the caller already holds, and counts what the
caller asked for however the engine chooses to fulfil it.

The same reasoning rules out installing listeners on the engine. Two concurrent
scopes share one engine, so its listeners cannot tell their commands apart, and
each scope is charged for the other's work.

## Python call counts

Function calls are not a sink. `cProfile` and `sys.monitoring` are both
process-global and single-owner, so a per-command profiler would attribute one
command's calls to whichever command happened to overlap it on the event loop.
Count them for a whole run instead, alongside the sink that counts tmux work:

```python
>>> import cProfile
>>> from libtmux.experimental.engines import MockEngine, instrument
>>> from libtmux.experimental.engines.base import CommandRequest
>>> from libtmux.experimental.engines.instrumentation import CountingSink
>>> counts = CountingSink()
>>> engine = instrument(MockEngine(), counts)
>>> profile = cProfile.Profile()
>>> profile.enable()
>>> for _ in range(3):
... _ = engine.run(CommandRequest.from_args("list-panes", "-a"))
>>> profile.disable()
>>> calls = sum(entry.callcount for entry in profile.getstats())
>>> counts.requests, calls > counts.requests
(3, True)
```

Pairing them is the point: tmux commands say what the server was asked to do,
and call counts say what Python spent getting there. A change that lowers one
while raising the other has moved cost rather than removed it.

## Overhead

The uninstrumented path is unchanged code, which is the argument for composing
rather than installing. An engine that carried its own event registry would
check whether anyone is listening on every call; an engine that carried wrapper
hooks would build a context object per call whether or not one is used. Wrapping
adds one Python call plus each sink's own work, and only for programs that ask
for it.

Two tests hold the claim: one asserts wrapping adds no attribute to the engine
it wraps, and one measures an instrumented call against the millisecond scale of
a real tmux round trip.
```console
$ uv run pytest tests/experimental/engines/test_instrumentation.py
```

For measurements of the transports themselves rather than of observation,
see {doc}`orchestration-benchmark`.
Loading
Loading