Skip to content

fix(vsock): kick() must not arm the RX gate on a plain vm resume - #14

Merged
nikhilunni merged 1 commit into
engram/live-migrationfrom
engram/fix-vsock-rx-gate-plain-resume
Jul 7, 2026
Merged

nikhilunni merged 1 commit into
engram/live-migrationfrom
engram/fix-vsock-rx-gate-plain-resume

Conversation

@nikhilunni

Copy link
Copy Markdown
Contributor

Bug

kick() runs on every vm resume (Vmm::resume_vm → kick_virtio_devices), not just the restore-resume it was written for, and it unconditionally armed pending_event_ack (the RX-delivery gate added by the "gate RX delivery on TRANSPORT_RESET ack" fix).

On a plain PATCH /vm Paused → Resumed cycle no TRANSPORT_RESET was queued (prepare_save never ran), so the guest gets an empty evq interrupt, has nothing to ack, and the gate never clears:

  • every host→guest vsock delivery black-holes (process_rx short-circuits forever), and
  • the event loop busy-spins at 100% CPU on the backend's undeliverable pending-RX (level-triggered epoll).

Hit in engrams production by ADR 0074 rung-2 parked-paused eviction (un-pause = plain pause/resume): prompts were "delivered" into an intact-but-gated connection and never reached the harness; in-guest exec hung; fc_vcpu threads idle-ticked while the FC main thread burned a full core. Also latent on the ADR 0045 admin pause/resume and migration-abort resume paths since the gate landed.

Fix

kick() signals (evq + TX-notification replay) only when pending_event_ack is already armed, and never arms it itself. The two paths that actually queue a reset own the arming:

  • prepare_save — in-process, covers the diff-checkpoint source resume (unchanged);
  • restore() — cross-snapshot, now re-arms from the saved virtio_state.activated flag (the gate flag itself is not persisted, and an activated snapshot always carries a queued reset because prepare_save runs on every save; a never-activated snapshot queued none, so arming would gate RX forever).

Tests

  • test_kick_plain_resume_unarmed_is_a_noop — the regression: active device, unarmed → kick leaves the gate open, signals nothing, RX still flows.
  • test_kick_when_armed_keeps_gate_and_suppresses_rx (rewrite of test_kick_when_active_arms_pending_event_ack) — armed kick re-signals and keeps gating.
  • test_kick_replays_tx_notification_only — now arms first (restore semantics preserved).
  • test_restore_arms_rx_gate_iff_snapshot_was_activated — the persist half.

cargo test -p vmm --lib -- vsock on the KVM dev VM: 67 passed, 2 failed — test_muxer_killq (parallelism flake, passes in isolation) and test_set_vsock_device fail identically at the base commit (pre-existing, environment-specific); everything this PR touches passes.

🤖 Generated with Claude Code

https://claude.ai/code/session_014jJi2vqAaxt3Q5UKxbe4Gx

kick() runs on EVERY vm resume (Vmm::resume_vm -> kick_virtio_devices),
not just the restore-resume it was written for, and it unconditionally
armed pending_event_ack. On a PLAIN pause->resume (no snapshot anywhere
in between) no TRANSPORT_RESET was queued, so the guest has nothing to
ack and the gate never clears: every host->guest vsock delivery
black-holes (process_rx short-circuits) and the event loop busy-spins
at 100% CPU on the backend's undeliverable pending-RX (level-triggered
epoll). Hit in production by ADR 0074 rung-2 parked-paused eviction,
whose un-pause is exactly a plain PATCH /vm Paused -> Resumed cycle;
also latent on the ADR 0045 admin pause/resume and migration-abort
resume paths.

Fix: kick() signals (evq + TX-notification replay) only when
pending_event_ack is ALREADY armed, and never arms it itself — the two
paths that queue a reset own the arming: prepare_save (in-process, the
diff-checkpoint source resume) and now restore() (cross-snapshot,
re-armed from the saved activation flag since the gate flag itself is
not persisted; an activated snapshot always carries a queued reset
because prepare_save runs on every save).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014jJi2vqAaxt3Q5UKxbe4Gx
@nikhilunni
nikhilunni merged commit 33c8fba into engram/live-migration Jul 7, 2026
3 checks passed
nikhilunni added a commit to cortexapps/engrams that referenced this pull request Jul 7, 2026
… plain resume (+ outbox/park hardening) (#596)

* fix(outbox): revert the #594 self-reattach window; defers bump attempts; terminal sessions drop their rows

Three outbox-delivery corrections from the rung-2 root-cause pass:

- Revert #594's NotFound await-in-guest-self-reattach window. Its
  premise was disproven twice over: an ADR 0074 rung-2 un-pause keeps
  the harness vsock connection INTACT (delivery succeeds; the NotFound
  arm never runs there), and a genuine harness-unbound desync has
  nothing to wait for (the harness only re-dials on a dropped link or
  agentd's SIGUSR1). It was also unreachable-fallback dead code -- see
  next point -- so a real desync deferred forever instead of
  self-healing via start_agent. NotFound now reattaches immediately
  again (e35ed1f).

- outbox_defer increments attempts. Only mark_delivered bumped it, so a
  row failing BEFORE the forward (ensure_active error, NotFound) sat at
  the floor failure_backoff forever and read attempts=0 in every
  investigation -- and #594's attempts<3 gate was always true.

- ensure_active maps Completed/Failed to Gone (410), not Conflict (409):
  a 409 reads as retry-later to every caller, and the outbox driver
  deferred a completed session's un-acked rows every backoff tick
  forever (observed live: 3 rows spinning the prod driver for hours).
  Gone routes them to the driver's Terminal drop arm and gives
  exec/upload/relay callers an honest 410.

Verified: coordinator unit suite + outbox_live_pg (extended with the
defer-bumps-attempts property) against live PG.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014jJi2vqAaxt3Q5UKxbe4Gx

* fix(park): parked-paused uniformly means Evicting; headroom resolves via PG; ascent clears park_rung on un-pause

Three ADR 0074 rung-2 hardening fixes found during the root-cause pass:

- The park branch now transitions an Active entry (the admin EvictIdle
  path) to Evicting after the pause lands, so park_rung=2 uniformly
  implies Evicting -- the admin path used to leave an ACTIVE session
  advertised over a frozen VM. On bookkeeping failure the VM is
  un-paused and the pipeline falls through to the full eviction. The
  ensure_active Active-arm un-park stays as the backstop for the crash
  window between the pause and the transition.

- try_cancel_nominated_eviction clears park_rung the moment the
  un-pause lands, not in the transition's Ok arm: the backstop case is
  already Active, and Active->Active is a same-state Conflict -- tying
  the clear to the Ok arm left park_rung=2 advertised forever over a
  running VM (observed live on prod).

- host_has_memory_headroom resolves the host via PG sessions.host_id
  instead of the in-memory host_registry: the registry is per-replica,
  so the RPC landing on the pod without the cached bind failed closed
  and silently degraded every park into a full eviction -- a coin flip
  in a 2-replica deployment (observed live: identical EvictIdle calls
  parked on one attempt and captured on the next).

- evict_idle_core reports the pipeline's actual outcome (parked /
  skipped / idle) instead of a hardcoded "idle".

Verified: coordinator unit suite (park tests updated to stamp the mock
session's host_id for the PG-resolved lookup).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014jJi2vqAaxt3Q5UKxbe4Gx

* test(fc): pause->resume vsock-delivery regression test + ADR 0074 incident record

The rung-2 un-pause black-hole was a VMM bug: upstream FC v1.16
(inherited by the fork) arms the vsock RX-delivery gate
(pending_event_ack) from kick() on EVERY resume_vm, but a plain
pause->resume queues no TRANSPORT_RESET for the guest to ack, so the
gate never clears -- all host->guest vsock delivery black-holes and the
FC event loop spins at 100% CPU. Fixed in the fork
(cortexapps/firecracker#14): kick() signals only when the gate is
already armed; prepare_save (in-process) and restore() (cross-snapshot,
from the saved activation flag) own the arming.

The new test proves the property end-to-end on the fork binary
(ENGRAM_FC_FORK_BIN, same gating + artifact plumbing as
stock_fork_snapshot_compat): boot a real microVM with the in-guest
agent, exec over vsock, plain pause+resume, exec again within a bounded
budget. Wired into ci.yml's FC-lane --test list. Bookend on the KVM dev
VM: base fork binary FAILS ("exec dial blocked after plain
pause->resume", pinned by the outer timeout -- the prod wedge
reproduced), fixed binary PASSES in 16s.

ADR 0074's divergence log gains the full incident record (the two wrong
theories, the actual mechanism, the engrams-side hardening, and the
file-upstream follow-up).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014jJi2vqAaxt3Q5UKxbe4Gx

* chore(fc-fork): bump third_party/firecracker — vsock RX gate must not arm on a plain resume

Picks up cortexapps/firecracker#14: kick() signals the vsock evq (and
replays the TXQ notification) only when pending_event_ack is ALREADY
armed, and never arms it itself; prepare_save and restore() (from the
saved activation flag) own the arming. A plain PATCH /vm
Paused->Resumed cycle no longer closes the RX-delivery gate with
nothing for the guest to ack -- the ADR 0074 rung-2 un-pause black-hole
(and the latent ADR 0045 pause/resume + migration-abort variants).
Snapshot byte-format unchanged (stock<->fork compat preserved).

bake-images' build-firecracker publishes the fork artifact keyed by
this gitlink SHA; the FC lane's pause_resume_vsock regression test
exercises it via ENGRAM_FC_FORK_BIN.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014jJi2vqAaxt3Q5UKxbe4Gx

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant