Skip to content

fix(resume): an unusable memory image recovers by disk-only cold boot, not Dead - #1613

Merged
nikhilunni merged 4 commits into
mainfrom
fix/resume-memory-image-unusable
Oct 8, 2026
Merged

nikhilunni merged 4 commits into
mainfrom
fix/resume-memory-image-unusable

Conversation

@nikhilunni

@nikhilunni nikhilunni commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Problem

After the ADR 0112 phase 2b roll, an Idle session with a memory snapshot from before the roll could not resume. That snapshot has no swap manifest. Every host refuses it (swap restore has no manifest; refusing a stale device). The resume verb counted each refusal as a terminal failure, and after five (RESUME_FAILURE_STREAK_BUDGET) it set the session Dead ("fork it to continue"). The disk was intact the whole time.

A host-agent that adopts a running guest from an older host-agent writes snapshots of the same shape. The guest still uses a raw swap file, and the new capture runs no swapoff, so those memory images can reference swap pages that no snapshot holds. The two kinds look the same in PG. A migration or a host-side rule cannot separate the safe ones from the unsafe ones.

Fix

Refuse the memory image, keep the disk:

  • engram-core: new SandboxError::MemoryImageUnusable. The refusal is deterministic (every host answers the same), and the disk is unaffected.
  • engram-host-agent: the three swap restore refusals in pooled_backend.rs return the new variant. grpc_server.rs sends it as a failed_precondition marker (the HarnessSpawn precedent), and engram-protocol maps it back.
  • WIRE_VERSION 34 → 35. An older host refuses the same snapshot with an untyped error, which would still count toward the Dead budget. The exact-match gate keeps such hosts out of placement during the roll, so resumes wait for upgraded hosts.
  • engram-coordinator resume: resume_from_fc_snapshot catches the variant and calls resume_disk_only_cold_boot on the newer of live_disk_manifest and the snapshot's disk_manifest. The cold boot now carries the image's memory and CPU budgets, so placement checks that it fits. After success, the coordinator appends a fenced resumed_from_disk event, and engram_session_resume_memory_image_fallback_total increments.
  • engram-coordinator teleport: a live move whose destination refuses the memory image aborts the export and rolls back. Before, it retried three times and then fail_move destroyed the source guest. A captured teleport already rolled back.
  • web: resumed_from_disk renders as a marker that says the files are intact and running processes stopped. The orchestrator frame taxonomy lists the new kind.
  • ADR 0112: a dated addendum line records the gap and the fallback.

The guest loses its processes and keeps its disk. An operator who upgrades from v0.10.0 needs no data migration: pre-upgrade memory snapshots resume by disk-only cold boot.

Not changed:

  • The refused snapshot row stays recoverable. Demoting it would unpin, from chunk GC, the snapshot disk that the cold boot mounts when the session has no live manifest.
  • A retained Idle binding is cleared together with live_disk_manifest. A retry after that point boots the snapshot's older disk. This is older than the PR and also affects the existing disk-only branch. It is left for a separate change.

Test

  • crates/engram-dst/tests/memory_image_fallback.rs: the sim host refuses the memory image once. The resume reaches Active with exactly one cold boot on the live disk and one resumed_from_disk event. The test fails with the coordinator arm disabled.
  • teleport_live_pg scenario refused_memory_image_rolls_back_live_move: one restore attempt, the export is aborted, the source is not destroyed, and the session is Active. It fails without the teleport arm.
  • Marker round trip and gRPC mapping unit tests in engram-protocol and engram-host-agent; the wire golden pins 35.
  • just check: 2767 passed, 317 skipped. cargo clippy --target aarch64-unknown-linux-musl -p engram-host-agent --all-targets -D warnings: clean.
  • web: format:check, lint, buildMessages.test.ts (79 passed), build. orchestrator: typecheck, frame-taxonomy.test.ts (5 passed).

🤖 Generated with Claude Code

nikhilunni and others added 4 commits October 7, 2026 18:23
…, not Dead

A memory snapshot taken before the ADR 0112 phase 2b roll has no swap
manifest, so every host refuses to restore it. The resume verb counted
each refusal as a terminal failure and set the session Dead after five,
although its disk was intact. A host-agent that adopts a running guest
from an older host-agent writes snapshots of the same shape, and those
can reference swap pages that no snapshot holds. PG cannot tell the two
kinds apart.

- engram-core: new SandboxError::MemoryImageUnusable. The refusal is
  deterministic, and the disk is unaffected.
- host-agent: the three swap restore refusals return it. It crosses
  gRPC as a failed_precondition marker (the HarnessSpawn precedent), so
  WIRE_VERSION does not change.
- coordinator: resume_from_fc_snapshot catches it and does a disk-only
  cold boot on the newer of the live and snapshot disk lineages. The
  guest loses its processes and keeps its disk. A new counter,
  engram_session_resume_memory_image_fallback_total, counts each case.
- engram-dst: a sim host can refuse memory images and records each
  create's root disk. The new memory_image_fallback test fails without
  the coordinator change.
- ADR 0112: a dated addendum line records the gap and the fallback.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…refusals

Before this branch, a host refused a pre-roll memory snapshot with an
untyped Snapshot error. The new coordinator reads that as a generic
failure and counts it toward the five-failure Dead budget, and snapshot
affinity keeps sending the resume to that same host. The exact-match
wire gate now fences such hosts out of placement during the roll, so
resumes wait for upgraded hosts instead of failing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… rolls back

finish_live_restore sent every untyped error through three retries and
then fail_move, which destroys the source guest. A memory image refusal
is deterministic, so a retry cannot help. Handle it like
postcopy-never-loaded: abort the export and roll back, so the source
keeps running. The new scenario fails without this branch.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…cesses were lost

- resume_disk_only_cold_boot passes the image's memory and CPU budgets to
  placement. Before, it passed none, so a cold boot could land on a host
  without room.
- After a successful fallback, the coordinator appends a fenced
  resumed_from_disk event. The web renders it as a marker that says the
  files are intact and running processes stopped. The orchestrator frame
  taxonomy lists the new kind.
- The fallback counter increments only after the cold boot succeeds.
- The sim host counts refusals, and the DST test asserts one refusal and
  one resumed_from_disk event. The cold boot can no longer come from the
  resume skipping the snapshot.
- ADR 0112: the addendum line covers the wire bump, the live teleport
  rollback, and why the refused row stays recoverable (demoting it would
  unpin the disk the cold boot can mount).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@nikhilunni
nikhilunni merged commit 8b2bbd6 into main Oct 8, 2026
30 checks passed
@nikhilunni
nikhilunni deleted the fix/resume-memory-image-unusable branch October 8, 2026 02:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant