Skip to content

fix(sftp): a tar stream stranded on a dead link resumes per file (#495) - #560

Merged
kipavy merged 1 commit into
devfrom
fix/tar-stall-watchdog-495
Oct 7, 2026
Merged

kipavy merged 1 commit into
devfrom
fix/tar-stall-watchdog-495

Conversation

@kipavy

@kipavy kipavy commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Closes #495. Closes #554.

Per-file copies already run under a stall watchdog (unless_lost). Tar streams (stream::upload / download / relay) did not. With SSH keepalive off, a half-open link hung a folder transfer forever instead of falling back to per-file resume.

What changed

after_tar now starts the stream itself and runs it under the watchdog. The watchdog's progress counter is the byte count the stream already keeps in its Progress. Both tar entry points (stream_or, relay_or_per_file) go through it.

After a stall of 15 s or more, the stream is kept while revive waits for the link (the UI shows "waiting"). It is only cut once it is stranded:

  • its SSH connection closed, or
  • a reconnect replaced that connection, and the old connection no longer answers an SSH-level probe (ssh_answers).
After a ≥15 s stall Before Now
Healthy link (nothing stalls) streams streams, no change
Probe fails, then the stream's own connection answers (Wi-Fi blip, frozen link) stream carries on stream kept and carries on (#554 can't happen)
Reconnect while the old connection still answers stream carries on stream kept and carries on
Stream's connection closed, or replaced and silent (#495) hangs forever cut, then resumed per file
Link never comes back hangs forever keeps waiting, shown as "waiting"; cancel works as before

A cut stream is stopped by cancelling its own child token, not by dropping it. It runs the same bounded unwind as a user cancel: it closes the remote channel, stops the status task and the local packer/unpacker. It gets UNWIND (5 s) for that before the per-file run starts. The watchdog never fails a transfer on its own.

Plumbing:

  • unless_lost takes a Stuck check (a small async_trait, so the future stays Send) that counts as dead at a stall. Per-file copies pass &false, so they are unchanged.
  • StartedOn records each TarHost's SSH handle when the stream starts and implements Stuck.
  • The _via functions take the cut token and the Progress. Job::with_progress builds their job, and Job::new is now test-only.
  • resume_per_file is shared by both fallback paths.

The one case left

A connection that stops answering, is replaced by a reconnect, and then comes back to life later could deliver its last buffered bytes to the remote tar after the per-file run has started. Before this PR that transfer hung, and finished only if the old connection revived. A reconnect normally happens because the old flow is gone for good (Wi-Fi roam, NAT reset). Ruling this out completely would mean killing the remote tar by PID over the new connection, which changes the remote command on every host type. This PR doesn't do that.

Tests

New unit tests (paused clock; tokio test-util added to dev-deps only):

  • a_tar_stream_stranded_on_a_dead_link_resumes_per_file_instead_of_hanging
  • a_tar_stream_left_on_a_replaced_connection_resumes_per_file
  • a_cut_tar_stream_runs_its_own_cancel_before_the_fallback
  • a_stalled_tar_stream_whose_own_connection_answers_again_is_kept: a probe fails once, the stream finishes by itself, it is never cut, and there is no per-file run.

Each was seen failing first. The "kept" test fails with cut-on-any-probe logic. The "replaced connection" test hangs without the Stuck check.

New live tests (#[ignore = "needs docker"]). They throttle the SSH container (--cpus 0.25) and SIGSTOP part of a running upload:

  • a_tar_upload_whose_link_freezes_and_thaws_finishes_on_its_own_stream: the tar's sshd is stopped for 25 s, then SIGCONT. The upload finishes on its stream, "waiting" was seen, there is no per-file run, and the tree is identical.
  • a_tar_upload_stranded_by_a_reconnect_resumes_per_file_on_the_new_link: the tar's sshd is stopped and a fresh connection is swapped into the session. The upload goes per file and the tree is identical.
  • a_tar_upload_whose_session_reconnects_while_its_link_lives_finishes_on_it: the remote tar itself is stopped (its link still answers), a fresh connection is swapped in, and after 30 s it gets SIGCONT. The upload finishes on its stream with no per-file run.

The live tests compile but have not been run. The build container here has no Docker access. To run them:
cargo test --lib commands::sftp::tar::tests::live -- --ignored --test-threads=1

Verification:

  • cargo fmt --check: clean.
  • cargo clippy --workspace --all-targets -D warnings: clean.
  • cargo test --lib: 785 passed, 3 failed. The 3 failures are the commands::plugins tests that need pnpm build:plugins output, the same as on dev.

🤖 Generated with Claude Code

Per-file copies run under a stall watchdog; tar streams (upload,
download, relay) had none, so with SSH keepalive off a half-open link
hung a folder transfer forever.

after_tar now runs the stream under unless_lost, fed by the bytes the
stream counts into its Progress. A stalled stream is kept while the link
comes back, and cut (by cancelling its own token, so it closes its
remote end as a cancel does) only once it is stranded: its connection
closed, or a reconnect replaced it and the old one no longer answers.
A cut stream resumes per file. A link that answers again keeps the
stream, and the watchdog never fails a transfer by itself (#554).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@kipavy
kipavy force-pushed the fix/tar-stall-watchdog-495 branch from 7df5b4e to 4a9d641 Compare October 7, 2026 10:45
@kipavy
kipavy merged commit 66ac6b4 into dev Oct 7, 2026
4 checks passed
@kipavy

kipavy commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Live tests run against a real SSH container (--network host + Docker socket), on the merged code (4a9d641c):

Live test Result
a_tar_upload_whose_link_freezes_and_thaws_finishes_on_its_own_stream ✅ 55 s. The watchdog fired ("waiting" seen), the stream was kept, there was no per-file run, and the tree is identical
a_tar_upload_stranded_by_a_reconnect_resumes_per_file_on_the_new_link ✅ 57 s. The stream was cut and resumed per file on the new link, and the tree is identical
a_tar_upload_whose_session_reconnects_while_its_link_lives_finishes_on_it ✅ 66 s. The stream was kept on its live connection, with no per-file run
a_tree_bigger_than_a_small_tmp_relays_downloads_and_uploads (existing) ✅
a_full_destination_is_named_in_the_error (existing) ✅

Check that the "reconnect while alive" test can fail: with the looser rule (replaced ⇒ stranded) it fails, because the per-file fallback ran on a live stream. So it guards the tightening.

@kipavy
kipavy deleted the fix/tar-stall-watchdog-495 branch October 7, 2026 12:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant