Skip to content

preserve trailing external events across continue-as-new - #163

Merged
Tomer Rosenthal (torosent) merged 3 commits into
mainfrom
torosent-continue-as-new-event-safety
Sep 29, 2026
Merged

Tomer Rosenthal (torosent) merged 3 commits into
mainfrom
torosent-continue-as-new-event-safety

Conversation

@torosent

@torosent Tomer Rosenthal (torosent) commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

The problem

An orchestrator can use WhenAny to wait for either an external event or a timer. If the timer wins, it can call ContinueAsNew and return while the event wait is still pending.

The same batch can contain a matching event after the timer. The SDK delivered that event to the old wait—even though the old run had finished and could never resume. The event was lost instead of being kept for the next run, despite WithKeepUnprocessedEvents being enabled.

We reproduced the loss on both the DTS emulator and live DTS. With the fix, the next run received all six test events once and in order.

The fix

  • Stop delivering events to the old run's pending waits after it successfully finishes with ContinueAsNew.
  • Keep those events for the next run when WithKeepUnprocessedEvents is enabled, preserving their arrival order.
  • Do not carry events forward if the orchestrator fails or is terminated instead.

What this means for event handling

Calling ContinueAsNew does not immediately end the current run. The orchestrator can still receive events before it returns.

Only undelivered events are kept. An event already delivered to a live wait is not carried forward, even if application code never awaits that wait. Without WithKeepUnprocessedEvents, undelivered events are dropped as before.

WhenAny still does not cancel its losing waits. Cancel a losing event wait explicitly, or use EventChannel with Select, if later code should receive that event.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@torosent
Tomer Rosenthal (torosent) marked this pull request as ready for review September 28, 2026 20:10
Copilot AI lite review requested due to automatic review settings September 28, 2026 20:10
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The unrelated public error-message and rewind changes must be split out or documented in the PR description.

Review effort: Balanced
Findings: 1 Low severity

Open (1)

Comment thread api/management.go
@andystaples

Copy link
Copy Markdown
Contributor

We found the same trailing-event loss in Python and opened microsoft/durabletask-python#281. That PR is deliberately draft while we discuss the cross-SDK contract. The Go fix looks sound; this is a non-blocking semantics discussion, not a request to hold the immediate fix.

Shared reproducer and consumption edge cases

A timer wins WhenAny, the orchestrator continues as new and returns, and the same work-item batch contains two matching external events. The first event satisfies the abandoned losing wait and disappears; only the second is buffered. The Python proposal preserves both trailing events when save_events=True, corresponding to Go's WithKeepUnprocessedEvents.

The important distinctions are:

  • Delivered is not necessarily observed. Both proposals treat delivery to a live registered wait as consumption, even if application code never awaits/reads that task. An event delivered to a losing WhenAny wait before the terminal boundary is not resurrected for the next execution.
  • WhenAny does not cancel its losers. Explicit cancellation is still necessary if later code should receive an event instead of an outstanding losing wait.
  • Retention remains opt-in. Without WithKeepUnprocessedEvents / save_events=True, undelivered events are not carried forward. Previously delivered events are not recovered by enabling retention.

Where Python deliberately differs, and why

Edge case This Go proposal Python draft
Continue-as-new is requested but the orchestrator has not returned The call records intent. Live waits can still receive events until successful finalization. Python already sets _is_complete at the call. The proposed dispatch guard uses that existing boundary, so subsequent ordinary events are buffered rather than delivered to pending waits. We avoided importing Go's deferred-intent model because that would be a broader Python lifecycle change.
Code or cleanup runs after the call Normal-return defers can still schedule and await durable work before finalization. This is distinct from the forced-unload protection in #164. The guard is not universal cancellation or a guarantee that execution stops: synchronous post-call code, synchronous consumption of already-buffered events, and finally cleanup can still run. Other history/resume paths are unchanged. The draft recommends returning immediately rather than relying on the call to interrupt Python control flow.
Failure/termination versus a prior continue-as-new request Carryover attaches only to actual successful CONTINUED_AS_NEW, not an execution that ultimately fails or is terminated. We retain Python's existing lifecycle: an exception exits history processing, and failure after the continue-as-new call does not replace the status already set by that call. Changing that precedence is deliberately outside this narrow loss fix.

The Python guard applies only to ordinary external-event delivery; entity call/lock responses retain their earlier routing. It neither cancels every pending task nor stops processing all trailing history.

We also reproduced cross-name reordering (a:1, b:2, a:3 becoming a:1, a:3, b:2) and fixed it separately in microsoft/durabletask-python#278. This Go PR already preserves global arrival order. The Python split keeps buffer ordering independent of the lifecycle/consumption decision under discussion.

Cross-SDK questions

  1. Should delivery to a live wait or observation by orchestrator code define consumption, especially for losing waits?
  2. Should continue-as-new become terminal at the call or at successful return/finalization, and what should subsequent failure or termination mean?
  3. What post-call work is supported, and how should we distinguish ordinary cleanup from forced unload and from stopping future event delivery?

It would be useful to document the intended observable event-retention contract across Durable, while allowing language-specific execution/cleanup mechanics where necessary. Python's draft and regression coverage make the current differences explicit rather than silently declaring a new standard.

@torosent
Tomer Rosenthal (torosent) merged commit 57e93dd into main Sep 29, 2026
9 checks passed
@torosent
Tomer Rosenthal (torosent) deleted the torosent-continue-as-new-event-safety branch September 29, 2026 18:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants