fix(envoy-client): declare actors lost before the engine can reallocate them - #5823
Conversation
|
Stack for rivet-dev/rivet Current stack:
Dependencies: Get stack: change nkyokown |
|
Review The change is small and the shared Likely bug: stale
Pings are about 3s apart, so there is a window for this on every reconnect after a lost event. A fix is to reset Smaller points
|
| let iter_start = crate::time::Instant::now(); | ||
| #[allow(unused_assignments)] | ||
| let mut branch: &'static str = "unknown"; | ||
| let ping_silence_wait = engine_ping_silence_wait(&ctx); |
There was a problem hiding this comment.
🔴 High · Ping liveness is not synchronized with the select loop or connection session
last_ping_ts is updated in forward_to_envoy without waking this loop. On a fresh connection, command replay can start actors before the ping task sends its first ping; this iteration therefore builds None, and subsequent pings do not arm a deadline until another envoy message or the 15-second KV cleanup tick. With the default 15-second lost threshold, a link that goes half-open in that window can let the engine expire the envoy before this branch runs, defeating the single-writer protection this change is meant to add. The timestamp also survives reconnects, so a reconnect after the old deadline immediately loses still-running or replayed actors before the new connection's first ping. Make ping reception/session changes an event observed by this loop (for example, a watch channel carrying the current session's last-ping value), reset it when a connection is established, and derive/restart the deadline from that event.
2b3bd85 to
8d4a713
Compare
26a3077 to
6a97303
Compare
No description provided.