CI flake: TailEndToEndTest.kill_restart_no_loss_no_duplicates_offline_scrollback 120s timeout (run #138) #56
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Build run #138 (
b7ef236, roadmap-only commit — cannot be code-caused) failed with TimeoutCancellationException after 120s in the tail sample's Synapse e2e; #136 was green two commits earlier. Same deflake family as #55: find the slow leg (container pull? initial sync await?), don't just bump the timeout.Likely root cause identified while fixing #66: the kill/restart starvation signature matches a dead sliding pos. Synapse can reject a resumed connection position with M_UNKNOWN_POS, and SyncLoop treated that as a generic error — backoff, retry the same dead pos, forever — so run2 never surfaced the markers and the 120s/240s timeout fired.
5579d06(D34 addendum) makes SlidingSyncSource restart the connection on M_UNKNOWN_POS (to-device token kept). Suggest keeping this open until a few CI runs confirm the flake is gone, then closing against5579d06.First CI run containing the M_UNKNOWN_POS restart (
45f6c20, run for task 1018) is green — build + spec-lint, tail kill/restart e2e included. Leaving open per the earlier note until a few more runs confirm the starvation signature is gone.Closing against
5579d06(D34 addendum: SlidingSyncSource restarts the connection on M_UNKNOWN_POS, to-device token kept).Evidence the starvation signature is gone — four consecutive green build runs since the fix, each verifiably executing :samples:katrix-tail:test (task line + job success in the logs; passing test names are not echoed):
45f6c20) — the first post-fix run, already noted above43b9e1e)b904e2a)165b510)Also re-confirmed the original failure matches the diagnosis: run #138's test task started 19:51:07 and failed at 19:53:13 with TimeoutCancellationException after exactly one starved 120s await — run2 spinning on a dead sliding pos, never surfacing the markers.