Repository navigation
4822 scheduled promotion after slot handoff - #4823
Merged
jeremydmiller merged 2 commits intoOct 5, 2026
Merged
jeremydmiller merged 2 commits into
jeremydmiller merged 2 commits into
Conversation
…sperFxGH-4822) A scheduled message to a global partition parks in the inbox at the external slot's address. A node that previously owned the slot keeps its stopped listening agent registered, so the scheduled poller handed the promoted envelopes to a listener without a receiver. That threw after the rows were committed as Incoming and owned by a live node, stranding them and every other destination in the same batch. EnqueueDirectlyAsync now settles slot ownership for the slot address itself and forwards to the slot when this node does not own it. Each destination group is also isolated: a failing group is logged and its envelopes are released back to any node so recovery retries them.
…GH-4822) When a promoted group fails part-way, envelopes that were already forwarded have had their inbox rows retired. Releasing them as well let recovery run them a second time, so only the envelopes not yet handed over are released now. The release also matched on the envelope's live destination, which the partition slot forward rewrites to the slot before sending. For a row parked at the companion local queue that matched nothing and left it owned by this live node. Stand-in envelopes now carry the parked address.
Member
|
@BlackChepo Thank you! That's for a JasperFx client, so now I feel a little bad, but I'm grateful for help like this! |
pltknttn
pushed a commit
to pltknttn/wolverine
that referenced
this pull request
Oct 6, 2026
…ng its inbox row A scheduled envelope that comes due on a node which does not listen to its destination is forwarded through a sending agent. JasperFxGH-4645 (JasperFx#4656) made that path delete the forwarded envelope's inbox row, which fixed the orphan. It left the send first and stored nothing as outgoing, which is still two losses in 6.46.0: - A poll by the owning node between the send and the delete. The listener's anti-duplicate probe (JasperFxGH-4316) finds the inbox row beside the new queue row and deletes the queue row, with no status filter. The message is not handled and not dead lettered. - A failed first send. EnqueueOutgoingAsync posts to the agent's RetryBlock and stores nothing, and RetryBlock.PostAsync awaits the first attempt and queues the item on an exception. The inbox row is deleted anyway, so until a retry succeeds no table holds the message and a process that stops in that interval loses it. For a DURABLE agent the order is now store the outgoing row, retire the inbox row, then send: - the inbox row is gone before any queue row exists, so the probe cannot match it; - a failed send is the ordinary outbox case the agent already retries from, and outbox recovery picks it up if this node stops; - a stop between the store and the delete leaves both rows, which can deliver the message twice and cannot lose it. An agent with no outbox behind it keeps the original order: there is nothing to move the message into, so reordering would only widen the window in which no table holds it at all. ISendingAgent.TryStoreOutgoingAsync is a DEFAULT interface member, so an external implementation keeps compiling and keeps today's behaviour. Only DurableSendingAgent overrides it, reusing its own DurableWriteRetry store path (JasperFxGH-4662) rather than a second copy; SendingAgent.setDefaults became protected so the outbox row carries the same Status/OwnerId/ReplyUri it would have had. Both call sites are fixed, not just the reported one: forwardToPartitionSlotAsync (JasperFxGH-4700) ran the identical send-then-delete pair, and shares the new helper. Its parked-address handling is unchanged -- the delete still names the address the row was written under, because received_at is part of the identity. Reported with a three-scenario reproduction (audit triggers recording every queue insert, queue delete and inbox delete, with the statement that ran; the loss observed in 4 of 4 runs and the no-table window in 2 of 2) by @smoqmilus, who also identified the fix. JasperFx#4823 landed first, as the issue said it needed to. Gates: wolverine.slnx Release -f net9.0 0 errors / 0 warnings; CoreTests 3289 green, including Bug_4700 and Bug_4822, whose stub transports are non-durable and so still exercise the original order. Co-Authored-By: smoqmilus <smoqmilus@users.noreply.github.com> Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VDUrBeB4tTnKj4AExCS1nj
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #4822
Problem
With
GlobalPartitioned+ sharded database queues, a scheduled message cascaded from a handler parks in the inbox at the external slot's address (by design since GH-4673, so ownership is decided when the message comes due).When a node joins and the leader moves a slot away before the message is due, the scheduled poller on the node that gave up the slot can win the advisory lock. That node still has the slot's stopped
ListeningAgentregistered, so:GlobalPartitionSlotFordoes not apply (the slot address is not a companion queue),FindListenerCircuitreturns the stopped agent,ListeningAgent.EnqueueDirectlyAsyncthrowsThere is no active, local queue for this listening endpoint.The poller had already committed the rows as
Incomingowned by this live node, so inbox recovery never picks them up. The exception also aborts the remaining destinations in the batch, which is why messages for the slot the node kept were lost too.A node that never owned the slot is not affected: it has no agent for the address and already forwards through the sender branch.
Changes
IEndpointCollection.IsGlobalPartitionSlot(Uri)(cached in anImHashMaplike its neighbours).EnqueueDirectlyAsyncnow treats a destination that is a slot the same way Global partitioning: a scheduled retry runs on a node that does not own the slot #4700 treats a companion queue: if this node does not own it, forward to the slot.AnyNodeso recovery retries them (same remedy as CircuitBreakingTests: durable_and_not_parallel.the_circuit_breaker_should_trip_and_restart is unreliable on CI #3680 on the recovery side).received_at), not the live destination, which the slot forward rewrites to the slot before sending.Tests
PostgresqlTests/Transport/Bug_4822_scheduled_promotion_after_slot_handoff: end-to-end through the real scheduled poller on a Solo host, with one slot's exclusive listener stopped to create a genuine ex-owner (same technique as Global partitioning: a scheduled retry runs on a node that does not own the slot #4700 / Global partitioning: a replayed dead letter runs on a node that does not own the slot #4776).Incoming/owner_id = 0,IsGlobalPartitionSlotagrees with the URIs the topology actually stamps.CoreTests/Bugs/Bug_4822_failed_promotion_releases_the_parked_rows: a slot forward that fails on the second envelope releases only the un-forwarded envelopes, at the companion queue address.Both bug tests failed before the fix with exactly the reported state (rows
Incoming, owned by the live node). Existing GH-3413 / GH-4645 / GH-4700 / GH-4776 / partitioning / scheduled suites pass, andwolverine.slnxbuilds clean with-c Release -f net9.0.Known limitations (not addressed here)
PostgresqlQueueListener.TryPopDurablyAsynccan delete the freshly forwarded queue row. This change routes the ex-owner case into that existing path. Before, these messages were stranded entirely; now they are exposed to a narrow window. I'd suggest tracking this separately.listener.EnqueueDirectlyAsynchands a group over as a whole, so a failure part-way through releases the whole group, and already-enqueued envelopes could run twice. Same trade-off as CircuitBreakingTests: durable_and_not_parallel.the_circuit_breaker_should_trip_and_restart is unreliable on CI #3680; the realistic failure (a listener with no receiver) throws before the first envelope.