Skip to content

4822 scheduled promotion after slot handoff - #4823

Merged
jeremydmiller merged 2 commits into
JasperFx:mainfrom
BlackChepo:gh-4822-scheduled-promotion-after-slot-handoff
Oct 5, 2026
Merged

jeremydmiller merged 2 commits into
JasperFx:mainfrom
BlackChepo:gh-4822-scheduled-promotion-after-slot-handoff

Conversation

@BlackChepo

Copy link
Copy Markdown
Contributor

Fixes #4822

Problem

With GlobalPartitioned + sharded database queues, a scheduled message cascaded from a handler parks in the inbox at the external slot's address (by design since GH-4673, so ownership is decided when the message comes due).

When a node joins and the leader moves a slot away before the message is due, the scheduled poller on the node that gave up the slot can win the advisory lock. That node still has the slot's stopped ListeningAgent registered, so:

  1. GlobalPartitionSlotFor does not apply (the slot address is not a companion queue),
  2. FindListenerCircuit returns the stopped agent,
  3. ListeningAgent.EnqueueDirectlyAsync throws There is no active, local queue for this listening endpoint.

The poller had already committed the rows as Incoming owned by this live node, so inbox recovery never picks them up. The exception also aborts the remaining destinations in the batch, which is why messages for the slot the node kept were lost too.

A node that never owned the slot is not affected: it has no agent for the address and already forwards through the sender branch.

Changes

Tests

  • PostgresqlTests/Transport/Bug_4822_scheduled_promotion_after_slot_handoff: end-to-end through the real scheduled poller on a Solo host, with one slot's exclusive listener stopped to create a genuine ex-owner (same technique as Global partitioning: a scheduled retry runs on a node that does not own the slot #4700 / Global partitioning: a replayed dead letter runs on a node that does not own the slot #4776).
    • ex-owner forwards the due message to the slot; it runs exactly once after the listener comes back, and the kept slot's message in the same batch still runs,
    • a destination that cannot accept promoted envelopes no longer strands its row or the rest of the batch; its row ends up Incoming / owner_id = 0,
    • IsGlobalPartitionSlot agrees with the URIs the topology actually stamps.
  • CoreTests/Bugs/Bug_4822_failed_promotion_releases_the_parked_rows: a slot forward that fails on the second envelope releases only the un-forwarded envelopes, at the companion queue address.

Both bug tests failed before the fix with exactly the reported state (rows Incoming, owned by the live node). Existing GH-3413 / GH-4645 / GH-4700 / GH-4776 / partitioning / scheduled suites pass, and wolverine.slnx builds clean with -c Release -f net9.0.

Known limitations (not addressed here)

…sperFxGH-4822)

A scheduled message to a global partition parks in the inbox at the
external slot's address. A node that previously owned the slot keeps its
stopped listening agent registered, so the scheduled poller handed the
promoted envelopes to a listener without a receiver. That threw after
the rows were committed as Incoming and owned by a live node, stranding
them and every other destination in the same batch.

EnqueueDirectlyAsync now settles slot ownership for the slot address
itself and forwards to the slot when this node does not own it. Each
destination group is also isolated: a failing group is logged and its
envelopes are released back to any node so recovery retries them.
…GH-4822)

When a promoted group fails part-way, envelopes that were already
forwarded have had their inbox rows retired. Releasing them as well let
recovery run them a second time, so only the envelopes not yet handed
over are released now.

The release also matched on the envelope's live destination, which the
partition slot forward rewrites to the slot before sending. For a row
parked at the companion local queue that matched nothing and left it
owned by this live node. Stand-in envelopes now carry the parked address.
@jeremydmiller

Copy link
Copy Markdown
Member

@BlackChepo Thank you! That's for a JasperFx client, so now I feel a little bad, but I'm grateful for help like this!

@jeremydmiller
jeremydmiller merged commit 6244beb into JasperFx:main Oct 5, 2026
45 checks passed
pltknttn pushed a commit to pltknttn/wolverine that referenced this pull request Oct 6, 2026
…ng its inbox row

A scheduled envelope that comes due on a node which does not listen to its
destination is forwarded through a sending agent. JasperFxGH-4645 (JasperFx#4656) made that path
delete the forwarded envelope's inbox row, which fixed the orphan. It left the send
first and stored nothing as outgoing, which is still two losses in 6.46.0:

- A poll by the owning node between the send and the delete. The listener's
  anti-duplicate probe (JasperFxGH-4316) finds the inbox row beside the new queue row and
  deletes the queue row, with no status filter. The message is not handled and not
  dead lettered.
- A failed first send. EnqueueOutgoingAsync posts to the agent's RetryBlock and
  stores nothing, and RetryBlock.PostAsync awaits the first attempt and queues the
  item on an exception. The inbox row is deleted anyway, so until a retry succeeds
  no table holds the message and a process that stops in that interval loses it.

For a DURABLE agent the order is now store the outgoing row, retire the inbox row,
then send:

- the inbox row is gone before any queue row exists, so the probe cannot match it;
- a failed send is the ordinary outbox case the agent already retries from, and
  outbox recovery picks it up if this node stops;
- a stop between the store and the delete leaves both rows, which can deliver the
  message twice and cannot lose it.

An agent with no outbox behind it keeps the original order: there is nothing to move
the message into, so reordering would only widen the window in which no table holds
it at all.

ISendingAgent.TryStoreOutgoingAsync is a DEFAULT interface member, so an external
implementation keeps compiling and keeps today's behaviour. Only DurableSendingAgent
overrides it, reusing its own DurableWriteRetry store path (JasperFxGH-4662) rather than a
second copy; SendingAgent.setDefaults became protected so the outbox row carries the
same Status/OwnerId/ReplyUri it would have had.

Both call sites are fixed, not just the reported one: forwardToPartitionSlotAsync
(JasperFxGH-4700) ran the identical send-then-delete pair, and shares the new helper. Its
parked-address handling is unchanged -- the delete still names the address the row
was written under, because received_at is part of the identity.

Reported with a three-scenario reproduction (audit triggers recording every queue
insert, queue delete and inbox delete, with the statement that ran; the loss observed
in 4 of 4 runs and the no-table window in 2 of 2) by @smoqmilus, who also identified
the fix. JasperFx#4823 landed first, as the issue said it needed to.

Gates: wolverine.slnx Release -f net9.0 0 errors / 0 warnings; CoreTests 3289 green,
including Bug_4700 and Bug_4822, whose stub transports are non-durable and so still
exercise the original order.

Co-Authored-By: smoqmilus <smoqmilus@users.noreply.github.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VDUrBeB4tTnKj4AExCS1nj
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Global partitioning: scheduled messages cascaded from a handler are never handled when a node joins before they are due

2 participants