Skip to content

fix(dart-worker): cooperative cancellation via isTaskCancelled (#66) - #68

Merged
vietnguyentuan2019 merged 9 commits into
mainfrom
fix/issue-66-dart-worker-cancellation
Sep 10, 2026
Merged

vietnguyentuan2019 merged 9 commits into
mainfrom
fix/issue-66-dart-worker-cancellation

Conversation

@vietnguyentuan2019

@vietnguyentuan2019 vietnguyentuan2019 commented Sep 9, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #67 (tracking issue, filed from discussion #66).

What

Cancelling a task did not interrupt a running DartWorker callback, and there was no way for a running callback to even find out it had been cancelled. Dart has no API to preemptively abort an already-running Future, so this can only be cooperative: NativeWorkManager.isTaskCancelled(taskId), polled between chunks of work.

'longSync': (input) async {
  final taskId = input?['__taskId'] as String?;
  for (var i = 1; i <= 100; i++) {
    if (taskId != null && await NativeWorkManager.isTaskCancelled(taskId)) {
      return false; // bail out — do not keep working
    }
    await processChunk(i);
  }
  return true;
},

Wired end-to-end via a new DartTaskCancellationRegistry (Kotlin + Swift):

  • Android: FlutterEngineManager.executeDartCallback marks the registry the moment CoroutineWorker cancellation is observed.
  • iOS foreground/simulator: every activeTasks[taskId]?.cancel() call site (cancel/cancelAll/cancelByTag, notification-driven cancel) now also marks the registry.
  • iOS true background: BGTaskSchedulerManager's expiration handler now marks the registry and actually cancels the running Task — previously it only called activeWorker?.stop() (a no-op for DartCallbackWorker), so an expired task's work kept running past the task's own completion.

Two bugs found along the way (same branch, per the CLAUDE.md second-pass rule)

  1. Android: an external CancellationException in executeDartCallback bypassed dispose()/scheduleDisposalCheck() entirely — a cancelled DartWorker task either leaked the ~50MB headless engine forever, or disposed it while an orphaned callback was still running.
  2. iOS: BGTask expiration never cancelled the Task backing execution, only called a no-op stop() for Dart workers.

Tests

  • test/unit/issue_66_dart_worker_cancellation_test.dart — Dart consumer test (mocked channel): forwards taskId, returns native reply, empty-taskId guard, no-handler-doesn't-throw guard.
  • android/.../engine/DartTaskCancellationRegistryTest.kt — registry correctness including concurrent access from 16 threads.
  • example/integration_test/device_integration_test.dart — issue_66: cancelling a running DartWorker task is observable via isTaskCancelled and issue_69: cancelling a background-session download actually aborts the transfer (iOS). Device-verified: Pixel 6 Pro (Android), iOS simulator, and a 3-device Firebase Test Lab spot-check (Samsung Galaxy S24, OPPO A79 5G, Google virtual baseline).
  • channel_method_parity_test.dart — exemption added for isTaskCancelled (same shape as reportProgress/dartReady).

Verified

  • flutter analyze: clean.
  • Full Dart test suite: green (1990 tests, 1 pre-existing skip).
  • Android Gradle unit tests: green (:native_workmanager:testDebugUnitTest).
  • iOS: the Swift changes compile via a real flutter build ios --simulator --no-codesign (a standalone xcodebuild/swift build of the bare SPM package fails on two pre-existing, unrelated issues — BGTask unavailable on macOS and the Flutter module unresolved outside a real Flutter build context — reproduced identically on unmodified main, so not caused by this change).

Pre-publish review update (2026-09-10)

Went through this branch end-to-end before a 1.7.0 release: docs, tests, sub-package
versions, demo app, pana.

  • Real CI bug found and fixed, unrelated to How to stop Dart execution when cancelling a running task? #66/iOS: cancelling a background-session HTTP download/upload never actually stops the transfer #69 in cause but blocking this PR:
    flutter-version-file: .flutter-version was never a format subosito/flutter-action
    actually supports (only pubspec.yaml/.fvmrc/.fvm/fvm_config.json) — it silently
    fell back to the floating channel: stable default, so CI ran whatever Flutter was
    newest that day and reformatted two files differently than local, turning this PR red
    for reasons that had nothing to do with its own diff. Switched to .fvmrc. CI is green
    as of the current HEAD.
  • Extended channel_method_parity_test.dart to actually check the
    dev.brewkits/dart_worker_channel methods (isTaskCancelled included) that were
    previously just exempted with a comment and no real check — red-then-green verified
    against a real regression.
  • native_workmanager_gen had no analysis_options.yaml of its own and was silently
    analyzing against the root plugin's unresolvable flutter_lints config; fixed, which
    surfaced (and fixed) two small pre-existing lint issues in its test suite.
  • Added a demo card + callback for isTaskCancelled to the example app — the API had
    zero demo coverage before. Device-verified the underlying mechanism via the
    Cancellation test group above; app boots clean on iOS simulator with no crash.
  • Version bump to 1.7.0 (both packages, lockstep), CHANGELOG [Unreleased] →
    [1.7.0], all doc install snippets updated from ^1.6.1.
  • pana: 160/160 on both native_workmanager and native_workmanager_gen.
  • Found, not fixed here: a pre-existing, cross-platform bug where, when a
    DartWorker's timeoutMs fires, no terminal event reaches NativeWorkManager.events
    — and the late result from the abandoned callback, once it does finish, is dropped
    too. Worker completions with no timeout race are unaffected — the stress test's 12
    cases show a perfect correlation between "no event ever arrived" and "timeoutMs beat
    delayMs". Reproduced on main at v1.6.1 via a worktree comparison — unrelated to and
    not blocking this PR, but real. Full write-up in CHANGELOG.md under "Known Issues".
    Also masks itself in the existing issue_30 stress test in stress_and_system_test.dart,
    which treats "no event ever arrived" the same as "correctly failed" — worth a look
    alongside the underlying bug.

Not fixed here

executeDartWorkerViaMethodChannel's foreground branch is a raw withCheckedContinuation around channel.invokeMethod — cancellation is only observed on the callback's next isTaskCancelled() poll, same as every other path. This is inherent to Dart's cooperative-only cancellation model, not a bug on top of it.

Claude-Session: https://claude.ai/code/session_011aaZsf2gDsriP8hfMXbgG8

Cancelling a task via NativeWorkManager.cancel()/cancelAll(), or the OS
reclaiming background time, does not interrupt a running DartWorker
callback — Dart has no API to preemptively abort a Future that is
already executing. A running callback had no way to even find out it
had been cancelled.

Adds NativeWorkManager.isTaskCancelled(taskId): a callback doing
long-running work polls it cooperatively between chunks and returns
early once it turns true. Wired end-to-end on both platforms via a new
DartTaskCancellationRegistry (Kotlin + Swift):

- Android: FlutterEngineManager.executeDartCallback marks the registry
  the moment CoroutineWorker cancellation is observed.
- iOS foreground/simulator: every activeTasks[taskId]?.cancel() call
  site (cancel/cancelAll/cancelByTag, notification-driven cancel) now
  also marks the registry.
- iOS true-background: BGTaskSchedulerManager's expirationHandler now
  marks the registry AND actually cancels the running Task — previously
  it only called activeWorker?.stop() (a no-op for DartCallbackWorker),
  so an expired task's work kept running past the task's own
  completion.

Along the way, fixed a real bug found while tracing the Android
cancellation path: external CancellationException in
executeDartCallback bypassed dispose()/scheduleDisposalCheck() entirely
(caught only by the outer rethrow-and-propagate catch), so a cancelled
DartWorker task either leaked the ~50MB headless engine forever, or —
if another task's idle timer was already armed — could dispose the
engine while the orphaned callback was still running.

Tests:
- test/unit/issue_66_dart_worker_cancellation_test.dart (Dart consumer
  test, mocked channel — round-trip + empty-taskId + no-handler guards)
- android/.../engine/DartTaskCancellationRegistryTest.kt (registry incl.
  concurrent access from many threads)
- example/integration_test/device_integration_test.dart: issue_66
  cancel-while-running test (not yet run on a physical device — no
  device/emulator was available in this session)
- channel_method_parity_test.dart exemption added for isTaskCancelled

Verified: flutter analyze clean, full Dart test suite green (1990
tests), Android Gradle unit tests green, iOS Swift changes compile via
a real `flutter build ios --simulator` (a standalone `xcodebuild`/`swift
build` of the SPM package fails on unrelated pre-existing issues:
BGTask unavailable on macOS, Flutter module unresolved outside a real
Flutter build context — reproduced identically on unmodified main).

Claude-Session: https://claude.ai/code/session_011aaZsf2gDsriP8hfMXbgG8
@vietnguyentuan2019

Copy link
Copy Markdown
Contributor Author

Update: ran the device tests on an iOS 26.5 simulator (iPhone 17) — no physical device or Android emulator was available in this session, but this covers the foreground DartWorker cancellation path end-to-end.

Cancellation issue_66: cancelling a running DartWorker task is observable via isTaskCancelled  PASSED
[DartWorker] dit_cancel_poll: observed cancellation at iteration 4

Cancelled ~600ms in (task polls every 200ms), callback observed it and returned at iteration 4 instead of running all 50 — confirms isTaskCancelled() actually short-circuits execution, not just that cancel() doesn't crash.

Also ran the adjacent groups that touch the files this PR changed, to check for regressions:

  • Cancellation (all 3, including the two pre-existing tests) — pass
  • Issue #38/#39 – DartWorker progress and persisted status — pass
  • DartWorker Constraint and Delay Enforcement — pass
  • Task Chains (including "Chain cancel after first step stops remaining steps") — pass

Still open: Android side (FlutterEngineManager.kt / DartCallbackWorker.kt changes) is only unit-tested (Gradle, DartTaskCancellationRegistryTest), not device-verified — no emulator/device available here. Someone with an Android device handy should run issue_66 from device_integration_test.dart before merging.

Claude-Session: https://claude.ai/code/session_011aaZsf2gDsriP8hfMXbgG8

…d the transfer (#69)

Found auditing for bugs similar to #66/#67 — same "cancel signal
doesn't reach the actual operation" shape, this time on native (non-
Dart) workers.

HttpDownloadWorker.downloadWithBackgroundSession and
HttpUploadWorker.uploadWithBackgroundSession (both gated by
useBackgroundSession: true, iOS-only) registered their
URLSessionDownloadTask/URLSessionUploadTask with BackgroundSessionManager
under a throwaway random id ("download-<uuid>" / "upload-<uuid>")
instead of the real plugin task id. NativeWorkManager.cancel()/
cancelAll()/cancelByTag() all call BackgroundSessionManager.cancel(taskId:
<real id>) to reach these — which always missed, since the transfer
was registered under a different, never-exposed id. The download/
upload just kept running regardless of cancel().

Fix: both functions now accept the real task id (extracted from
__taskId in the input JSON, same pattern HttpDownloadWorker.doWork
already used for progress reporting) and register with
BackgroundSessionManager under that id. Falls back to a random id only
when none is available. No change needed to BackgroundSessionManager
itself — the existing cancel(taskId:) calls already wired into every
cancel path now find the right task.

Verified red-then-green on a real iOS simulator: the new device test
(issue_69) reproduces the bug on the pre-fix code — cancelling 1s into
a request to a 6s-delayed endpoint still writes the destination file —
then passes with the fix. Also reran the full Cancellation and All
Workers groups on-device: no regressions.

Android has no equivalent code path (WorkManager itself is the
background-survival mechanism there, no separate background-session
concept), so no Android-side fix is needed.

Claude-Session: https://claude.ai/code/session_011aaZsf2gDsriP8hfMXbgG8
@vietnguyentuan2019

Copy link
Copy Markdown
Contributor Author

Follow-up: while auditing for bugs similar to this one, found and fixed a second, unrelated instance of the same shape — cancelling a useBackgroundSession: true HttpDownloadWorker/HttpUploadWorker on iOS never actually stopped the transfer (both registered their URLSessionTask under a throwaway random id instead of the real task id, so BackgroundSessionManager.cancel(taskId:) always missed). Filed as #69, fixed in the latest commit here.

Verified red-then-green on a real iOS simulator (not just "it passes now" — confirmed the new issue_69 test actually fails on the pre-fix code first):

# before fix:
Expected: false
  Actual: <true>
issue_69: a cancelled background-session download must not still write its destination file

# after fix:
Cancellation issue_69: cancelling a background-session download actually aborts the transfer (iOS)  PASSED

Also reran the full Cancellation and All Workers groups on-device after this change — 16/16 and 4/4 pass, no regressions in the untouched HTTP download/upload/foreground paths.

Claude-Session: https://claude.ai/code/session_011aaZsf2gDsriP8hfMXbgG8

…ure was a no-op on real hardware

Device-verified on a Pixel 6 Pro (no emulator/device was available in
the prior commits on this branch — this is the actual on-device
verification the PR called out as still missing).

The registry was cleared from a `finally` around executeDartCallback's
own coroutine, which runs the instant THIS coroutine unwinds after
observing external cancellation. But the whole premise of the feature
is that the orphaned Dart callback keeps running and polling
isTaskCancelled() for a while *after* that — Dart has no way to be
interrupted. So the sequence was: markCancelled() → (this coroutine
returns, finally fires) → clear() → every subsequent poll from the
still-running callback sees false again. Confirmed via logcat: the
poll immediately after the mark still returned false, and the
counterFile kept climbing 10 more iterations in the following 2s —
cancellation was never actually observed by the callback on a real
device, even though every unit test (channel mocked, no real
coroutine-cancellation race) stayed green.

Every unit test in the previous commits mocked the channel and never
exercised the real timing race between "this coroutine's cancellation
unwinds" and "the orphaned Dart isolate's next poll" — so nothing
caught this until the device test that was supposed to prove the fix
actually ran on real hardware.

Fix: clear the registry from the three MethodChannel.Result overrides
on the executeCallback invocation (success/error/notImplemented) —
which fire whenever that specific invocation truly finishes, including
when it finishes late, well after this coroutine already unwound from
cancellation — not from a finally tied to this coroutine's own
lifetime.

Verified on the Pixel 6 Pro: issue_66 now passes, and logcat shows the
first poll after markCancelled() correctly returns true. Reran
Cancellation, Issue #38/#39, Task Chains, and All Workers device-test
groups (no regressions) plus the full Gradle unit test suite (green).

Claude-Session: https://claude.ai/code/session_011aaZsf2gDsriP8hfMXbgG8
@vietnguyentuan2019

Copy link
Copy Markdown
Contributor Author

Important update — the Android side did not actually work, and this is exactly the gap I flagged as unverified. A Pixel 6 Pro became available and running issue_66 on it immediately failed:

Expected: <10>
  Actual: <20>
issue_66: iteration count must not still be climbing 2s later

logcat showed the real bug: DartTaskCancellationRegistry.markCancelled(taskId) fired correctly on external cancellation, but the registry entry was cleared again in the very same breath — from a finally around executeDartCallback's own coroutine, which unwinds immediately after observing the cancellation. The whole premise of the feature is that the orphaned Dart callback keeps running (and polling) for a while after that, since Dart can't be interrupted — so the very next poll from the still-running callback already saw false again. isTaskCancelled() was a no-op on real hardware the entire time this PR existed. Every unit test stayed green because they mock the channel and never hit this real coroutine-cancellation timing race.

Fixed by moving the clear() to the three MethodChannel.Result overrides on the executeCallback invocation (success/error/notImplemented) — those fire when that specific invocation truly finishes, including finishing late, well after the cancelling coroutine already unwound.

Re-verified on the Pixel 6 Pro:

channel isTaskCancelled poll taskId=... -> false   (before markCancelled)
markCancelled called for taskId=...
channel isTaskCancelled poll taskId=... -> true    (immediately after — fixed)

issue_66 now passes on-device. Reran Cancellation, Issue #38/#39, Task Chains, and All Workers device-test groups on the same Pixel (no regressions) plus the full Gradle unit test suite.

This is now device-verified on both platforms (iOS simulator + Pixel 6 Pro). The lesson for next time is in the memory file for this project already — a device test that mocks the channel doesn't exercise the real cancellation race, so "green on device" still needs a real cancel to actually happen mid-callback, not just a mocked reply.

Claude-Session: https://claude.ai/code/session_011aaZsf2gDsriP8hfMXbgG8

…#69

Adds the pieces needed to run integration_test files as Android
instrumentation tests on Firebase Test Lab, which the repo didn't have
(the weekly firebase-benchmark.yml workflow has never actually run —
its own comment notes the required secrets were never configured):

- example/android/app/src/androidTest/.../MainActivityTest.kt — the
  standard Flutter FlutterTestRunner wrapper (from the flutter/packages
  integration_test example), drives whichever integration_test Dart
  entrypoint was baked into the app APK.
- testInstrumentationRunner + androidx.test:runner/rules deps in
  app/build.gradle.kts, pinned to 1.2.0 to match what the
  integration_test plugin's own Android dependency already resolves to
  on the main runtime classpath (Gradle's consistent-resolution check
  fails otherwise).
- integration_test/issue_66_69_ftl_test.dart — a small, self-contained
  entrypoint covering just the two cancellation regressions (issue #66,
  issue #69), NOT the full device_integration_test.dart suite. That
  suite is long-running and has known hang risks on periodic-trigger
  tests (see project memory) — unsuitable for a quick, cheap Test Lab
  spot-check, which is what this is for.
- scripts/firebase-ftl-cancellation.sh — builds both APKs, submits to
  Test Lab, and polls for the result via the Testing API's REST
  endpoint directly (this environment's gcloud CLI has no `matrices
  describe` subcommand under `firebase test android`).

Claude-Session: https://claude.ai/code/session_011aaZsf2gDsriP8hfMXbgG8
@vietnguyentuan2019

Copy link
Copy Markdown
Contributor Author

Ran a targeted Firebase Test Lab spot-check on top of the Pixel 6 Pro + iOS simulator verification, to catch any OEM-specific timing difference the two local devices couldn't show. Built a minimal instrumentation entrypoint (integration_test/issue_66_69_ftl_test.dart — just the two regressions, not the full 100+ test suite, which has known hang risks on periodic triggers) and added the Android instrumentation wrapper this repo didn't have yet (MainActivityTest.kt + gradle config, see scripts/firebase-ftl-cancellation.sh).

3 real device/OS combos, chosen deliberately for OEM battery-deferral behavior (the class of device most likely to time this differently than a Pixel):

Device API Result
Google MediumPhone (virtual, stock) 34 ✅ 2/2 pass
Samsung Galaxy S24 36 ✅ 2/2 pass
OPPO A79 5G 34 ✅ 2/2 pass

0 failures, 0 errors across all three. issue_66 completed in 97–338ms on every device (well inside the ~600ms-to-cancel / 2s-wait window), consistent with the Pixel 6 Pro run.

This is now verified on 5 distinct real/virtual environments across the fix's lifetime: Pixel 6 Pro (where the registry-clear-timing bug was actually caught), an iOS simulator, and these 3 Firebase Test Lab devices. Ready for review.

Claude-Session: https://claude.ai/code/session_011aaZsf2gDsriP8hfMXbgG8

subosito/flutter-action's `channel` input defaults to 'stable' even when
the key is omitted from the workflow yaml entirely, and a non-empty
channel wins over flutter-version-file. The existing fix (just omitting
`channel:`) left that default in place, so .flutter-version was still
being silently ignored on every job — confirmed today: the pin says
3.41.9, but every job in this PR actually ran 3.47.2, and its `dart
format` disagreed with local (3.41.9), turning an unrelated PR red.

Fix: explicitly set `channel: ''` alongside flutter-version-file so
there is no non-empty default left for the action to prefer.
…ted format

subosito/flutter-action's flutter-version-file only parses pubspec.yaml,
.fvmrc, or .fvm/fvm_config.json. This repo pointed it at a plain-text
.flutter-version file, which the action can't parse in any supported
format — it silently fell back to the channel default ('stable', which
is non-empty even when the key is omitted from this yaml), so every job
ran whatever Flutter happened to be newest that day regardless of the
pin. Confirmed today: pin said 3.41.9, CI ran 3.47.2, then 3.47.3 the
next run, each formatting lib/src/native_work_manager.dart and
test/unit/issue_66_dart_worker_cancellation_test.dart differently than
local — the actual cause of this PR's unrelated CI failures. `channel`
is also explicitly emptied as defense in depth, in case both ever
resolve non-empty again. The project already manages Flutter locally
via `fvm`, so .fvmrc is also the natural single source of truth for
local `fvm use`.

Also:
- native_workmanager_gen had no analysis_options.yaml of its own, so
  `dart analyze` walked up to the root plugin's, which includes
  `package:flutter_lints/flutter.yaml` — unresolvable against a pure-Dart
  package that only depends on `lints`. That silently broke analysis for
  the whole package (every run just warned and skipped), hiding an
  unused import and two leading-underscore locals in its test suite.
  Gave it its own analysis_options.yaml (package:lints/recommended.yaml)
  and fixed the newly-surfaced issues.
- Extended test/unit/channel_method_parity_test.dart to cover
  dev.brewkits/dart_worker_channel (dartReady/reportProgress/
  isTaskCancelled on both platforms' handlers), which the existing
  parity test explicitly exempted from any automated check — the exact
  "Dart calls a method no platform registers" defect class this repo has
  shipped five times, just on the channel the guard didn't reach yet.
  Red-then-green verified against a real regression in
  FlutterEngineManager.kt.
… 160/160

- Bump native_workmanager, native_workmanager_gen, and the example app to
  1.7.0 in lockstep. CHANGELOG headers moved from [Unreleased] to
  [1.7.0] - 2026-09-10 in both packages.
- Updated every doc install snippet (README, GETTING_STARTED,
  MIGRATION_GUIDE, MIGRATION_TOOL_README, ANDROID_SETUP) from ^1.6.1 to
  ^1.7.0.
- Added a "1c. Dart Worker — Cooperative Cancellation (Issue #66)" demo
  card to the example app's comprehensive demo page, plus its backing
  `cancellableTaskCallback` in main.dart — the new isTaskCancelled API
  had zero demo coverage before this. Device-verified the underlying
  mechanism via the existing issue_66/issue_69 device_integration_test.dart
  Cancellation group (4/4 passing on iOS simulator).
- CHANGELOG "Known Issues" entry documenting a pre-existing, cross-platform
  DartWorker.timeoutMs event-delivery gap found while auditing the stress
  suite for this release — reproduced on main at v1.6.1 via a worktree
  comparison, so unrelated to and not blocking this release, but real and
  worth its own investigation. See CHANGELOG for the full write-up.
- pana: 160/160 on both native_workmanager and native_workmanager_gen.
The 12-case stress-test log shows a perfect correlation: every case
where the callback finished naturally (no timeout race) delivered its
terminal event; only the cases where timeoutMs actually fired dropped
it. The original wording claimed natural completion could be affected
too — that's not what the evidence shows and overstates the blast
radius. Narrowed to what's actually demonstrated: the timeout-fires
path drops both the immediate failure event and the late result from
the abandoned invocation.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

DartWorker cancellation doesn't stop a running callback (no onTaskStopped / cancellation token)

1 participant