Repository navigation
fix(dart-worker): cooperative cancellation via isTaskCancelled (#66) - #68
Conversation
Cancelling a task via NativeWorkManager.cancel()/cancelAll(), or the OS reclaiming background time, does not interrupt a running DartWorker callback — Dart has no API to preemptively abort a Future that is already executing. A running callback had no way to even find out it had been cancelled. Adds NativeWorkManager.isTaskCancelled(taskId): a callback doing long-running work polls it cooperatively between chunks and returns early once it turns true. Wired end-to-end on both platforms via a new DartTaskCancellationRegistry (Kotlin + Swift): - Android: FlutterEngineManager.executeDartCallback marks the registry the moment CoroutineWorker cancellation is observed. - iOS foreground/simulator: every activeTasks[taskId]?.cancel() call site (cancel/cancelAll/cancelByTag, notification-driven cancel) now also marks the registry. - iOS true-background: BGTaskSchedulerManager's expirationHandler now marks the registry AND actually cancels the running Task — previously it only called activeWorker?.stop() (a no-op for DartCallbackWorker), so an expired task's work kept running past the task's own completion. Along the way, fixed a real bug found while tracing the Android cancellation path: external CancellationException in executeDartCallback bypassed dispose()/scheduleDisposalCheck() entirely (caught only by the outer rethrow-and-propagate catch), so a cancelled DartWorker task either leaked the ~50MB headless engine forever, or — if another task's idle timer was already armed — could dispose the engine while the orphaned callback was still running. Tests: - test/unit/issue_66_dart_worker_cancellation_test.dart (Dart consumer test, mocked channel — round-trip + empty-taskId + no-handler guards) - android/.../engine/DartTaskCancellationRegistryTest.kt (registry incl. concurrent access from many threads) - example/integration_test/device_integration_test.dart: issue_66 cancel-while-running test (not yet run on a physical device — no device/emulator was available in this session) - channel_method_parity_test.dart exemption added for isTaskCancelled Verified: flutter analyze clean, full Dart test suite green (1990 tests), Android Gradle unit tests green, iOS Swift changes compile via a real `flutter build ios --simulator` (a standalone `xcodebuild`/`swift build` of the SPM package fails on unrelated pre-existing issues: BGTask unavailable on macOS, Flutter module unresolved outside a real Flutter build context — reproduced identically on unmodified main). Claude-Session: https://claude.ai/code/session_011aaZsf2gDsriP8hfMXbgG8
|
Update: ran the device tests on an iOS 26.5 simulator (iPhone 17) — no physical device or Android emulator was available in this session, but this covers the foreground DartWorker cancellation path end-to-end. Cancelled ~600ms in (task polls every 200ms), callback observed it and returned at iteration 4 instead of running all 50 — confirms Also ran the adjacent groups that touch the files this PR changed, to check for regressions:
Still open: Android side ( Claude-Session: https://claude.ai/code/session_011aaZsf2gDsriP8hfMXbgG8 |
…d the transfer (#69) Found auditing for bugs similar to #66/#67 — same "cancel signal doesn't reach the actual operation" shape, this time on native (non- Dart) workers. HttpDownloadWorker.downloadWithBackgroundSession and HttpUploadWorker.uploadWithBackgroundSession (both gated by useBackgroundSession: true, iOS-only) registered their URLSessionDownloadTask/URLSessionUploadTask with BackgroundSessionManager under a throwaway random id ("download-<uuid>" / "upload-<uuid>") instead of the real plugin task id. NativeWorkManager.cancel()/ cancelAll()/cancelByTag() all call BackgroundSessionManager.cancel(taskId: <real id>) to reach these — which always missed, since the transfer was registered under a different, never-exposed id. The download/ upload just kept running regardless of cancel(). Fix: both functions now accept the real task id (extracted from __taskId in the input JSON, same pattern HttpDownloadWorker.doWork already used for progress reporting) and register with BackgroundSessionManager under that id. Falls back to a random id only when none is available. No change needed to BackgroundSessionManager itself — the existing cancel(taskId:) calls already wired into every cancel path now find the right task. Verified red-then-green on a real iOS simulator: the new device test (issue_69) reproduces the bug on the pre-fix code — cancelling 1s into a request to a 6s-delayed endpoint still writes the destination file — then passes with the fix. Also reran the full Cancellation and All Workers groups on-device: no regressions. Android has no equivalent code path (WorkManager itself is the background-survival mechanism there, no separate background-session concept), so no Android-side fix is needed. Claude-Session: https://claude.ai/code/session_011aaZsf2gDsriP8hfMXbgG8
|
Follow-up: while auditing for bugs similar to this one, found and fixed a second, unrelated instance of the same shape — cancelling a Verified red-then-green on a real iOS simulator (not just "it passes now" — confirmed the new Also reran the full Claude-Session: https://claude.ai/code/session_011aaZsf2gDsriP8hfMXbgG8 |
…ure was a no-op on real hardware Device-verified on a Pixel 6 Pro (no emulator/device was available in the prior commits on this branch — this is the actual on-device verification the PR called out as still missing). The registry was cleared from a `finally` around executeDartCallback's own coroutine, which runs the instant THIS coroutine unwinds after observing external cancellation. But the whole premise of the feature is that the orphaned Dart callback keeps running and polling isTaskCancelled() for a while *after* that — Dart has no way to be interrupted. So the sequence was: markCancelled() → (this coroutine returns, finally fires) → clear() → every subsequent poll from the still-running callback sees false again. Confirmed via logcat: the poll immediately after the mark still returned false, and the counterFile kept climbing 10 more iterations in the following 2s — cancellation was never actually observed by the callback on a real device, even though every unit test (channel mocked, no real coroutine-cancellation race) stayed green. Every unit test in the previous commits mocked the channel and never exercised the real timing race between "this coroutine's cancellation unwinds" and "the orphaned Dart isolate's next poll" — so nothing caught this until the device test that was supposed to prove the fix actually ran on real hardware. Fix: clear the registry from the three MethodChannel.Result overrides on the executeCallback invocation (success/error/notImplemented) — which fire whenever that specific invocation truly finishes, including when it finishes late, well after this coroutine already unwound from cancellation — not from a finally tied to this coroutine's own lifetime. Verified on the Pixel 6 Pro: issue_66 now passes, and logcat shows the first poll after markCancelled() correctly returns true. Reran Cancellation, Issue #38/#39, Task Chains, and All Workers device-test groups (no regressions) plus the full Gradle unit test suite (green). Claude-Session: https://claude.ai/code/session_011aaZsf2gDsriP8hfMXbgG8
|
Important update — the Android side did not actually work, and this is exactly the gap I flagged as unverified. A Pixel 6 Pro became available and running logcat showed the real bug: Fixed by moving the Re-verified on the Pixel 6 Pro:
This is now device-verified on both platforms (iOS simulator + Pixel 6 Pro). The lesson for next time is in the memory file for this project already — a device test that mocks the channel doesn't exercise the real cancellation race, so "green on device" still needs a real cancel to actually happen mid-callback, not just a mocked reply. Claude-Session: https://claude.ai/code/session_011aaZsf2gDsriP8hfMXbgG8 |
…#69 Adds the pieces needed to run integration_test files as Android instrumentation tests on Firebase Test Lab, which the repo didn't have (the weekly firebase-benchmark.yml workflow has never actually run — its own comment notes the required secrets were never configured): - example/android/app/src/androidTest/.../MainActivityTest.kt — the standard Flutter FlutterTestRunner wrapper (from the flutter/packages integration_test example), drives whichever integration_test Dart entrypoint was baked into the app APK. - testInstrumentationRunner + androidx.test:runner/rules deps in app/build.gradle.kts, pinned to 1.2.0 to match what the integration_test plugin's own Android dependency already resolves to on the main runtime classpath (Gradle's consistent-resolution check fails otherwise). - integration_test/issue_66_69_ftl_test.dart — a small, self-contained entrypoint covering just the two cancellation regressions (issue #66, issue #69), NOT the full device_integration_test.dart suite. That suite is long-running and has known hang risks on periodic-trigger tests (see project memory) — unsuitable for a quick, cheap Test Lab spot-check, which is what this is for. - scripts/firebase-ftl-cancellation.sh — builds both APKs, submits to Test Lab, and polls for the result via the Testing API's REST endpoint directly (this environment's gcloud CLI has no `matrices describe` subcommand under `firebase test android`). Claude-Session: https://claude.ai/code/session_011aaZsf2gDsriP8hfMXbgG8
|
Ran a targeted Firebase Test Lab spot-check on top of the Pixel 6 Pro + iOS simulator verification, to catch any OEM-specific timing difference the two local devices couldn't show. Built a minimal instrumentation entrypoint ( 3 real device/OS combos, chosen deliberately for OEM battery-deferral behavior (the class of device most likely to time this differently than a Pixel):
0 failures, 0 errors across all three. This is now verified on 5 distinct real/virtual environments across the fix's lifetime: Pixel 6 Pro (where the registry-clear-timing bug was actually caught), an iOS simulator, and these 3 Firebase Test Lab devices. Ready for review. Claude-Session: https://claude.ai/code/session_011aaZsf2gDsriP8hfMXbgG8 |
subosito/flutter-action's `channel` input defaults to 'stable' even when the key is omitted from the workflow yaml entirely, and a non-empty channel wins over flutter-version-file. The existing fix (just omitting `channel:`) left that default in place, so .flutter-version was still being silently ignored on every job — confirmed today: the pin says 3.41.9, but every job in this PR actually ran 3.47.2, and its `dart format` disagreed with local (3.41.9), turning an unrelated PR red. Fix: explicitly set `channel: ''` alongside flutter-version-file so there is no non-empty default left for the action to prefer.
…ted format
subosito/flutter-action's flutter-version-file only parses pubspec.yaml,
.fvmrc, or .fvm/fvm_config.json. This repo pointed it at a plain-text
.flutter-version file, which the action can't parse in any supported
format — it silently fell back to the channel default ('stable', which
is non-empty even when the key is omitted from this yaml), so every job
ran whatever Flutter happened to be newest that day regardless of the
pin. Confirmed today: pin said 3.41.9, CI ran 3.47.2, then 3.47.3 the
next run, each formatting lib/src/native_work_manager.dart and
test/unit/issue_66_dart_worker_cancellation_test.dart differently than
local — the actual cause of this PR's unrelated CI failures. `channel`
is also explicitly emptied as defense in depth, in case both ever
resolve non-empty again. The project already manages Flutter locally
via `fvm`, so .fvmrc is also the natural single source of truth for
local `fvm use`.
Also:
- native_workmanager_gen had no analysis_options.yaml of its own, so
`dart analyze` walked up to the root plugin's, which includes
`package:flutter_lints/flutter.yaml` — unresolvable against a pure-Dart
package that only depends on `lints`. That silently broke analysis for
the whole package (every run just warned and skipped), hiding an
unused import and two leading-underscore locals in its test suite.
Gave it its own analysis_options.yaml (package:lints/recommended.yaml)
and fixed the newly-surfaced issues.
- Extended test/unit/channel_method_parity_test.dart to cover
dev.brewkits/dart_worker_channel (dartReady/reportProgress/
isTaskCancelled on both platforms' handlers), which the existing
parity test explicitly exempted from any automated check — the exact
"Dart calls a method no platform registers" defect class this repo has
shipped five times, just on the channel the guard didn't reach yet.
Red-then-green verified against a real regression in
FlutterEngineManager.kt.
… 160/160 - Bump native_workmanager, native_workmanager_gen, and the example app to 1.7.0 in lockstep. CHANGELOG headers moved from [Unreleased] to [1.7.0] - 2026-09-10 in both packages. - Updated every doc install snippet (README, GETTING_STARTED, MIGRATION_GUIDE, MIGRATION_TOOL_README, ANDROID_SETUP) from ^1.6.1 to ^1.7.0. - Added a "1c. Dart Worker — Cooperative Cancellation (Issue #66)" demo card to the example app's comprehensive demo page, plus its backing `cancellableTaskCallback` in main.dart — the new isTaskCancelled API had zero demo coverage before this. Device-verified the underlying mechanism via the existing issue_66/issue_69 device_integration_test.dart Cancellation group (4/4 passing on iOS simulator). - CHANGELOG "Known Issues" entry documenting a pre-existing, cross-platform DartWorker.timeoutMs event-delivery gap found while auditing the stress suite for this release — reproduced on main at v1.6.1 via a worktree comparison, so unrelated to and not blocking this release, but real and worth its own investigation. See CHANGELOG for the full write-up. - pana: 160/160 on both native_workmanager and native_workmanager_gen.
The 12-case stress-test log shows a perfect correlation: every case where the callback finished naturally (no timeout race) delivered its terminal event; only the cases where timeoutMs actually fired dropped it. The original wording claimed natural completion could be affected too — that's not what the evidence shows and overstates the blast radius. Narrowed to what's actually demonstrated: the timeout-fires path drops both the immediate failure event and the late result from the abandoned invocation.
Fixes #67 (tracking issue, filed from discussion #66).
What
Cancelling a task did not interrupt a running
DartWorkercallback, and there was no way for a running callback to even find out it had been cancelled. Dart has no API to preemptively abort an already-runningFuture, so this can only be cooperative:NativeWorkManager.isTaskCancelled(taskId), polled between chunks of work.Wired end-to-end via a new
DartTaskCancellationRegistry(Kotlin + Swift):FlutterEngineManager.executeDartCallbackmarks the registry the momentCoroutineWorkercancellation is observed.activeTasks[taskId]?.cancel()call site (cancel/cancelAll/cancelByTag, notification-driven cancel) now also marks the registry.BGTaskSchedulerManager's expiration handler now marks the registry and actually cancels the runningTask— previously it only calledactiveWorker?.stop()(a no-op forDartCallbackWorker), so an expired task's work kept running past the task's own completion.Two bugs found along the way (same branch, per the CLAUDE.md second-pass rule)
CancellationExceptioninexecuteDartCallbackbypasseddispose()/scheduleDisposalCheck()entirely — a cancelledDartWorkertask either leaked the ~50MB headless engine forever, or disposed it while an orphaned callback was still running.Taskbacking execution, only called a no-opstop()for Dart workers.Tests
test/unit/issue_66_dart_worker_cancellation_test.dart— Dart consumer test (mocked channel): forwards taskId, returns native reply, empty-taskId guard, no-handler-doesn't-throw guard.android/.../engine/DartTaskCancellationRegistryTest.kt— registry correctness including concurrent access from 16 threads.example/integration_test/device_integration_test.dart—issue_66: cancelling a running DartWorker task is observable via isTaskCancelledandissue_69: cancelling a background-session download actually aborts the transfer (iOS). Device-verified: Pixel 6 Pro (Android), iOS simulator, and a 3-device Firebase Test Lab spot-check (Samsung Galaxy S24, OPPO A79 5G, Google virtual baseline).channel_method_parity_test.dart— exemption added forisTaskCancelled(same shape asreportProgress/dartReady).Verified
flutter analyze: clean.:native_workmanager:testDebugUnitTest).flutter build ios --simulator --no-codesign(a standalonexcodebuild/swift buildof the bare SPM package fails on two pre-existing, unrelated issues —BGTaskunavailable on macOS and theFluttermodule unresolved outside a real Flutter build context — reproduced identically on unmodifiedmain, so not caused by this change).Pre-publish review update (2026-09-10)
Went through this branch end-to-end before a 1.7.0 release: docs, tests, sub-package
versions, demo app, pana.
flutter-version-file: .flutter-versionwas never a format subosito/flutter-actionactually supports (only
pubspec.yaml/.fvmrc/.fvm/fvm_config.json) — it silentlyfell back to the floating
channel: stabledefault, so CI ran whatever Flutter wasnewest that day and reformatted two files differently than local, turning this PR red
for reasons that had nothing to do with its own diff. Switched to
.fvmrc. CI is greenas of the current HEAD.
channel_method_parity_test.dartto actually check thedev.brewkits/dart_worker_channelmethods (isTaskCancelledincluded) that werepreviously just exempted with a comment and no real check — red-then-green verified
against a real regression.
native_workmanager_genhad noanalysis_options.yamlof its own and was silentlyanalyzing against the root plugin's unresolvable
flutter_lintsconfig; fixed, whichsurfaced (and fixed) two small pre-existing lint issues in its test suite.
isTaskCancelledto the example app — the API hadzero demo coverage before. Device-verified the underlying mechanism via the
Cancellationtest group above; app boots clean on iOS simulator with no crash.[Unreleased]→[1.7.0], all doc install snippets updated from^1.6.1.native_workmanagerandnative_workmanager_gen.DartWorker'stimeoutMsfires, no terminal event reachesNativeWorkManager.events— and the late result from the abandoned callback, once it does finish, is dropped
too. Worker completions with no timeout race are unaffected — the stress test's 12
cases show a perfect correlation between "no event ever arrived" and "timeoutMs beat
delayMs". Reproduced on
mainat v1.6.1 via a worktree comparison — unrelated to andnot blocking this PR, but real. Full write-up in
CHANGELOG.mdunder "Known Issues".Also masks itself in the existing
issue_30 stresstest instress_and_system_test.dart,which treats "no event ever arrived" the same as "correctly failed" — worth a look
alongside the underlying bug.
Not fixed here
executeDartWorkerViaMethodChannel's foreground branch is a rawwithCheckedContinuationaroundchannel.invokeMethod— cancellation is only observed on the callback's nextisTaskCancelled()poll, same as every other path. This is inherent to Dart's cooperative-only cancellation model, not a bug on top of it.Claude-Session: https://claude.ai/code/session_011aaZsf2gDsriP8hfMXbgG8