Skip to content

Commit e49d29c

Browse files
authored
Merge pull request #26 from CSSFrancis/feat/webgpu-2d-images
feat: WebGPU 2-D image rendering, tiled display, binary pixel transport
2 parents 2186b0b + a2fdb4e commit e49d29c

40 files changed

Lines changed: 5811 additions & 183 deletions

‎LARGE_IMAGE_WEBGPU_PLAN.md‎

Lines changed: 223 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,223 @@
1+
# Large-image WebGPU 2D rendering + binary transport — scoping document
2+
3+
Status: **Phase 0 + core Phase 1/2 DONE, hardware-verified** (2026-07-06). Branch:
4+
`feat/webgpu-2d-images`. Owner: @CSSFrancis
5+
6+
## DONE (verified on an NVIDIA Pascal GPU via the SpyDE Electron consumer)
7+
- **WebGPU 2-D image render path**: a `gpuCanvas` below `plotCanvas`; the normalized
8+
uint8 frame → R8 texture, the 256-entry colormap → a 256×1 RGBA LUT texture, a
9+
fullscreen-quad WGSL fragment shader re-stretches by clim (dmin/dmax uniform) and
10+
samples the LUT. Replaces the 64M-iteration Canvas2D atob+LUT loop with one GPU draw.
11+
- **Reuses the 3D contract**: the `_gpuDevice()` singleton, first-frame-canvas-then-
12+
async-swap, `device.lost` → permanent Canvas2D fallback (extended to 2-D panels).
13+
- **API**: `imshow(..., gpu="auto"|True|False)` → `gpu_mode`; `GPU_IMAGE_THRESHOLD`
14+
(~1 Mpx) auto-gate; `plot.gpu_active` echo via the `gpu_status` event.
15+
- **Fallback intact**: RGB images, sub-threshold images, `gpu=False`, and no-device
16+
all render on Canvas2D (anyplotlib suite green; the DOM keeps `plotCanvas` first so
17+
`querySelector('canvas')` still resolves the image canvas).
18+
- **Correctness (review-hardened)**: the shader is a plain IDENTITY LUT lookup —
19+
`_buildLut32` already bakes clim + scale_mode into the LUT, so the shader must NOT
20+
re-apply the window (an earlier version double-applied clim; correct only at
21+
full-range, wrong for any narrowed contrast/log/symlog — fixed). A zoomed/panned
22+
view falls back to Canvas2D so the base image stays registered with the axes/
23+
overlays (the GPU quad is full-extent). Nearest sampling on both textures matches
24+
Canvas2D's `imageSmoothingEnabled=false` pixel-for-pixel.
25+
- **Zoom/pan v-window fix + headless GPU parity tests (2026-07-09)**: the shader
26+
applied a global `1 - v` flip AFTER interpolating the `[v0,v1]` window, sampling
27+
the MIRRORED window `[1-v1, 1-v0]` — correct only when v0+v1==1 (rest / centred
28+
zoom), so it survived the readback tests but inverted pan-y and detached the
29+
image from markers/overlays on any vertically off-centre view (base AND detail
30+
passes; fixed by interpolating v from v1 at the screen bottom to v0 at the top).
31+
Guarded by `tests/test_plot2d/test_gpu_parity_playwright.py`: GPU-vs-Canvas2D
32+
PNG parity (rest/zoom/pan/markers/widgets/detail-tile) on REAL WebGPU in
33+
headless Chromium — `channel="chromium"` (the full build; the default headless
34+
shell has no `navigator.gpu`) + `--enable-unsafe-webgpu`, skipped when no
35+
adapter. NB `page.screenshot()` DOES capture the WebGPU canvas there; the
36+
"swapchain reads black" caveat below applies to Electron offscreen capture.
37+
- **Verification**: `__apl_gpuReadback` renders the active panel to an OFFSCREEN
38+
texture (the live swapchain reads black under automation) and copies it to CPU.
39+
On a real 4k movie frame: min 0 / max 255 / 96% non-black / correct gray values.
40+
A **narrowed clim [0.3,0.7]** matches the numpy windowed-colormap to meanDiff
41+
0.65/255 (regression guard for the double-apply bug; `tests/_gpu_clim_check.cjs`).
42+
Movie scrub + playback re-upload the texture per frame and stay correct (5/5
43+
scrub, 5 distinct played frames). GPU resources are freed on panel close / figure
44+
dispose / device loss (no leak on repeated large-image open/close).
45+
46+
## DEFERRED (documented; not blockers now)
47+
- **Mipmaps** (smooth downscale-on-zoom): the sampler uses `minFilter:'linear'`
48+
without a mip chain. Marginal here because LOD already caps the uploaded texture
49+
near display size; matters for deep zoom-out. Needs an R8 render-based mip chain.
50+
- **Binary pixel transport** (base64-in-JSON → binary buffer): the Phase-0 headline,
51+
but LOD decimation already cut the shipped payload ~35× (≤1536 px, ~2 MB not
52+
~85 MB) and the GPU shader removed the render cost, so this is now an incremental
53+
transport optimization spanning Jupyter/Pyodide/standalone/Electron — do it with
54+
the multi-environment verification it needs.
55+
56+
Prerequisite reading: `WEBGPU_PLAN.md` (the 3D points/voxels WebGPU path this extends),
57+
`anyplotlib/FIGURE_ESM.md` (the `figure_esm.js` section map), `AGENTS.md` (repo conventions).
58+
Prerequisite reading: `WEBGPU_PLAN.md` (the 3D points/voxels WebGPU path this extends),
59+
`anyplotlib/FIGURE_ESM.md` (the `figure_esm.js` section map), `AGENTS.md` (repo conventions).
60+
61+
This extends the repo's existing hardware-verified `WEBGPU_PLAN.md` (3D instanced points +
62+
voxels) to **2D large images**, deliberately lifting that doc's "No 2D pipeline changes"
63+
non-goal (§2). Motivating consumer: a SpyDE in-situ movie viewer that must scrub/play through
64+
8k×8k image frames smoothly.
65+
66+
## 1. Goal
67+
68+
Render **large 2D image frames** (up to 8k×8k) interactively — smooth scrub/playback of a
69+
frame stream and smooth zoom — by moving the 2D image path from Canvas2D to **WebGPU** (texture
70+
upload + WGSL colormap LUT + mipmap downscale) and moving the pixel bytes off **base64-in-JSON**
71+
onto a **binary transport**.
72+
73+
| Workload | Today (Canvas2D) | Target (WebGPU) |
74+
|---|---|---|
75+
| 8k×8k frame colormap+draw | ~64M-iter JS LUT loop → OffscreenCanvas | shader LUT, **<5 ms** |
76+
| Scrub (new frame/tick) | ~85 MB base64/frame + full rebuild | binary uint8 + texture upload, **≥15–30 fps** |
77+
| Zoom-in on a still | re-blit from OffscreenCanvas | GPU **mipmap**, **60 fps** |
78+
79+
## 2. Non-goals
80+
81+
- **Not** replacing Canvas2D for images — it stays the universal baseline, the fallback, the
82+
small-image path, and the fully-CI-tested path. Default behaviour for small images is
83+
byte-identical to today.
84+
- No WebGL2, no three.js, no bundler — raw WebGPU + inline WGSL, mirroring `WEBGPU_PLAN.md`.
85+
- 1D lines / bars stay Canvas2D. Only the 2D **image** path is affected.
86+
87+
## 3. Coverage & the fallback contract
88+
89+
Inherited verbatim from `WEBGPU_PLAN.md` §3: WebGPU is a progressive enhancement. `navigator.gpu`
90+
present → `requestAdapter()` resolves → device created → *then* a panel may switch. Any failure at
91+
any point (including mid-session device loss) lands on the Canvas2D path silently and permanently
92+
for that session. **A figure must never render nothing because GPU was attempted.**
93+
94+
## 4. Reuse (do NOT rebuild)
95+
96+
- **Device singleton + fallback state machine**: `_gpuDevice()` (`figure_esm.js` ~L1955,
97+
module-level `_gpuDevicePromise`), the `p._gpu ∈ {pending,active,unavailable}` per-panel state,
98+
the "first frame always Canvas2D → swap on device resolve" pattern, and `device.lost` → permanent
99+
per-session Canvas2D. All proven by the 3D path — the 2D image path plugs into the exact same
100+
machinery.
101+
- **The `gpuCanvas`-below-`plotCanvas` split**: decorations (axes, ticks, colorbar, scale bar,
102+
overlay mask, markers, widgets) keep drawing on the 2D `plotCanvas` **verbatim** — only the image
103+
raster moves to `gpuCanvas`. Mirrors `WEBGPU_PLAN.md` §4.3.
104+
- **`gpu_mode` / `_gpu_active` plumbing**: the state field + echo already exist (`st.gpu_mode`,
105+
~L1998). Add an image-megapixel `GPU_IMAGE_THRESHOLD` alongside the existing point threshold.
106+
- **The zoom/letterbox model**: `_imgFitRect` (~L1262) — the GPU draw honors the same fit-rect /
107+
`zoom` semantics; no new zoom math.
108+
- **The colormap LUT**: `_build_colormap_lut` (`_utils.py:118`) → `st.colormap_data`
109+
([[r,g,b]×256]) already ships to JS. Upload it as a 256×1 texture; the shader samples it.
110+
- **The normalized payload**: `set_data` already produces `img_u8` (single-channel uint8 via
111+
`_normalize_image`, `_utils.py:82`) + `display_min/max` (clim) + `raw_min/max`. The WebGPU path
112+
uploads that **same uint8** (8× smaller than raw float64; keeps current visual behaviour) as an
113+
**R8 texture**; the shader does the clim re-stretch + LUT — exactly what the JS loop at
114+
`draw2d` L1379–1381 does today, but on the GPU.
115+
116+
## 5. The function being replaced
117+
118+
`draw2d` (`figure_esm.js` ~L1344): today it `atob`-decodes `image_b64`, runs a per-pixel LUT loop
119+
(L1379–1381, 64M iters at 8k) into an `OffscreenCanvas` 2D context, caches the bitmap in
120+
`blitCache`, and `_blit2d`-down-blits to the panel. The WebGPU path replaces the decode+LUT+blit
121+
with: upload R8 texture → WGSL fragment shader (clim uniform + LUT texture) → mipmapped textured
122+
quad over the `_imgFitRect`. `blitCache`'s "bytes unchanged → reuse" logic maps to "texture
123+
unchanged → skip re-upload". Canvas2D `draw2d` remains as the fallback when `p._gpu !== 'active'`.
124+
125+
## 6. Binary transport (generalize `WEBGPU_PLAN.md` §4.7 to images)
126+
127+
Today `set_data` → `_encode_bytes(img_u8)` → `image_b64` string on the `panel_<id>_geom` trait
128+
(base64-in-JSON). Change: send the raw `img_u8` **bytes** as an anywidget **binary buffer**
129+
(`_repr_utils._widget_state` already handles `bytes`; the geom-trait split already isolates heavy
130+
keys — `FIGURE_ESM.md` ~L237), with a small JSON header (`image_width/height`, `display_min/max`,
131+
`raw_min/max`, dtype, an optional `lod` level). Keep base64 for **small images and the
132+
standalone/Pyodide/`save_html` paths** (no binary channel there). Keep `FigureBridge` (`embed.py`)
133+
transport-agnostic so an Electron embed can supply an even faster channel (shared memory / a
134+
transferable `ArrayBuffer`) without the library caring. JS: pick up the binary buffer, upload to
135+
the GPU texture; on the Canvas2D fallback, build the `ImageData` from the same bytes (no atob).
136+
137+
## 7. API surface
138+
139+
- `Axes.imshow(..., gpu="auto"|True|False)` and `Plot2D.set_data(..., gpu=...)` — mirror
140+
`scatter3d`/`voxels`. Default `"auto"`: attempt WebGPU only above `GPU_IMAGE_THRESHOLD`
141+
(initial ~4 megapixels — below it Canvas2D is already instant). `True` forces an attempt (still
142+
falls back); `False` never attempts.
143+
- `plot.gpu_active` — bool echo after first render (reuse `_gpu_active`).
144+
- Optional `set_data(..., lod=k)` affordance so a consumer can mark a frame as decimated (scrub)
145+
vs full-res (settle); at minimum, an uploaded full-res texture gets **free** GPU downscale-on-zoom
146+
via mipmaps, so LOD-on-zoom needs no consumer work.
147+
148+
## 8. Phases (with decision gates, mirroring WEBGPU_PLAN.md style)
149+
150+
### Phase 0 — Binary transport + minimal GPU texture-blit (risk-retire; do first)
151+
152+
Add the binary image trait + header (keep b64 fallback). Add a minimal WebGPU 2D image pipeline:
153+
R8 texture + 256×1 LUT texture + clim uniform + a textured quad over `_imgFitRect`, **no mipmaps
154+
yet**. Verify a single static 8k image renders identically to Canvas2D via **offscreen-texture
155+
readback** (the 3D path's proven test method — the WebGPU swapchain doesn't snapshot reliably under
156+
automation, per `FIGURE_ESM.md` ~L256).
157+
158+
**Gate A:** GPU image matches Canvas2D within tolerance on a real GPU, and `gpu=False`/adapter-absent
159+
renders identically via Canvas2D (automated).
160+
161+
### Phase 1 — Mipmaps + zoom + scrub
162+
163+
Generate mipmaps on upload; sample with trilinear filtering so zoom-out/downscale is smooth and
164+
zoom-in honors `_imgFitRect`. Wire the scrub path: new bytes → texture re-upload (reuse
165+
`blitCache`-style "unchanged → skip").
166+
167+
**Acceptance:** 8k still zooms at 60 fps; a frame stream scrubs ≥15–30 fps on a real GPU; benchmark
168+
`js_gpu_image_8k` recorded.
169+
170+
### Phase 2 — API + fallback hardening + tests
171+
172+
`gpu="auto"|True|False` on `imshow`/`set_data`, `GPU_IMAGE_THRESHOLD`, `plot.gpu_active`, optional
173+
`lod=`. Canvas2D-parity tests in the normal suite (assert `_gpu_active` false + pixel parity when
174+
GPU absent); flagged headless-GPU smoke job (skip-on-null-adapter); `upcoming_changes/*.rst`
175+
towncrier fragment.
176+
177+
**Gate B:** full fallback matrix green (adapter-absent, mid-session device loss via forced
178+
`device.destroy()`, `gpu=False`) + real-GPU image benchmark, before this is released.
179+
180+
### Phase 3 (gated) — raw-float precision path
181+
182+
If uint8 quantization proves visibly lossy for scientific contrast adjustment, add an optional R32F
183+
texture upload (raw floats, 4× bytes) with the clim applied in-shader over true data range. Only
184+
build on a concrete need (Gate C).
185+
186+
## 9. Biggest risks / do-not-break
187+
188+
- **The 3D WebGPU path and the Canvas2D 2D path must not regress** — the 2D image GPU path is
189+
additive and plugs into the shared `_gpuDevice()` singleton; keep the device/fallback code shared,
190+
not forked.
191+
- **Canvas2D stays the default + fallback forever** — no figure may render nothing because GPU was
192+
attempted (the `WEBGPU_PLAN.md` §3 contract).
193+
- **Automation can't snapshot the WebGPU swapchain** — test via offscreen-texture readback or state
194+
echo, not screenshot diffing of the live canvas (per `FIGURE_ESM.md`).
195+
- **Small-image / standalone / Pyodide paths unchanged** — binary transport is large-image-only;
196+
b64 remains for the rest.
197+
- Repo norms (`AGENTS.md`): OO API only, state in `_state` dicts (no new traitlets on Plot2D — the
198+
Figure adds panel traits dynamically), end every `_state` mutation with `_push()`, `uv run
199+
pytest`, `uv run playwright install chromium` first, add the towncrier fragment.
200+
201+
## 10. Verify
202+
203+
`uv run pytest` (Canvas2D parity + fallback in the normal suite); the flagged GPU smoke job;
204+
offscreen-readback comparison GPU-vs-Canvas2D on a large image; `js_gpu_image_8k` benchmark on a
205+
real-GPU machine; a manual Electron-embed pass (the real consumer) once SpyDE pins the new version.
206+
207+
## 11. API sketch
208+
209+
```python
210+
# Python
211+
plot = ax.imshow(frame, gpu="auto") # auto-attempt WebGPU above GPU_IMAGE_THRESHOLD
212+
plot.set_data(next_frame, clim=(lo, hi)) # scrub: binary bytes → texture re-upload
213+
plot.set_data(decimated, lod=2) # scrub-time LOD hint (optional)
214+
plot.gpu_active # bool, after first render echo
215+
```
216+
217+
```js
218+
// JS internals (figure_esm.js)
219+
_gpuDevice() // existing module singleton → Promise<GPUDevice|null>
220+
p._gpu // 'pending' | 'active' | 'unavailable' (existing)
221+
_buildImagePipeline(device, p) // NEW: R8 sampled texture + LUT texture + clim uniform
222+
_drawGpu2d(p) // NEW: image raster on gpuCanvas; decorations still 2D via draw2d
223+
```

‎anyplotlib/_binary_frame.py‎

Lines changed: 78 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,78 @@
1+
"""
2+
_binary_frame.py — a tiny length-prefixed binary wire format for shipping raw
3+
image pixels to an Electron host without base64-in-JSON.
4+
5+
Used ONLY by the Electron transport (``_electron.py``); Jupyter / Pyodide /
6+
standalone keep the base64-in-JSON path untouched. The format is a single
7+
self-describing frame written to a byte stream (stdout), interleaved with the
8+
existing text ``PLOTAPP:{json}\\n`` lines:
9+
10+
PLOTBIN:<header_len>:<payload_len>\\n<header_json_bytes><payload_bytes>
11+
12+
- ``PLOTBIN:`` is the ASCII marker (distinct from ``PLOTAPP:``) so a host reading
13+
the raw stream can tell a binary frame from a text line.
14+
- ``<header_len>`` / ``<payload_len>`` are ASCII decimal byte counts.
15+
- ``<header_json_bytes>`` is UTF-8 JSON: ``{fig_id, key, ...metadata}`` (dims,
16+
dtype, clim — everything the renderer needs EXCEPT the pixels).
17+
- ``<payload_bytes>`` is the raw payload (e.g. the single-channel uint8 image),
18+
exactly ``payload_len`` bytes, NO base64, NO JSON.
19+
20+
Both the Python producer and the JS host use this one definition, so the wire
21+
format has a single source of truth and can be unit-tested in isolation (the
22+
producer's ``encode_frame`` round-trips through ``decode_frame`` without any
23+
Electron / stdout involved).
24+
"""
25+
from __future__ import annotations
26+
27+
import json
28+
29+
MARKER = b"PLOTBIN:"
30+
31+
32+
def encode_frame(fig_id: str, key: str, header: dict, payload: bytes) -> bytes:
33+
"""Serialise one binary frame to bytes (marker + lengths + header + payload).
34+
35+
``header`` is arbitrary JSON-able metadata; ``fig_id`` and ``key`` are merged
36+
in (so the decoder always recovers them). ``payload`` is the raw bytes."""
37+
hdr = dict(header or {})
38+
hdr["fig_id"] = fig_id
39+
hdr["key"] = key
40+
hdr_bytes = json.dumps(hdr, default=str).encode("utf-8")
41+
prefix = MARKER + f"{len(hdr_bytes)}:{len(payload)}\n".encode("ascii")
42+
return prefix + hdr_bytes + bytes(payload)
43+
44+
45+
def decode_frame(buf: bytes):
46+
"""Parse ONE complete frame from ``buf`` (as produced by ``encode_frame``).
47+
48+
Returns ``(header_dict, payload_bytes, n_consumed)`` where ``n_consumed`` is
49+
the number of bytes the frame occupied. Raises ``ValueError`` if ``buf`` does
50+
not begin with a complete frame. (A streaming host uses ``parse_prefix`` +
51+
incremental reads instead; this is the whole-buffer convenience for tests.)"""
52+
if not buf.startswith(MARKER):
53+
raise ValueError("not a PLOTBIN frame")
54+
nl = buf.find(b"\n")
55+
if nl < 0:
56+
raise ValueError("incomplete PLOTBIN prefix (no newline)")
57+
prefix = buf[len(MARKER):nl].decode("ascii")
58+
hdr_len_s, pay_len_s = prefix.split(":")
59+
hdr_len, pay_len = int(hdr_len_s), int(pay_len_s)
60+
start = nl + 1
61+
end_hdr = start + hdr_len
62+
end_pay = end_hdr + pay_len
63+
if len(buf) < end_pay:
64+
raise ValueError("incomplete PLOTBIN frame (truncated body)")
65+
header = json.loads(buf[start:end_hdr].decode("utf-8"))
66+
payload = buf[end_hdr:end_pay]
67+
return header, payload, end_pay
68+
69+
70+
def parse_prefix(line: bytes):
71+
"""Parse a ``PLOTBIN:<hlen>:<plen>`` prefix line (the bytes up to, not
72+
including, the newline). Returns ``(header_len, payload_len)``. For a
73+
streaming host that reads the prefix line, then exactly ``header_len +
74+
payload_len`` more bytes."""
75+
if not line.startswith(MARKER):
76+
raise ValueError("not a PLOTBIN prefix")
77+
hdr_len_s, pay_len_s = line[len(MARKER):].decode("ascii").split(":")
78+
return int(hdr_len_s), int(pay_len_s)

0 commit comments

Comments
 (0)