| Filename | Latest commit message | Latest commit date |
|---|---|---|
Replace the legacy Emscripten JS-exceptions model with native wasm-EH (legacy encoding) across the whole build, keeping Asyncify coroutines working via a from-source Binaryen --hoist-cpp-catches pre-pass. Net result: native-EH is the only build mode, the 3D viewer is on by default, and pcbnew shrinks substantially. Highlights: - Binaryen submodule everywhere + --hoist-cpp-catches integration in apply-asyncify; post-link Asyncify covers every app wasm (not just standalone test wasm). - Build deps (incl. OpenCASCADE without OCC_CONVERT_SIGNALS) and all KiCad apps with -fwasm-exceptions; emscripten_sleep added to the post-link asyncify-imports. - libcontext fiber entry wired under native exceptions; while-loop main loop + currData shim injected into all wx apps. - Native-EH collab apply fixed: DEBUG-define the embind TU + match all out-of-CMake C++ TUs' ABI flags to the core, fixing the vtable-layout skew / mis-dispatch. - 3D viewer enabled by default (real raytracer linked, not the stub). - Retire the EH-spike scaffolding; flip the asyncify-races ablation pins to shim-redundancy pins (native-EH stays clean with the legacy shims ablated). - Fix the asyncify-races quiescence check to not require Asyncify.currData==0: under the native-EH per-frame-yield top loop the main stack is asyncify-suspended every frame, so currData legitimately churns (a freed-but-not-yet-nulled buffer, not a leak). Refresh the pcbnew toolbar screenshot baseline for the new kicad. - CI: drop the obsolete binaryen_version input/env (the build uses the binaryen submodule fork's wasm-opt, not a version download); key the wasm-output cache on the binaryen submodule SHA instead. Bumps the wxwidgets + binaryen submodules to their squashed feature commits. Validated green: all 7 apps native-EH (real 3D in pcbnew); KiCad e2e 63/63 Firefox + Chromium (3D viewer renders); wx 336; coroutine 34/34 both engines; asyncify 7/7 both engines. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
| .. | ||
| README.md | ||
Cross-browser performance: why Firefox > Chrome > Safari, and how to close the gap
Research notes, 2026-06-18. The KiCad WASM port runs fastest in Firefox, slower in Chrome, slowest in Safari. This document explains why at the browser-engine level and lays out a ranked, build-specific plan to speed up Chrome and Safari. Web claims are dated and linked in Sources; codebase claims carry
file:linerefs. Companion work lives in../async/(Asyncify) and../wasm-exceptions/(the-fwasm-exceptionsmigration).
TL;DR
The Firefox lead is not a Firefox trick. Our binary is dominated by Asyncify instrumentation, and Firefox's compilers simply tolerate Asyncify's pathological code far better than Chrome's or Safari's do. So the highest-leverage work for Chrome and Safari is to shrink/attack the Asyncify footprint, plus a handful of cheap, orthogonal wins.
There are two independent axes, and both need attention:
- WASM compile/execute — Asyncify-dominated. This explains the Firefox > Chrome > Safari ordering.
- WebGL rendering — Safari's Metal/ANGLE overhead. This is extra Safari slowness on top of axis 1, and several fixes are one-liners.
Ranked levers
| # | Lever | Axis | Effort | Impact | Where |
|---|---|---|---|---|---|
| 1 | Confirm/force instantiateStreaming + Content-Type: application/wasm + stable URL/ETag |
startup | hours | ~1.5–1.8× cold start (FF); arms V8 cache | web/standalone/src/wasm/boot.ts |
| 2 | Brotli instead of gzip-9 on R2 | startup | hours | ~15–25% smaller transfer | R2 / edge config |
| 3 | powerPreference: 'high-performance' + context-lost handlers |
WebGL (Safari/Chrome) | hours | discrete GPU instead of integrated | wxwidgets/src/wasm/glcanvas.cpp:524-535 |
| 4 | Audit GAL shaders for the flat qualifier |
WebGL (Safari) | hours–days | up to seconds/frame in worst case | kicad/common/gal/shaders/ |
| 5 | Remove glGetError() from the render loop |
WebGL (Safari) | hours | avoids per-call Metal flush | GAL compositor |
| 6 | Test antialias: false |
WebGL (Safari) | hours | cuts MSAA resolve cost | glcanvas.cpp:524-535 |
| 7 | Enable -msimd128 |
WASM exec (all) | days | 1.5–2.5× geometry/render hot loops | build flags |
| 8 | ASYNCIFY_ADVISE → ASYNCIFY_IGNORE_INDIRECT + extend REMOVE |
WASM exec (all, esp. Chrome/Safari) | days | smaller binary + faster tier-up | scripts/common/apply-asyncify.sh |
| 9 | -fwasm-exceptions (size) |
WASM (all) | weeks | 64.5 → 36 MB gz | tracked — see § Structural bets |
| 10 | JSPI (delete Asyncify) | WASM (all, esp. Safari) | weeks | ~40–50% smaller, removes JIT pressure | tracked — see § Structural bets |
| — | wasm-split, WebGPU GAL backend | startup / WebGL | weeks+ | deferred (see § Deferred) | — |
Current build (the baseline)
Verified from the build scripts and runtime glue:
| Knob | Value | Location |
|---|---|---|
| Asyncify | -sASYNCIFY=1, ASYNCIFY_STACK_SIZE=65536 |
scripts/kicad/build-kicad-target.sh:~400 |
| Exceptions | legacy -fexceptions (not -fwasm-exceptions) |
build-kicad-target.sh:240-255 |
| SIMD | none (-msimd128 absent) |
— |
| Threads | -sUSE_PTHREADS=1, pool = navigator.hardwareConcurrency (+ COOP/COEP) |
build-kicad-target.sh, web/.../preflight/capabilities.ts |
| Memory | INITIAL_MEMORY=256MB, MAXIMUM_MEMORY=4GB, ALLOW_MEMORY_GROWTH=1 |
build-kicad-target.sh |
| Opt | clang -O2 (release); link -O0 then host wasm-opt -O2 after --asyncify |
apply-asyncify.sh:88-157 |
| WebGL | WebGL2 (-sMAX_WEBGL_VERSION=2), antialias:true, powerPreference:DEFAULT |
glcanvas.cpp:524-535 |
| Loading | Emscripten script-glue; streaming not confirmed; gzip-9, no Brotli | boot.ts:145-294 |
| Artifact | pcbnew 186 MB raw / 64 MB gzip; eeschema 99/34; pl_editor 52/17; gerbview 49/16 | output/ |
Note: all three modern browsers support SharedArrayBuffer/threads under COOP+COEP
(the app demonstrably runs in each) — capability gating is in
capabilities.ts, not UA sniffing.
Why the ordering exists (engine internals)
The villain: Asyncify
Asyncify rewrites every instrumented function with unwind/rewind state checks and
saves/restores all locals to linear memory. That expands each local's live range
across the whole function, producing a nearly fully-connected interference graph
— exactly the input that is catastrophic for optimizing register allocators.
Asyncify's own docs warn: "VMs may also limit compilation to the baseline tier on
such pathological code." Result: ~+70% binary, giant functions, and the
186 MB-raw pcbnew. See ../async/02-asyncify-internals.md.
How each engine copes
| Engine | Baseline tier | Optimizing tier | On Asyncify's giant functions |
|---|---|---|---|
| Firefox / SpiderMonkey | Rabaldr, ~25 ns/byte, eager whole-module, multithreaded (30–60 MB/s) | Ion — 75× large-function fix, Oct 2024 (sorted live ranges, Semi-NCA dominators, sparse bitsets) targeting exactly the huge-CFG/high-vreg shape Asyncify creates (ONNX: 5 min → 3.9 s) | Best. Whole module baseline-compiled before download finishes; Ion swallows the big functions. No OSR gap. |
| Chrome / V8 | Liftoff, ~50 ns/byte (½ Firefox) | TurboFan — chokes on huge fns (a 1.96 MB fn → 95 s, 7.4 GB RAM, 87% in regalloc); falls back to mid-tier allocator or skips optimization | Middle. V8 has no OSR for wasm — a function in a long loop (Asyncify rewind/unwind loops!) finishes that whole call in Liftoff; only the next call gets TurboFan. |
| Safari / JSC | Lazy everything: IPInt (interpreter) → BBQ → OMG. Nothing eager. | OMG (B3) — did not get Ion's 2024 large-fn treatment | Worst. First run executes at interpreter speed; Asyncify ~doubles fn count → huge OMG backlog → documented 300–400% CPU spike for 30 s+ after a workload. No persistent compiled-code cache, so it re-pays every session; above ~10 MB it switches to a slower JIT mode. |
Two corollaries that bite us specifically
- Chrome's V8 wasm code cache is effectively unavailable. It only caches
modules under ~150 MB compiled, and compiled code is 5–7× the
.wasm. Our 186 MB pcbnew → ~1 GB compiled — far over the ceiling. So Chrome re-runs TurboFan on every cold load today. Shrinking the binary (levers 7–10) is the only way to get Chrome's repeat-load cache back. See V8 wasm code caching. - Benchmark trap: with DevTools open, V8 tiers all wasm down to Liftoff. Never measure Chrome speed with DevTools open (except via an actual Performance recording, which forces tier-up). This likely makes Chrome look worse than it is in casual testing.
The ranked plan
Tier 1 — cheap, do now (days, low risk)
1. Confirm + force streaming instantiation and cache headers. The loader
injects the Emscripten JS glue via <script> (boot.ts); whether the runtime
then streams the .wasm depends on serving conditions. In DevTools → Network,
confirm the .wasm returns Content-Type: application/wasm with no console
"falling back to ArrayBuffer instantiation" warning. Streaming is ~1.5–1.8× faster
cold-start on Firefox and is the only path that arms V8's code cache. Serve the
.wasm from a stable URL (no content-hash in the path; use a stable alias)
with ETag/304. With -pthread, confirm the module is compiled once and shared
to workers (Emscripten does this via the shared WebAssembly.Module), not
recompiled per worker.
2. Brotli instead of gzip-9 on R2. Brotli is ~15–25% smaller on wasm (our
64 MB pcbnew → ~50 MB). Verify Cloudflare actually Brotli-compresses it at the
edge — it often skips large binary types — and if not, precompress and serve
with Content-Encoding: br + Content-Type: application/wasm. Smaller transfer
also means less to compile, so it compounds with everything below.
3. powerPreference: 'high-performance' for WebGL. We default to
EM_WEBGL_POWER_PREFERENCE_DEFAULT (glcanvas.cpp:524-535). Safari (and Chrome
on dual-GPU Macs) defaults WebGL to the integrated GPU. Requesting
high-performance switches to the discrete GPU — often the single biggest
GPU-bound framerate win on MacBook Pros. Caveat: Safari only honors it if you also
register webglcontextlost/webglcontextrestored handlers.
4. Audit GAL shaders for the flat interpolation qualifier. This is the big
Safari sleeper. flat triggers a provoking-vertex workaround in Safari's
Metal/ANGLE backend that has cost real apps seconds per frame. PCB renderers
commonly use flat for per-primitive net/layer colors. Grep the GAL shaders
(kicad/common/gal/shaders/, source GLSL 1.20 before convert_glsl_es3.py); if
present, replace with regular interpolation or restructure. Potentially a massive
Safari-only win.
5. Remove glGetError() from the render loop. On Safari each call forces a
Metal pipeline flush. Restrict to init/debug builds only. (Note: the WebGL
compositor already drains stale glGetError() once before draws — that's fine;
the concern is per-call error checks inside the hot path.)
6. Test antialias: false. We default MSAA on (antialias:true). KiCad's GAL
does much of its line AA in-shader (SMAA) and has its own AA setting; if MSAA is
redundant, dropping it cuts Metal's resolve cost on Safari. Quality/perf tradeoff —
A/B it on a dense board; consider exposing it as a setting.
Tier 2 — medium effort, high impact
7. Enable -msimd128. Expect 1.5–2.5× on the geometry/render hot loops
(polygon booleans in shape_poly_set, DRC overlap checks, vertex-buffer fills) via
LLVM autovectorization at -O2+. Safe on all three engines (Chrome 91 / FF 89 /
Safari 16.4). Helps absolute Chrome and Safari speed. Verify v128.* actually
appears in the disassembly for the hot functions, and prefer pmin/pmax over
min/max (the SSE→wasm emulation table has slow paths). Do not ship Relaxed SIMD
yet (Safari still flags it). Minor interaction to watch: SIMD slightly grows
per-function size, which feeds the Asyncify/locals pressure — measure after the
wasm-opt -O2 pass.
8. Shrink the Asyncify surface. This directly attacks the root cause for Chrome
and Safari. Run ASYNCIFY_ADVISE to see which functions get instrumented and
why — it surfaces the biggest instrumented functions (the JIT pressure points).
Then:
ASYNCIFY_IGNORE_INDIRECT=1is the high-impact one: our wxWidgets/GAL code is vtable-heavy, and Asyncify conservatively instruments every indirect call site, which is why instrumentation spreads everywhere. With our understanding of the suspend paths (the park-throw work), we may be able to assert no indirect call is on the suspend stack and add specific ones back viaASYNCIFY_ADD.- Extend the existing 12-function
ASYNCIFY_REMOVElist (apply-asyncify.sh:92-104) with cold/startup-only large functions ADVISE flags. - Smaller instrumented set → smaller functions → better tier-up in all engines and a smaller Safari OMG backlog.
- ⚠️ Error-prone (wrong config = silent runtime breakage). Gate behind the
red/green Asyncify harness in
../asyncify-arbiter/redgreen.md/tests/asyncify/.
Structural bets (track / prototype)
9. -fwasm-exceptions (we're on legacy -fexceptions). Biggest size lever —
../wasm-exceptions/02-measurements.md
puts pcbnew at 64.5 → 36 MB gzip, which would also start bringing Chrome back
under the code-cache ceiling and fix the unreliable catch/destructor landing-pad
behavior we've documented. Two blockers to track before committing:
- a Safari 26.0 startup regression for
-fwasm-exceptionsmodules (emscripten #25365) — verify whether it's fixed in a 26.x / Safari 27 beta before shipping, since Safari is the browser we're trying to help; - the asyncify-EH unwind-from-catch interaction
(
../wasm-exceptions/05-asyncify-fork-design.md).
10. JSPI — the eventual real fix for Safari, because it deletes Asyncify entirely (no instrumentation → no giant functions → no OMG backlog → the 300–400% Safari spike goes away) and cuts ~40–50% of binary size. Status:
| Engine | JSPI status |
|---|---|
| Chrome / V8 | shipped, Chrome 137 (May 2025) |
| Firefox / SpiderMonkey | Firefox 153 intent-to-ship (June 2026); Nightly now, stable ~late summer/fall 2026 |
| Safari / JSC | Safari 27 beta (WWDC26), enabled by default; stable Fall 2026 |
Don't migrate wholesale yet:
- unresolved ~350× regression on the
JS→C→JSre-entry pattern (emscripten #21081) — exactly what a GUI event loop hits constantly; - static-init
SuspendError(emscripten #24302); invoke_*over-tagging under our legacy exceptions.
Recommended path: prototype JSPI on a small tool (calculator or pl_editor)
behind feature detection ('Suspending' in WebAssembly) with Asyncify fallback,
profile the re-entry pattern, and watch #21081. By the time it's safe, all three
engines will support it.
Deferred / not now
wasm-split/-sSPLIT_MODULE— the secondary module can't be loaded lazily and asynchronously on the main thread, which is incompatible with our main-thread Asyncify model unless we move to-sPROXY_TO_PTHREAD.- WebGPU GAL backend — the structural exit from Safari's Metal/ANGLE overhead
(Safari 26 ships WebGPU), but it's a multi-week GLSL→WGSL port with no upstream
KiCad support. See the GAL history in
../archive/webgl/. - Global
-O3— its inlining bloats the binary and makes the Chrome cache/compile problem worse. If anything, compile cold/utility units at-Os.
What to do first
Quick, visible, near-zero-risk wins this week:
- Safari: #3 (high-performance GPU) + #4 (
flataudit) + #5 (glGetError). - Startup everywhere: #2 (Brotli) + #1 (streaming/headers).
- Runtime everywhere: #7 (
-msimd128).
Then invest in #8 (Asyncify ADVISE + IGNORE_INDIRECT) as the real lever against the Chrome/Safari gap, and keep #9 / #10 on a tracking list.
Sources
Engine internals
- 75× faster: optimizing the Ion compiler backend — SpiderMonkey, Oct 2024
- Understanding WebAssembly code generation throughput — wingolog, 2020
- V8 WebAssembly compilation pipeline · Dynamic tiering · Liftoff
- Code caching for WebAssembly developers — V8
- Introducing the JetStream 3 Benchmark Suite — WebKit, 2024 (IPInt/BBQ/OMG)
- Pause and Resume WebAssembly with Binaryen's Asyncify — kripken, 2019
Startup / size / SIMD
- Optimizing WebAssembly Startup Time — Nutrient
- MDN: WebAssembly.instantiateStreaming
- Module Splitting — Emscripten
- Using SIMD with WebAssembly — Emscripten · V8 SIMD · caniuse wasm-simd
- Asynchronous Code (Asyncify settings) — Emscripten
Safari WASM + WebGL
- WebGL Performance on Safari & Apple Vision Pro — Wonderland Engine (
flat, UBO timing) - WebKit features in Safari 26.0 · News from WWDC26 — Safari 27 beta
- emscripten #25365 — Safari 26.0 wasm-exceptions regression
- emscripten #26027 — Safari + Asyncify + unaligned-load leak
- ONNX Runtime #26827 — Safari WebKit 26 OMG CPU loop
JSPI
- V8: Introducing the WebAssembly JavaScript Promise Integration API · new API
- caniuse: JSPI · Chrome 137 release notes
- Mozilla dev-platform: Intent to Ship JSPI (Fx153), June 2026
- emscripten #21081 — JSPI 350× slower for JS→C→JS · #24302 — JSPI static-init SuspendError
Local cross-references
../async/— Asyncify internals, single-slot contention, park-throw.../wasm-exceptions/—-fwasm-exceptionsmeasurements, toolchain status, asyncify-EH fork design.../asyncify-arbiter/redgreen.md— the harness to gate Asyncify-surface changes.../archive/webgl/— GAL → WebGL2 history (context attrs, compositor, shaders).- Build:
scripts/kicad/build-kicad-target.sh,scripts/common/apply-asyncify.sh. - WebGL context:
wxwidgets/src/wasm/glcanvas.cpp:489-535. Loader:web/standalone/src/wasm/boot.ts.