pcbjam/docs/features/web-e2e-rot/01-editor-lib-bridge-flows.md
Viktor Vaczi 63ed1f3c1f e2e/CI: dual-engine suites, per-engine screenshots, SwiftShader retired, prod web suite, CI-coverage gate
Squash of experiment/ff-big-modules vs main.

Big-module routing removed: native-EH shrank kicad_editor below
SpiderMonkey's x86-64 code budget (runs 29355049705/29356152413 green on
stock Firefox), so BIG_MODULE_SPECS routing and the baseline-only-JIT
crutch are gone — kicad-firefox and kicad-chromium both run the full
suite, with the module compiled the way real users' browsers compile it.

Per-engine screenshots end to end: stableShot/shotPath write
test-results/<engine>/<name>.png; baselines move to
baseline-screenshots/{chromium,firefox}/ and the whole tools/screenshots
pipeline (compare/promote/manifest/spec-map/changelog/Discord) keys on
<engine>/<name>. Previously Firefox and Chromium renders of one spec
overwrote each other and Firefox renders were never actually gated.
Seeded from CI run 29421380806 (92 new firefox baselines, +24 chromium
web-suite shots); manifest generated from the baseline tree.

One merged playwright.config.ts (kicad/asyncify/coroutine/perf as
projects); ~25 dead npm scripts dropped. The web suite is gated in CI for
the first time ever (4 rotted specs fixed, 5 broken lib-bridge specs
triaged as fixme in docs/features/web-e2e-rot/); cheap lint step after
npm ci; last 26 blind-sleep violations fixed.

SwiftShader retired: CI Chromium renders WebGL on ANGLE → Mesa llvmpipe
(--use-gl=angle --use-angle=gl --ignore-gpu-blocklist; the blocklist flag
is mandatory — llvmpipe is blocklisted and WebGL is silently unavailable
without it) in BOTH configs. Under WORKERS=4 congestion SwiftShader
transiently failed the first post-board-load draw and the recovery
cascade ended in a silent permanent Cairo fallback — that engine flip was
the "~1.2% changedRatio both directions" occ-export baseline flake.
Validated 160/160 across two 80-repeat rigs; full analysis in
docs/features/wx-parity-bugs/occ-export-context-eviction.md. Chromium
baselines shift slightly on llvmpipe — promote once from the first green
run. Deflakes the new coverage exposed: presence baselines settle before
capture; presence fixtures declare current file formats; perf gets its
own outputDir so CI evidence survives; occ-export settles the board paint
before the export dialog; menu-item waits (waitForRenderedByLabel before
clickMenuItem) in 4 specs + the TESTING.md rule.

Web suite runs the PROD build, in parallel: webServer becomes backend
`start` + the standalone's e2e:preview (build-preview.mjs: link-wasm →
stash the public/wasm symlink aside during vite build, build-demo.mjs's
move — then vite preview as the persistent server). The wasm middleware
serves /wasm/* in preview and emits COOP/COEP/CORP itself (a pthread
worker script's own response must carry COEP or Chrome kills it with
ERR_BLOCKED_BY_RESPONSE). VITE_* flags bake at build time;
VITE_ALLOW_USER_OVERRIDE joins turbo globalEnv. fullyParallel + default
workers: 5.2m → 1.4m. Determinism fixes the parallel run exposed:
shared-page specs become serial groups; locks.spec grabs alice's exact
item via the new kicadCollabTestSelectByUuid hook (cross-tab "first
footprint" order is not a ysync invariant); quit specs poll page.url()
(quit supersedes its own navigation — NS_BINDING_ABORTED on Firefox).
Suite: 51 passed / 12 skipped / 0 failed in 1.6m.

CI-coverage gate (lint:ci-coverage): every tests/**/*.spec.ts must be
reachable from the npm scripts the workflows invoke — scraped from
.github/workflows/, resolved through package.json, coverage asked from
playwright --list itself. Rules: uncovered-spec + orphan-project (with a
documented LOCAL_ONLY_PROJECTS allowlist). Gating next to
lint:determinism; 138 spec files / 13 projects accounted for.

Product fixes kept from the investigations (reachable on real GPUs too):
wx 7799fd1be5 — paint flags clear before dispatch + Invalidate always
propagates; kicad 3dcfea5e45 — SwiftShader pass-boundary flush +
per-instance font texture + first-frame GL-error drain (GAL recovery
recovers instead of falling back to Cairo) + the user-facing eeschema
switch navigates again under __EMSCRIPTEN__ (project-sync's
FaceRegistered gate had rerouted it into the hidden sync player; caught
by the newly-gated web suite).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018eUxiPApHgGiu9NFyQfhAq
2026-07-17 12:21:54 +02:00

3.4 KiB

Web-suite rot: footprint/symbol editor library-bridge flows (5 specs)

Surfaced 2026-07-15 by wiring the web suite into CI for the first time (branch experiment/ff-big-modules, runs 29414003275 / 29415027412). These specs were green when written and have silently rotted since — the web suite never ran in CI, so nothing caught it.

Affected specs (marked test.fixme — flip back when fixed)

spec first-broken symptom (local, web-chromium) CI symptom
web/footprint-browse-remote.spec.ts boots, lib tree lists libs, but expanding Resistor_SMD never produces child rows (FootprintEnumerate yields nothing) #canvas never visible in 180 s
web/footprint-write-remote.spec.ts boots; New Footprint + Ctrl+S never lands an item in the backend (/api/scopes/default/libs/<lib>/items stays empty 30 s) same boot timeout
web/footprint-write-spike.spec.ts boots (route fixed); window.__pcbjamSaved never captures a body after New Footprint + save (?fpwrite=1 spike provider silent) same boot timeout
web/symbol-write-remote.spec.ts same family as footprint-write-remote, symbol domain same boot timeout
web/symbol-write-spike.spec.ts same family as footprint-write-spike (?libwrite=1) same boot timeout

Evidence pointing at one root cause

All five live in the same domain: window.kicadLibs.request(...) traffic from the footprint/symbol EDITOR tools (enumerate / save). Probing a booted /default/projects/demo/-/footprint_editor locally:

  • the lib tree lists the origin libs (Capacitor_SMD, Diode_SMD, LED_SMD, My Symbols, Resistor_SMD) — the list path works (pre-sync/IDB);
  • __libsCalls (a wrapper capturing every kicadLibs.request) records ZERO calls during boot and ZERO on expanding a lib — the per-lib enumerate/get/save traffic never happens;
  • one [pageerror] __name is not defined fires at boot — esbuild's keep-names helper missing in whatever context executes a bundled callback; prime suspect for the provider dying silently on first use.

By contrast the SCHEMATIC editor's bridge works (eeschema-fp-selector records index calls), and eeschema/pcbnew tool pages boot and pass their specs — the rot is specific to the fp/sym editor lib flows, not the bridge as a whole.

Separately fixed while triaging (NOT part of this bug): dead /p/<slug>/<tool>/ routes in 3 of these specs (router grammar is /:scope/projects/:name/-/:tool since the scope/kind/name change), the tool-switch overlay race + /p/ URL asserts, and read-only-editor's cross-tab SelectFirst-order assumption.

Why fixme and not expected-fail

The ysync convention (expected-fail repros) is right for fast unit-level repros. These five die in 180 s boot timeouts on CI — as test.fail() they would burn ~30 min of CI per run across two engines for zero extra signal. test.fixme keeps them visible in every report as skipped-with-reason; remove the marker (and delete this table row) when the bridge flow is fixed.

CI-only sibling (kept RUNNING, conditional expected-fail)

web/eeschema-fp-selector.spec.ts completes its flow but records a [pageerror] index out of bounds wasm trap on CI (BOTH engines, llvmpipe/SwiftShader) that does not reproduce on a Mac with real GL — see test.fail(!!process.env.CI, …) in the spec. Likely the same software-GL render-path family as the presence ghost-wipe (SwiftShader pass-boundary flush, webgl_gal.cpp). Needs its own investigation on a CI-like rig.