pcbjam/tests/kicad/eeschema-sim-recovery.spec.ts
Istvan Matejcsok c421d724b0 findings(E-10..E-22): fix the defects a code review found in the E-1..E-9 work
A review of the group-E fixes found 13 further defects; ten were introduced by
those fixes, two pre-existed and were merely relocated, one is deferred.

Services / transport
  E-10  retireWorker synthesized no bg/exit frame, so sharedspice's s_bgRunning
        mirror stayed latched true after a mid-run worker death: Run stayed
        disabled and the promised fresh-worker restart was unreachable for the
        whole session. Retirement now dispatches a synthetic controlled-exit
        straight to the installed handler (never through dispatchEvt — a
        fabricated frame must not touch the credit ledger). Driving the repro
        exposed two further defects, both fixed here: a replacement worker
        trapped on pre-init engine reads, and the rerun's cm_input_path/circ hit
        that uninitialized engine before KiCad's validate() re-init (the native
        flow assumes a crashed engine survives in-process — true for the dll,
        false for a dead worker). Reads now answer their empty shapes pre-init,
        writes lazy-init, and init is idempotent per worker engine.
  E-19  dispatchEvt acked only AFTER handler(evt) returned, and the sharedspice
        client deliberately rethrows non-trap errors — so each throw leaked one
        unit of the 64-frame credit window until the stream died with a
        misattributed "transport exceeded". The ack moves to a finally in both
        service copies; the throw still propagates (the trap machinery needs it).
  E-20  the oversize-line path promises to transfer the accepted prefix, but
        with the window full that flush only DEFERS, and stopEventStream wiped
        the deferred queue — losing the diagnostics that explain the failure.
        The terminal notice now carries them as pendingEvents; both hosts
        deliver them in order, unacked (the fatal frame is outside the credit
        protocol).
  E-21  the 30s prefetch deadline discarded every model already collected and
        reported nothing. A caller-owned progress sink ships the partials and
        the omission reaches the export report. (Awaiting the aborted collection
        was rejected: an in-flight source fetch is not abortable — E-4's
        original disease.) Plus a serving-candidate memo, so a .wrl ref served
        by its .step fallback stops re-probing the miss on every export.

Scheduler
  E-14  _terminalizeNativeTrap classified by message substring, so any plain JS
        error QUOTING 'Aborted(' or 'out of bounds' permanently bricked a
        healthy instance. Now structural only: instanceof RuntimeError plus a
        duck-typed name check (verified in this build's glue that abort() throws
        a genuine RuntimeError both pre- and post-runtime-init). Module.onAbort
        now latches the gate — the authoritative notification, previously
        ignored.
  E-15  the shim half: _pumpResume gates on terminal (catching wakes already
        queued at latch time) and resolveWait refuses on terminal WITHOUT
        consuming the entry, so a frame stays visibly parked rather than
        resuming inside a trapped module.
  E-16  the E-5 handler read the realm-global scheduler at dispatch instead of
        its installing module's; also frees the per-line buffer on the non-trap
        rethrow path.
  E-11  get_vec trusted the worker's res.length over the transferred arrays.
        Observed death shape: a 4 GiB std::vector threw an unhandled
        std::length_error that exited the editor's main loop. Now clamped, with
        the buffers freed on every failure path.

Guardrails (replacing two deferred refactors: e2e→production-code injection and
collapsing the four copies of the worker-lifecycle machinery)
  E-18  the source contract asserted comment-string counts — rewording failed
        CI while moving a guard outside its #ifdef passed. It now parses the
        #ifdef regions and asserts on code.
        service-stub-parity.ts pins what the four lifecycle copies must share:
        credit-window equality parsed from source, the finally-ack, boot
        deadlines, terminal-notice consumption. The transport numbers are now
        single-sourced from the worker.
        CI actually runs the gates: the web/standalone vitest suites (which had
        NEVER run in CI), the reducer, the source contract and the parity tool —
        with a NON_PLAYWRIGHT_GATES check so deleting a step re-fails the lint.
  E-22  the e2e occ stub's 60s boot watchdog, deleted in a66e109, is restored in
        the ngspice-stub shape with a wedgeNextBoot() repro hook.

Every behavioral fix has red-then-green evidence (the reds were captured first).
E-17 (a stale RUNNING cross-stamping the next run's generation under E-6's
transport deferral) is DEFERRED with its analysis recorded — a real fix needs
run identity on the bg frames.

Test hygiene: the dwell lint now requires the mandated ": <why>" and all 47 bare
markers carry their reason; three export-report dwells became modal-lease polls;
exact-ledger assertions became relative deltas; the dead data-wx-dom-id branch,
an unused fault hook and unused receipt plumbing are gone; abort scans, wx
dialog drivers, the sim harness and the vitest FakeWorker are each one copy now.

Bumps kicad and wxwidgets to their findings-group-e tips.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 18:19:16 +02:00

183 lines
9.4 KiB
TypeScript

import { test, expect } from './fixtures';
import { clickByTooltip, waitForEditorReady } from '../e2e/utils/element-tracker';
import { FATAL_WASM_PATTERNS, findNativeFailure } from './utils/native-failure';
import {
loadRectifier,
openSimulator,
runSimulation,
waitForRunToolEnabled,
} from './utils/sim-harness';
/**
* eeschema simulator worker-death recovery (findings E-10/E-12/E-11): the
* promise of the out-of-process engine is that a worker death settles
* everything in flight and the next Run transparently boots a fresh worker.
* These specs kill (or corrupt) the service at exact points and assert the
* simulator UI actually recovers:
*
* - E-10: a mid-run worker death must unlatch the client's s_bgRunning
* mirror (via the service's synthetic controlled-exit) — otherwise the
* Run tool's ENABLE(!simRunning) holds "running" forever and the promised
* fresh-worker restart is unreachable for the whole session.
* - E-12: a run whose transport dies between launch acceptance and its
* RUNNING transition delivers its crash-exit completion — the wasm-only
* unowned-event drop must not swallow an owned run's only IDLE.
* - E-11: a corrupted worker's oversized get_vec length must be clamped to
* the actually-transferred arrays — not copied into the editor heap as a
* multi-gigabyte read that traps the instance.
*/
test.describe('eeschema simulator worker-death recovery', () => {
test.setTimeout(300000);
test('a mid-run worker death re-enables Run and a rerun succeeds (E-10)', async ({ page, testLogger }) => {
await page.goto('/kicad/eeschema.html');
await waitForEditorReady(page);
await loadRectifier(page);
await openSimulator(page);
await waitForRunToolEnabled(page);
const checkpoint = await page.evaluate(() => {
const hooks = (globalThis as any).__ngspiceServiceTestHooks;
return hooks.appliedGenerationCheckpoint() as number;
});
// Start a run and inject the worker death while it is live. The check
// and the retirement happen in ONE page.evaluate — frames dispatch on
// the same main thread, so no finish frame can interleave between the
// "still running" check and the kill. Event scans are scoped past any
// frame-open activity (workbook plot restoration).
const eventFloor = await page.evaluate(
() => ((window as any).__ngspiceEvents as unknown[]).length);
expect(await clickByTooltip(page, 'Run Simulation', { elementType: 'tool' }),
'Run tool').toBe(true);
// Run accepted: the worker's bg started frame arrived.
await expect.poll(
() => page.evaluate((floor: number) =>
((window as any).__ngspiceEvents as Array<{ kind: string; finished?: boolean }>)
.slice(floor)
.some((e) => e.kind === 'bg' && e.finished === false), eventFloor),
{ message: 'the run must report bg started', timeout: 60000 },
).toBe(true);
const injected = await page.evaluate((floor: number) => {
const events = ((window as any).__ngspiceEvents as Array<{
kind: string; finished?: boolean }>).slice(floor);
const finishSeen = events.some((e) => e.kind === 'bg' && e.finished === true);
const retired = (globalThis as any).__ngspiceServiceTestHooks
.forceRetire('E-10 repro: worker death mid-run');
return { finishSeen, retired };
}, eventFloor);
expect(injected.finishSeen,
'repro window: the run must still be live when the fault is injected').toBe(false);
expect(injected.retired, 'the active generation was retired').toBe(true);
// THE E-10 oracle: without the synthetic controlled-exit the
// s_bgRunning mirror stays latched true; and without the worker's
// pre-init read guard the crash-recovery finish parks on a vector
// pull into the trapped replacement engine — either way this poll
// times out with the Run tool disabled forever.
await waitForRunToolEnabled(page);
// The crashed run's completion was delivered (cursor/finish body ran).
const crashReceipt = await page.evaluate(async (after: number) => {
const hooks = (globalThis as any).__ngspiceServiceTestHooks;
return await hooks.waitForAppliedGenerationAfter(after, 60000);
}, checkpoint);
expect(crashReceipt.generation, 'the crashed run applied its completion')
.toBeGreaterThan(checkpoint);
// The synthetic exit is visible in the event record.
const exitSeen = await page.evaluate((floor: number) =>
((window as any).__ngspiceEvents as Array<{ kind: string }>)
.slice(floor)
.some((e) => e.kind === 'exit'), eventFloor);
expect(exitSeen, 'a controlled-exit event reached the client').toBe(true);
// And the promised transparent restart: a full rerun on a fresh
// worker generation succeeds end to end.
const rerunGeneration = await runSimulation(page);
expect(rerunGeneration).toBeGreaterThan(crashReceipt.generation);
const generations = await page.evaluate(() =>
(globalThis as any).__ngspiceServiceTestHooks.snapshot());
expect(generations.retiredGenerations, 'the killed generation was retired')
.toContain(1);
expect(findNativeFailure([...testLogger.consoleLogs, ...testLogger.errors]),
'no wasm abort during the recovery').toBeUndefined();
});
test('a launch that dies before RUNNING still delivers its completion (E-12)', async ({ page, testLogger }) => {
await page.goto('/kicad/eeschema.html');
await waitForEditorReady(page);
await loadRectifier(page);
await openSimulator(page);
await waitForRunToolEnabled(page);
// Arm: the transport dies on the bg_run launch itself — after the
// native side published its run generation, before any RUNNING
// transition could fire. The retirement's synthetic exit then
// delivers this run's ONLY completion. (The arm keys on bg_run
// specifically, so frame-open plot restoration cannot consume it.)
const checkpoint = await page.evaluate(() => {
const hooks = (globalThis as any).__ngspiceServiceTestHooks;
hooks.dieOnNextBgRun();
return hooks.appliedGenerationCheckpoint() as number;
});
expect(await clickByTooltip(page, 'Run Simulation', { elementType: 'tool' }),
'Run tool').toBe(true);
// THE E-12 oracle: on the unfixed build the crash-exit IDLE carries
// generation 0 (its RUNNING never fired) and is deleted — the owned
// run's completion never applies and this receipt times out.
const receipt = await page.evaluate(async (after: number) => {
const hooks = (globalThis as any).__ngspiceServiceTestHooks;
return await hooks.waitForAppliedGenerationAfter(after, 60000);
}, checkpoint);
expect(receipt.generation, 'the dead launch applied its crash completion')
.toBeGreaterThan(checkpoint);
// Recovery stays intact: a rerun on the replacement generation works.
const rerunGeneration = await runSimulation(page);
expect(rerunGeneration).toBeGreaterThan(receipt.generation);
expect(findNativeFailure([...testLogger.consoleLogs, ...testLogger.errors]),
'no wasm abort during the recovery').toBeUndefined();
});
test('a corrupted get_vec length is clamped, not copied out of bounds (E-11)', async ({ page, testLogger }) => {
await page.goto('/kicad/eeschema.html');
await waitForEditorReady(page);
await loadRectifier(page);
await openSimulator(page);
await waitForRunToolEnabled(page);
// Arm BEFORE the run: the next vector pull reports a ~5e8-element
// length while its arrays stay ~101 elements. (Frame-open plot
// restoration also pulls vectors; whichever pull the arm hits, the
// corrupted answer flows through the same client prepare.)
await page.evaluate(() => {
(globalThis as any).__ngspiceServiceTestHooks.corruptNextGetVec();
});
// THE E-11 oracle: on the unfixed build the client copies v_length
// doubles from the small buffer. Observed death shape on this build:
// the 4 GiB std::vector throws an UNHANDLED std::length_error that
// exits the editor's main loop — the scheduler shuts down and the
// whole session is dead (an OOB trap is the sibling shape). Fixed,
// the length clamps to the transferred arrays and the run completes.
await runSimulation(page);
const scheduler = await page.evaluate(() => ({
dead: (globalThis as any).__wxScheduler?.dead === true,
terminal: (globalThis as any).__wxScheduler?.terminal === true,
}));
expect(scheduler.dead,
'the corrupted vector must not exit the editor main loop').toBe(false);
expect(scheduler.terminal, 'the editor instance must not be terminal').toBe(false);
const fatal = findNativeFailure([...testLogger.consoleLogs, ...testLogger.errors]);
expect(fatal, `no wasm trap from the corrupted vector (patterns: ${
FATAL_WASM_PATTERNS.join(', ')})`).toBeUndefined();
});
});