Make the next e2e hang report itself, and add a soak to provoke one #157

Merged
cruelacid merged 1 commit from worktree-nec-89-hang-report into main 2026-09-24 20:35:18 +01:00
Owner

Task: NEC-89

The e2e hang has struck six CI runs. Nothing has captured it yet: every diagnostic read went through the renderer that had stopped answering. This makes the next occurrence report itself from the host, and adds a soak to provoke one.

  • packages/e2e/test/multi/hang-watch.ts, installed via startVault:
    • Every executeObsidian is bounded at 120 s and deleteSession at 30 s.
    • On a timeout it prints a NEC-89 hang report. The report holds the renderer heartbeat files, each vault's diag.log read from disk, /proc process state and CPU, /dev/shm, and the Chromium log.
    • It then kills the hung vault, so that afterAll still closes the server.
  • ws-server.ts: the Client disconnected line carries cid, close code, reason and lifetime. Both e2e vaults connect as the same user, so without these the two sockets can't be told apart.
  • soak.mjs and e2e-soak.yml loop external-edit and repro-runaway on the runner, from a fresh launch each time. The workflow runs on workflow_dispatch, or on a branch by bumping its marker. It is not a gate.
  • README: the re-examination of every failed e2e job. All six hangs share one signature: one socket closes 1–3 ms after the second vault's opens and never reconnects. No passing control shows it.

Mutation-checked locally:

  • A renderer spun in for(;;) is reported at state R, with its heartbeat stopped while the other vault's continues.
  • A renderer that calls process.crash() surfaces at once as tab crashed, not as a hang.
  • A normal run prints no report.
  • The soak exits 1 on a hang and 0 on clean iterations.
  • The full test:e2e:multi suite passes locally: 69 passed, 2 skipped.

Changelog

NONE

🤖 Generated with Claude Code

https://claude.ai/code/session_01Q6gkmHZ3pYvsZCCpiA9qe9

Task: NEC-89 The e2e hang has struck six CI runs. Nothing has captured it yet: every diagnostic read went through the renderer that had stopped answering. This makes the next occurrence report itself from the host, and adds a soak to provoke one. - **`packages/e2e/test/multi/hang-watch.ts`**, installed via `startVault`: - Every `executeObsidian` is bounded at 120 s and `deleteSession` at 30 s. - On a timeout it prints a `NEC-89 hang report`. The report holds the renderer heartbeat files, each vault's `diag.log` read from disk, `/proc` process state and CPU, `/dev/shm`, and the Chromium log. - It then kills the hung vault, so that `afterAll` still closes the server. - **`ws-server.ts`**: the `Client disconnected` line carries `cid`, close code, reason and lifetime. Both e2e vaults connect as the same user, so without these the two sockets can't be told apart. - **`soak.mjs` and `e2e-soak.yml`** loop `external-edit` and `repro-runaway` on the runner, from a fresh launch each time. The workflow runs on `workflow_dispatch`, or on a branch by bumping its marker. It is not a gate. - **README**: the re-examination of every failed e2e job. All six hangs share one signature: one socket closes 1–3 ms after the second vault's opens and never reconnects. No passing control shows it. Mutation-checked locally: - A renderer spun in `for(;;)` is reported at state R, with its heartbeat stopped while the other vault's continues. - A renderer that calls `process.crash()` surfaces at once as `tab crashed`, not as a hang. - A normal run prints no report. - The soak exits 1 on a hang and 0 on clean iterations. - The full `test:e2e:multi` suite passes locally: 69 passed, 2 skipped. ## Changelog NONE 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01Q6gkmHZ3pYvsZCCpiA9qe9
Make the next e2e hang report itself, and add a soak to provoke one
Some checks failed
Release note / release-note (pull_request) Successful in 11s
CI / e2e (pull_request) Successful in 5m5s
CI / build (pull_request) Successful in 5m5s
CI / promote (pull_request) Has been skipped
e2e-soak / soak (push) Failing after 28m59s
f74623151b
Six CI runs have hung the same way (NEC-89): one vault stops answering
WebDriver at startup and the run waits out 300 s in silence. Re-examining
every failed e2e job found one signature across all six (runs 468, 504,
568, 611, 615, 692): the first execute of the test hangs, and one socket
closes 1-3 ms after the second vault's opens and never reconnects. No
passing control shows that. Everything the suite had for diagnosing it
read through the renderer that had stopped answering, so six occurrences
produced no data.

- hang-watch.ts bounds executeObsidian at 120 s and deleteSession at 30 s,
  via startVault, so every multi file gets it. On a timeout it prints a
  "NEC-89 hang report" read entirely from the host: each renderer's
  heartbeat file, each vault's diag.log from disk, process state and CPU
  from /proc, /dev/shm, and Chromium's log. It writes to fd 2 because
  vitest swallowed console output in a local run, report included.
- The server logs cid, close code, reason and lifetime on disconnect, so
  a socket the plugin closed (1000/1001) can be told from one that
  dropped (1006), and paired with its connect.
- soak.mjs and e2e-soak.yml loop the two files from a fresh launch each
  time, on the runner, and dump the evidence in full on a hang.

Mutation-checked locally: a renderer spun with for(;;) is reported at
state R with its heartbeat stopped while the other vault's continues, and
killed at teardown; process.crash() surfaces at once as "tab crashed",
not as a hang; a normal run prints nothing; the soak stops at a hang and
exits 1, and exits 0 on clean iterations. The server test fails with the
old log line.

Task: NEC-89

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q6gkmHZ3pYvsZCCpiA9qe9
cruelacid force-pushed worktree-nec-89-hang-report from f74623151b
Some checks failed
Release note / release-note (pull_request) Successful in 11s
CI / e2e (pull_request) Successful in 5m5s
CI / build (pull_request) Successful in 5m5s
CI / promote (pull_request) Has been skipped
e2e-soak / soak (push) Failing after 28m59s
to 8884a58c9e
All checks were successful
Release note / release-note (pull_request) Successful in 15s
CI / build (pull_request) Successful in 4m35s
CI / e2e (pull_request) Successful in 4m50s
CI / promote (pull_request) Has been skipped
CI / build (push) Successful in 5m8s
CI / e2e (push) Successful in 5m23s
CI / promote (push) Successful in 30s
e2e-soak / soak (push) Successful in 19m8s
2026-09-24 20:30:15 +01:00
Compare
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
Nectenda/nectenda!157
No description provided.