No flake ledger: per-scenario e2e history does not survive a run #93

Closed
opened 2026-09-22 17:38:16 +01:00 by cruelacid · 1 comment
Owner

Why

This is the concrete mechanism for work this repo has already specified.
packages/e2e/README.md § "Stability is not yet established" lists five things
that would establish suite stability; items 2, 3 and 5 are exactly what a
flake ledger delivers:

Item 2 — Record per-scenario duration, not just pass or fail. A scenario
drifting from 800ms to 8s is a race narrowing before it starts failing, and
today nothing would notice until it failed.

Item 3 — Keep a run history rather than reading the latest. "Which
scenarios have ever failed, and how often" is not answerable right now without
paging through run logs by hand, which means in practice nobody will.

Item 5 — Decide the policy before it is needed. There is currently no
answer to "the nightly went red, is it the change or the suite?"

docs/plan-master.md carries the same item as Phase 9's one outstanding piece.

It is live, not theoretical. The README records CI run 468 on 22 September
2026: external-edit.test.ts hung for five minutes and timed out at 300s, taking
a second test down with it, against three commits that touched nothing under
packages/. Not reproducible — four local passes immediately after. That is the
second observation at an unknown rate, and the README says it plainly: "Two
observations at an unknown rate is still not a measurement."

What exists today

Nightly cron at 03:00 in .forgejo/workflows/e2e.yml, and nothing
machine-readable comes out of it
: no vitest JSON reporter, no outputFile, no
artifact upload. A run's per-scenario result exists only inside a zstd job log on
the Forgejo host (/data/gitea/actions_log/..., keyed by task id — see
CLAUDE.md). So "has recovery.test.ts ever failed, and how often" is a
by-hand log-paging exercise, which means nobody will do it.

The total-duration band (162s / 166s / 163s across three identical-code runs) is
currently our best signal, and the README notes it works this way by accident.
The per-scenario version should be deliberate.

How Relay does it

Read 22 September 2026 from No-Instructions/Relay,
.github/workflows/e2e-burnin.yml (887 lines, public and readable — unlike their
unit tests, see #92).

Their header states the purpose exactly:

"every night's per-plan pass/fail is appended to a persistent R2 ledger
(s3://system3-ci/burnin/ledger.jsonl) so flake RATES accumulate across nights
rather than resetting with each run's artifacts."

The structural insight is that last clause. CI artifacts are ephemeral and
scoped to one run, so a flake rate cannot live in them. The ledger has to sit
outside CI entirely.

Mechanics worth copying

Their design Detail
Append-only JSONL in object storage One row per plan per night. S3 has no append, so the job pulls the whole file, appends, and puts it back. Crude and completely adequate.
Three verdicts, not two pass, fail, infra. Infra misses are excluded from the rate in both directions — "they are not a product signal". Without this the rate measures your CI provider, not your suite.
An expected-red map const EXPECTED_RED = { tp058: 'a fork created mid-outage never reconciles' } — plans with an open tracked cause are surfaced but excluded from the rate, with the comment "Keep this map current as bugs resolve."
Isolation per scenario One freshly-leased VM per plan, one matrix leg each, "so a flake in one plan cannot contaminate another's verdict".
Rolling rate across all nights Summary table of Plan / Pass rate / Runs / Nights / Last / Note, rendered into the run summary and published at a stable URL.
A graduation gate Slow, timing-sensitive or render-heavy plans "must prove flake-free here BEFORE they enter the per-push suite". The ledger is the evidence for promotion.
Membership from one manifest A shared suite manifest flags burnIn: true; both e2e workflows resolve their set from it rather than each listing tests.

Worth noting what they don't do: no retries, and no auto-quarantine. The ledger
informs a human decision; it does not silently hide anything.

What ours should be

Smaller than theirs. We have 34 scenarios in 10 files and one runner, not a VM
pool — the ledger is the valuable half, and the fleet management is not.

  • Emit machine-readable results. A vitest JSON reporter with outputFile in
    e2e.yml. Nothing downstream is possible until this exists, and it is the
    cheapest step.
  • One row per scenario per run, not per file: scenario id, verdict, duration
    in ms
    , run id, commit, Obsidian version, date. Duration is item 2 of the
    README list and is the part their design does not emphasise — for us it is
    the early-warning signal, because every timing bug here has been a window that
    is narrow on a developer machine and wide somewhere else.
  • Adopt pass / fail / infra and the expected-red map. We already have a
    case for the latter: the README notes the Obsidian-version job "tracks
    Obsidian's release schedule and can go red with no code change."
  • Somewhere durable that is not a CI artifact. We already run S3-compatible
    storage (packages/server/src/blob-store.ts, and Backblaze for backups), so a
    bucket is the path of least resistance.
  • Answer item 5 in writing: what to do when the nightly is red. The README is
    unambiguous about why this matters — "a suite people re-run on failure has
    stopped being evidence of anything."

Check first: Forgejo may already hold half of this

action_run and action_task in Forgejo's SQLite already carry per-run status
and timestamps, and CLAUDE.md documents how to query them. That gives job-level
history for free — worth confirming before building anything.

What it cannot give is per-scenario granularity, which is the part that
matters. So the likely answer is: derive run-level history from Forgejo, emit
scenario-level rows ourselves. Check before building either.

The honest counterweight

The README's own datum argues against urgency: "of eight red CI runs to date,
every single one was a real fault in the code, the pipeline or the test's own
assumptions, and none was noise."
Our base rate of genuine failures is high, so
the re-run habit has not started.

The ledger is how that claim stays checkable as run volume grows — and run 468 is
the first data point that does not fit it. It is also, per the README, second in
value to simply letting the nightly accumulate runs, which costs only
patience. Do not let building this delay that.

Verification

  • Inverting a verdict must change the computed rate — mutation-check it in both
    directions, per CLAUDE.md. A rate that cannot move is a number, not a
    measurement.
  • Assert the ledger survives a run whose artifacts have expired. That is the
    whole point of it, and it is the one property an artifact-based version would
    silently fail.
## Why **This is the concrete mechanism for work this repo has already specified.** `packages/e2e/README.md` § *"Stability is not yet established"* lists five things that would establish suite stability; items **2, 3 and 5** are exactly what a flake ledger delivers: > **Item 2 — Record per-scenario duration, not just pass or fail.** A scenario > drifting from 800ms to 8s is a race narrowing before it starts failing, and > today nothing would notice until it failed. > > **Item 3 — Keep a run history rather than reading the latest.** "Which > scenarios have ever failed, and how often" is not answerable right now without > paging through run logs by hand, which means in practice nobody will. > > **Item 5 — Decide the policy before it is needed.** There is currently no > answer to "the nightly went red, is it the change or the suite?" `docs/plan-master.md` carries the same item as Phase 9's one outstanding piece. **It is live, not theoretical.** The README records CI run 468 on 22 September 2026: `external-edit.test.ts` hung for five minutes and timed out at 300s, taking a second test down with it, against three commits that touched nothing under `packages/`. Not reproducible — four local passes immediately after. That is the second observation at an unknown rate, and the README says it plainly: *"Two observations at an unknown rate is still not a measurement."* ## What exists today Nightly cron at 03:00 in `.forgejo/workflows/e2e.yml`, and **nothing machine-readable comes out of it**: no vitest JSON reporter, no `outputFile`, no artifact upload. A run's per-scenario result exists only inside a zstd job log on the Forgejo host (`/data/gitea/actions_log/...`, keyed by task id — see `CLAUDE.md`). So "has `recovery.test.ts` ever failed, and how often" is a by-hand log-paging exercise, which means nobody will do it. The total-duration band (162s / 166s / 163s across three identical-code runs) is currently our best signal, and the README notes it *works this way by accident*. The per-scenario version should be deliberate. ## How Relay does it Read 22 September 2026 from `No-Instructions/Relay`, `.github/workflows/e2e-burnin.yml` (887 lines, public and readable — unlike their unit tests, see #92). Their header states the purpose exactly: > *"every night's per-plan pass/fail is appended to a persistent R2 ledger > (`s3://system3-ci/burnin/ledger.jsonl`) so flake RATES accumulate across nights > rather than resetting with each run's artifacts."* **The structural insight is that last clause.** CI artifacts are ephemeral and scoped to one run, so a flake rate cannot live in them. The ledger has to sit outside CI entirely. ### Mechanics worth copying | Their design | Detail | |---|---| | **Append-only JSONL in object storage** | One row per plan per night. S3 has no append, so the job pulls the whole file, appends, and puts it back. Crude and completely adequate. | | **Three verdicts, not two** | `pass`, `fail`, **`infra`**. Infra misses are excluded from the rate in both directions — *"they are not a product signal"*. Without this the rate measures your CI provider, not your suite. | | **An expected-red map** | `const EXPECTED_RED = { tp058: 'a fork created mid-outage never reconciles' }` — plans with an open tracked cause are surfaced but excluded from the rate, with the comment *"Keep this map current as bugs resolve."* | | **Isolation per scenario** | One freshly-leased VM per plan, one matrix leg each, *"so a flake in one plan cannot contaminate another's verdict"*. | | **Rolling rate across all nights** | Summary table of `Plan / Pass rate / Runs / Nights / Last / Note`, rendered into the run summary and published at a stable URL. | | **A graduation gate** | Slow, timing-sensitive or render-heavy plans *"must prove flake-free here BEFORE they enter the per-push suite"*. The ledger is the evidence for promotion. | | **Membership from one manifest** | A shared suite manifest flags `burnIn: true`; both e2e workflows resolve their set from it rather than each listing tests. | Worth noting what they *don't* do: no retries, and no auto-quarantine. The ledger informs a human decision; it does not silently hide anything. ## What ours should be Smaller than theirs. We have 34 scenarios in 10 files and one runner, not a VM pool — the ledger is the valuable half, and the fleet management is not. - **Emit machine-readable results.** A vitest JSON reporter with `outputFile` in `e2e.yml`. Nothing downstream is possible until this exists, and it is the cheapest step. - **One row per scenario per run**, not per file: scenario id, verdict, **duration in ms**, run id, commit, Obsidian version, date. Duration is item 2 of the README list and is the part their design does *not* emphasise — for us it is the early-warning signal, because every timing bug here has been a window that is narrow on a developer machine and wide somewhere else. - **Adopt `pass` / `fail` / `infra`** and the expected-red map. We already have a case for the latter: the README notes the Obsidian-version job *"tracks Obsidian's release schedule and can go red with no code change."* - **Somewhere durable that is not a CI artifact.** We already run S3-compatible storage (`packages/server/src/blob-store.ts`, and Backblaze for backups), so a bucket is the path of least resistance. - **Answer item 5 in writing**: what to do when the nightly is red. The README is unambiguous about why this matters — *"a suite people re-run on failure has stopped being evidence of anything."* ### Check first: Forgejo may already hold half of this `action_run` and `action_task` in Forgejo's SQLite already carry per-run status and timestamps, and `CLAUDE.md` documents how to query them. That gives **job-level** history for free — worth confirming before building anything. What it cannot give is **per-scenario** granularity, which is the part that matters. So the likely answer is: derive run-level history from Forgejo, emit scenario-level rows ourselves. Check before building either. ## The honest counterweight The README's own datum argues against urgency: *"of eight red CI runs to date, every single one was a real fault in the code, the pipeline or the test's own assumptions, and none was noise."* Our base rate of genuine failures is high, so the re-run habit has not started. The ledger is how that claim stays checkable as run volume grows — and run 468 is the first data point that does not fit it. It is also, per the README, second in value to simply **letting the nightly accumulate runs**, which costs only patience. Do not let building this delay that. ## Verification - Inverting a verdict must change the computed rate — mutation-check it in both directions, per `CLAUDE.md`. A rate that cannot move is a number, not a measurement. - Assert the ledger survives a run whose artifacts have expired. That is the whole point of it, and it is the one property an artifact-based version would silently fail.
Author
Owner

Moved to the Vikunja board as NEC-69: https://projectron.nerchure.com/tasks/69

Moved to the Vikunja board as **NEC-69**: https://projectron.nerchure.com/tasks/69
Sign in to join this conversation.
No milestone
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
Nectenda/nectenda#93
No description provided.