No flake ledger: per-scenario e2e history does not survive a run #93
Labels
No labels
area:docs
area:identity
area:ops
area:plugin
area:server
channel:community
channel:direct
channel:owned
channel:press
channel:social
e2ee-constrained
gate:at-ga
gate:pre-ga
marketing
parity
relay:absent
relay:planned
relay:requested
relay:supported
risk:additive
risk:contract
risk:none
usability
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
Nectenda/nectenda#93
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Why
This is the concrete mechanism for work this repo has already specified.
packages/e2e/README.md§ "Stability is not yet established" lists five thingsthat would establish suite stability; items 2, 3 and 5 are exactly what a
flake ledger delivers:
docs/plan-master.mdcarries the same item as Phase 9's one outstanding piece.It is live, not theoretical. The README records CI run 468 on 22 September
2026:
external-edit.test.tshung for five minutes and timed out at 300s, takinga second test down with it, against three commits that touched nothing under
packages/. Not reproducible — four local passes immediately after. That is thesecond observation at an unknown rate, and the README says it plainly: "Two
observations at an unknown rate is still not a measurement."
What exists today
Nightly cron at 03:00 in
.forgejo/workflows/e2e.yml, and nothingmachine-readable comes out of it: no vitest JSON reporter, no
outputFile, noartifact upload. A run's per-scenario result exists only inside a zstd job log on
the Forgejo host (
/data/gitea/actions_log/..., keyed by task id — seeCLAUDE.md). So "hasrecovery.test.tsever failed, and how often" is aby-hand log-paging exercise, which means nobody will do it.
The total-duration band (162s / 166s / 163s across three identical-code runs) is
currently our best signal, and the README notes it works this way by accident.
The per-scenario version should be deliberate.
How Relay does it
Read 22 September 2026 from
No-Instructions/Relay,.github/workflows/e2e-burnin.yml(887 lines, public and readable — unlike theirunit tests, see #92).
Their header states the purpose exactly:
The structural insight is that last clause. CI artifacts are ephemeral and
scoped to one run, so a flake rate cannot live in them. The ledger has to sit
outside CI entirely.
Mechanics worth copying
pass,fail,infra. Infra misses are excluded from the rate in both directions — "they are not a product signal". Without this the rate measures your CI provider, not your suite.const EXPECTED_RED = { tp058: 'a fork created mid-outage never reconciles' }— plans with an open tracked cause are surfaced but excluded from the rate, with the comment "Keep this map current as bugs resolve."Plan / Pass rate / Runs / Nights / Last / Note, rendered into the run summary and published at a stable URL.burnIn: true; both e2e workflows resolve their set from it rather than each listing tests.Worth noting what they don't do: no retries, and no auto-quarantine. The ledger
informs a human decision; it does not silently hide anything.
What ours should be
Smaller than theirs. We have 34 scenarios in 10 files and one runner, not a VM
pool — the ledger is the valuable half, and the fleet management is not.
outputFileine2e.yml. Nothing downstream is possible until this exists, and it is thecheapest step.
in ms, run id, commit, Obsidian version, date. Duration is item 2 of the
README list and is the part their design does not emphasise — for us it is
the early-warning signal, because every timing bug here has been a window that
is narrow on a developer machine and wide somewhere else.
pass/fail/infraand the expected-red map. We already have acase for the latter: the README notes the Obsidian-version job "tracks
Obsidian's release schedule and can go red with no code change."
storage (
packages/server/src/blob-store.ts, and Backblaze for backups), so abucket is the path of least resistance.
unambiguous about why this matters — "a suite people re-run on failure has
stopped being evidence of anything."
Check first: Forgejo may already hold half of this
action_runandaction_taskin Forgejo's SQLite already carry per-run statusand timestamps, and
CLAUDE.mddocuments how to query them. That gives job-levelhistory for free — worth confirming before building anything.
What it cannot give is per-scenario granularity, which is the part that
matters. So the likely answer is: derive run-level history from Forgejo, emit
scenario-level rows ourselves. Check before building either.
The honest counterweight
The README's own datum argues against urgency: "of eight red CI runs to date,
every single one was a real fault in the code, the pipeline or the test's own
assumptions, and none was noise." Our base rate of genuine failures is high, so
the re-run habit has not started.
The ledger is how that claim stays checkable as run volume grows — and run 468 is
the first data point that does not fit it. It is also, per the README, second in
value to simply letting the nightly accumulate runs, which costs only
patience. Do not let building this delay that.
Verification
directions, per
CLAUDE.md. A rate that cannot move is a number, not ameasurement.
whole point of it, and it is the one property an artifact-based version would
silently fail.
Moved to the Vikunja board as NEC-69: https://projectron.nerchure.com/tasks/69