KPI dashboard: Grafana on clconsole fed by read-only /api/kpi routes #184

Merged
cruelacid merged 1 commit from worktree-kpi-dashboard into main 2026-09-26 16:50:38 +01:00
Owner

Task: NEC-122

A private KPI dashboard: Grafana on clconsole, fed every 15 minutes by a collector that reads new read-only GET /api/kpi routes on the identity service and each shard.

What changes on the fleet

  • identity and server each gain GET /api/kpi, guarded by a new optional KPI_TOKEN. Leave it unset and the route answers 404, so this ships inert until the token is set on the hosts. The route refuses ADMIN_TOKEN, the directory token and user sessions. Both routes only read existing tables: no schema change and no new retention.
  • The payloads are pseudonymous: people appear as the identity service's random UUID, with no email, name, device label, document name (not even its HMAC), folder name, ciphertext or key material. Tests plant canaries for each and mutation-check the leak assertion.
  • /api/admin/users?id= now resolves a dashboard id to a person, using the admin token.
  • docs/security-model.md gains a "Usage metrics, which we keep internally" bullet.

Two things that weren't obvious

  • Sign-in method comes from auth_flows.method. completeFlow already wrote it, but nothing read it. A finished flow is kept for about an hour after it expires, so the collector polls every 15 minutes. A flow counts as a sign-up only when it is the user's first flow and the user was created while it was open, because a second sign-in within 10 minutes of signing up would otherwise count as one too. Sign-ups from before collection are inferred: a provider counts as the sign-up method if it was linked within 5 s of the account's creation.
  • Folder activity is the per-document max seq summed per folder, not the number of doc_updates rows. Snapshots delete the updates they absorb, so a row count undercounts. seq is max+1 across the log and the snapshot, so it counts every update ever made. A test compacts, then checks the total.

deploy/clconsole/kpi

Postgres 18 with Grafana 13.0.2 (OSS, pinned by digest), a run-to-completion collector container, systemd user timers, and a nightly pg_dump kept for 30 days. Every metric is defined as a view in collector/schema.sql. The dashboard JSON is generated by grafana/build-dashboard.mjs, and a test fails if the committed file drifts from what the generator produces. Install steps are in deploy/clconsole/README.md.

Verified locally

  • pnpm test, pnpm -r typecheck, pnpm lint and build-mirror --check all pass.
  • End to end: a real identity service and a real shard, seeded with email and Google sign-ups, a returning sign-in and a compacted folder, with the stack in Docker.
    • Two collector runs left identical row counts.
    • A run with a bad token wrote only its failure.
    • All 36 panel queries ran through Grafana's read-only role.
    • The dashboard renders.
    • The download backfill read 60 commits from obsidian-releases and wrote the plugin's 8 listed days. The total, 212, matched the plugin page.
  • Mutation checks, each confirmed to fail: widening the 5 s inference window, accepting the admin token, leaking doc_name, counting rows instead of seq, and dating active days by the run instead of by last-seen.

Changelog

NONE

🤖 Generated with Claude Code

https://claude.ai/code/session_01MY63BZ4UVtMHFcADNr1jFg

Task: NEC-122 A private KPI dashboard: Grafana on clconsole, fed every 15 minutes by a collector that reads new read-only `GET /api/kpi` routes on the identity service and each shard. ## What changes on the fleet - **identity** and **server** each gain `GET /api/kpi`, guarded by a new optional `KPI_TOKEN`. Leave it unset and the route answers 404, so this ships inert until the token is set on the hosts. The route refuses `ADMIN_TOKEN`, the directory token and user sessions. Both routes only read existing tables: **no schema change and no new retention**. - The payloads are pseudonymous: people appear as the identity service's random UUID, with no email, name, device label, document name (not even its HMAC), folder name, ciphertext or key material. Tests plant canaries for each and mutation-check the leak assertion. - `/api/admin/users?id=` now resolves a dashboard id to a person, using the admin token. - `docs/security-model.md` gains a "Usage metrics, which we keep internally" bullet. ## Two things that weren't obvious - **Sign-in method** comes from `auth_flows.method`. `completeFlow` already wrote it, but nothing read it. A finished flow is kept for about an hour after it expires, so the collector polls every 15 minutes. A flow counts as a sign-up only when it is the user's first flow and the user was created while it was open, because a second sign-in within 10 minutes of signing up would otherwise count as one too. Sign-ups from before collection are inferred: a provider counts as the sign-up method if it was linked within 5 s of the account's creation. - **Folder activity** is the per-document max `seq` summed per folder, not the number of `doc_updates` rows. Snapshots delete the updates they absorb, so a row count undercounts. `seq` is max+1 across the log and the snapshot, so it counts every update ever made. A test compacts, then checks the total. ## deploy/clconsole/kpi Postgres 18 with Grafana 13.0.2 (OSS, pinned by digest), a run-to-completion collector container, systemd user timers, and a nightly `pg_dump` kept for 30 days. Every metric is defined as a view in `collector/schema.sql`. The dashboard JSON is generated by `grafana/build-dashboard.mjs`, and a test fails if the committed file drifts from what the generator produces. Install steps are in `deploy/clconsole/README.md`. ## Verified locally - `pnpm test`, `pnpm -r typecheck`, `pnpm lint` and `build-mirror --check` all pass. - End to end: a real identity service and a real shard, seeded with email and Google sign-ups, a returning sign-in and a compacted folder, with the stack in Docker. - Two collector runs left identical row counts. - A run with a bad token wrote only its failure. - All 36 panel queries ran through Grafana's read-only role. - The dashboard renders. - The download backfill read 60 commits from obsidian-releases and wrote the plugin's 8 listed days. The total, 212, matched the plugin page. - Mutation checks, each confirmed to fail: widening the 5 s inference window, accepting the admin token, leaking `doc_name`, counting rows instead of `seq`, and dating active days by the run instead of by last-seen. ## Changelog NONE 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01MY63BZ4UVtMHFcADNr1jFg
Add a KPI dashboard: Grafana on clconsole, fed by read-only /api/kpi routes
All checks were successful
Release note / release-note (pull_request) Successful in 13s
CI / build (pull_request) Successful in 4m42s
CI / e2e (pull_request) Successful in 5m16s
CI / promote (pull_request) Has been skipped
Deploy site / deploy (push) Successful in 49s
CI / build (push) Successful in 4m57s
CI / e2e (push) Successful in 5m17s
CI / promote (push) Successful in 32s
51f1e395e3
A private dashboard for the numbers that steer the product: plugin downloads,
sign-ups and sign-ins by method, passkey adoption, active users, a 30-day
activity leaderboard, organisation sizes, retention cohorts, platforms,
plugin versions, feature use and revenue.

The identity service and each shard gain GET /api/kpi, guarded by a new
KPI_TOKEN that opens that route and nothing else, because the collector runs
on clconsole, outside the fleet. Both read existing tables only. No schema
changes and no new retention. People are the identity service's random ids;
no email, name, label, document name or key material leaves either service.

- Sign-in method comes from auth_flows.method, which completeFlow already
  wrote and nothing read. The collector polls every 15 minutes, inside the
  hour a finished flow is kept. Sign-ups from before collection are inferred
  from whether the first provider link was made with the account.
- Folder activity is each folder's lifetime update count (sum of per-document
  max seq), because snapshots compact the update log and a row count
  undercounts.

deploy/clconsole/kpi/ is the stack: Postgres, Grafana 13, a run-to-completion
collector, systemd user timers, a nightly pg_dump, and a generated dashboard
whose metric definitions all live as views in schema.sql.
docs/security-model.md says what is copied there.

Task: NEC-122

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MY63BZ4UVtMHFcADNr1jFg
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
Nectenda/nectenda!184
No description provided.