Compare commits

..

110 Commits

Author SHA1 Message Date
Ben Stull 242831a4a0 Merge pull request 'fix(§6/auth): reconcile by email in provision_user so OTC→OAuth doesn't 500 (v0.54.1)' (#50) from fix-provision-reconcile into main 2026-06-09 12:01:49 +00:00
Ben Stull f0533dc073 fix(§6/auth): reconcile by email in provision_user so OTC→OAuth doesn't 500 (v0.54.1)
The Gitea-OAuth callback's `auth.provision_user` matched an existing `users` row
only by `gitea_id`. A human who signed in first via OTC owns an email-only row
(`gitea_id` NULL); a later Gitea-OAuth sign-in with the same email missed that
row and INSERTed a new one, colliding on the `idx_users_email` case-insensitive
unique index → IntegrityError → 500 callback.

Fix (mirror of `otc.provision_or_link_user`): when no `gitea_id` row exists,
reconcile by email and link the OAuth identity (gitea_id/gitea_login/profile)
onto the existing email row. Owner-zero (§6.1) bootstrap applied on link
(matching the fresh-insert path); `permission_state` preserved so linking never
changes admission status as a side effect.

One §9-surfaced framework fragility of two. backend 680 green (+3).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 05:01:02 -07:00
Ben Stull f7b93d797c Merge pull request 'feat(§8.12/§8.3): main-view Ask cuts an edit branch and runs the question there (v0.54.0)' (#49) from fix-main-view-ask into main 2026-06-09 11:20:30 +00:00
Ben Stull 24596842ea feat(§8.12/§8.3): main-view Ask cuts an edit branch and runs the question there (v0.54.0)
The AI "Ask" affordance (selection tooltip + prompt bar) had no response
surface on an entry's canonical `main` view: RFCView renders the human-
discussion panel there, not the chat panel, so a chat turn posted from the
reading view had nowhere to render — asking-while-reading silently did
nothing. AI chat is an editing activity (a turn can emit <change> proposals)
and only runs on an edit branch.

Option B: when Ask is invoked from main, transparently cut an edit branch via
the same dispatch as Start Contributing (promote-to-branch for active,
start-edit-branch for super-draft; idempotent — reuse an existing edit
branch), navigate onto it so the chat panel mounts, and run the question
(text + selected quote) as the branch's first chat turn once its view and
main_thread_id resolve. The turn fires from the message-load effect's
continuation (not a parallel effect) so a late message load can't clobber the
optimistic turn; a live ref keeps the latest submitChatTurn closure in reach.

Degrades gracefully: signed-out → login; a viewer who can't contribute has the
cut rejected server-side and the error surfaced (no spurious branch); asking
from a branch already, and Flag on main (human discussion thread), unchanged.

Frontend-only; the branch chat endpoints already work on a branch. SPEC §8.3
records the behavior; new RFCView unit tests + an e2e spec cover the flow.
backend 677 / frontend 66 / e2e 5 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 04:19:20 -07:00
Ben Stull ff88be2e91 Merge pull request 'fix(§22/G-15): three-tier-aware branch/edit/body subsystem (v0.53.0)' (#47) from fix-g15-three-tier-write-paths into main 2026-06-09 05:40:59 +00:00
Ben Stull 4e7410f90b fix(§22/G-15): make the branch/edit/body subsystem three-tier aware (v0.53.0)
§22 migrated the READ/catalog path to the three-tier (project/collection)
model but the WRITE/branch/body subsystem still hardcoded the default
project's default collection — resolving every meta-resident entry to the
default content repo at rfcs/<slug>.md, ignoring the entry's project (its own
content repo) and collection (a <subfolder>/rfcs/ prefix). An entry outside
the default collection rendered a blank canonical body and every edit/PR/
body-write path hit the wrong file.

- New single resolver: projects.content_repo_for_collection() +
  projects.entry_location(config, cid, slug) -> (org, repo, md_path)
  (collection -> project -> content_repo + subfolder; falls back to the
  default repo for a legacy/unknown collection).
- Every entry write/read path resolves via it: api_branches (body GET + all
  branch write paths), api_prs (pr-draft/open/merge/withdraw/review/
  resolution-branch), api_graduation (graduate/claim/retire/unretire +
  orchestrator + state-flip), api.py mark-reviewed + idea-PR merge/decline/
  withdraw + proposal preview, api_metadata. bot.open_metadata_pr and
  bot.mark_entry_reviewed gained a file_path param. refresh_meta_branches,
  the webhook corpus-refresh dispatch, and the hygiene branch-delete now
  span every project's content repo, not just the default.
- New additive collection-scoped body-read routes:
  GET /api/projects/{pid}/collections/{cid}/rfcs/{slug}/main and
  .../branches/{branch} disambiguate a slug across collections (G-5) and read
  the entry's own repo. Slug-only routes kept (now collection-aware via the
  cached row). Frontend getRFCMain/getBranch take optional pid+cid; RFCView
  threads them (mirrors the v0.52.1 entry-detail fix).

No migration, no config change. Existing default-collection entries
unaffected. Tests: backend resolver + collection-scoped branch/body + graduate-
in-subfolder write path; frontend api unit. backend 677 / frontend 60 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 22:39:46 -07:00
Ben Stull afa8d26378 Merge pull request 'fix(routing): 404 page for unmatched routes (v0.52.3)' (#45) from fix-404-unmatched-routes into main 2026-06-09 03:14:39 +00:00
Ben Stull 948ee88160 fix(routing): render a 404 page for unmatched routes (v0.52.3)
The top-level catch-all silently redirected unknown paths to "/", and the
nested /admin/*, /p/:projectId/*, and /docs/* route groups had no catch-all
(invalid subpaths rendered a blank pane). Add a shared NotFound component and
wire it into all four route groups so a bad/typo'd URL shows a clear 404 with
a link home. Client-side 404 UI (SPA still serves HTTP 200).

Patch 0.52.2 → 0.52.3; CHANGELOG updated. Caught by the operator on PPE.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 20:14:25 -07:00
Ben Stull 7bcf784d06 Merge pull request 'fix(admin): absolute sidebar nav links (v0.52.2)' (#44) from fix-admin-nav-relative-links into main 2026-06-09 03:08:25 +00:00
Ben Stull 8e207a60e6 fix(admin): absolute sidebar nav links so /admin URLs don't accumulate (v0.52.2)
The /admin left-rail NavLinks used relative targets (to="users", etc.). Under
the /admin/* nested route a relative link resolves against the current URL, so
each click appended a segment — Users→Allowlist→Graduation yielded
/admin/users/allowlist/graduation instead of /admin/graduation. Use absolute
targets (/admin/<tab>) + `end` for exact active matching.

Patch 0.52.1 → 0.52.2; CHANGELOG updated. Caught by the operator on PPE.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 20:07:44 -07:00
Ben Stull fbaa975b5c Merge pull request 'fix(§22.4a): scope RFCView entry-detail fetch to its collection (v0.52.1)' (#43) from fix-collection-scoped-entry-detail into main 2026-06-08 13:45:42 +00:00
Ben Stull ba37da927a fix(§22.4a): scope RFCView entry-detail fetch to its collection (v0.52.1)
The §9 deployed-environment E2E harness (0.52.0), run against a PPE host
with per-collection-isolated content, surfaced a latent multi-collection
bug: RFCView computed the collection id from the route but called
getRFC(pid, slug) without it, so a named-collection entry was always
fetched via the project default-collection route — which 404s for an entry
that exists only in a named collection ("Error: Not found"; metadata panel
absent). Local/Tier-1 stacks masked it (same slug also reachable via the
default collection). Thread cid through all three getRFC call sites; re-run
the load effect on collection change.

Harness/test-infra (not in the deployed artifact):
- e2e: pre-record cookie consent via addInitScript (lib/fixtures.js) so the
  bottom-fixed consent banner can't intercept catalog row-select clicks on
  the slower deployed edge.
- testing/seed-ppe.sh: fail loudly on any non-2xx Gitea response (a
  swallowed 403 org-repo create had reached the deploy as a 502).
- testing/ppe-deploy-and-test.sh: seed via the Keychain admin token
  (write:organization needed to create the PPE repos); store the E2E secret
  newline-free; read EXPECT_VERSION from VERSION.

Patch bump 0.52.0 → 0.52.1; CHANGELOG updated.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 06:45:01 -07:00
Ben Stull 9c8035bdbd test(e2e): one-shot PPE deploy+E2E resume script
Runs the whole §9 PPE stage non-interactively after the operator's gcloud
reauth: bot-token-seed the PPE repos -> ensure E2E secret -> deploy ->
wait for health=0.52.0 + bdd-collection sync -> run metadata.spec.js
against the deployed host. Idempotent; secrets fetched from SM, never
echoed. Test/ops infra (not in the deployed artifact).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 00:01:19 -07:00
Ben Stull fd123da6a3 test(e2e): harden harness for deployed runs (retries, serialize, banner)
Test-infra only (e2e/ is not in the deployed artifact). Refines the
v0.52.0 deployed-env harness:
- retries:2 + on-first-retry trace now meaningful (was dead: retries
  defaulted to 0); de-risks timing flakes over the public edge.
- workers:1 — the metadata specs run in order and write real commits;
  parallel workers would race on shared state and concurrent bot pushes.
- deployed timeouts bumped (60s/20s) when E2E_TEST_AUTH_SECRET is set.
- dismissCookies forcibly removes any lingering consent-banner node so it
  can't intercept catalog-footer checkbox clicks (the recurring flake).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 23:56:32 -07:00
Ben Stull 96e2214213 Merge pull request 'v0.52.0 — deployed-environment E2E harness + gated test-auth' (#42) from ppe-e2e-deployed-harness into main 2026-06-08 06:36:18 +00:00
Ben Stull 83eafe72ee feat(e2e): deployed-environment E2E harness + gated test-auth (v0.52.0)
The §9 pipeline's PPE+E2E stage was unreachable: e2e/metadata.spec.js was
bound to Tier-1-only scaffolding (docker-seeded faceted collection,
SQLite-injected owner, Mailpit OTC sink). This makes the same suite run
against a deployed host.

- backend: POST /auth/test/login — fail-closed, secret-gated, single-
  identity owner test-login (404 unless E2E_TEST_AUTH_SECRET +
  E2E_TEST_AUTH_EMAIL both set; constant-time compare; loud startup warn).
  6 vertical tests; backend 665 green. Documented in backend/.env.example.
- e2e/lib/auth.js: branch on E2E_TEST_AUTH_SECRET (deployed test-login vs
  local Mailpit OTC); OWNER_EMAIL from E2E_OWNER_EMAIL. Spec unchanged so
  the localhost Tier-1 path keeps working.
- testing/seed-ppe.sh: seed a dedicated, prod-untouching PPE registry +
  content repo (faceted bdd collection) — real OHM content never touched.
- docs/design/2026-06-07-deployed-env-e2e-harness.md; CHANGELOG; VERSION
  + frontend/package.json -> 0.52.0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 23:35:53 -07:00
Ben Stull dd9ceff69e Merge pull request 'v0.51.1 — collection-id divergence fix + faceted-catalog scoping fix + metadata E2E' (#41) from fix-mig029-collection-divergence into main 2026-06-08 05:30:47 +00:00
Ben Stull 52f465b4dd release: v0.51.1 — collection-id divergence + faceted-catalog scoping fixes
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 22:30:09 -07:00
Ben Stull 2fc7029bd9 test(e2e): §22-current Tier-1 harness + metadata E2E (SLICE-3/4/5)
Modernize the Tier-1 stack to the three-tier app and add browser coverage for
the §22.4a metadata UI, closing the E2E gap deferred across SLICE-3/4/5:
- seed-gitea.sh: create a REGISTRY_REPO with projects.yaml + a faceted named
  collection (.collection.yaml fields: priority enum + tags) seeded with three
  metadata-bearing entries; register content+registry webhooks; self-guarding
  (skip if a prior token still works) so a dependency-triggered re-run can't
  remint and invalidate the backend's token.
- .env.tier1: REGISTRY_REPO/DEFAULT_PROJECT_ID; disable OTC cooldown + lift the
  per-IP auth limiter for the single-IP test runner.
- docker-compose: pin backend image; backend-seed inserts a granted owner the
  OTC path can sign in as (write paths need contributor+).
- Makefile: two-phase tier1-up (seed to completion, then create backend so it
  reads the populated token env); robust down; e2e-fresh = down+up+e2e (the
  canonical run, since the edit/bulk specs mutate the seeded corpus).
- metadata.spec.js: SLICE-3 faceted filter (anon), SLICE-4 edit panel (owner),
  SLICE-5 bulk bar (owner). 4 passed against a fresh stack.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 22:29:11 -07:00
Ben Stull 9e1b7ce34f feat(ratelimit): env-tunable per-IP auth limiter budgets
Add RATELIMIT_OTC_REQUEST_MAX / RATELIMIT_VERIFY_MAX / RATELIMIT_CHECK_MAX
(default to the existing secure values; non-positive/unparseable → default).
Lets a test/PPE stack that drives the auth endpoints repeatedly from one IP
lift the budget; production leaves them unset. Mirrors the existing
OTC_REQUEST_COOLDOWN_SECONDS knob.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 22:28:57 -07:00
Ben Stull cbaba76345 fix(§22): catalog scopes to the named collection in the URL (E2E-caught)
useCollectionId read useParams().collectionId, but the Catalog renders at
/p/:projectId/* — outside the c/:collectionId route — so it always fell back
to the default collection. The faceted filter (SLICE-3) and bulk action bar
(SLICE-5) therefore never scoped to a named (fields-bearing) collection in the
deployed app; the Catalog unit test had mocked useCollectionId, hiding it.
Resolve the /c/<id>/ segment from the pathname when the route param isn't in
scope. Caught by the new Playwright metadata suite.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 22:28:57 -07:00
Ben Stull 620926b834 fix(§22): heal mig-029 vs registry-mirror collection-id divergence
On a deployment that already had >=2 projects when migration 029 ran, 029
seeds the default project's collection id as the project id (e.g. 'ohm'),
but the registry mirror expects 'default' -> it inserted a duplicate empty
'default' collection, orphaning the entries.

Adds projects.reconcile_default_collection_id() -- the collection-grain twin
of restamp_default_project -- run at startup BEFORE the mirror so it merges
onto the canonical 'default' collection instead of duplicating. Renames the
divergent collection id and cascades collection_id across all keyed tables
(FK-off atomic rename + foreign_key_check). Idempotent; no-op on fresh /
single-project / already-aligned deploys. Seam tests show the duplicate forms
without the fix and merges with it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 21:54:52 -07:00
Ben Stull 4ac3955056 Merge pull request 'v0.51.0 — SLICE-5: bulk tag/untag metadata (§22.4a PUC-2)' (#35) from slice5-bulk-metadata into main 2026-06-08 04:04:21 +00:00
Ben Stull 7886840362 fix(slice5): validate raw metadata input + prune stale bulk selections
Code-review follow-ups:
- Validate the raw submitted value before apply_values coerces it, in both
  the single-edit and bulk endpoints — a scalar set onto a tags field now
  rejects (422 / rejected) instead of silently char-splitting into a list.
- Catalog prunes bulk selections to the currently-visible (filtered) entries
  after each list fetch, so the bulk bar can't act on rows the user has
  filtered out of view.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 21:03:29 -07:00
Ben Stull 8eee907893 docs(slice5): record SLICE-5 bulk edit shipped in SPEC §22.4a status
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 20:52:48 -07:00
Ben Stull a2b55f94ce release(slice5): v0.51.0 — bulk tag/untag metadata (§22.4a PUC-2 SLICE-5)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 20:52:16 -07:00
Ben Stull edbf68909a feat(slice5): catalog multi-select + bulk action bar (§22.4a PUC-2)
Faceted catalog rows are selectable for contributors; a sticky bulk bar
(BulkActionBar) drives set/add/remove gestures from the collection fields
schema, calls bulkEntryMeta (one commit), toasts applied/skipped counts,
and re-fetches. Non-contributors and legacy (no-fields) collections see
no selection (INV-5).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 20:51:48 -07:00
Ben Stull 282706d7ef feat(slice5): POST .../meta/bulk one-commit bulk metadata edit (§22.4a PUC-2)
set/add/remove ops reusing the SLICE-4 sidecar write-through; per-entry
partial-rejection; contributor+ gated (INV-4); validated at the write
boundary. Tests cover one-commit, add/remove, partial reject, authz,
and op/field/empty guards.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 20:48:54 -07:00
Ben Stull eaf69cd05c Merge pull request 'v0.50.0 — SLICE-4: single-entry metadata edit (§22.4a)' (#34) from worktree-metadata-slice4 into main 2026-06-08 02:18:21 +00:00
Ben Stull 0d2fdfacf2 release(slice4): v0.50.0 — single-entry metadata edit + sidecar-aware writes (§22.4a SLICE-4)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 19:17:25 -07:00
Ben Stull abd17a6cc8 feat(slice4): schema-driven metadata edit panel in RFCView (§22.4a PUC-1 §5.2)
MetadataFieldsPanel renders one control per declared field (enum→select,
tags→chips, text→input), read-only without contribute access, saving changed
values via saveEntryMeta (direct sidecar commit). Wired into the canonical
main view, gated on the collection declaring fields (INV-5). saveEntryMeta
API client added.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 19:14:13 -07:00
Ben Stull 46c957cff5 feat(slice4): Owner-gated collection migrate endpoint (§22.4a PUC-5)
POST .../collections/{cid}/migrate — now safe to ship since all write paths
are sidecar-aware. Idempotent, one commit per collection.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 19:01:18 -07:00
Ben Stull d687a65470 feat(slice4): POST .../meta single-entry edit endpoint + GET RFC meta/can_edit_meta (§22.4a PUC-1)
Direct-commit to the sidecar (D7), contributor+ gated (INV-4), schema-validated
at the write boundary, lazy-migrates a legacy entry, re-ingests. GET RFC now
returns the per-entry meta mapping and a can_edit_meta capability.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 19:00:40 -07:00
Ben Stull aee9b582e5 fix(slice4): make all entry write paths sidecar-aware (§22.4a carried from SLICE-1)
graduate, claim, retire/unretire, _read_meta_entry, mark_entry_reviewed,
body extract/wrap (api_branches + api_prs replay) now dual-read and write
metadata to the sidecar via write_entry_files + bot.commit_entry_files/
open_entry_pr — a migrated body-only .md no longer crashes entry.parse or
re-grows frontmatter; legacy entries lazy-migrate on first metadata write.
Existing tests updated to assert the sidecar (INV-2 clean docs).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 18:57:54 -07:00
Ben Stull 734290f344 feat(slice4): bot.commit_entry_files + open_entry_pr multi-file primitives (§22.4a)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 18:46:09 -07:00
Ben Stull 49981e2d6e feat(slice4): metadata sidecar git read/write helpers — dual-read + lazy-migrate ops (§22.4a)
apply_values, EntryGitState, read_entry_from_git, write_entry_files, sidecar_path_for.
Plus the SLICE-4 implementation plan.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 18:44:14 -07:00
Ben Stull 7ece6d348b Merge pull request 'v0.49.0 — SLICE-3: faceted left-pane filtering (§22.4a)' (#33) from worktree-metadata-slice3-facets into main 2026-06-08 00:44:54 +00:00
Ben Stull 6bb6d654fa release(slice3): v0.49.0 — faceted left-pane filtering (§22.4a SLICE-3) 2026-06-07 17:43:56 -07:00
Ben Stull 276a625997 test(slice3): Catalog faceted-pane component test (§22.4a PUC-3) 2026-06-07 17:41:04 -07:00
Ben Stull ae3afe2f48 feat(slice3): faceted catalog pane gated on collection fields (§22.4a §5.1) 2026-06-07 17:40:29 -07:00
Ben Stull ae083bfcaa feat(slice3): FacetGroups faceted-pane component (§22.4a §5.1) 2026-06-07 17:38:35 -07:00
Ben Stull bcce40d2cb feat(slice3): listRFCs accepts facet selections, returns {items,facets} (§22.4a) 2026-06-07 17:38:16 -07:00
Ben Stull 7b269e11c4 test(slice3): integration tests for faceted list endpoint (§22.4a PUC-3) 2026-06-07 17:37:46 -07:00
Ben Stull 27061c30b0 feat(slice3): collection list endpoint returns facets + honours filters (§22.4a) 2026-06-07 17:35:26 -07:00
Ben Stull 14ea3c0cce feat(slice3): pure facet field-set + filter/count helper (§22.4a) 2026-06-07 17:34:33 -07:00
Ben Stull 644bf35d89 feat(slice3): persist entry metadata to cached_rfcs.meta_json at ingest (§22.4a) 2026-06-07 17:33:47 -07:00
Ben Stull 3636fa5afd feat(slice3): migration 034 — cached_rfcs.meta_json for facet values (§22.4a) 2026-06-07 17:33:11 -07:00
Ben Stull 1bcf8aa77e Merge pull request 'v0.48.0 — SLICE-2: collection field schema + central validation (§22.4a)' (#32) from worktree-metadata-slice2-schema into main 2026-06-08 00:06:41 +00:00
Ben Stull e336e31812 v0.48.0 — SLICE-2: collection field schema + central validation (§22.4a)
Collections declare a `fields:` schema in `.collection.yaml`; entries carry
typed metadata (enum/tags/text). New central `metadata_schema` module parses
the schema leniently (INV-3) and validates entry values — advisory at read
(corpus mirror flags violations as `metadata_malformed` without hard-failing),
the enforcement point for the write boundary (edit endpoints land SLICE-4/5).

- app/metadata_schema.py: parse_fields (lenient/normalizing) + validate
- registry.parse_collection_manifest reads `fields:` into collection config
- collections.get_collection unpacks `fields`; served by the collection API
- cache._refresh_collection_corpus validates each entry advisory-only

Non-breaking, opt-in: no `fields:` → unchanged (INV-5); default `document`
collection declares none (N=1 unchanged). No DB migration — schema rides in
collections.config_json. SLICE-2 of
docs/design/2026-06-06-configurable-collection-metadata.md §7.2.

Backend suite green (601 passed).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 17:06:04 -07:00
Ben Stull 98c276a662 Merge pull request 'v0.47.0 — SLICE-1: metadata sidecars (storage + dual-read + migration tool)' (#31) from worktree-metadata-slice1-sidecar into main 2026-06-07 15:24:55 +00:00
Ben Stull f05ee59763 v0.47.0 — SLICE-1: metadata sidecars — storage + dual-read + migration tool
§22.4a SLICE-1 of docs/design/2026-06-06-configurable-collection-metadata.md
(§7.2). Entry metadata can live in a per-entry `<slug>.meta.yaml` sidecar with
the `.md` kept as pure prose (INV-2). Additive and non-breaking — with no
sidecars present every corpus stays on the legacy frontmatter path,
byte-identical (N=1 unchanged).

- Dual-read (app/metadata.py `read_entry`) — sidecar-else-legacy-frontmatter,
  identical records (INV-6); unknown/forward-compat keys ride along through
  parse→serialize and migration (INV-7, `Entry.extra`). A degenerate sidecar
  (malformed/empty/slug-less) never drops the entry — slug backstopped from the
  filename stem, flagged not lost (INV-3).
- Migration tool (`metadata.migrate_collection`) — idempotent, one ChangeFiles
  commit per collection (new `gitea.change_files`). Tested as a function; its
  Owner-gated operator trigger is DEFERRED to SLICE-4 (write paths must become
  sidecar-aware first — see the design's SLICE-4 note + INV-8). No production
  trigger ships here, so no corpus is rewritten.
- Malformed flag — migration 033 adds `cached_rfcs.metadata_malformed`
  (additive); the corpus mirror derives it; catalog list + entry-detail APIs
  surface `metadata_malformed`.
- INV-7 at graduation — graduation now carries `Entry.extra` through the rebuild
  instead of dropping forward-compat keys.

Gate: backend 575 passed (28 new: test_metadata / _migration / _cache +
graduation extra-preservation). Frontend untouched. CHANGELOG 0.47.0 +
upgrade-steps; VERSION + frontend/package.json -> 0.47.0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 08:24:27 -07:00
Ben Stull b1acc2382d Merge pull request 'v0.46.2 — SLICE-0: §22.4a contract amendment (entry metadata is collection-configured, not type-driven)' (#30) from worktree-metadata-slice0-spec into main 2026-06-07 14:33:21 +00:00
Ben Stull 5cb5f4a4a2 v0.46.2 — SLICE-0: §22.4a contract amendment (entry metadata is collection-configured, not type-driven)
Reframes binding SPEC.md §22.4a so a collection's entry metadata schema is
collection-configured — a `fields:` schema in `.collection.yaml` plus per-entry
`<slug>.meta.yaml` sidecars — rather than a frontmatter schema hard-wired to the
collection's `type`. Item 3's type-specific surfaces (release planning;
bdd scenario/coverage views) are deferred to a future design; bdd coverage is
recorded as a future `ref`-field surface rendered as hyperlinks (no cross-
collection corpus fusion). `type` still selects terminology (entry noun,
v0.45.0) + default initial_state/review posture (§22.4b-c).

SLICE-0 of docs/design/2026-06-06-configurable-collection-metadata.md (§7.2);
supersedes the per-type-surfaces draft (D11). Doc-only: per-type frontmatter
validation was never implemented, so no operator action, no schema/behavior
change. Sidecar storage + validation + UI arrive in SLICE-1+.

- SPEC.md §22.4a reframed; document/specification/bdd bullets updated; metadata
  amendment blockquote added.
- SPEC.md §2 and §22 forward-pointer blockquotes: "type-dependent frontmatter
  schema" -> "collection-configured, not type-driven (§22.4a, as amended)".
- CHANGELOG 0.46.2; VERSION + frontend/package.json -> 0.46.2.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-07 07:32:27 -07:00
Ben Stull 36cb6187eb Merge pull request 'preview.env.example: flotilla-core overlay set (retire shim ref)' (#29) from fix/preview-env-flotilla-core into main 2026-06-07 05:52:03 +00:00
Ben Stull 677c5eb72f preview.env.example: flotilla-core overlay set (not the retired per-app shim)
The example comment referenced 'ohm-rfc-app-flotilla overlay set' — the retired
per-app shim. flotilla-core is the operator CLI for all deployments.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 22:44:09 -07:00
Ben Stull 077563ea47 Merge pull request 'docs(spec): Corpus Tree — universal directory-tree left pane (Solution Design)' (#28) from docs/corpus-tree-spec into main 2026-06-07 00:55:02 +00:00
Ben Stull dc5345cef4 docs(spec): Corpus Tree — universal directory-tree left pane (Solution Design)
Solution Design for making the left pane a universal git-directory tree:
host existing documentation repos as path-addressed corpora, full
governance lifecycle per file, dual-mode (structure/flat) pane replacing
the §7 Catalog, zero-migration onboarding ("no record = active").

Discovery output of ohm session 0081.0. Conforms to handbook §3.3
Solution Design standard. Sliced SLICE-1..4 in §7.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 17:47:24 -07:00
Ben Stull ee74a39b62 Merge pull request 'v0.46.1 — fix migration 029 for the §22.13 re-stamp aftermath (OHM deploy fault)' (#27) from fix/migration-029-orphan-collection into main 2026-06-06 19:03:20 +00:00
Ben Stull 2f507e5721 v0.46.1 — fix migration 029 for the §22.13 re-stamp aftermath (OHM deploy fault)
Deploying the three-tier series onto OHM crash-looped on migration 029:
`NOT NULL constraint failed: cached_branches__new.collection_id`, then a UNIQUE
collision. Root cause: the v0.39.0 default→ohm re-stamp updated
cached_rfcs.project_id but NOT the entry-satellite tables, so ~1.3k
cached_branches rows were stranded at project_id='default' — which 029's
per-project collection backfill can't map (NULL), some of which also duplicate
freshly-re-mirrored 'ohm' rows (UNIQUE), and some of which reference entries
that no longer exist.

Fix — a repair prologue at the top of 029 (no schema change):
- drop stale rows that duplicate an already-correctly-stamped row (keep the
  fresh copy) for the branch-keyed tables;
- re-derive each satellite's project_id from its entry (cached_rfcs, by slug);
- drop rows whose entry no longer exists (stale cache, rebuildable from gitea).
A no-op on clean/fresh deployments (empty or consistent satellites).

Validated against a snapshot of the live OHM DB: 029→032 apply cleanly, zero
NULL collection_id, FK check clean, watches/RFCs/branch_visibility preserved,
cached_branches 1397 → 1291 (−67 no-RFC, −39 dups). Fresh-install path
unchanged: existing 029 suite green + a new regression test for the
stale/dup/orphan shape. Full backend suite 547 passed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 12:02:53 -07:00
Ben Stull 9785782532 Merge pull request 'Discovery spec: configurable collection metadata (clean-doc tagging)' (#26) from docs/collection-metadata-design into main 2026-06-06 18:43:26 +00:00
Ben Stull 43a002c6aa Spec v0.1.6: §7.1 execution convention (one slice, one session, just-in-time)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 11:19:28 -07:00
Ben Stull 8ce3e5792d Spec v0.1.5: define Business Actors (§1.3) before Problem/Pain reference them
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 08:12:20 -07:00
Ben Stull 1be4a2edbf Spec v0.1.4: two-part structure — Business Context (§1) + Solution Proposal (§2)
- §1 Business Context holds the whole business lens (1.1–1.9), solution-agnostic
- §2 Solution Proposal introduces the chosen approach (and justifies build over
  a manual alternative); may be non-software in general
- Renumber product/engineering lenses to §§3–7

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 08:02:33 -07:00
Ben Stull 561cd73760 Spec v0.1.3: supersede per-type-surfaces; split personas; harvest patterns
- Supersede docs/design/2026-06-06-per-type-surfaces.md (banner added there);
  harvest validation seam, malformed-metadata flag, unknown-fields-ride-along,
  engine-unchanged invariant, N=1 document backcompat
- bdd coverage kept as a future per-type surface over a generic `ref` field
- Schema model = pure collection-config (not type-driven)
- SLICE-0 added: amend binding SPEC.md §22.4a (frontmatter→sidecar,
  type-driven→collection-configured, defer surfaces)
- Split personas: §6 Business Actors/Roles (solution-agnostic) + §10 Product
  Personas (mapped to business roles); renumber accordingly

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 07:34:15 -07:00
Ben Stull 3c910e89ab Revise spec to template: value-only summary, Pain Points, business framing
- Executive Summary → value-only (no solution)
- New §4 Pain Points (PP-1..PP-7)
- Business Outcomes → §5, restated as business outcomes (adoption/diversity),
  not solution outputs
- Business Use Cases rewritten as solution-agnostic actor-goal scenarios
- Renumber per the Solution Design template revision

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 07:15:35 -07:00
Ben Stull b7e23a01f8 Merge pull request '§22 S6 (last item): spec pass for per-type surfaces (§22.4a items 1 & 3)' (#25) from spec/s7-per-type-surfaces into main 2026-06-06 09:16:07 +00:00
Ben Stull 281dd29e62 §22 S6 (last item): spec pass for per-type surfaces (§22.4a items 1 & 3)
The per-type frontmatter schemas (item 1) and type-specific surfaces (item 3:
specification release-planning; bdd scenario/coverage views) were flagged at
v0.45.0 as wanting a discovery/spec pass first (they lacked BDD scenarios in
Part C). This is that pass — a design doc specifying the schemas, the surfaces,
their data model / API / frontend shape, BDD-style acceptance scenarios, and a
three-slice delivery plan (S7a schemas, S7b releases, S7c bdd surfaces), all
additive and engine-preserving per §22.4a. Doc-only; no code, no version bump.

- docs/design/2026-06-06-per-type-surfaces.md — the new spec pass.
- three-tier design doc S6 bullet: forward-pointer + S6 shipped/spec'd status.

Leaves implementation to the S7a–S7c coding slices; S7a (schemas) is unblocked,
S7b/S7c carry open product questions for the operator/discovery to settle.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 02:15:45 -07:00
Ben Stull 1c17fecea3 Merge pull request '§22 S6: request-to-join + cross-collection inbox (§22.8) — v0.46.0' (#24) from feat/s6-join-requests into main 2026-06-06 09:12:00 +00:00
Ben Stull fcc3c84d76 §22 S6: request-to-join + cross-collection inbox (§22.8) — v0.46.0
Ships the request side of joining a gated scope, completing the §22.8 pair
(S4 shipped the invite half). A user who knows a project/collection exists
asks to join it naming a desired role; the request fans out to that scope's
Owners across the subtree (the cross-collection inbox, §22.11), who accept
(writing the memberships row via memberships.grant) or decline. Built by
analogy to §28 contribution_requests + the S4 memberships surface.

Backend
- migration 032: join_requests (scope_type ∈ {project,collection}, scope_id,
  requester, requested_role, message, status, granted_role); one-open-per
  (scope, requester) partial unique index. Additive — no rebuild.
- api_join_requests.py: GET join-target / POST join-requests / POST
  {id}/accept / {id}/decline under /api/scopes/{scope_type}/{scope_id}/.
  Accept grants via memberships.grant; the request POST does not require the
  scope be readable (that is how one joins a gated scope).
- notify: fan_out_join_request (subtree-Owner enumeration via
  _scope_owner_user_ids), notify_join_decided, 3 render_summary cases.
- auth.effective_role_at_scope — scope-grain twin of effective_scope_role,
  folding global → project for a project target.
- api_collections: viewer.can_request_join on the project + collection blocks.

Frontend
- api.js join verbs; JoinRequestModal; "Request to join" affordance in the
  collection directory + catalog footer; JoinRequestRow in the inbox.

Tests: backend test_join_requests_vertical (11) + test_migration_032 (5);
frontend api.joinrequests + CollectionDirectory cases. 546 backend / 36
frontend green.

Per docs/design/2026-06-05-three-tier-projects-collections.md Part E (S6) and
SPEC.md §22.8 / §22.11. Closes the request-to-join item flagged open at
0.45.0; per-type surfaces (§22.4a items 1 & 3) remain the last S6 item.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 02:11:16 -07:00
Ben Stull e86fc65643 Discovery spec: configurable collection metadata (clean-doc tagging)
Solution Design for a generic per-collection metadata system — typed fields
(enum/tags/text) declared in .collection.yaml, per-entry values in a clean
<slug>.meta.yaml sidecar, schema-derived faceted left-pane filtering, and
single/bulk direct-commit tag/untag. Restores the retired BDD Release Planner's
corpus-annotation half inside the framework; release planning stays downstream.

Output of discovery session OHM-0079.0. Proposed as §23; flag for reconciliation
with §22 S6 type-modules at the SPEC merge.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 01:52:26 -07:00
Ben Stull 014015014b Merge pull request '§22 S6: SPEC merge + per-collection model universe + type noun — v0.45.0' (#23) from feat/s6-three-tier-final into main 2026-06-06 08:44:25 +00:00
Ben Stull b392fa923c §22 S6: release v0.45.0 — two-project/multi-collection test + §20.4 changelog
The release wrap for the S6 scope shipped this slice (SPEC merge, per-collection
enabled_models, type-driven entry noun).

- test_s6_two_project_multicollection.py: an integrative pass proving the new
  per-collection knobs (§22.12 enabled_models, §22.4a noun) resolve independently
  per collection across two projects with no cross-project/-collection bleed.
- CHANGELOG.md 0.45.0: the §20.4 entry + upgrade-steps (migration 031 is
  additive + automatic; per-collection enabled_models and typed collections are
  optional). Notes the two OPEN S6 items carried to a follow-up slice (per-type
  surfaces; request-to-join + cross-collection inbox).
- VERSION + frontend/package.json → 0.45.0.

Gate: backend 530 passed, frontend 30 passed, build green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 01:43:09 -07:00
Ben Stull 839404da0c §22 S6: type-driven entry noun (§22.4a) — backend source of truth + chrome
The displayed noun for an entry is a framework concept keyed on the collection's
immutable type — document→"RFC", specification→"Spec", bdd→"Feature" — not
deployment content. The chrome reads it from the API instead of hardcoding "RFC".

- collections.ENTRY_NOUN + entry_noun(type) (unknown type → generic "RFC").
- Surfaced as entry_noun on get_collection, list_collections items, the
  /api/deployment directory items, and GET /api/projects/:id.
- Frontend: Catalog reads entry_noun for the "+ Propose New <noun>" control;
  ProposeModal fetches the active collection's noun for its title + field copy.
- test_s6_entry_noun_vertical.py: map + API surfacing (collection + directory).

Scope note: this is §22.4a item (2) — terminology. The deeper type-module work
(item 1 per-type frontmatter schemas; item 3 specification release-planning +
bdd scenario/coverage surfaces) is flagged OPEN in the design doc and is handed
off to a follow-up spec+slice.

Backend 528 passed; frontend 30 passed + build green. Part of v0.45.0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 01:35:57 -07:00
Ben Stull 79a27a946b §22 S6: per-collection enabled_models — the §22.12 model-universe chain
A collection's .collection.yaml may carry an enabled_models list that NARROWS
its project's universe, which narrows the deployment ENABLED_MODELS. Resolution
(extending §6.6/§6.7): funder ∩ per-entry models ∩ collection ∩ project, with
the operator providers as the ceiling.

- migration 031: additive collections.config_json (parallels projects.config_json).
- registry.parse_collection_manifest: read enabled_models into CollectionEntry.config
  (absent = inherit, present incl. [] = narrow; [] opts the collection out of AI);
  reject a non-list. _upsert_named_collection persists config_json. The default
  collection (from projects.yaml) leaves it NULL — it inherits the project.
- models_resolver: _scope_narrowed_universe narrows the operator universe by the
  entry's project then collection enabled_models; the funder universe is bounded
  by it too. A collection cannot widen its project (narrowing from the ceiling).
- collections.get_collection surfaces enabled_models; GET
  /api/projects/:id/collections/:cid returns it.
- test_s6_collection_models_vertical.py: 9 cases (parse, the narrowing chain
  incl. opt-out + cannot-widen, API surfacing).

Backend: 525 passed. Per SPEC §22.12. Releases as part of v0.45.0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 01:30:16 -07:00
Ben Stull 26f3680197 §22 S6: merge three-tier model into SPEC.md + registry format in DEPLOYMENTS
Write the canonical §22 (deployment → project → RFC collection) into the
binding spec, applying Part A of the design doc and the Part D amendments in
place. Lands the S3 keystone reinterpretation of §B.1/§B.3: a plain granted
account is a granted account, not a global write role; "global RFC Contributor"
is an explicit memberships(scope_type='global') grant; the implicit-public
write baseline is grandfathered onto the migration-seeded default collection
only (N=1 preserved).

- SPEC.md §22.1–§22.14: tiers + isolation, the registry + .collection.yaml
  manifests, content-repo-per-project, per-collection slug identity, collection
  type/initial_state/unreviewed, two-tier visibility (narrow-only), the unified
  {owner, contributor} role vocabulary at {global, project, collection}, the
  four-layer most-permissive union, discovery/joining (invite + request-to-join),
  runtime branding, /p/<project>/c/<collection>/ routing, one inbox,
  per-collection model universe, the default-project+collection migration, and
  §22.14 consolidating the §§1–21 amendments.
- Forward-pointer amendment notes (house style) at §1, §2, §5, §6 routing
  readers to §22.
- docs/DEPLOYMENTS.md: the registry (REGISTRY_REPO / projects.yaml) and
  .collection.yaml manifest formats, the N=1 default-project+collection upgrade,
  and in-app create-project/collection.

Per docs/design/2026-06-05-three-tier-projects-collections.md Parts A/B/D/E (S6).
Docs-only; no code or migration. Releases as part of v0.45.0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 01:21:39 -07:00
Ben Stull b0737380cd §22 S5: in-app create-project + global directory — v0.44.0 (@S5) (#22) 2026-06-06 08:03:12 +00:00
Ben Stull 33212c71e4 §22 S5: in-app create-project + global-directory empty states (@S5 C3.1–C3.2)
Adds the global-Owner create-project action and the role-aware deployment
directory, completing slice S5 of the three-tier refactor (release v0.44.0,
minor/non-breaking — no migration). Per
docs/design/2026-06-05-three-tier-projects-collections.md Part E (S5),
Part C.3 (C3.1–C3.2).

Backend:
- auth.can_create_project — global-Owner gate (deployment owner/admin or an
  explicit scope_type='global' Owner grant).
- bot.create_project — provision the Gitea content repo (seed README so main
  exists), read+append+commit projects.yaml in the registry repo, audit-log
  (create_project). The bot stays the only git writer (§1).
- POST /api/projects (api_deployment) — global-Owner gated; validates the id
  (slug, not 'default'), name, type, visibility, content_repo; commits via the
  bot, then re-mirrors the registry so projects + default collection rows flow
  from git (§22.2). GET /api/deployment gains viewer.can_create_project and
  default_project_readable. make_router now takes gitea + bot.

Frontend:
- api.createProject; DeploymentProvider surfaces viewer + defaultProjectReadable
  + refresh; DeploymentLanding redirects into the default only when readable
  (gated/absent default falls through to the directory, no 404 bounce).
- Directory.jsx role-aware empty states (C3.1 "Create your first project" CTA;
  C3.2 "Nothing has been shared with you yet") + Owner-only "New project"
  control + CreateProjectModal.

Tests: backend test_create_project_vertical.py (vertical + gates + the
deployment empty-state signals); frontend Directory.test.jsx empty-state cases.

Also: ignore the session-local .superpowers/ tooling dir.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 01:02:16 -07:00
Ben Stull 2696e64ff5 Merge pull request '§22 S4: invitation surfaces + role-aware empty states — v0.43.0 (@S4)' (#21) from feat/s4-invitation-surfaces-empty-states into main 2026-06-06 07:31:01 +00:00
Ben Stull ff54632657 §22 S4: release v0.43.0 — invitation surfaces + role-aware empty states (@S4)
Bump VERSION + frontend/package.json to 0.43.0, add the CHANGELOG entry
(minor, non-breaking — additive endpoints + UI, no migration, no change to
existing authz outcomes), and mark Part E slice S4 shipped in the design doc.

Completes @S4 (C2.1–C2.7 invitation + C3.3–C3.5 role-aware empty states).
Next: S5 (in-app create-project + the global directory, C3.1–C3.2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 00:30:12 -07:00
Ben Stull 93cf506059 §22 S4: invitation modal + role-aware empty states (@S4 C.2, C.3)
The frontend surfaces for the S4 backend (Part E S4 / Part C.2, C.3):

- ScopeMembersModal — the Owner-only scope-role invitation surface (sibling of
  the per-RFC InvitationsModal): grant {owner, contributor} by email at the
  project or a single collection, with a scope picker bounded to the inviter's
  reach (no parent-grant-child-exclude option, C.2.5), plus a current-members
  list with revoke. Opened from the directory's owner-only "Members" control;
  contributors never see it (C.2.4).
- CreateCollectionModal — surfaces the S2 create-collection endpoint (was
  UI-less), gated on the viewer's can_create_collection capability.
- CollectionDirectory — reads the new `viewer` capability block to render the
  role-aware empty states: a project Owner sees "Create your first collection"
  (C3.3); a contributor without create rights sees the bare empty directory
  (C3.4); an Owner with management reach sees the "Members" control.
- Catalog — the empty collection shows "Propose the first entry" to a viewer
  who may contribute here (C3.5), and the propose control is gated on the
  collection's can_contribute flag (anon keeps the sign-in prompt, S2).
- api.js — getCollection, listScopeMembers, grantScopeMember, revokeScopeMember.

Frontend builds clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 00:27:49 -07:00
Ben Stull c9fd1c535e §22 S4: scope-role invitation backend + capability flags (@S4 C.2)
The Owner-only scope-role grant surface (Part E S4 / Part C.2). An Owner
grants {owner, contributor} at a scope their reach covers — the project or
one collection within it — to an existing account looked up by email; the
grant writes a `memberships` row immediately and §15-notifies the grantee
(direct grant, no accept round-trip — the C.2 scenarios name existing
accounts and write the row directly).

- auth.can_invite_at_project / can_invite_at_collection — the Owner-reach
  invite gates (is_project_superuser / is_collection_superuser).
- memberships.py — grant (with the C.2.6 broader-supersedes-narrower prune,
  preserving a stronger child grant — no negative override), revoke, list,
  user_by_email.
- api_memberships.py — GET/POST/DELETE /api/projects/:id/members, the single
  POST keying on optional collection_id so the invite UI's one control maps
  to one endpoint; reach bounded by the inviter's Owner reach (C.2.3);
  contributors refused (C.2.4); a pending grantee's row is recorded but
  confers no write (C.2.7, the §6 floor).
- notify.notify_scope_role_granted + render_summary — the §15 personal-direct
  notification naming the project and role.
- api_collections — surface viewer capabilities (can_create_collection,
  can_invite, can_contribute, role) on the project/collection GETs to drive
  the C.3 role-aware empty states.

Acceptance: test_s4_invitations_vertical.py covers C.2.1–C.2.7 over the HTTP
surface + resolver, plus the C.3 (@S4) capability flags. Full suite green
(504 passed).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 23:47:28 -07:00
benstull e6bd69f132 Merge pull request '§22 S3: scope-role enforcement + collection-grain visibility — v0.42.0 (@S3)' (#1) from feat/s3-scope-roles-collection-visibility into main 2026-06-06 01:08:54 +00:00
Ben Stull c2f566512a §22 S3: scope-role enforcement + collection-grain visibility (@S3) — v0.42.0
Implement slice S3 of the §22 three-tier refactor: the four-layer
most-permissive scope-role resolver (§B.2) over {owner, contributor}
grants at {global, project, collection}, with the §22.5 visibility gate
enforced at the collection grain.

- migration 030: memberships.scope_type += 'global' (the global RFC
  Contributor tier; sentinel scope_id '*').
- auth.effective_scope_role folds global → project → collection,
  most-permissive, no negative override; can_read_collection /
  can_contribute_in_collection / is_collection_superuser /
  can_create_collection gate reads, writes, admin, and create.
- collection-grain visibility: a gated collection is hidden from the
  public (404, omitted from the directory) yet visible+listed for a
  scope-role holder; a collection may be set only as strict or stricter
  than its project (public < unlisted < gated), validated at create and
  clamped at the mirror.
- entry-scoped authority (mark-reviewed, graduate, branch read/contribute,
  PR/discussion/contribution moderation) re-pointed from the project grain
  to the entry's collection.
- create-collection authority widened to a project/global-scope grant
  holder (§B.1), not only a deployment owner/admin.

Keystone reconciliation (session 0076): a plain granted account is a
granted *account*, not a write-everywhere global role; the implicit-public
write baseline is grandfathered onto the migration-seeded `default`
collection only, so the N=1 deployment loses no capability. Reinterprets
§B.1/§B.3 literally — flagged for the SPEC merge (S6).

Completes @S3 (C1.1–C1.8). Tests: test_s3_scope_roles_vertical.py (8 C.1
scenarios + visibility/strictness), test_migration_030_global_scope.py.
Full backend suite 493 passed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 18:07:51 -07:00
Ben Stull 39ce54fbcc Merge pull request '§22 S2: create & navigate a second collection — v0.41.0 (@S2)' (#20) from feat/s2-second-collection into main 2026-06-05 20:19:25 +00:00
Ben Stull 55d04ce4ca §22 S2: release v0.41.0 — create & navigate a second collection (@S2)
Minor, non-breaking: named collections via .collection.yaml, create-collection
endpoint, collection-scoped serve/propose, and the /p/<project>/ collection
directory. Completes acceptance @S2 (C3.6). 478 backend + 26 frontend green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 13:17:31 -07:00
Ben Stull bd6dc6524a §22 S2: @S2 acceptance — anonymous empty public collection catalog (C3.6)
Anonymous reader of an empty public collection gets a 200 empty catalog and no
propose action (the propose route rejects anonymous); the Catalog footer's
'Sign in to propose' prompt is the existing anonymous affordance.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 13:14:52 -07:00
Ben Stull 17bdd5fd9a §22 S2: collection directory at /p/<project>/ (1 → redirect, 2+ → list)
Replace DefaultCollectionRedirect with a CollectionDirectory that lists the
project's visible collections, or redirects into the sole one when there is
exactly one (preserving the S1 C3.7/C3.8 single-collection UX). The
create-first-collection empty state is S4.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 13:14:18 -07:00
Ben Stull 2b32e124ab §22 S2: Catalog + propose scoped to the active collection
Catalog reads the /c/:collectionId/ segment via useCollectionId and fetches the
collection-scoped catalog, building entry/proposal links with the active
collection; the propose modal threads the active collection so a propose from a
named collection targets it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 13:12:45 -07:00
Ben Stull 98eea3e2d6 §22 S2: collection-scoped frontend path + API helpers
useCollectionId() hook; listRFCs/getRFC/proposeRFC take an optional collection
id and target the /collections/<cid>/ routes; add listCollections +
createCollection.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 13:11:27 -07:00
Ben Stull 0c654b173d §22 S2: create-collection endpoint (bot commit + registry refresh)
POST /api/projects/<id>/collections, owner/admin-gated, commits a
.collection.yaml to the content repo main via bot.create_collection, then
re-mirrors the registry so the collections row appears (§22.2). Adds GET
list/one collection routes. Extends FakeGitea to model directory listings so
the mirror's content-repo walk is exercised end to end.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 13:10:09 -07:00
Ben Stull f57d4080dc §22 S2: collection-scoped list/get/propose endpoints
Refactor the project-scoped serve/propose internals into collection-grained
helpers (_list_rfcs_for_collection, _get_rfc_for_collection,
_propose_into_collection); add routes under
/api/projects/<id>/collections/<cid>/rfcs[/<slug>|/propose]. Propose writes
the entry under the target collection's <subfolder>/rfcs via a new rfcs_dir
param on bot.open_idea_pr. Default-collection routes preserved as wrappers.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 13:05:58 -07:00
Ben Stull 91b0fb358c §22 S2: corpus mirror reads each collection's <subfolder>/rfcs/
refresh_meta_repo now iterates a project's collections and keys cached_rfcs by
collection_id; the default collection (subfolder '') keeps the shipped rfcs/
root path. N=1 default path unchanged (469 green).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 13:02:59 -07:00
Ben Stull 868391870c §22 S2: collection read helpers (list/get/subfolder)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 13:00:28 -07:00
Ben Stull 74476423ba §22 S2: registry mirror reads .collection.yaml manifests
Discover named collections by walking each project's content-repo root for
<subdir>/.collection.yaml; parse + upsert with immutable-type enforcement
(§22.4a) and project-visibility inheritance. The default collection still
flows from projects.yaml.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 12:59:25 -07:00
Ben Stull 599e7018f6 §22 S2: implementation plan — create & navigate a second collection
Plan for slice S2 of the three-tier (deployment→project→collection) refactor.
Completes acceptance @S2 (C3.6). See
docs/design/2026-06-05-three-tier-projects-collections.md Part E.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 12:55:56 -07:00
Ben Stull 4ffff6b677 Merge pull request '§22 S1: three-tier collection grain — migration 029 + threading + 308 redirect (v0.40.0)' (#19) from feat/s1-collection-grain into main 2026-06-05 15:34:07 +00:00
Ben Stull aaf7b09bbe §22 S1: release v0.40.0 — three-tier collection grain (breaking URL + migration 029)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 08:31:43 -07:00
Ben Stull 4f72aa31e0 §22 S1: frontend /c/<collection>/ route layer + C3.7 + legacy-URL redirects
- entryPaths builders carry the /c/<collection>/ segment (default collection in S1)
- /p/:projectId/* gains c/:collectionId/ corpus routes; serving stays project-scoped
- DefaultCollectionRedirect (C3.7: project landing -> default collection)
- LegacyCorpusRedirect (v0.35.0 /p/<p>/e/<slug> bookmarks -> /c/default/, query preserved)
- entryPaths unit test; build + vitest green (18 tests)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 08:30:18 -07:00
Ben Stull 9ca07a3f81 §22 S1: thread collection_id through backend + update tests (N=1 unchanged, 454 green)
- collections.py resolution helpers (default_collection_id, type, initial_state)
- registry mirror writes project grouping fields + default-collection corpus fields
- auth.project_of_rfc joins collections; project_member_role reads memberships
- cache/api_*/funder writers+readers re-keyed to collection_id (cached_prs + denormalised
  project_id tags unchanged); api_deployment reads type/initial_state from the default collection
- projects.py restamp detects bootstrap via collections; initial_state via collection
- tests updated to the three-tier schema; test_migration_028 retired (superseded by 029)
- add @S1 acceptance test (collection grain + N=1 serving)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 08:26:14 -07:00
Ben Stull 867f2504d6 §22 S1: migration 029 — collections grain, field move-down, 13-table re-key, memberships
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 07:58:39 -07:00
Ben Stull 08bdea8539 §22 S1 plan: three-tier collection grain (migration 029 + threading + redirect)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 07:53:41 -07:00
Ben Stull 0de91fe35c Merge pull request '§22 refactor (spec): three tiers — deployment → project → RFC collection' (#18) from spec/three-tier-projects-collections into main 2026-06-05 14:38:48 +00:00
Ben Stull 2f5d09aef5 §22 spec: re-cut Part E into usable BDD-tagged slices (S1-S6)
Per operator (session 0072): every slice must end in a usable deployment and
declare which Part C scenarios it makes pass. Tag all 23 Gherkin scenarios with
@S<n> (the slice that completes them) and re-cut Part E from layer-by-layer
(N1-N6) to usable increments (S1-S6) with a slice->scenario index table. S1
bundles the coupled migration 029 + threading + redirect as one right-sized
first session; S2 second collection; S3 role enforcement; S4 invitation;
S5 create-project + directory; S6 types + SPEC merge.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 06:42:12 -07:00
Ben Stull 87279fc545 §22 three-tier spec Part E: decided migration-029 strategy (collection grain beneath project)
Operator chose (session 0072) to add the collection grain beneath today's
project: projects table stays the group tier (keeps content_repo), a new
collections table holds the per-corpus fields, entries re-key to
(collection_id, slug), one default collection per project on migrate, breaking
/p/<project>/e/<slug> -> /p/<project>/c/<collection>/e/<slug> with 308s.
Re-sloted slices N1-N6. Correction banners updated from pending -> decided.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 03:46:51 -07:00
Ben Stull 9c0e3b60ac Correct §22 three-tier spec: two-tier model already shipped (v0.39.0)
Re-checked code vs the stale memory: migration 028 (slug PK -> (project_id,
slug)), v0.35.0 /p/<project>/ routing, and v0.37/0.38 per-project read+propose
are all shipped to main. The 'fold into not-yet-shipped Plan B + M3-frontend'
premise is false. Neutralize the wrong claims in §0/§A.3 and flag Part E's
sequencing as pending re-decision; structural model (Parts A-D) unaffected.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 03:43:44 -07:00
Ben Stull 31d680be54 §22 three-tier refactor spec: project → RFC collection + unified roles + BDDs
Splits the original §22 two-tier model (deployment → project=corpus) into
three tiers (deployment → project → RFC collection). Project owns one
content repo; collections are typed subfolders declared by .collection.yaml
manifests (git-truth). Reconciles the accumulated role vocabulary onto one
{owner, contributor} enum attached at {global, project, collection}, with
downward additive inheritance and no negative override. Adds BDD scenarios
(Part C) for role usage, invitation, and empty states. Re-slots the roadmap
to fold the tier into the not-yet-shipped Plan B (mig 028) + M3-frontend.

Session 0072 (spec). Revises docs/design/multi-project-spec.md §22.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 03:39:04 -07:00
Ben Stull f758fe072f Merge pull request '§22.13 step 1: default-project-id re-stamp (v0.39.0)' (#17) from feat/m3-planb-restamp into main 2026-06-04 14:45:21 +00:00
Ben Stull 33c67ccc09 §22.13 step 1: default-project-id re-stamp (v0.39.0)
projects.restamp_default_project(config): at startup after the registry mirror,
renames project_id from the M1 bootstrap 'default' to the configured
DEFAULT_PROJECT_ID across every project-scoped table (discovered by column) and
drops the stale 'default' projects row, so a deployment's original corpus lands
at a meaningful /p/<id>/ and 'default' is never a public URL. FK off for the
rename (parent+children move together) + foreign_key_check backstop. Idempotent;
no-op unless DEFAULT_PROJECT_ID is set to a non-'default' value.

test_restamp_default_project.py (3 tests). 450 backend green.

This is the last framework piece for OHM's clean /p/ohm/ cutover.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 07:45:10 -07:00
Ben Stull 508a8cb6d0 Merge pull request '§22 Plan B write path (propose): project-scoped propose (v0.38.0)' (#16) from feat/m3-planb-write into main 2026-06-04 14:02:29 +00:00
Ben Stull fec51bdbb6 §22 M3-backend Plan B (write path, propose): project-scoped propose (v0.38.0)
A new entry can be proposed into a specific project; it lands in that project's
content repo and shows under that project's proposals. A non-default project is
no longer read-only.

- api.py: POST /api/projects/{pid}/rfcs/propose (propose body extracted into a
  project-parameterized helper; unscoped /api/rfcs/propose kept as default
  compat). Slug uniqueness, idea-PR reservation, landing state, and the
  proposed_use_cases row scoped to the target project. GET
  /api/projects/{pid}/proposals.
- cache.py: refresh_meta_pulls loops every project's content_repo, stamping
  cached_prs.project_id; projects.content_repo(pid) helper.
- frontend: proposeRFC(projectId,…)/listProposals(projectId); ProposeModal
  takes projectId; App resolves current project from the /p/<id>/ URL; Catalog
  lists that project's proposals.
- tests: test_project_scoped_propose.py (lands scoped + gated 404). 447 backend
  + 11 Vitest green; clean build.

Known limitation: branch/PR/graduation edit flows + default-id re-stamp not yet
scoped (next slice). Per docs/superpowers/specs/2026-06-04-m3-backend-planb-design.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 07:02:19 -07:00
148 changed files with 22899 additions and 887 deletions
+1
View File
@@ -26,3 +26,4 @@ data/
# Claude Code (per-machine settings only; shared config under .claude/ is committed)
.claude/settings.local.json
.superpowers/
+995
View File
File diff suppressed because it is too large Load Diff
+15 -1
View File
@@ -1,9 +1,18 @@
.PHONY: tier1-up tier1-down tier1-logs fe-unit e2e e2e-install
.PHONY: tier1-up tier1-down tier1-logs fe-unit e2e e2e-install e2e-fresh
# Two-phase: run the Gitea seed to completion FIRST so it writes the bot token /
# OAuth creds into generated/.env.tier1.generated, THEN create the backend/web —
# compose snapshots env_file at container-create time, so the backend must be
# created after the seed has populated it. The touch seeds an empty placeholder
# for compose's up-front env_file existence check on a clean checkout.
tier1-up:
touch testing/generated/.env.tier1.generated
docker compose -f testing/docker-compose.yml up --build -d gitea-seed
docker compose -f testing/docker-compose.yml wait gitea-seed
docker compose -f testing/docker-compose.yml up --build -d
tier1-down:
touch testing/generated/.env.tier1.generated
docker compose -f testing/docker-compose.yml down -v
tier1-logs:
@@ -17,3 +26,8 @@ e2e-install:
e2e:
cd e2e && BASE_URL=$${BASE_URL:-http://localhost:8080} MAILSINK_URL=$${MAILSINK_URL:-http://localhost:8025} npm run e2e
# Canonical run: the metadata specs mutate the seeded corpus (edit/bulk write
# real commits), so they assume a freshly-seeded stack. This brings the stack
# down, back up (re-seeds), and runs the suite once — the shape CI uses.
e2e-fresh: tier1-down tier1-up e2e
+594 -1
View File
@@ -65,6 +65,18 @@ in `rfcs/`" is the whole mental model a new deployer needs.
> and its `wiggleverse/rfc-0001-human` repo archived (see §13.6). The
> decision record is OHM ROADMAP #36.
> **Three-tier change (v0.45.0 — supersedes the single-corpus topology;
> see §22).** A deployment is no longer one corpus in one repo. It has a
> **registry** (§22.2) naming N **projects**, each owning **one content
> repo** (§22.3) that holds N typed **RFC collections** as subfolders.
> "This single repository is its content repository" now reads "each
> *project* names one content repository; the deployment's registry lists
> them." The bot and app-owned-authorization paragraphs below are
> unchanged and now read **org-wide** across every content repo and the
> registry repo. The single-corpus deployment is the N=1 case and keeps
> running unchanged via a generated default project + default collection
> (§22.13). §22 is the binding model.
All Git operations on the meta repository are performed by a single **bot
service account** in Gitea. Real human users do not have meaningful Gitea
permissions on the repo itself; their accounts exist for OAuth identity
@@ -107,6 +119,20 @@ That's the entirety of the meta repo. App-level permission state, user
accounts, chat history, audit logs, and branch visibility grants do **not**
live in the meta repo — they live in the app database (see §5).
> **Three-tier amendment (v0.45.0 — see §22).** Slugs are unique **per
> collection** (§22.4): `model/intro` and `specs/intro` coexist. The entry
> **metadata schema is collection-configured** (§22.4a, as amended by
> v0.46.2) — each collection declares a `fields:` schema in its
> `.collection.yaml` and stores per-entry values in a `<slug>.meta.yaml`
> sidecar; the fields below are the `document` baseline, **not**
> type-driven. The §2.3 `RFC-NNNN` `max+1` allocation is
> **removed** (the slug is the identity); pre-change `id` values survive as
> frozen legacy labels. New `active`-entry frontmatter: `unreviewed` (bool)
> and the `reviewed_at` / `reviewed_by` provenance pair (§22.4c).
> Collection configuration (`type`, `visibility`, `initial_state`) lives in
> a `.collection.yaml` manifest in the content repo, mirrored like entry
> frontmatter (§22.2).
### 2.1 Entry file format
```markdown
@@ -298,6 +324,18 @@ the natural path is SQLite FTS5 indexed off the reconciler.
These are the tables that are app-owned (not cached from Gitea). Names
and exact columns are illustrative; the implementing session can adjust.
> **Three-tier amendment (v0.45.0 — see §22).** Every entry-scoped table
> below carries the corpus grain, which is now the **collection**: the
> column is **`collection_id`** (not `project_id`), and `cached_rfcs` is
> keyed `(collection_id, slug)`. A denormalized `project_id` rides the
> high-churn cache/notification rows for filtering. New app-owned tables:
> **`projects`** (one content repo, project settings), **`collections`**
> (the immutable `type`, `subfolder`, `initial_state`, `visibility`,
> mirrored from `.collection.yaml`, §22.2), and **`memberships`**
> (`scope_type ∈ {global, project, collection}`, the unified
> `{owner, contributor}` roles — §22.6, replacing M2's `project_members`).
> `users.role` is the **deployment** admission tier only (§22.7).
- `users``id`, `email`, `display_name`, `gitea_login`, `role` (one
of `owner` / `admin` / `contributor`), `muted` (bool — the §6.2
app-wide write-mute, distinct from the per-RFC and per-user
@@ -435,6 +473,20 @@ merge with no data movement.
Authorization is owned by the app. Gitea sees only the bot account.
> **Three-tier amendment (v0.45.0 — see §22.6–§22.7).** The deployment
> roles described in this section are the **global** tier of a four-layer
> most-permissive union (global → project → collection → per-entry).
> Scope roles use one vocabulary — **Owner** and **RFC Contributor**
> attached at `{global, project, collection}` via the `memberships` table
> and inheriting downward with no negative override. A plain granted
> `contributor` account is **not** an implicit global write role: it can
> sign in and read, but writing an explicitly-created collection requires
> an explicit `memberships` grant. The one carve-out is the N=1
> default-collection baseline, where the pre-multi-project implicit-public
> write capability is grandfathered (§22.6 keystone note). The §6.3
> per-RFC delegated-authority idea is now **Owner** at project/collection
> scope, sitting above per-entry authority in the union.
Authentication has three paths, in the order a visitor encounters
them:
@@ -845,7 +897,13 @@ a hierarchy on the user that gets in the way of finding by title.
- **Filter chip strip** — multi-select, AND-combined. Chips:
`State: super-draft | active | withdrawn`, `My RFCs` (I'm an owner
or arbiter), `Has open PRs`, `Unclaimed` (super-drafts with empty
`owners:`), `Tag: …`.
`owners:`), `Tag: …`. A collection that declares a metadata field
schema (§22.4a) replaces this chip strip with **faceted filter
groups** — one collapsible group per `enum`/`tags` field plus state,
each showing per-value result counts and multi-select checkboxes
(OR within a field, AND across fields); see the Configurable
Collection Metadata design. A collection with no `fields:` schema
keeps the chip strip described here unchanged.
### 7.2 The list rows
@@ -964,6 +1022,20 @@ the user on it in contribute mode. New-branch naming defaults to an
auto-generated value (user-renamable); the exact format is an
implementation detail.
The AI chat is an editing activity — a turn can emit `<change>`
proposals — so it runs on an edit branch, not on read-only main, whose
right column is the human-discussion surface (§8.12 is branch-scoped).
Invoking the AI **Ask** affordance (the selection tooltip's prompt, or
the prompt bar) from main therefore transparently cuts an edit branch
via the same dispatch as "Start Contributing" (promote-to-branch for an
active RFC, start-edit-branch for a super-draft per §9.5; idempotent, so
an existing edit branch is reused rather than a second one cut), lands
the user on it, and runs the question — text plus any selected quote —
as the branch's first chat turn. A viewer who cannot contribute has the
branch cut rejected and the error surfaced (no branch is created); a
signed-out viewer keeps the §8.7 read-only path. **Flag** from main is
unaffected — it opens a human discussion thread on main, not a branch.
Discuss vs. contribute is an *intent* affordance, not a *permission*
affordance. A user without contribute access to a branch sees the
toggle disabled, with a sign-in or request-access path (see §8.7).
@@ -4912,3 +4984,524 @@ existing consenters. The cleanest moment to do this is the next
material privacy-policy revision; the conventions in §21.4 hold
in the interim.
---
## 22. Three tiers: deployment → project → RFC collection
A **deployment** hosts one or more **projects**; each project owns one
content repository and holds one or more **RFC collections**; each
collection is a typed corpus of **entries**. This is the three-tier model
that supersedes the original single-corpus framing of §§121 and the
two-tier (deployment → project) draft that preceded it.
```
deployment (= "global" in the UI) one Gitea org, one bot, one account
│ system, one inbox, one running process;
│ the surface a visitor first lands on.
└─ project a named grouping + project settings;
│ owns exactly ONE content repo. No type.
└─ RFC collection a typed corpus: type, slug namespace,
│ catalog, philosophy, initial_state,
│ unreviewed flag, members.
└─ entry an RFC / spec / feature, identified by
its slug within the collection.
```
Everything §§121 describe about *a corpus* is now *an RFC collection*.
Everything they describe about *a deployment* that is not corpus-specific —
accounts, the §6 admission gate, the §15 inbox, the §1 bot — stays at the
deployment level and is shared. A grouping layer, the **project**, sits
between: it owns the content repo and project-wide settings, and groups the
collections beneath it. The numbered sections that assume a single corpus are
amended in §22.14; **§22 is the binding model they defer to.**
> **Three-tier change (v0.40.0 → v0.45.0 — supersedes the single-corpus and
> the two-tier models).** §1 originally said "this single repository is its
> content repository," and an earlier draft of this section said "a deployment
> hosts N projects, each a corpus." The current model is three tiers: a
> deployment has a **registry** (§22.2) naming N **projects**, each project
> owns **one content repo** (§22.3) holding N **RFC collections** as typed
> subfolders (§22.2). The single-corpus deployment is the **N=1 case** and
> continues to run after migration via a generated default project carrying a
> generated default collection (§22.13); no deployment is forced to adopt more
> than one of either. Where earlier sections say "the meta repo," "the corpus,"
> or "the project," read "the collection's content repo" and "the collection's
> corpus." The slices that delivered this are S1S6 (the design record is
> `docs/design/2026-06-05-three-tier-projects-collections.md`).
### 22.1 The tiers and isolation
One deployment, N projects (N ≥ 1); one project, N collections (N ≥ 1). A
collection belongs to exactly one project; a project to exactly one
deployment; neither moves. **Isolation (§22.5) holds at the collection
grain:** an RFC, branch, thread, star, or watch belongs to exactly one
collection, and no app surface joins across collections except the
per-account ones the deployment owns (the §15 inbox, the §6 account roster,
sign-in).
- **Deployment / "global."** The unchanged top tier. Owns accounts, the §6
admission gate, the §15 inbox, the §1 bot, and the landing directory. Its
management surface is **projects + global settings**.
- **Project.** Belongs to one deployment; never moves. Owns one content repo
(§22.3) and carries project settings (name, tagline, theme, visibility,
model universe). Has **no `type`** of its own. Its management surface is
**RFC collections + project settings**.
- **RFC collection.** A typed subfolder of its project's content repo (§22.2).
Carries everything the original draft pinned on a "project": the immutable
`type` (§22.4a), the per-collection slug namespace (§22.4), `initial_state`
(§22.4b), the `unreviewed` flag (§22.4c), catalog, philosophy.
- **Entry.** Unchanged (§2). Identified by its slug **within its collection**.
### 22.2 The registry and the collection manifests — git is still truth
Project and collection configuration is declared in git and mirrored into
cache tables (`projects`, `collections`) exactly the way content is mirrored
into `cached_rfcs` (§4). There are **two git sources**, both read by the bot:
1. **The registry repo** declares **projects**. A `projects.yaml` at the root
of a dedicated **registry repo** under the deployment's Gitea org lists each
project's `id`, `name`, `content_repo`, `visibility`, `theme`, and
`enabled_models`. The framework learns the registry repo's location from a
required env var (`REGISTRY_REPO`, the successor to `META_REPO`); the repo's
*name* is the deployment's choice per the separation-of-concerns rule, and
the framework fails loudly at startup if the var is unset. `content_repo`
lives on the **project** (one repo per project), not the collection.
2. **Each project's content repo** declares its **collections** as typed
subfolders, each carrying a **`.collection.yaml` manifest** (the
collection's `type`, `visibility`, `initial_state`, `name`, and optional
`enabled_models`). The registry mirror walks the content repo and reads
these manifests, so collection configuration is git-truth and survives a
cache rebuild — exactly as entry frontmatter does.
```yaml
# projects.yaml (registry repo root)
deployment:
name: Wiggleverse # deployment display name (replaces VITE_APP_NAME)
tagline: ... # deployment landing deck (§22.10)
projects:
- id: ohm # url-stable slug, unique within the deployment
name: Open Human Model
content_repo: ohm-content # ONE repo under the org; collections live inside it
visibility: public # gated | public | unlisted (§22.5)
theme: { accent: "#5b5bd6" } # optional per-project token overrides (§22.9)
enabled_models: [claude, gemini] # optional; falls back to deployment ENABLED_MODELS
```
```yaml
# ohm-content/model/.collection.yaml (one per collection subfolder)
type: document # document | specification | bdd — immutable (§22.4a)
visibility: gated # defaults to the project's, may only narrow (§22.5)
initial_state: super-draft # super-draft | active — defaults from type (§22.4b)
name: The Model
# enabled_models: [claude] # optional; narrows the project's universe (§22.12)
```
```
ohm-content/
model/
.collection.yaml # type: document
rfcs/intro.md
specs/
.collection.yaml # type: specification
rfcs/runtime.md
features/
.collection.yaml # type: bdd
rfcs/login.md
```
**Creation is in-app, wrapping a bot commit, at both tiers.** *+ New project*
(a global-Owner action, §22.6) has the bot create a Gitea content repo under
the org and commit a project entry to `projects.yaml`. *+ New collection* (a
project Owner / RFC-Contributor-with-create action) has the bot commit a new
subfolder + `.collection.yaml` to the project's content repo. The in-app
button is a thin convenience over a git write; nothing becomes app state that
git cannot rebuild. `projects` and `collections` cache rows are never written
from user actions directly — they flow from the mirror only. **Membership**
(§22.6) remains app state, as `rfc_collaborators` always has been — it churns
at user speed and is not document state.
### 22.3 Content repositories — one per project
Each project names one content repo under the deployment's single Gitea org
(convention `<project-id>-content`). Collections are **subfolders** within it
(§22.2). The §1 bot service account operates org-wide across every content
repo and the registry repo; nothing about the bot, the §6 app-owned
authorization, or the "app is the only contribution surface" stance changes.
There are no per-project Gitea orgs and no per-project bot accounts.
### 22.4 The slug namespace is per-collection; the slug is the identity
An entry's slug (§2) is unique **within its collection**: `model/intro` and
`specs/intro` coexist. The fully-qualified identity is `(project, collection,
slug)` — there is no type prefix and **no numeric ID**. The §22.4 retirement
of `RFC-NNNN` allocation stands: the slug is the identity, and graduation
(§13) flips state without allocating a number. The displayed *noun* around a
slug ("RFC", "Spec", "Feature") is a presentation concern driven by the
collection's `type` (§22.4a), not part of the identity.
**Legacy numbers.** Entries graduated *before* this change keep their existing
`id` (`RFC-NNNN`) in frontmatter as a **frozen, non-identity legacy label**
preserved and shown so external "RFC-0001"-style citations still resolve, but
never used for routing or lookup. New entries are never assigned one.
### 22.4a Collection type
> **Metadata amendment (v0.46.2 — Configurable Collection Metadata).** Item 1
> below originally made the **entry metadata schema type-driven** — a
> `document`/`specification`/`bdd` frontmatter schema baked into a per-type
> module. That is **superseded**: entry metadata is **collection-configured**,
> not type-driven. Each collection declares a `fields:` schema in its
> `.collection.yaml`, and per-entry values live in a `<slug>.meta.yaml`
> **sidecar** (the `.md` body stays pure prose; a parser reads the sidecar
> else legacy top-of-doc frontmatter). See the
> [Configurable Collection Metadata](./docs/design/2026-06-06-configurable-collection-metadata.md)
> design, which supersedes the
> [per-type-surfaces](./docs/design/2026-06-06-per-type-surfaces.md) draft
> (D11). Item 3's **type-specific surfaces** (release planning for
> `specification`; scenario/coverage views for `bdd`) are **deferred** to a
> future design; the `bdd` **coverage** capability is recorded there as a
> future `ref`-field surface that maps features to the spec entries they
> verify **as hyperlinks**, honoring the §22 rule against fusing corpora
> across collections. What `type` still selects is the **terminology** (item 2,
> the entry noun, shipped v0.45.0) and the **default `initial_state` / review
> posture** (§22.4bc).
>
> **Shipped status.** The design lands incrementally: sidecar storage +
> dual-read (SLICE-1, v0.47.0), the collection `fields:` schema + central
> validation (SLICE-2, v0.48.0), faceted left-pane filtering (SLICE-3, v0.49.0),
> and **single-entry metadata edit (SLICE-4, v0.50.0)** — the direct-commit
> `POST …/rfcs/{slug}/meta` editor (contributor+, validated at the write
> boundary), the schema-driven detail panel, all entry **write paths made
> sidecar-aware** (a migrated body-only `.md` never re-grows frontmatter), and
> the Owner-gated `POST …/collections/{cid}/migrate` endpoint. **Bulk
> tag/untag (SLICE-5, v0.51.0)** completes the design: the
> `POST …/collections/{cid}/meta/bulk` endpoint (`{slugs, op: set|add|remove,
> field, value}`) applies one field change to many entries' sidecars in a
> **single commit** (D7), validated per entry at the write boundary with
> partial-rejection reporting (`{applied, rejected}`), plus the catalog's
> row multi-select + sticky bulk action bar. See §9.5 for the edit-metadata
> write-through this reuses.
Every collection declares a `type` in its `.collection.yaml` manifest
(§22.2), chosen at creation and **immutable**: one of `document`,
`specification`, or `bdd`. Type does not change the engine — every type uses
the same content repo (§22.3), the same propose→branch→PR→discuss→graduate
lifecycle (§§913), the same threads, flags, and chat. Type selects:
1. the **terminology** the chrome uses for an entry (the §8.1 noun, catalog
labels) — the entry noun, shipped v0.45.0;
2. the **default `initial_state`** a new entry lands in, and its review
posture (§22.4b, §22.4c).
Entry **metadata** is **not** selected by type — it is **collection-configured**
(a `.collection.yaml` `fields:` schema + per-entry `<slug>.meta.yaml`
sidecars; see the amendment above, the Configurable Collection Metadata
design, and the §2 baseline). **Type-specific surfaces** layered on the shared
§7 catalog are **deferred** to a future design.
Type-specific behavior, where it exists, is implemented as a per-type module
the framework selects on `collection.type`; the engine itself treats every
entry as markdown + a metadata sidecar regardless of type. `type` is an
**open set** in shape — a future type is a new module plus a new allowed enum
value, no schema rebuild. The type names and their behavior are framework
concepts (like role names), not deployment content: a deployment picks which
type each collection is, but does not define or rename types.
- **`document`** — long-form normative prose (OHM: a model of principles and
definitions). Metadata is the §2 baseline; no collection-configured `fields:`
are required. The §22.13 generated default collection is a `document`
collection with no `fields:`, so the N=1 case is unchanged.
- **`specification`** — a versioned technical specification (this framework's
own `SPEC.md` is the archetype). A deployment that wants spec metadata
(`version`, lifecycle `status` of draft/active/superseded, `supersedes`)
declares those as collection `fields:`. **Release planning** — grouping
entries into versioned releases with a changelog + §20-style upgrade-steps —
is a **deferred** type surface.
- **`bdd`** — behavior-driven feature specs: each entry states a feature as
Given/When/Then scenarios with acceptance criteria. Feature metadata is
declared as collection `fields:`. The **scenario/acceptance view** and a
**coverage view** (mapping features to the spec entries they verify via a
future `ref` field, rendered as hyperlinks — never fusing corpora across
collections) are **deferred** type surfaces.
### 22.4b Initial state of a new entry
A collection sets the **landing state** a new entry takes when its creating
idea-PR merges (§2.4) — the `initial_state` manifest field, one of the §2.4
entry-states:
- **`super-draft`** (default for `document` and `specification`) — a new entry
lands as a super-draft and must be explicitly graduated (§13) to reach
`active`. This is today's flow.
- **`active`** (default for `bdd`) — a new entry lands `active` on the idea-PR
merge, with the **`unreviewed` flag set** (§22.4c). The §13 graduate gate is
not surfaced — the entry is already active — but because nothing reviewed
it, the flag marks it as not-yet-vetted until an owner clears it.
The default comes from the collection's **type** (§22.4a), but `initial_state`
is an independent knob. It changes only the landing state and whether
graduation is required; the underlying engine is unchanged. The §22.13 default
collection keeps `super-draft`, preserving the N=1 flow.
### 22.4c The `unreviewed` flag
An `active` entry carries an **`unreviewed`** boolean, orthogonal to its
`state`, recording whether a human gate has vetted it. An entry that reaches
`active` by the normal **graduate** path (§13) is never flagged — the graduate
action *is* the review. An entry that skips straight to `active` via
`initial_state: active` lands `unreviewed = true`. A collection **Owner**
(§22.6) clears it with a **mark-reviewed** action (§17), stamping
`reviewed_at`/`reviewed_by` for provenance. The flag is git-truth
(frontmatter, §2 amendment) and survives a cache rebuild. The §7 catalog gains
an **unreviewed filter** — the owner's worklist for the action.
### 22.5 Visibility applies at both project and collection
Visibility is `gated` | `public` | `unlisted`, and is carried at **both** the
project and the collection tier:
- **`gated`** — invisible to non-members: not shown in the directory (§22.10),
returns 404 to non-members, and reading or writing requires a scope role
(§22.6).
- **`public`** — any visitor may read under the §6.1 anonymous-read contract;
appears in the directory; contributing still requires a grant (subject to
the N=1 baseline, §22.6).
- **`unlisted`** — readable by anyone with a direct link, but not shown in the
directory and not enumerated by the deployment/project listing.
**A collection defaults to its project's visibility and may only narrow it**
(`public` < `unlisted` < `gated`; a collection may be as strict or stricter
than its project, never looser). Reading or writing a collection requires
passing **both** gates — the stricter of project and collection wins. This is
validated at create-collection (422 on a looser setting) and clamped at the
registry mirror. Project/collection visibility does not relax the §11
per-branch `read_public` controls *within* a collection.
### 22.6 Roles: one vocabulary, attached at a scope
There is **one role enum — `{owner, contributor}`** — displayed as **Owner**
and **RFC Contributor**. A grant *attaches that role at a scope*: **global**,
**project**, or **collection**. "Owner at all levels, RFC Contributor at all
levels" is literal — the same two words at every tier.
| Role | Capabilities within its scope's subtree |
|---|---|
| **Owner** | Superuser: manage settings and membership; create child projects/collections; act on any entry (merge on behalf, graduate, mark-reviewed, withdraw/reopen, set branch visibility). |
| **RFC Contributor** | Propose entries, create branches, open PRs, claim unclaimed super-drafts, participate in discussion. At **project** (or global) scope this additionally includes **creating collections** in that project. (A *collection*-scope grant cannot create sibling collections — creating one is a project-level action.) |
**Schema.** `users.role` continues to carry the **deployment admission** tier
(`owner` / `admin` / `contributor`, the §6 gate). Scope grants live in a single
polymorphic **`memberships(scope_type ∈ {global, project, collection},
scope_id, user_id, role, granted_by, granted_at)`** table; the role enum is
`{owner, contributor}`. The prior `project_members` three-role set
(`viewer`/`contributor`/`admin`) collapsed: `admin → owner`, `contributor →
contributor`, and `viewer` is deferred (a read grant folded into visibility,
not a membership role this pass). When the richer set returns it **re-splits
out of** Owner / re-adds a tier; the unified roles are not aliases.
> **Keystone reinterpretation (v0.42.0, S3 — reconciles the role mapping).**
> An earlier draft equated "deployment `contributor`" with "global RFC
> Contributor." That contradicted the open-by-default baseline. The binding
> reading: a **plain granted account** (`users.role='contributor'`, no
> membership row) is a granted *account* — it can sign in and read — **not** a
> write-everywhere global role. **"Global RFC Contributor"** is an **explicit
> `memberships(scope_type='global')` grant** (the cleo case, §22.6a). The one
> carve-out preserving N=1: the pre-multi-project **implicit-public write
> baseline is grandfathered onto the migration-seeded `default` collection
> only** — a granted `contributor` keeps its historical write capability there
> with no membership row. Every **explicitly created** collection (and any
> second project) requires an explicit scope grant to write. Deployment
> `owner`/`admin` remain superusers everywhere.
Membership is still gated by the deployment-level
`users.permission_state='granted'` (§6): a pending account has no write
capability at any scope regardless of its `memberships` rows.
### 22.6a Role & invitation scenarios
The behavioral spec for role usage, invitation, and empty-state experiences is
the BDD scenario set in
`docs/design/2026-06-05-three-tier-projects-collections.md` Part C (C.1 role
usage / inheritance / most-permissive union; C.2 invitation — who may invite
whom, at which scope; C.3 empty states at each tier). They are written so they
can also seed a `bdd`-type collection (the framework dogfooding its own model).
Each scenario carries the slice tag (`@S1``@S6`) that makes it pass.
### 22.7 How the four tiers compose
Effective authority on an entry is the **most permissive** union of four
layers, inheriting **downward**, **additive**, with **no negative override**:
```
effective authority on an entry =
global role (memberships scope_type='global'; + users.role owner/admin)
project role (membership at the entry's project)
collection role (membership at the entry's collection)
per-entry authority (owners / arbiters / rfc_collaborators — §6.3, §12)
then minus §6.2 write-mute and §22.5 visibility (subtractive, as today)
```
- A grant at **global** covers every project and collection in the deployment.
- A grant at **project** covers every collection in that project — including
collections added later, with no new grant.
- A grant at **collection** covers just that collection.
- You **cannot** grant at a parent scope and revoke at a child; resolution
never subtracts a parent grant.
**Per-entry authority is a distinct, finer layer — not a synonym.** `owners` /
`arbiters` / `rfc_collaborators` apply to *one specific entry* (§6.3, §12);
the three named scopes apply to a *subtree*. `arbiter` is narrower than Owner
(one entry, not a subtree) and stays distinct. `users.role` now means
deployment level only; no schema change demotes an existing owner/admin —
their powers read as "superuser in every project and collection."
### 22.8 Discovery and joining
Because a gated project or collection is invisible to non-members, joining is
by one of:
- **Invite** — an Owner (at the target scope or any scope above it) grants a
user a role directly, writing a `memberships` row and fanning a §15
notification. The grant may name any scope at or beneath the inviter's reach;
the **broader-scope-supersedes** rule prunes membership rows the new grant
subsumes (a project grant removes subsumed collection rows of same-or-lower
rank; a global grant removes subsumed project + collection rows; a *stronger*
child grant survives).
- **Request to join** — a user who knows a scope exists requests membership
naming a desired role; the request is recorded and surfaced to the scope's
Owners across the subtree (the cross-collection inbox), who accept or decline.
Accepting writes the `memberships` row.
A `public` project/collection needs neither for read; the existing §6 / §12
contribute-grant paths cover write.
### 22.9 Branding is resolved at runtime
`VITE_APP_NAME` is **deprecated** (§20 amendment): a single build-time name
cannot serve N projects. Deployment, project, and collection identity are
served at runtime — `GET /api/deployment` (deployment `name`, `tagline`, the
visible projects), `GET /api/projects/:id` (the project's settings + visible
collections), `GET /api/projects/:id/collections/:cid` (the collection's
settings incl. `type`). The frontend reads these instead of
`import.meta.env.VITE_APP_NAME`. Three chrome layers result: **deployment
chrome** (directory, switcher, shared inbox), **project chrome** (the
collection directory, project settings), and **collection chrome** (the §7
catalog, the §8 entry view, the §14 philosophy).
### 22.10 Routing and the landing surfaces
The canonical route gains a collection segment:
```
/p/<project>/c/<collection>/e/<slug>
```
The `c/` segment keeps collection ids from colliding with reserved
project-level segments. Reserved **collection-level** siblings (`proposals`,
`philosophy`) sit under `/p/<project>/c/<collection>/…`. The displayed entry
noun is the collection type's label (§22.4a), not part of the path.
- `/` is the **deployment landing**: a directory of the projects the visitor
can see (§22.5). Redirects to the sole visible project when there is exactly
one (the N=1 case).
- `/p/<project>/` is the **project landing**: a directory of the collections
the visitor can see. Redirects to its sole visible collection when there is
exactly one.
**Backcompat.** The shipped `/p/<project>/e/<slug>` URLs (v0.35.0)
**308-redirect** to `/p/<project>/c/<default>/e/<slug>`, and the
pre-multi-project `/rfc/<slug>` / `/proposals/<n>` redirect to their
`/p/<default-project>/c/<default-collection>/…` equivalents. Both are handled
in the migration (§22.13).
### 22.11 Notifications span the deployment, one inbox
Accounts are deployment-wide, so the §15 inbox is one inbox across all the
caller's collections. Entry-scoped notification rows carry the entry's
`collection_id` (and a denormalized `project_id`) so the inbox filters by
collection or project and a user can mute an entire collection. Quiet hours,
digest cadence, and email preferences stay per-account at the deployment level
(§5, §15). The **cross-collection inbox** (§22.8) surfaces join requests to
the Owners of the scope they target, aggregated across the subtree.
### 22.12 Per-collection model universe
A collection's `enabled_models` (its `.collection.yaml` manifest, §22.2)
narrows its **project's** `enabled_models` (registry, §22.2), which in turn
overrides the deployment `ENABLED_MODELS` (§18). Resolution order is **funder
universe ∩ §6.6 per-entry list ∩ collection universe ∩ project universe**,
with the collection universe substituting for the deployment universe at the
outermost step. A collection's universe may only narrow, never widen, its
project's; the project's may only narrow the deployment's.
### 22.13 Migration — the default project and default collection (N=1)
A deployment on the shipped two-tier schema (v0.39.0) is migrated so it keeps
running unchanged:
1. The existing `projects` row **stays as the project** (it already owns
`content_repo` and its config-derived `id` from the §22.13 re-stamp).
2. A **default collection** (`id='default'`, `subfolder` = repo root) is
created per project, inheriting that project's `type` / `initial_state` /
visibility; those per-corpus fields are then dropped from `projects`.
3. Every entry-scoped row is re-keyed `(project_id, slug)`
`(collection_id, slug)` via the migration-028 rebuild pattern.
4. `project_members` rows migrate to `memberships(scope_type='collection')` on
the default collection, role-collapsed (§22.6).
5. **308 redirects:** the shipped `/p/<project>/e/<slug>`
`/p/<project>/c/<default>/e/<slug>`, and the pre-multi-project `/rfc/<slug>`
/ `/proposals/<n>` → their `/p/<project>/c/<default>/…` equivalents.
Until a second collection is added, the deployment is functionally identical
to before, with one extra path segment. This is the §20.4 upgrade-steps
content for the release.
### 22.14 Amendments to §§121 (applied in place)
The single-corpus sections defer to §22; the load-bearing reinterpretations:
- **§1 Repository topology.** Each *project* names one content repo; the
deployment's registry (§22.2) lists them; collections are subfolders within
a project's repo (§22.3). The bot and app-owned-authorization paragraphs are
unchanged and now read org-wide.
- **§2 Schema / §2.3 IDs.** Slugs are unique **per collection**; the entry
**metadata schema is collection-configured**, not type-driven (§22.4a, as
amended by v0.46.2) — a `.collection.yaml` `fields:` schema + per-entry
`<slug>.meta.yaml` sidecars. The `RFC-NNNN` `max+1` allocation is
**removed** — the slug is the identity (§22.4). New `active`-entry fields:
`unreviewed` (bool) and the `reviewed_at`/`reviewed_by` provenance pair
(§22.4c).
- **§2.4 State machine.** The `(no entry) ─[idea-PR merged]→` transition
targets the collection's `initial_state` (§22.4b); a new `active
─[mark-reviewed, Owner]→ active` self-transition clears the `unreviewed`
flag (§22.4c).
- **§5 Data model.** `project_id` becomes **`collection_id`** on every
entry-scoped row (the corpus grain is now the collection); a separate
`project_id` exists only on the `collections` table and project-scoped rows.
`cached_rfcs` PK → `(collection_id, slug)` and mirrors the `unreviewed`
frontmatter flag. New tables: `projects`, `collections` (carrying the
immutable `type`, §22.4a), and `memberships` (§22.6, replacing
`project_members`). `users.role` is annotated deployment-scope (§22.7).
- **§6 Permission model.** Deployment roles are the global tier of the §22.7
four-layer union; a plain `contributor` has no implicit write at any
explicitly-created scope until a `memberships` grant gives it one (the N=1
default-collection baseline is the sole carve-out, §22.6 keystone note).
`project_admin`'s delegation idea is now **Owner** at project/collection
scope (§22.6), sitting above per-RFC authority (§6.3).
- **§7 / §8.1 / §13.3.** The catalog is per-collection under
`/p/<project>/c/<collection>/`; the project collection-directory and the
deployment directory (§22.10) sit above it; the §7 catalog gains the
unreviewed filter (§22.4c). The §8.1 breadcrumb gains leading project +
collection segments. §13.3 graduation operates on the collection's content
subfolder, allocates no number, and is a no-op (replaced by mark-reviewed)
for collections whose `initial_state` is `active`.
- **§14.1 / §17 / §18 / §20.** The landing splits into deployment directory,
project collection-directory, and per-collection philosophy/deck. §17 routes
gain the `/p/<project>/c/<collection>/` scoping plus `GET /api/deployment`,
`GET /api/projects/:id`, `GET /api/projects/:id/collections[/:cid]`, the
`memberships` management + request-to-join endpoints, and the mark-reviewed
endpoint. `ENABLED_MODELS` is the deployment fallback under §22.12.
`VITE_APP_NAME` is deprecated and `REGISTRY_REPO` is a required env var
(§20.3); `META_REPO` is legacy, consulted only by the §22.13 migration.
+1 -1
View File
@@ -1 +1 @@
0.37.0
0.54.1
+13
View File
@@ -130,3 +130,16 @@ CLOUDFLARE_TURNSTILE_SECRET=
# config drift surfaces as a loud 500 rather than a silent abuse-
# defense disablement.
TURNSTILE_REQUIRED=false
# --- Deployed-environment E2E test auth (v0.52.0) ---
# DANGER: NEVER set these on a production deployment. Together they
# enable `POST /auth/test/login`, which mints an authenticated OWNER
# session for the one configured email without any OTC/email round trip
# — it exists only to run the Playwright E2E suite against a deployed
# pre-prod (PPE) host that has no Mailpit sink. The route is fail-closed:
# it returns 404 unless BOTH vars below are set, requires the caller to
# present E2E_TEST_AUTH_SECRET in the `X-Test-Auth-Secret` header
# (constant-time compare), and only ever mints the single configured
# email (any other → 403). Leave BOTH unset everywhere except PPE.
# E2E_TEST_AUTH_EMAIL=e2e-owner@example.test
# E2E_TEST_AUTH_SECRET= # a Secret Manager ref on real deployments; never a literal here
+275 -72
View File
@@ -21,14 +21,19 @@ from pydantic import BaseModel, Field
from . import (
api_admin,
api_branches,
api_collections,
api_contributions,
api_deployment,
api_discussion,
api_graduation,
api_invitations,
api_join_requests,
api_memberships,
api_metadata,
api_notifications,
api_prs,
auth,
collections as collections_mod,
projects as projects_mod,
db,
device_trust as device_trust_mod,
@@ -37,6 +42,7 @@ from . import (
docs_specs,
entry as entry_mod,
cache,
facets,
funder,
health,
notify,
@@ -125,6 +131,8 @@ def make_router(
router.include_router(api_prs.make_router(config, gitea, bot, providers))
# Slice 5: §13 graduation + §13.1 claim.
router.include_router(api_graduation.make_router(config, gitea, bot))
# §22.4a SLICE-4/5: entry metadata edit + Owner-gated collection migrate.
router.include_router(api_metadata.make_router(config, gitea, bot))
# Slice 6: §15 notifications surface (inbox, watches, prefs,
# quiet hours, per-user mute, email unsubscribe, bounce webhook).
router.include_router(api_notifications.make_router(config))
@@ -150,7 +158,15 @@ def make_router(
router.include_router(api_contributions.make_router())
# §22.9/§22.10 (M3): runtime deployment + per-project config (replaces
# VITE_APP_NAME) + the old-URL 308 redirects.
router.include_router(api_deployment.make_router(config))
router.include_router(api_deployment.make_router(config, gitea, bot))
router.include_router(api_collections.make_router(config, gitea, bot))
# §22 S4 (C.2): the scope-role invitation surface — Owners grant
# {owner, contributor} at project/collection scope to existing accounts.
router.include_router(api_memberships.make_router())
# §22.8 S6: request-to-join + the cross-collection inbox — a user asks into a
# scope (naming a role); the scope's Owners across the subtree accept (writing
# the membership row) or decline.
router.include_router(api_join_requests.make_router())
# ---------------------------------------------------------------
# §17: /api/health — unauthenticated post-flight probe.
@@ -637,13 +653,14 @@ def make_router(
unreviewed_clause = " AND unreviewed = 1 AND state = 'active'"
rows = db.conn().execute(
f"""
SELECT slug, title, state, rfc_id, repo,
owners_json, arbiters_json, tags_json,
last_main_commit_at, last_entry_commit_at, updated_at
FROM cached_rfcs
WHERE state IN ('super-draft', 'active')
AND project_id IN ({placeholders}){unreviewed_clause}
ORDER BY COALESCE(last_main_commit_at, last_entry_commit_at) DESC
SELECT r.slug, r.title, r.state, r.rfc_id, r.repo,
r.owners_json, r.arbiters_json, r.tags_json,
r.metadata_malformed,
r.last_main_commit_at, r.last_entry_commit_at, r.updated_at
FROM cached_rfcs r JOIN collections c ON c.id = r.collection_id
WHERE r.state IN ('super-draft', 'active')
AND c.project_id IN ({placeholders}){unreviewed_clause}
ORDER BY COALESCE(r.last_main_commit_at, r.last_entry_commit_at) DESC
""",
params,
).fetchall()
@@ -672,6 +689,7 @@ def make_router(
"last_active_at": r["last_main_commit_at"] or r["last_entry_commit_at"] or r["updated_at"],
"starred_by_me": r["slug"] in starred,
"has_open_prs": False, # wired in Slice 2 when per-RFC repos exist
"metadata_malformed": bool(r["metadata_malformed"]),
}
)
return {"items": items}
@@ -685,8 +703,8 @@ def make_router(
raise HTTPException(404, "Not found")
viewer = auth.current_user(request)
# §22.5 visibility gate (subtractive, §22.7): a gated project's entries
# 404 to non-members.
auth.require_project_readable(viewer, row["project_id"])
# 404 to non-members. Recover the project via the entry's collection.
auth.require_project_readable(viewer, auth.project_of_rfc(slug))
# §13.7: a retired entry is removed from every browsing surface. The
# sole exception is a site owner, so the un-retire affordance has
# somewhere to live; everyone else gets a plain 404.
@@ -706,6 +724,9 @@ def make_router(
(slug,),
).fetchone()
payload["proposed_use_case"] = uc["use_case"] if uc else None
# §22.4a SLICE-4: contributor+ on the entry's collection may edit metadata.
payload["can_edit_meta"] = bool(
auth.can_contribute_in_collection(viewer, auth.collection_of_rfc(slug)))
return payload
# ---------------------------------------------------------------
@@ -716,40 +737,75 @@ def make_router(
# second project's corpus renders under /p/<id>/.
# ---------------------------------------------------------------
@router.get("/api/projects/{project_id}/rfcs")
async def list_project_rfcs(
project_id: str, request: Request, unreviewed: str | None = None
def _require_collection_in_project(collection_id: str, project_id: str) -> None:
# §22 S2: a collection-scoped route 404s when the collection does not
# belong to the project in the path (shape matches an unknown id).
if collections_mod.project_of_collection(collection_id) != project_id:
raise HTTPException(404, "Not found")
def _list_rfcs_for_collection(
collection_id: str, viewer, unreviewed: str | None,
query_params=None,
) -> dict[str, Any]:
viewer = auth.current_user(request)
# §22.5 read gate: a gated project's catalog 404s to a non-member.
auth.require_project_readable(viewer, project_id)
viewer_id = viewer.user_id if viewer else None
# §22.4a SLICE-3: the collection's declared field schema drives the facet
# set (None when undeclared → no facets, INV-5).
col = collections_mod.get_collection(collection_id)
fields_schema = (col or {}).get("fields") or None
# Parse + validate filter selections from the query string. Unknown
# field → 400 (§6.4). `unreviewed` keeps its existing meaning; an
# empty-valued selection is ignored, not an error (plan decision 6).
selections: dict[str, set[str]] = {}
only_malformed = False
if query_params is not None:
allowed = facets.allowed_filter_keys(fields_schema)
facet_names = {n for n, _ in facets.facet_fields(fields_schema)}
for key in query_params.keys():
if key not in allowed:
raise HTTPException(400, f"unknown filter field {key!r}")
if (query_params.get("malformed") or "").lower() in ("1", "true", "yes"):
only_malformed = True
for name in facet_names:
vals = {v for v in query_params.getlist(name) if v != ""}
if vals:
selections[name] = vals
unreviewed_clause = ""
if unreviewed is not None and unreviewed.lower() in ("1", "true", "yes"):
unreviewed_clause = " AND unreviewed = 1 AND state = 'active'"
rows = db.conn().execute(
f"""
SELECT slug, title, state, rfc_id, repo,
owners_json, arbiters_json, tags_json,
owners_json, arbiters_json, tags_json, metadata_malformed,
meta_json,
last_main_commit_at, last_entry_commit_at, updated_at
FROM cached_rfcs
WHERE state IN ('super-draft', 'active')
AND project_id = ?{unreviewed_clause}
AND collection_id = ?{unreviewed_clause}
ORDER BY COALESCE(last_main_commit_at, last_entry_commit_at) DESC
""",
(project_id,),
(collection_id,),
).fetchall()
starred = set()
if viewer_id is not None:
starred = {
r["rfc_slug"]
for r in db.conn().execute(
"SELECT rfc_slug FROM stars WHERE user_id = ? AND project_id = ?",
(viewer_id, project_id),
"SELECT rfc_slug FROM stars WHERE user_id = ? AND collection_id = ?",
(viewer_id, collection_id),
)
}
items = [
{
# Build entry dicts the facet helper understands (state + malformed +
# parsed meta), preserving SQL order.
entries = []
for r in rows:
try:
meta = json.loads(r["meta_json"]) if r["meta_json"] else {}
except (TypeError, ValueError):
meta = {}
entries.append({
"slug": r["slug"],
"title": r["title"],
"state": r["state"],
@@ -761,18 +817,19 @@ def make_router(
"last_active_at": r["last_main_commit_at"] or r["last_entry_commit_at"] or r["updated_at"],
"starred_by_me": r["slug"] in starred,
"has_open_prs": False,
}
for r in rows
]
return {"items": items}
"metadata_malformed": bool(r["metadata_malformed"]),
"meta": meta,
})
@router.get("/api/projects/{project_id}/rfcs/{slug}")
async def get_project_rfc(project_id: str, slug: str, request: Request) -> dict[str, Any]:
viewer = auth.current_user(request)
auth.require_project_readable(viewer, project_id)
filtered, facet_counts = facets.filter_and_count(
entries, fields_schema, selections, only_malformed=only_malformed
)
return {"items": filtered, "facets": facet_counts}
def _get_rfc_for_collection(collection_id: str, slug: str, viewer) -> dict[str, Any]:
row = db.conn().execute(
"SELECT * FROM cached_rfcs WHERE project_id = ? AND slug = ?",
(project_id, slug),
"SELECT * FROM cached_rfcs WHERE collection_id = ? AND slug = ?",
(collection_id, slug),
).fetchone()
if row is None:
raise HTTPException(404, "Not found")
@@ -782,14 +839,64 @@ def make_router(
uc = db.conn().execute(
"""
SELECT use_case FROM proposed_use_cases
WHERE scope = 'rfc' AND rfc_slug = ? AND project_id = ?
WHERE scope = 'rfc' AND rfc_slug = ? AND collection_id = ?
ORDER BY id DESC LIMIT 1
""",
(slug, project_id),
(slug, collection_id),
).fetchone()
payload["proposed_use_case"] = uc["use_case"] if uc else None
# §22.4a SLICE-4: contributor+ on the collection may edit metadata (INV-4).
payload["can_edit_meta"] = bool(
auth.can_contribute_in_collection(viewer, collection_id))
return payload
@router.get("/api/projects/{project_id}/rfcs")
async def list_project_rfcs(
project_id: str, request: Request, unreviewed: str | None = None
) -> dict[str, Any]:
viewer = auth.current_user(request)
# §22.5 read gate: a gated project's catalog 404s to a non-member.
auth.require_project_readable(viewer, project_id)
# §22 S1: the project-scoped route serves the default collection.
collection_id = collections_mod.default_collection_id(project_id)
return _list_rfcs_for_collection(
collection_id, viewer, unreviewed, query_params=request.query_params
)
@router.get("/api/projects/{project_id}/rfcs/{slug}")
async def get_project_rfc(project_id: str, slug: str, request: Request) -> dict[str, Any]:
viewer = auth.current_user(request)
auth.require_project_readable(viewer, project_id)
collection_id = collections_mod.default_collection_id(project_id)
return _get_rfc_for_collection(collection_id, slug, viewer)
# §22 S2: collection-scoped serve + propose. The catalog/entry views read
# these under /p/<project>/c/<collection>/; the project-scoped routes above
# stay as the default-collection compat surface.
@router.get("/api/projects/{project_id}/collections/{collection_id}/rfcs")
async def list_collection_rfcs(
project_id: str, collection_id: str, request: Request,
unreviewed: str | None = None,
) -> dict[str, Any]:
viewer = auth.current_user(request)
auth.require_project_readable(viewer, project_id)
_require_collection_in_project(collection_id, project_id)
# §22.5 (S3): a hidden/gated collection 404s to a non-scope-role viewer.
auth.require_collection_readable(viewer, collection_id)
return _list_rfcs_for_collection(
collection_id, viewer, unreviewed, query_params=request.query_params
)
@router.get("/api/projects/{project_id}/collections/{collection_id}/rfcs/{slug}")
async def get_collection_rfc(
project_id: str, collection_id: str, slug: str, request: Request
) -> dict[str, Any]:
viewer = auth.current_user(request)
auth.require_project_readable(viewer, project_id)
_require_collection_in_project(collection_id, project_id)
auth.require_collection_readable(viewer, collection_id)
return _get_rfc_for_collection(collection_id, slug, viewer)
# ---------------------------------------------------------------
# §22.4c: mark-reviewed — clear an active entry's `unreviewed` flag
# ---------------------------------------------------------------
@@ -797,27 +904,33 @@ def make_router(
@router.post("/api/projects/{project_id}/rfcs/{slug}/mark-reviewed")
async def mark_reviewed(project_id: str, slug: str, request: Request) -> dict[str, Any]:
"""§22.4c — clear an active entry's `unreviewed` flag. Authority is the
§22.7 project superuser (project_admin or deployment owner/admin)."""
§B.2 collection Owner (a collection/project/global Owner or deployment
owner/admin reaching the entry's collection)."""
viewer = auth.require_user(request)
auth.require_project_readable(viewer, project_id)
if not auth.is_project_superuser(viewer, project_id):
raise HTTPException(403, "Only a project owner/admin can mark an entry reviewed")
collection_id = collections_mod.default_collection_id(project_id)
if not auth.is_collection_superuser(viewer, collection_id):
raise HTTPException(403, "Only a collection owner can mark an entry reviewed")
row = db.conn().execute(
"SELECT state, unreviewed FROM cached_rfcs WHERE slug = ? AND project_id = ?",
(slug, project_id),
"SELECT state, unreviewed FROM cached_rfcs WHERE slug = ? AND collection_id = ?",
(slug, collection_id),
).fetchone()
if row is None:
raise HTTPException(404, "Not found")
if row["state"] != "active" or not row["unreviewed"]:
raise HTTPException(409, "Entry is not an unreviewed active entry")
# §22/G-15: write to the entry's project content_repo + collection
# subfolder, not the deployment default.
org, meta_repo, md_path = projects_mod.entry_location(config, collection_id, slug)
try:
await bot.mark_entry_reviewed(
viewer.as_actor(),
org=config.gitea_org,
meta_repo=(projects_mod.default_content_repo(config) or ""),
org=org,
meta_repo=meta_repo,
slug=slug,
reviewed_by=viewer.gitea_login,
reviewed_at=entry_mod.today(),
file_path=md_path,
)
except GiteaError as e:
raise HTTPException(502, f"Gitea: {e.detail}")
@@ -870,6 +983,35 @@ def make_router(
]
}
@router.get("/api/projects/{project_id}/proposals")
async def list_project_proposals(project_id: str, request: Request) -> dict[str, Any]:
# §22.4/§22.5: the pending idea-PRs scoped to one project.
viewer = auth.current_user(request)
auth.require_project_readable(viewer, project_id)
rows = db.conn().execute(
"""
SELECT rfc_slug, pr_number, title, description, opened_by, opened_at, state
FROM cached_prs
WHERE pr_kind = 'idea' AND state = 'open' AND project_id = ?
ORDER BY opened_at DESC
""",
(project_id,),
).fetchall()
return {
"items": [
{
"slug": r["rfc_slug"],
"pr_number": r["pr_number"],
"title": r["title"],
"description": r["description"],
"opened_by": r["opened_by"],
"opened_at": r["opened_at"],
"proposed_use_case": _proposal_use_case(r["pr_number"]),
}
for r in rows
]
}
@router.get("/api/proposals/{pr_number}")
async def get_proposal(pr_number: int, request: Request) -> dict[str, Any]:
"""§9.3 pending-idea view data.
@@ -893,7 +1035,19 @@ def make_router(
# Read the proposed entry file from the head branch.
slug = row["rfc_slug"]
head = row["head_branch"]
result = await gitea.read_file(config.gitea_org, (projects_mod.default_content_repo(config) or ""), f"rfcs/{slug}.md", ref=head)
# §22/G-15: the proposal lives in its project's content_repo under its
# collection's `<subfolder>/rfcs/`. cached_prs carries project_id but not
# collection_id, so resolve the repo from the project and locate the file
# by trying each of the project's collection subfolders (default first).
repo = (projects_mod.content_repo(row["project_id"])
or projects_mod.default_content_repo(config) or "")
result = None
for col in collections_mod.list_collections(row["project_id"], include_unlisted=True):
sub = col["subfolder"] or ""
cand = f"{sub}/rfcs/{slug}.md" if sub else f"rfcs/{slug}.md"
result = await gitea.read_file(config.gitea_org, repo, cand, ref=head)
if result:
break
entry_payload: dict[str, Any] | None = None
if result:
text, _sha = result
@@ -922,15 +1076,21 @@ def make_router(
# §9.1: propose a new RFC
# ---------------------------------------------------------------
@router.post("/api/rfcs/propose")
async def propose_rfc(payload: ProposeBody, request: Request) -> dict[str, Any]:
user = auth.require_contributor(request)
# §22.6/§22.7: proposing a new entry requires project-level contribute
# standing. Through M2 every entry lands in the default project; M3's
# routing carries the target project. On the public default project the
# implicit-public baseline preserves the pre-multi-project flow.
if not auth.can_contribute_in_project(user, auth.DEFAULT_PROJECT_ID):
raise HTTPException(403, "You do not have contribute access to this project")
async def _propose_into_project(project_id: str, payload: ProposeBody, user) -> dict[str, Any]:
# Default-collection wrapper (§22 S1/S2): resolve the project's default
# collection and delegate. Keeps the project-scoped propose routes intact.
return await _propose_into_collection(
project_id, collections_mod.default_collection_id(project_id), payload, user
)
async def _propose_into_collection(
project_id: str, collection_id: str, payload: ProposeBody, user
) -> dict[str, Any]:
# §B.2 (S3): proposing a new entry requires contribute standing in the
# *target collection* — the four-layer scope-role union, with the
# grandfathered implicit-public baseline on the default collection.
if not auth.can_contribute_in_collection(user, collection_id):
raise HTTPException(403, "You do not have contribute access to this collection")
slug = payload.slug.strip().lower()
if not entry_mod.is_valid_slug(slug):
raise HTTPException(422, "Slug must be lowercase letters, digits, and dashes")
@@ -940,22 +1100,24 @@ def make_router(
# on every keystroke, since a concurrent submission could land
# between dialog-open and submit.
clash = db.conn().execute(
"SELECT 1 FROM cached_rfcs WHERE slug = ?", (slug,)
"SELECT 1 FROM cached_rfcs WHERE slug = ? AND collection_id = ?", (slug, collection_id)
).fetchone()
if clash:
raise HTTPException(409, f"Slug `{slug}` is already taken")
idea_clash = db.conn().execute(
"SELECT 1 FROM cached_prs WHERE pr_kind = 'idea' AND state = 'open' AND rfc_slug = ?",
(slug,),
"SELECT 1 FROM cached_prs WHERE pr_kind = 'idea' AND state = 'open' "
"AND rfc_slug = ? AND project_id = ?",
(slug, project_id),
).fetchone()
if idea_clash:
raise HTTPException(409, f"Slug `{slug}` is already reserved by an open proposal")
# §22.4b: the target project's landing state. Through Plan A every
# entry lands in the default project; M3-frontend routing carries a
# non-default target later.
target_project = projects_mod.resolved_default_id(config)
landing_state = "active" if projects_mod.project_initial_state(target_project) == "active" else "super-draft"
# §22.4b: the target collection's landing state (the per-corpus field
# moved down to the collection in migration 029).
landing_state = (
"active" if collections_mod.collection_initial_state(collection_id) == "active"
else "super-draft"
)
entry = entry_mod.Entry(
slug=slug,
@@ -986,15 +1148,19 @@ def make_router(
f"**Topic:** {entry.title}\n\n"
f"{payload.pitch.strip()}"
)
# §22 S2: write the entry under the target collection's <subfolder>/rfcs.
subfolder = collections_mod.subfolder_of(collection_id)
rfcs_dir = f"{subfolder}/rfcs" if subfolder else "rfcs"
try:
pr = await bot.open_idea_pr(
user.as_actor(),
org=config.gitea_org,
meta_repo=(projects_mod.default_content_repo(config) or ""),
meta_repo=(projects_mod.content_repo(project_id) or ""),
slug=slug,
file_contents=contents,
pr_title=pr_title,
pr_description=pr_description,
rfcs_dir=rfcs_dir,
)
except GiteaError as e:
raise HTTPException(502, f"Gitea: {e.detail}")
@@ -1014,19 +1180,48 @@ def make_router(
if use_case:
db.conn().execute(
"""
INSERT INTO proposed_use_cases (scope, rfc_slug, pr_number, use_case)
VALUES ('rfc', ?, ?, ?)
ON CONFLICT(project_id, scope, pr_number) DO UPDATE SET use_case = excluded.use_case
INSERT INTO proposed_use_cases (scope, rfc_slug, pr_number, use_case, collection_id)
VALUES ('rfc', ?, ?, ?, ?)
ON CONFLICT(collection_id, scope, pr_number) DO UPDATE SET use_case = excluded.use_case
""",
(slug, pr["number"], use_case),
(slug, pr["number"], use_case, collection_id),
)
db.conn().execute(
"UPDATE cached_prs SET proposed_use_case = ? WHERE pr_kind = 'idea' AND pr_number = ?",
(use_case, pr["number"]),
"UPDATE cached_prs SET proposed_use_case = ? WHERE pr_kind = 'idea' AND pr_number = ? AND project_id = ?",
(use_case, pr["number"], project_id),
)
return {"pr_number": pr["number"], "slug": slug}
@router.post("/api/rfcs/propose")
async def propose_rfc(payload: ProposeBody, request: Request) -> dict[str, Any]:
# Default-project compat path (pre-multi-project clients).
user = auth.require_contributor(request)
return await _propose_into_project(projects_mod.resolved_default_id(config), payload, user)
@router.post("/api/projects/{project_id}/rfcs/propose")
async def propose_project_rfc(
project_id: str, payload: ProposeBody, request: Request
) -> dict[str, Any]:
# §22.4: propose a new entry into a specific project (read-gated first
# so a gated project 404s a non-member before the contribute check).
user = auth.require_contributor(request)
auth.require_project_readable(user, project_id)
return await _propose_into_project(project_id, payload, user)
@router.post("/api/projects/{project_id}/collections/{collection_id}/rfcs/propose")
async def propose_collection_rfc(
project_id: str, collection_id: str, payload: ProposeBody, request: Request
) -> dict[str, Any]:
# §22 S2: propose a new entry into a specific collection of a project.
user = auth.require_contributor(request)
auth.require_project_readable(user, project_id)
_require_collection_in_project(collection_id, project_id)
# §22.5 (S3): a hidden/gated collection 404s a non-scope-role viewer
# before the contribute check (existence is not revealed).
auth.require_collection_readable(user, collection_id)
return await _propose_into_collection(project_id, collection_id, payload, user)
# ---------------------------------------------------------------
# §9.1 Slice 2 (roadmap #27): Claude Haiku tag suggestions as the
# propose-RFC fields fill in. The modal debounce-posts the partial
@@ -1066,7 +1261,9 @@ def make_router(
await bot.merge_idea_pr(
user.as_actor(),
org=config.gitea_org,
meta_repo=(projects_mod.default_content_repo(config) or ""),
# §22/G-15: the idea PR lives in its project's content_repo.
meta_repo=(projects_mod.content_repo(row["project_id"])
or projects_mod.default_content_repo(config) or ""),
pr_number=pr_number,
slug=row["rfc_slug"],
)
@@ -1086,7 +1283,8 @@ def make_router(
await bot.decline_idea_pr(
user.as_actor(),
org=config.gitea_org,
meta_repo=(projects_mod.default_content_repo(config) or ""),
meta_repo=(projects_mod.content_repo(row["project_id"])
or projects_mod.default_content_repo(config) or ""),
pr_number=pr_number,
slug=row["rfc_slug"],
comment=body.comment,
@@ -1110,7 +1308,8 @@ def make_router(
await bot.withdraw_idea_pr(
user.as_actor(),
org=config.gitea_org,
meta_repo=(projects_mod.default_content_repo(config) or ""),
meta_repo=(projects_mod.content_repo(row["project_id"])
or projects_mod.default_content_repo(config) or ""),
pr_number=pr_number,
slug=row["rfc_slug"],
)
@@ -1162,12 +1361,12 @@ def make_router(
async def add_funder_consent(slug: str, request: Request) -> dict[str, Any]:
user = auth.require_contributor(request)
rfc = db.conn().execute(
"SELECT project_id FROM cached_rfcs WHERE slug = ?", (slug,)
"SELECT 1 FROM cached_rfcs WHERE slug = ?", (slug,)
).fetchone()
if rfc is None:
raise HTTPException(404, "RFC not found")
# §22.5 visibility gate (subtractive): gated → 404 to non-members.
auth.require_project_readable(user, rfc["project_id"])
auth.require_project_readable(user, auth.project_of_rfc(slug))
# §6.7: refuse consent from a user with no registered credentials
# — a consent without a universe would be inert and the surface
# should fail loudly rather than silently.
@@ -1216,6 +1415,10 @@ def _serialize_rfc(row) -> dict[str, Any]:
"arbiters": json.loads(row["arbiters_json"] or "[]"),
"tags": json.loads(row["tags_json"] or "[]"),
"body": row["body"] or "",
"metadata_malformed": bool(row["metadata_malformed"]),
# §22.4a SLICE-4: the full per-entry metadata mapping (known + custom
# fields) so the detail panel can render schema-driven controls.
"meta": json.loads(row["meta_json"] or "{}"),
}
+127 -55
View File
@@ -29,7 +29,7 @@ from fastapi import APIRouter, HTTPException, Request
from fastapi.responses import StreamingResponse
from pydantic import BaseModel, Field
from . import auth, cache, chat as chat_layer, db, entry as entry_mod, funder, models_resolver, projects as projects_mod
from . import auth, cache, chat as chat_layer, collections as collections_mod, db, entry as entry_mod, funder, metadata as metadata_mod, models_resolver, projects as projects_mod
from .bot import Bot
from .config import Config
from .gitea import Gitea, GiteaError
@@ -40,6 +40,37 @@ log = logging.getLogger(__name__)
RFC_FILE_PATH = "RFC.md"
# ---------------------------------------------------------------------------
# §22.4a SLICE-4: sidecar-aware body extract/wrap (pure, unit-testable)
# ---------------------------------------------------------------------------
def _extract_body_pure(rfc, file_contents: str, branch: str, *, is_meta: bool) -> str:
"""Editable body of an entry file. Meta-resident files carry a frontmatter
envelope (legacy) or are already body-only (migrated, §22.4a); per-RFC repo
files are body-only. Dual-read tolerant: a body-only `.md` returns as-is."""
if not is_meta:
return file_contents
return metadata_mod.strip_frontmatter(file_contents)
def _wrap_body_pure(rfc, prior_contents: str, new_body: str, branch: str, *, is_meta: bool) -> str:
"""Inverse of `_extract_body_pure`. Under §22.4a the body lives in the `.md`
and metadata in the sidecar, so wrapping is identity for body-only files —
frontmatter is never re-grown here. A legacy un-migrated meta file still has
its metadata in the `.md` frontmatter (no sidecar yet), so preserve it rather
than silently dropping it on a pure body edit; it is migrated to body-only on
its next *metadata* edit."""
nb = new_body if new_body.endswith("\n") else new_body + "\n"
if not is_meta:
return nb
if entry_mod.FRONTMATTER_RE.match(prior_contents):
e = entry_mod.parse(prior_contents)
e.body = nb
return entry_mod.serialize(e)
return nb
# ---------------------------------------------------------------------------
# Request bodies
# ---------------------------------------------------------------------------
@@ -139,10 +170,7 @@ def make_router(
# open edit branches, open meta-repo body-edit and metadata PRs.
# -------------------------------------------------------------------
@router.get("/api/rfcs/{slug}/main")
async def get_rfc_main(slug: str, request: Request) -> dict[str, Any]:
viewer = auth.current_user(request)
rfc = _require_rfc(slug, viewer)
async def _main_payload(rfc, slug: str, viewer) -> dict[str, Any]:
if rfc["state"] not in ("active", "super-draft"):
raise HTTPException(409, f"RFC is {rfc['state']}")
@@ -263,6 +291,23 @@ def make_router(
"pre_graduation_history": pre_grad,
}
@router.get("/api/rfcs/{slug}/main")
async def get_rfc_main(slug: str, request: Request) -> dict[str, Any]:
viewer = auth.current_user(request)
return await _main_payload(_require_rfc(slug, viewer), slug, viewer)
@router.get("/api/projects/{project_id}/collections/{collection_id}/rfcs/{slug}/main")
async def get_rfc_main_scoped(
project_id: str, collection_id: str, slug: str, request: Request,
) -> dict[str, Any]:
# §22/G-15: collection-scoped canonical-body read. Disambiguates a slug
# that exists in two collections (G-5) and resolves the entry's own
# content repo / subfolder via `_repo_for`/`_file_path_for`.
viewer = auth.current_user(request)
_check_collection_in_project(project_id, collection_id)
rfc = _require_rfc(slug, viewer, collection_id=collection_id)
return await _main_payload(rfc, slug, viewer)
# The bare `GET /api/rfcs/<slug>/branches/<branch>` is declared
# at the *bottom* of this router so the more-specific deeper GET
# routes — `branches/{branch:path}/threads` and
@@ -458,6 +503,7 @@ def make_router(
org=owner,
meta_repo=repo,
slug=slug,
file_path=path,
new_file_contents=new_content,
prior_sha=prior_sha,
pr_title=pr_title,
@@ -742,7 +788,7 @@ def make_router(
"""
INSERT INTO branch_visibility (rfc_slug, branch_name, read_public, contribute_mode)
VALUES (?, ?, ?, ?)
ON CONFLICT(project_id, rfc_slug, branch_name) DO UPDATE SET
ON CONFLICT(collection_id, rfc_slug, branch_name) DO UPDATE SET
read_public = excluded.read_public,
contribute_mode = excluded.contribute_mode
""",
@@ -896,7 +942,7 @@ def make_router(
"""
INSERT INTO branch_chat_seen (user_id, rfc_slug, branch_name, last_seen_message_id, seen_at)
VALUES (?, ?, ?, ?, datetime('now'))
ON CONFLICT(project_id, user_id, rfc_slug, branch_name) DO UPDATE SET
ON CONFLICT(collection_id, user_id, rfc_slug, branch_name) DO UPDATE SET
last_seen_message_id = excluded.last_seen_message_id,
seen_at = excluded.seen_at
""",
@@ -1017,10 +1063,7 @@ def make_router(
# else, including slashed branch names like `foo/bar`.
# -------------------------------------------------------------------
@router.get("/api/rfcs/{slug}/branches/{branch:path}")
async def get_branch_view(slug: str, branch: str, request: Request) -> dict[str, Any]:
viewer = auth.current_user(request)
rfc = _require_rfc_with_repo(slug, viewer)
async def _branch_view_payload(rfc, slug: str, branch: str, viewer) -> dict[str, Any]:
if not _can_read_branch(slug, branch, viewer):
raise HTTPException(403, "Branch is private")
@@ -1087,12 +1130,48 @@ def make_router(
"capabilities": capabilities,
}
@router.get("/api/rfcs/{slug}/branches/{branch:path}")
async def get_branch_view(slug: str, branch: str, request: Request) -> dict[str, Any]:
viewer = auth.current_user(request)
rfc = _require_rfc_with_repo(slug, viewer)
return await _branch_view_payload(rfc, slug, branch, viewer)
@router.get("/api/projects/{project_id}/collections/{collection_id}/rfcs/{slug}/branches/{branch:path}")
async def get_branch_view_scoped(
project_id: str, collection_id: str, slug: str, branch: str, request: Request,
) -> dict[str, Any]:
# §22/G-15: collection-scoped branch-body read (the canonical-body GET
# RFCView renders). Disambiguates a slug across collections (G-5) and
# reads the entry's own content repo / subfolder.
viewer = auth.current_user(request)
_check_collection_in_project(project_id, collection_id)
rfc = _require_rfc_with_repo(slug, viewer, collection_id=collection_id)
return await _branch_view_payload(rfc, slug, branch, viewer)
# ------------------------------------------------------------------
# Permission + state helpers (closures, share `config` etc.)
# ------------------------------------------------------------------
def _require_rfc(slug: str, viewer):
row = db.conn().execute("SELECT * FROM cached_rfcs WHERE slug = ?", (slug,)).fetchone()
def _check_collection_in_project(project_id: str, collection_id: str) -> None:
"""§22/G-15: a collection-scoped route 404s when the collection isn't in
the named project (matches api_metadata's guard)."""
if collections_mod.project_of_collection(collection_id) != project_id:
raise HTTPException(404, "Collection not in project")
def _require_rfc(slug: str, viewer, collection_id: str | None = None):
# §22/G-15: when a collection_id is supplied (the collection-scoped
# body-read routes), scope the lookup to that collection so a slug that
# exists in two collections resolves unambiguously (G-5); otherwise the
# legacy slug-only lookup picks the entry by slug alone.
if collection_id is not None:
row = db.conn().execute(
"SELECT *, (SELECT c.project_id FROM collections c WHERE c.id = cached_rfcs.collection_id) AS project_id "
"FROM cached_rfcs WHERE slug = ? AND collection_id = ?",
(slug, collection_id)).fetchone()
else:
row = db.conn().execute(
"SELECT *, (SELECT c.project_id FROM collections c WHERE c.id = cached_rfcs.collection_id) AS project_id "
"FROM cached_rfcs WHERE slug = ?", (slug,)).fetchone()
if row is None:
raise HTTPException(404, "RFC not found")
# §22.5 visibility gate (subtractive, §22.7): a gated project's entries
@@ -1100,13 +1179,13 @@ def make_router(
auth.require_project_readable(viewer, row["project_id"])
return row
def _require_rfc_with_repo(slug: str, viewer):
def _require_rfc_with_repo(slug: str, viewer, collection_id: str | None = None):
"""Used by every branch-scoped endpoint. Under the meta-only
topology (§1) the meta repo is the implicit target for every
entry — super-draft and active alike — so there is no per-RFC
repo check. The name is retained for call-site stability; a
withdrawn entry is still rejected."""
row = _require_rfc(slug, viewer)
row = _require_rfc(slug, viewer, collection_id)
if row["state"] == "withdrawn":
raise HTTPException(409, "RFC is withdrawn")
return row
@@ -1152,39 +1231,32 @@ def make_router(
return _is_meta_branch_name(branch)
def _repo_for(rfc, branch: str = "main") -> tuple[str, str]:
# §22/G-15: a meta-resident entry's repo is its COLLECTION's project
# content_repo (collection → project → content_repo), not the deployment
# default — so an entry in a non-default project reads/writes its own
# repo. `entry_location` falls back to the default repo for a legacy /
# unknown collection, preserving the single-corpus behaviour.
if _is_meta_target(rfc, branch):
return config.gitea_org, (projects_mod.default_content_repo(config) or "")
org, repo, _ = projects_mod.entry_location(config, rfc["collection_id"], rfc["slug"])
return org, repo
owner, repo = rfc["repo"].split("/", 1)
return owner, repo
def _file_path_for(rfc, branch: str = "main") -> str:
# §22/G-15: path is the collection's `<subfolder>/rfcs/<slug>.md`
# (repo root `rfcs/<slug>.md` for a default collection).
if _is_meta_target(rfc, branch):
return f"rfcs/{rfc['slug']}.md"
_, _, path = projects_mod.entry_location(config, rfc["collection_id"], rfc["slug"])
return path
return RFC_FILE_PATH
def _extract_body(rfc, file_contents: str, branch: str = "main") -> str:
"""For super-draft entries (and active-RFC pre-graduation reads
per §9.8) the file on disk is the full frontmatter+body envelope;
the editable body is entry.body. For active RFCs reading their
per-RFC repo the file is just RFC.md and the whole thing is body."""
if not _is_meta_target(rfc, branch):
return file_contents
try:
entry = entry_mod.parse(file_contents)
except Exception:
return file_contents
return entry.body
return _extract_body_pure(
rfc, file_contents, branch, is_meta=_is_meta_target(rfc, branch))
def _wrap_body(rfc, prior_contents: str, new_body: str, branch: str = "main") -> str:
"""Inverse of _extract_body: re-wrap a new body into the entry
envelope, preserving the prior frontmatter exactly."""
if not _is_meta_target(rfc, branch):
return new_body
entry = entry_mod.parse(prior_contents)
# Ensure exactly one trailing newline so the serializer's
# round-trip is stable.
entry.body = new_body if new_body.endswith("\n") else new_body + "\n"
return entry_mod.serialize(entry)
return _wrap_body_pure(
rfc, prior_contents, new_body, branch, is_meta=_is_meta_target(rfc, branch))
async def _refresh_cache_for(rfc) -> None:
if _is_meta_resident(rfc):
@@ -1264,11 +1336,11 @@ def make_router(
return row["on_behalf_of"] if row else None
def _can_read_branch(slug: str, branch: str, viewer) -> bool:
# §22.5 visibility gate first (subtractive, §22.7): in a gated project
# nothing — not even main or a read_public branch — is readable by a
# non-member.
pid = auth.project_of_rfc(slug)
if not auth.can_read_project(viewer, pid):
# §22.5 visibility gate first (subtractive, §B.2): in a hidden/gated
# collection nothing — not even main or a read_public branch — is
# readable by a non-scope-role viewer.
cid = auth.collection_of_rfc(slug)
if not auth.can_read_collection(viewer, cid):
return False
if branch == "main":
return True
@@ -1277,7 +1349,7 @@ def make_router(
return True
if viewer is None:
return False
if auth.is_project_superuser(viewer, pid):
if auth.is_collection_superuser(viewer, cid):
return True
creator = _branch_creator(slug, branch)
if creator and viewer.gitea_login == creator:
@@ -1310,12 +1382,12 @@ def make_router(
# legacy `repo:` is set (nothing, after the RFC-0001 fold-back).
if rfc["state"] == "active" and rfc["repo"] and _is_meta_branch_name(branch):
return False
pid = auth.project_of_rfc(slug)
cid = auth.collection_of_rfc(slug)
# §22.5 visibility gate (subtractive): no contribute in an unreadable
# project.
if not auth.can_read_project(viewer, pid):
# collection.
if not auth.can_read_collection(viewer, cid):
return False
if auth.is_project_superuser(viewer, pid):
if auth.is_collection_superuser(viewer, cid):
return True
owners = json.loads(rfc["owners_json"] or "[]")
arbiters = json.loads(rfc["arbiters_json"] or "[]")
@@ -1326,10 +1398,10 @@ def make_router(
return True
vis = _branch_vis(slug, branch)
if vis["contribute_mode"] == "any-contributor":
# "any contributor" means anyone with project-level write standing
# (§22.6/§22.7) — the implicit-public baseline on a public project,
# or an explicit project_contributor/admin elsewhere.
return auth.can_contribute_in_project(viewer, pid)
# "any contributor" means anyone with collection-level write standing
# (§B.2) — the grandfathered baseline on the public default
# collection, or an explicit scope grant reaching the collection.
return auth.can_contribute_in_collection(viewer, cid)
if vis["contribute_mode"] == "specific":
row = db.conn().execute(
"""
@@ -1352,7 +1424,7 @@ def make_router(
def _require_branch_owner(rfc, viewer, creator: str | None) -> None:
# §22.6: a project_admin is the per-RFC owner/arbiter authority lifted
# to project scope, so it (and a deployment owner/admin) clears here.
if auth.is_project_superuser(viewer, rfc["project_id"]):
if auth.is_collection_superuser(viewer, rfc["collection_id"]):
return
owners = json.loads(rfc["owners_json"] or "[]")
arbiters = json.loads(rfc["arbiters_json"] or "[]")
@@ -1368,7 +1440,7 @@ def make_router(
has no owners, so the set collapses to the superuser tier only —
sensible because admin oversight is the only path to canonicalizing
edits on an unclaimed entry."""
if auth.is_project_superuser(viewer, rfc["project_id"]):
if auth.is_collection_superuser(viewer, rfc["collection_id"]):
return True
owners = json.loads(rfc["owners_json"] or "[]")
arbiters = json.loads(rfc["arbiters_json"] or "[]")
@@ -1381,7 +1453,7 @@ def make_router(
"can_read": _can_read_branch(slug, branch, viewer),
"can_contribute": _can_contribute(rfc, slug, branch, viewer) if viewer else False,
"can_change_branch_settings": viewer is not None and (
auth.is_project_superuser(viewer, rfc["project_id"])
auth.is_collection_superuser(viewer, rfc["collection_id"])
or (creator is not None and viewer.gitea_login == creator)
or viewer.gitea_login in (owners + arbiters)
),
@@ -1427,7 +1499,7 @@ def make_router(
def _can_resolve_thread(rfc, thread, creator: str | None, viewer) -> bool:
if viewer is None:
return False
if auth.is_project_superuser(viewer, rfc["project_id"]):
if auth.is_collection_superuser(viewer, rfc["collection_id"]):
return True
owners = json.loads(rfc["owners_json"] or "[]")
arbiters = json.loads(rfc["arbiters_json"] or "[]")
+182
View File
@@ -0,0 +1,182 @@
"""§22 S2 — collection directory + create-collection.
GET /api/projects/:id/collections — list the project's visible collections.
GET /api/projects/:id/collections/:cid — one collection's settings.
POST /api/projects/:id/collections — create a collection. Authorized by a
deployment owner/admin (S2; scoped
{owner, contributor} roles at the
collection axis land in S3). The bot
commits a `.collection.yaml` to the
content repo, then the registry mirror
upserts the collections row — §22.2
keeps the registry the source of truth.
"""
from __future__ import annotations
import re
from typing import Any
import yaml
from fastapi import APIRouter, HTTPException, Request
from pydantic import BaseModel
from . import (
auth,
collections as collections_mod,
projects as projects_mod,
registry as registry_mod,
)
from .bot import Bot
from .config import Config
from .gitea import Gitea, GiteaError
_SLUG_RE = re.compile(r"^[a-z0-9](?:[a-z0-9-]*[a-z0-9])?$")
class CreateCollectionBody(BaseModel):
collection_id: str
type: str
name: str | None = None
visibility: str | None = None
initial_state: str | None = None
def _project_viewer_caps(viewer: Any, project_id: str) -> dict[str, Any]:
"""§22 S4: the viewer's project-grain capabilities for role-aware UI — may
they create a collection, may they manage membership (invite), and their
project role. `role` maps the §22.6 legacy strings back to the unified
`{owner, contributor}` vocabulary the frontend speaks."""
legacy = auth.project_member_role(viewer, project_id)
role = None
if viewer is not None and viewer.role in ("owner", "admin"):
role = "owner"
elif legacy == "project_admin":
role = "owner"
elif legacy == "project_contributor":
role = "contributor"
return {
"can_create_collection": auth.can_create_collection(viewer, project_id),
"can_invite": auth.can_invite_at_project(viewer, project_id),
# §22.8: a signed-in, granted account with no role at the project may ask
# to join it (the request-to-join affordance). Owners/members and
# not-yet-granted accounts don't see it.
"can_request_join": (
viewer is not None
and viewer.permission_state == "granted"
and auth.effective_role_at_scope(viewer, "project", project_id) is None
),
"role": role,
}
def make_router(config: Config, gitea: Gitea, bot: Bot) -> APIRouter:
router = APIRouter()
@router.get("/api/projects/{project_id}/collections")
async def list_cols(project_id: str, request: Request) -> dict[str, Any]:
viewer = auth.current_user(request)
# §22.5 read gate: a gated project 404s a non-member.
auth.require_project_readable(viewer, project_id)
# §22.5 (S3): the directory is viewer-aware — a hidden/gated collection
# is listed only for a scope-role holder who can read it; `unlisted` is
# omitted from enumeration for everyone (link-only).
items = [
c
for c in collections_mod.list_collections(project_id, include_unlisted=True)
if c["visibility"] != "unlisted" and auth.can_read_collection(viewer, c["id"])
]
# §22 S4: surface the viewer's project-level capabilities so the
# directory can render role-aware affordances (the create-first-
# collection CTA, the invite control) without a second round-trip.
return {"items": items, "viewer": _project_viewer_caps(viewer, project_id)}
@router.get("/api/projects/{project_id}/collections/{collection_id}")
async def get_col(project_id: str, collection_id: str, request: Request) -> dict[str, Any]:
viewer = auth.current_user(request)
auth.require_project_readable(viewer, project_id)
col = collections_mod.get_collection(collection_id)
if col is None or col["project_id"] != project_id:
raise HTTPException(404, "Not found")
# §22.5 (S3): a hidden/gated collection 404s a non-scope-role viewer.
auth.require_collection_readable(viewer, collection_id)
# §22 S4: the viewer's collection-level capabilities drive the
# propose-first empty state and the collection invite control.
col = dict(col)
col["viewer"] = {
"can_contribute": auth.can_contribute_in_collection(viewer, collection_id),
"can_invite": auth.can_invite_at_collection(viewer, collection_id),
# §22.8: a signed-in, granted account with no role reaching this
# collection may ask to join it.
"can_request_join": (
viewer is not None
and viewer.permission_state == "granted"
and auth.effective_scope_role(viewer, collection_id) is None
),
"role": auth.effective_scope_role(viewer, collection_id),
}
return col
@router.post("/api/projects/{project_id}/collections")
async def create_col(
project_id: str, body: CreateCollectionBody, request: Request
) -> dict[str, Any]:
# §B.1 (S3) authority: a deployment owner/admin or a project/global-scope
# grant holder (Owner or RFC Contributor) may create a collection. The
# read gate runs first so a gated project 404s a non-member.
user = auth.require_contributor(request)
auth.require_project_readable(user, project_id)
if not auth.can_create_collection(user, project_id):
raise HTTPException(403, "You may not create collections in this project")
cid = body.collection_id.strip().lower()
if not _SLUG_RE.match(cid) or cid == "default":
raise HTTPException(422, "collection id must be a slug and not 'default'")
if body.type not in registry_mod.VALID_TYPES:
raise HTTPException(422, f"invalid type {body.type!r}")
if body.visibility is not None:
if body.visibility not in registry_mod.VALID_VISIBILITY:
raise HTTPException(422, f"invalid visibility {body.visibility!r}")
# §22.5 (S3) strictness: a collection may be set only as strict or
# stricter than its project — never more public.
pvis = auth.project_visibility(project_id)
if auth.visibility_rank(body.visibility) < auth.visibility_rank(pvis):
raise HTTPException(
422,
f"collection visibility {body.visibility!r} is looser than "
f"the project's {pvis!r}; a collection may only narrow it",
)
if body.initial_state is not None and body.initial_state not in registry_mod.VALID_INITIAL_STATE:
raise HTTPException(422, f"invalid initial_state {body.initial_state!r}")
if collections_mod.get_collection(cid) is not None:
raise HTTPException(409, f"collection `{cid}` already exists")
content_repo = projects_mod.content_repo(project_id)
if not content_repo:
raise HTTPException(409, "project has no content repo")
manifest: dict[str, Any] = {"type": body.type}
if body.name:
manifest["name"] = body.name
if body.visibility:
manifest["visibility"] = body.visibility
if body.initial_state:
manifest["initial_state"] = body.initial_state
manifest_yaml = yaml.safe_dump(manifest, sort_keys=False)
try:
await bot.create_collection(
user.as_actor(),
org=config.gitea_org,
content_repo=content_repo,
collection_id=cid,
manifest_yaml=manifest_yaml,
)
except GiteaError as e:
raise HTTPException(502, f"Gitea: {e.detail}")
# §22.2: re-read the registry so the new manifest becomes a row.
await registry_mod.refresh_registry(config, gitea)
col = collections_mod.get_collection(cid)
if col is None:
raise HTTPException(500, "collection committed but not mirrored")
return col
return router
+3 -3
View File
@@ -62,7 +62,7 @@ def _require_super_draft(slug: str, viewer):
visibility gate is subtractive: a gated project's entries 404 to
non-members (§22.7)."""
row = db.conn().execute(
"SELECT slug, title, state, owners_json, proposed_by, project_id FROM cached_rfcs WHERE slug = ?",
"SELECT slug, title, state, owners_json, proposed_by, (SELECT c.project_id FROM collections c WHERE c.id = cached_rfcs.collection_id) AS project_id FROM cached_rfcs WHERE slug = ?",
(slug,),
).fetchone()
if row is None:
@@ -91,7 +91,7 @@ def _viewer_relationship(viewer, slug: str) -> str | None:
"""Why this viewer can't *request* to contribute — or None if they can.
Owners/admins already have the RFC; existing collaborators are already
in. Both get a clear 409 rather than a useless self-request."""
if auth.is_rfc_owner(viewer, slug) or auth.is_project_superuser(viewer, auth.project_of_rfc(slug)):
if auth.is_rfc_owner(viewer, slug) or auth.is_collection_superuser(viewer, auth.collection_of_rfc(slug)):
return "You already own or administer this RFC."
if auth.is_rfc_collaborator(viewer, slug):
return "You're already a collaborator on this RFC."
@@ -108,7 +108,7 @@ def make_router() -> APIRouter:
@router.get("/api/rfcs/{slug}/contribution-target")
async def contribution_target(slug: str, request: Request) -> dict[str, Any]:
row = db.conn().execute(
"SELECT slug, title, state, owners_json, proposed_by, project_id FROM cached_rfcs WHERE slug = ?",
"SELECT slug, title, state, owners_json, proposed_by, (SELECT c.project_id FROM collections c WHERE c.id = cached_rfcs.collection_id) AS project_id FROM cached_rfcs WHERE slug = ?",
(slug,),
).fetchone()
if row is None:
+174 -17
View File
@@ -1,28 +1,60 @@
"""§22.9 runtime deployment/project config (replaces VITE_APP_NAME) + §22.10
old-URL 308 redirects.
old-URL 308 redirects + §22 S5 in-app create-project.
GET /api/deployment — the deployment name/tagline + the projects the caller can
GET /api/deployment — the deployment name/tagline + the projects the caller can
see (§22.5: gated filtered by membership, unlisted omitted from enumeration),
plus the corpus-served `default_project_id` the M3-frontend guard keys on.
GET /api/projects/:id — one project's runtime config + optional theme overlay,
plus the corpus-served `default_project_id` the M3-frontend guard keys on, the
`viewer` capability block (S5: `can_create_project`), and
`default_project_readable` (whether the N=1 redirect target is reachable by this
viewer — drives the deployment-directory empty state vs the land-in-corpus
redirect).
POST /api/projects — §22 S5 create-project (global-Owner only). The bot
provisions a Gitea content repo and commits a project entry to `projects.yaml`;
the registry mirror then upserts the `projects` + default `collections` rows
(§22.2 keeps the registry the source of truth).
GET /api/projects/:id — one project's runtime config + optional theme overlay,
gated behind the §22.5 read gate (404 for a non-member of a gated project).
GET /rfc/{slug}, /proposals/{n} — §22.10 server-side 308s onto the new
GET /rfc/{slug}, /proposals/{n} — §22.10 server-side 308s onto the new
`/p/<default>/…` routes (the SPA no longer owns these paths; nginx proxies them
to the backend instead of serving index.html).
"""
from __future__ import annotations
import json
import re
from typing import Any
import yaml
from fastapi import APIRouter, HTTPException, Request
from fastapi.responses import RedirectResponse
from pydantic import BaseModel
from . import auth, db, projects as projects_mod
from . import (
auth,
collections as collections_mod,
db,
projects as projects_mod,
registry as registry_mod,
)
from .bot import Bot
from .config import Config
from .gitea import Gitea, GiteaError
# A project id is a slug (the §22.2 registry key + the `/p/<id>/` path segment).
_SLUG_RE = re.compile(r"^[a-z0-9](?:[a-z0-9-]*[a-z0-9])?$")
# A Gitea repo name: alphanumeric start, then alphanumerics / `-` / `_` / `.`.
_REPO_RE = re.compile(r"^[A-Za-z0-9][A-Za-z0-9._-]*$")
def make_router(config: Config) -> APIRouter:
class CreateProjectBody(BaseModel):
project_id: str
name: str
type: str
visibility: str | None = None
content_repo: str | None = None
def make_router(config: Config, gitea: Gitea, bot: Bot) -> APIRouter:
router = APIRouter()
@router.get("/api/deployment")
@@ -33,15 +65,39 @@ def make_router(config: Config) -> APIRouter:
).fetchone()
# §22.5: enumerate only public + (member-)gated; unlisted is never listed.
visible = set(auth.visible_project_ids(viewer))
# §22 three-tier: `type` is a per-corpus field on the (default) collection
# now; surface the default collection's type for each project.
rows = db.conn().execute(
"SELECT id, name, type, visibility FROM projects "
"SELECT id, name, visibility FROM projects "
"WHERE visibility != 'unlisted' ORDER BY name"
).fetchall()
projects = [
{"id": r["id"], "name": r["name"], "type": r["type"], "visibility": r["visibility"]}
{
"id": r["id"],
"name": r["name"],
"type": collections_mod.collection_type(
collections_mod.default_collection_id(r["id"])
),
"entry_noun": collections_mod.entry_noun(
collections_mod.collection_type(
collections_mod.default_collection_id(r["id"])
)
),
"visibility": r["visibility"],
}
for r in rows
if r["id"] in visible
]
# §22 S5: the N=1 land-in-corpus redirect targets the default project, but
# only when this viewer can actually read it. A `gated` default (C3.2) or
# an absent default (C3.1, a deployment with no projects) is *not* a valid
# redirect target — the frontend then falls through to the deployment
# directory's role-aware empty state instead of bouncing into a 404.
default_id = projects_mod.resolved_default_id(config)
default_exists = db.conn().execute(
"SELECT 1 FROM projects WHERE id = ?", (default_id,)
).fetchone() is not None
default_readable = default_exists and auth.can_read_project(viewer, default_id)
return {
"name": (dep["name"] if dep else "") or "",
"tagline": (dep["tagline"] if dep else "") or "",
@@ -49,8 +105,100 @@ def make_router(config: Config) -> APIRouter:
# serves the corpus for (the default, until Plan B serves per
# project). The frontend renders corpus routes only for this id and
# shows a "content not yet served" placeholder for any other.
"default_project_id": projects_mod.resolved_default_id(config),
"default_project_id": default_id,
"default_project_readable": default_readable,
"projects": projects,
# §22 S5 (C3.1/C3.2): role-aware deployment-directory affordances.
"viewer": {"can_create_project": auth.can_create_project(viewer)},
}
@router.post("/api/projects")
async def create_project(body: CreateProjectBody, request: Request) -> dict[str, Any]:
# §22 S5 / §A.2 / §B.1: "+ New project" is a global-Owner action.
user = auth.require_contributor(request)
if not auth.can_create_project(user):
raise HTTPException(403, "Only a global Owner may create projects")
pid = body.project_id.strip().lower()
if not _SLUG_RE.match(pid) or pid == "default":
raise HTTPException(422, "project id must be a slug and not 'default'")
name = (body.name or "").strip()
if not name:
raise HTTPException(422, "project name is required")
if body.type not in registry_mod.VALID_TYPES:
raise HTTPException(422, f"invalid type {body.type!r}")
# A project is created visible by default — the point of standing one up
# is for it to be seen; an Owner narrows it afterwards (or picks gated).
visibility = (body.visibility or "public").strip()
if visibility not in registry_mod.VALID_VISIBILITY:
raise HTTPException(422, f"invalid visibility {visibility!r}")
if db.conn().execute("SELECT 1 FROM projects WHERE id = ?", (pid,)).fetchone():
raise HTTPException(409, f"project `{pid}` already exists")
content_repo = (body.content_repo or f"{pid}-content").strip()
if not _REPO_RE.match(content_repo):
raise HTTPException(422, f"invalid content repo name {content_repo!r}")
if await gitea.get_repo(config.gitea_org, content_repo) is not None:
raise HTTPException(409, f"repo `{content_repo}` already exists")
# Read the current registry, append the project, recompose. Reads live in
# gitea.py and may be called anywhere; the bot owns the write back.
read = await gitea.read_file(
config.gitea_org, config.registry_repo, "projects.yaml", ref="main"
)
if read is None:
raise HTTPException(409, "registry projects.yaml not found")
text, sha = read
try:
doc = yaml.safe_load(text) or {}
except yaml.YAMLError as e:
raise HTTPException(500, f"registry projects.yaml is not valid YAML: {e}")
if not isinstance(doc, dict):
raise HTTPException(500, "registry projects.yaml is malformed")
projects = doc.get("projects")
if not isinstance(projects, list):
projects = []
if any(isinstance(p, dict) and str(p.get("id") or "") == pid for p in projects):
raise HTTPException(409, f"project `{pid}` already in the registry")
projects.append(
{
"id": pid,
"name": name,
"type": body.type,
"content_repo": content_repo,
"visibility": visibility,
}
)
doc["projects"] = projects
new_text = yaml.safe_dump(doc, sort_keys=False)
readme_text = f"# {name}\n\nContent repository for project `{pid}`.\n"
try:
await bot.create_project(
user.as_actor(),
org=config.gitea_org,
registry_repo=config.registry_repo,
content_repo=content_repo,
project_id=pid,
projects_yaml_new=new_text,
projects_yaml_sha=sha,
readme_text=readme_text,
)
except GiteaError as e:
raise HTTPException(502, f"Gitea: {e.detail}")
# §22.2: re-read the registry so the new entry becomes projects +
# default-collection rows.
await registry_mod.refresh_registry(config, gitea)
row = db.conn().execute(
"SELECT id, name, visibility FROM projects WHERE id = ?", (pid,)
).fetchone()
if row is None:
raise HTTPException(500, "project committed but not mirrored")
cid = collections_mod.default_collection_id(pid)
return {
"id": row["id"],
"name": row["name"],
"visibility": row["visibility"],
"type": collections_mod.collection_type(cid),
}
@router.get("/api/projects/{project_id}")
@@ -60,8 +208,7 @@ def make_router(config: Config) -> APIRouter:
# an unknown id). unlisted is readable by direct id.
auth.require_project_readable(viewer, project_id)
row = db.conn().execute(
"SELECT id, name, type, visibility, initial_state, config_json "
"FROM projects WHERE id = ?",
"SELECT id, name, visibility, config_json FROM projects WHERE id = ?",
(project_id,),
).fetchone()
if row is None:
@@ -71,13 +218,17 @@ def make_router(config: Config) -> APIRouter:
except (ValueError, TypeError):
cfg = {}
dep = db.conn().execute("SELECT tagline FROM deployment WHERE id = 1").fetchone()
# §22 three-tier: type + initial_state moved down to the (default)
# collection in migration 029.
cid = collections_mod.default_collection_id(row["id"])
return {
"id": row["id"],
"name": row["name"],
"tagline": (dep["tagline"] if dep else "") or "",
"type": row["type"],
"type": collections_mod.collection_type(cid),
"entry_noun": collections_mod.entry_noun(collections_mod.collection_type(cid)),
"visibility": row["visibility"],
"initial_state": row["initial_state"],
"initial_state": collections_mod.collection_initial_state(cid),
"theme": cfg.get("theme") or {},
}
@@ -86,21 +237,27 @@ def make_router(config: Config) -> APIRouter:
# is permanent, so external "RFC-0001" links and bookmarks land correctly.
# nginx routes /rfc/ and /proposals/ to the backend so these are reached
# before the SPA's index.html fallback.
# §22 three-tier (S1): the canonical entry route now carries the collection
# segment /p/<project>/c/<collection>/…. The legacy roots redirect through
# the default project's default collection.
@router.get("/rfc/{slug}")
async def redirect_old_rfc(slug: str) -> RedirectResponse:
default_id = projects_mod.resolved_default_id(config)
return RedirectResponse(url=f"/p/{default_id}/e/{slug}", status_code=308)
cid = collections_mod.default_collection_id(default_id)
return RedirectResponse(url=f"/p/{default_id}/c/{cid}/e/{slug}", status_code=308)
@router.get("/rfc/{slug}/pr/{pr_number}")
async def redirect_old_rfc_pr(slug: str, pr_number: int) -> RedirectResponse:
default_id = projects_mod.resolved_default_id(config)
cid = collections_mod.default_collection_id(default_id)
return RedirectResponse(
url=f"/p/{default_id}/e/{slug}/pr/{pr_number}", status_code=308
url=f"/p/{default_id}/c/{cid}/e/{slug}/pr/{pr_number}", status_code=308
)
@router.get("/proposals/{pr_number}")
async def redirect_old_proposal(pr_number: int) -> RedirectResponse:
default_id = projects_mod.resolved_default_id(config)
return RedirectResponse(url=f"/p/{default_id}/proposals/{pr_number}", status_code=308)
cid = collections_mod.default_collection_id(default_id)
return RedirectResponse(url=f"/p/{default_id}/c/{cid}/proposals/{pr_number}", status_code=308)
return router
+2 -2
View File
@@ -252,7 +252,7 @@ def _require_rfc_readable(slug: str, viewer):
entries refuse reads of every shape — same rule `_require_rfc_with_repo`
in `api_branches.py` follows."""
row = db.conn().execute(
"SELECT * FROM cached_rfcs WHERE slug = ?", (slug,)
"SELECT *, (SELECT c.project_id FROM collections c WHERE c.id = cached_rfcs.collection_id) AS project_id FROM cached_rfcs WHERE slug = ?", (slug,)
).fetchone()
if row is None:
raise HTTPException(404, "RFC not found")
@@ -312,7 +312,7 @@ def _ensure_discussion_thread(slug: str, viewer) -> int:
def _can_resolve(rfc, thread, viewer) -> bool:
if viewer is None:
return False
if auth.is_project_superuser(viewer, rfc["project_id"]):
if auth.is_collection_superuser(viewer, rfc["collection_id"]):
return True
owners = json.loads(rfc["owners_json"] or "[]")
arbiters = json.loads(rfc["arbiters_json"] or "[]")
+79 -72
View File
@@ -42,7 +42,7 @@ from fastapi import APIRouter, HTTPException, Request
from fastapi.responses import StreamingResponse
from pydantic import BaseModel, Field
from . import auth, cache, db, entry as entry_mod, projects as projects_mod
from . import auth, cache, db, entry as entry_mod, metadata as metadata_mod, projects as projects_mod
from .bot import Actor, Bot
from .config import Config
from .gitea import Gitea, GiteaError
@@ -218,7 +218,7 @@ def make_router(
can_merge = (
viewer is not None
and (
auth.is_project_superuser(viewer, rfc["project_id"])
auth.is_collection_superuser(viewer, rfc["collection_id"])
or viewer.gitea_login in owners
or viewer.gitea_login in arbiters
)
@@ -341,19 +341,18 @@ def make_router(
if _rfc_id_taken(rfc_id, excluding_slug=slug):
raise HTTPException(409, f"Integer ID {rfc_id} is already taken")
# Read the meta-repo entry once — we need the file's sha for the
# graduation PR's update_file call and the body to carry through
# Dual-read the meta-repo entry once (§22.4a sidecar-aware) — we need its
# git state for the graduation commit and the body to carry through
# unchanged (meta-only keeps the body in the entry, §13.3).
fetched = await gitea.read_file(
config.gitea_org, (projects_mod.default_content_repo(config) or ""), f"rfcs/{slug}.md", ref="main",
)
if fetched is None:
raise HTTPException(409, f"Meta entry rfcs/{slug}.md not found on main")
meta_text, meta_sha = fetched
try:
super_draft_entry = entry_mod.parse(meta_text)
except Exception as e:
raise HTTPException(500, f"Meta entry malformed: {e}")
# §22/G-15: resolve the repo+path from the entry's COLLECTION, not the
# deployment default, so an entry in a non-default project graduates in
# its own content repo / collection subfolder.
org, meta_repo, md_path = projects_mod.entry_location(
config, rfc["collection_id"], slug)
st = await metadata_mod.read_entry_from_git(gitea, org, meta_repo, md_path)
if st is None:
raise HTTPException(409, f"Meta entry {md_path} not found on main")
super_draft_entry = st.entry
arbiters = json.loads(rfc["arbiters_json"] or "[]") or owners[:1]
@@ -376,8 +375,12 @@ def make_router(
models=super_draft_entry.models,
funder=super_draft_entry.funder,
body=super_draft_entry.body,
# INV-7 (§22.4a): carry forward-compat / unknown frontmatter keys
# through graduation rather than dropping them on the rebuild.
extra=dict(super_draft_entry.extra),
)
graduated_contents = entry_mod.serialize(graduated_entry)
graduation_files = metadata_mod.write_entry_files(
md_path, graduated_entry, st)
state = _new_active(
slug, rfc_id=rfc_id, owners=owners, arbiters=arbiters,
@@ -397,8 +400,8 @@ def make_router(
coro = _orchestrate(
config=config, gitea=gitea, bot=bot,
actor=viewer.as_actor(), state=state,
graduated_contents=graduated_contents,
meta_file_sha=meta_sha,
meta_repo=meta_repo,
graduation_files=graduation_files,
)
if request.query_params.get("_sync") == "1":
await coro
@@ -478,26 +481,23 @@ def make_router(
if already:
raise HTTPException(409, f"A claim PR is already open: #{already['pr_number']}")
fetched = await gitea.read_file(
config.gitea_org, (projects_mod.default_content_repo(config) or ""), f"rfcs/{slug}.md", ref="main",
)
if fetched is None:
raise HTTPException(409, f"Meta entry rfcs/{slug}.md not found on main")
meta_text, meta_sha = fetched
try:
ent = entry_mod.parse(meta_text)
except Exception as e:
raise HTTPException(500, f"Meta entry malformed: {e}")
if viewer.gitea_login in ent.owners:
# §22/G-15: resolve repo+path from the entry's collection, not the default.
org, meta_repo, md_path = projects_mod.entry_location(
config, rfc["collection_id"], slug)
st = await metadata_mod.read_entry_from_git(gitea, org, meta_repo, md_path)
if st is None:
raise HTTPException(409, f"Meta entry {md_path} not found on main")
if viewer.gitea_login in st.entry.owners:
return {"ok": True, "noop": True}
ent.owners = ent.owners + [viewer.gitea_login]
new_contents = entry_mod.serialize(ent)
ent = metadata_mod.apply_values(
st.entry, {"owners": st.entry.owners + [viewer.gitea_login]})
files = metadata_mod.write_entry_files(md_path, ent, st)
try:
pr = await bot.open_claim_pr(
viewer.as_actor(),
org=config.gitea_org, meta_repo=(projects_mod.default_content_repo(config) or ""),
org=org, meta_repo=meta_repo,
slug=slug,
new_file_contents=new_contents, prior_sha=meta_sha,
files=files,
)
except GiteaError as e:
raise HTTPException(502, f"Gitea: {e.detail}")
@@ -524,11 +524,15 @@ def make_router(
403, "Only this RFC's owners or a site owner may retire it"
)
prior_state = rfc["state"]
entry, sha = await _read_meta_entry(slug)
entry.state = "retired"
# §22/G-15: resolve repo+path from the entry's collection, not the default.
org, meta_repo, md_path = projects_mod.entry_location(
config, rfc["collection_id"], slug)
st = await _read_meta_entry(org, meta_repo, md_path)
entry = metadata_mod.apply_values(st.entry, {"state": "retired"})
files = metadata_mod.write_entry_files(md_path, entry, st)
await _run_state_flip(
config=config, gitea=gitea, bot=bot, actor=viewer.as_actor(),
slug=slug, new_contents=entry_mod.serialize(entry), prior_sha=sha,
meta_repo=meta_repo, slug=slug, files=files,
verb="retire", target_state="retired",
)
_audit(
@@ -552,13 +556,17 @@ def make_router(
viewer = auth.require_user(request)
if viewer.role != "owner":
raise HTTPException(403, "Only a site owner may un-retire an RFC")
_require_retired(slug)
rfc = _require_retired(slug)
restored = _prior_state_before_retire(slug)
entry, sha = await _read_meta_entry(slug)
entry.state = restored
# §22/G-15: resolve repo+path from the entry's collection, not the default.
org, meta_repo, md_path = projects_mod.entry_location(
config, rfc["collection_id"], slug)
st = await _read_meta_entry(org, meta_repo, md_path)
entry = metadata_mod.apply_values(st.entry, {"state": restored})
files = metadata_mod.write_entry_files(md_path, entry, st)
await _run_state_flip(
config=config, gitea=gitea, bot=bot, actor=viewer.as_actor(),
slug=slug, new_contents=entry_mod.serialize(entry), prior_sha=sha,
meta_repo=meta_repo, slug=slug, files=files,
verb="unretire", target_state=restored,
)
_audit(
@@ -574,7 +582,7 @@ def make_router(
# -------------------------------------------------------------------
def _require_super_draft(slug: str, viewer):
row = db.conn().execute("SELECT * FROM cached_rfcs WHERE slug = ?", (slug,)).fetchone()
row = db.conn().execute("SELECT *, (SELECT c.project_id FROM collections c WHERE c.id = cached_rfcs.collection_id) AS project_id FROM cached_rfcs WHERE slug = ?", (slug,)).fetchone()
if row is None:
raise HTTPException(404, "RFC not found")
# §22.5 visibility gate (subtractive, §22.7): gated → 404 to non-members.
@@ -584,7 +592,7 @@ def make_router(
return row
def _require_retirable(slug: str):
row = db.conn().execute("SELECT * FROM cached_rfcs WHERE slug = ?", (slug,)).fetchone()
row = db.conn().execute("SELECT *, (SELECT c.project_id FROM collections c WHERE c.id = cached_rfcs.collection_id) AS project_id FROM cached_rfcs WHERE slug = ?", (slug,)).fetchone()
if row is None:
raise HTTPException(404, "RFC not found")
if row["state"] not in ("super-draft", "active"):
@@ -592,24 +600,22 @@ def make_router(
return row
def _require_retired(slug: str):
row = db.conn().execute("SELECT * FROM cached_rfcs WHERE slug = ?", (slug,)).fetchone()
row = db.conn().execute("SELECT *, (SELECT c.project_id FROM collections c WHERE c.id = cached_rfcs.collection_id) AS project_id FROM cached_rfcs WHERE slug = ?", (slug,)).fetchone()
if row is None:
raise HTTPException(404, "RFC not found")
if row["state"] != "retired":
raise HTTPException(409, f"RFC is {row['state']}, not retired")
return row
async def _read_meta_entry(slug: str) -> tuple[entry_mod.Entry, str]:
fetched = await gitea.read_file(
config.gitea_org, (projects_mod.default_content_repo(config) or ""), f"rfcs/{slug}.md", ref="main",
)
if fetched is None:
raise HTTPException(409, f"Meta entry rfcs/{slug}.md not found on main")
text, file_sha = fetched
try:
return entry_mod.parse(text), file_sha
except Exception as e:
raise HTTPException(500, f"Meta entry malformed: {e}")
async def _read_meta_entry(org: str, repo: str, md_path: str):
"""Dual-read an entry from meta-main → EntryGitState (sidecar-aware,
§22.4a). A migrated body-only `.md` reads cleanly; never raises on bad
metadata (INV-3). `org`/`repo`/`md_path` are collection-resolved by the
caller (§22/G-15) so a non-default project's entry reads its own repo."""
st = await metadata_mod.read_entry_from_git(gitea, org, repo, md_path)
if st is None:
raise HTTPException(409, f"Meta entry {md_path} not found on main")
return st
async def _refresh_catalog() -> None:
# Inline refresh so the catalog reflects the flip immediately; the
@@ -637,8 +643,8 @@ async def _orchestrate(
bot: Bot,
actor: Actor,
state: GraduationState,
graduated_contents: str,
meta_file_sha: str,
meta_repo: str,
graduation_files: list[dict],
) -> None:
"""Open the flip PR, then merge it. Two steps, no transaction:
@@ -656,10 +662,9 @@ async def _orchestrate(
try:
pr = await bot.open_graduation_pr(
actor,
org=config.gitea_org, meta_repo=(projects_mod.default_content_repo(config) or ""),
org=config.gitea_org, meta_repo=meta_repo,
slug=state.slug,
new_file_contents=graduated_contents,
prior_sha=meta_file_sha,
files=graduation_files,
rfc_id=state.rfc_id,
owners=state.owners,
)
@@ -676,14 +681,15 @@ async def _orchestrate(
try:
await bot.merge_graduation_pr(
actor,
org=config.gitea_org, meta_repo=(projects_mod.default_content_repo(config) or ""),
org=config.gitea_org, meta_repo=meta_repo,
pr_number=state.new_pr_number,
head_branch=state.graduation_branch or "",
slug=state.slug, rfc_id=state.rfc_id,
)
except GiteaError as e:
await _fail(state, "merge_pr", f"Gitea: {e.detail}")
await _cleanup_unmerged(config=config, bot=bot, actor=actor, state=state)
await _cleanup_unmerged(
config=config, bot=bot, actor=actor, state=state, meta_repo=meta_repo)
await _finish_failed(state, failed_at="merge_pr", on_behalf_of=actor.gitea_login)
return
await _done(state, "merge_pr", f"PR #{state.new_pr_number} merged")
@@ -726,7 +732,7 @@ async def _orchestrate(
async def _cleanup_unmerged(
*, config: Config, bot: Bot, actor: Actor, state: GraduationState,
*, config: Config, bot: Bot, actor: Actor, state: GraduationState, meta_repo: str,
) -> None:
"""A merge failure leaves the flip PR open on its `graduate-<slug>-<hex>`
branch. Close the PR and delete the branch so failed attempts don't
@@ -738,7 +744,7 @@ async def _cleanup_unmerged(
try:
await bot.close_graduation_pr(
actor,
org=config.gitea_org, meta_repo=(projects_mod.default_content_repo(config) or ""),
org=config.gitea_org, meta_repo=meta_repo,
pr_number=state.new_pr_number,
head_branch=state.graduation_branch or "",
slug=state.slug, reason="graduation merge failed",
@@ -751,7 +757,7 @@ async def _cleanup_unmerged(
await bot.delete_branch(
actor,
owner=config.gitea_org,
repo=(projects_mod.default_content_repo(config) or ""),
repo=meta_repo,
branch=branch_name,
slug=state.slug,
action_kind="delete_post_merge_branch",
@@ -795,7 +801,7 @@ def _can_graduate(rfc, viewer) -> bool:
if viewer is None:
return False
# §6.1 admin/owner or §22.6 project_admin OR §6.3 RFC owners/arbiters.
if auth.is_project_superuser(viewer, rfc["project_id"]):
if auth.is_collection_superuser(viewer, rfc["collection_id"]):
return True
owners = json.loads(rfc["owners_json"] or "[]")
arbiters = json.loads(rfc["arbiters_json"] or "[]")
@@ -846,13 +852,14 @@ async def _run_state_flip(
gitea: Gitea,
bot: Bot,
actor: Actor,
meta_repo: str,
slug: str,
new_contents: str,
prior_sha: str,
files: list[dict],
verb: str,
target_state: str,
) -> None:
"""§13.7: open + merge a retire / un-retire frontmatter flip PR. Runs
"""§13.7: open + merge a retire / un-retire state-flip PR. The flip is
written to the entry's metadata sidecar (§22.4a) via `files` ops. Runs
inline (no SSE — the flip is a single quick state change, unlike the
multi-step graduation that streams progress). On an open failure
nothing was created; on a merge failure the half-open PR/branch is
@@ -861,8 +868,8 @@ async def _run_state_flip(
try:
pr = await bot.open_retire_flip_pr(
actor,
org=config.gitea_org, meta_repo=(projects_mod.default_content_repo(config) or ""),
slug=slug, new_file_contents=new_contents, prior_sha=prior_sha,
org=config.gitea_org, meta_repo=meta_repo,
slug=slug, files=files,
verb=verb, target_state=target_state,
)
except GiteaError as e:
@@ -872,7 +879,7 @@ async def _run_state_flip(
try:
await bot.merge_retire_flip_pr(
actor,
org=config.gitea_org, meta_repo=(projects_mod.default_content_repo(config) or ""),
org=config.gitea_org, meta_repo=meta_repo,
pr_number=pr_number, head_branch=head_branch,
slug=slug, verb=verb,
)
@@ -881,12 +888,12 @@ async def _run_state_flip(
# accumulate (mirrors graduation's `_cleanup_unmerged`).
try:
await bot.close_graduation_pr(
actor, org=config.gitea_org, meta_repo=(projects_mod.default_content_repo(config) or ""),
actor, org=config.gitea_org, meta_repo=meta_repo,
pr_number=pr_number, head_branch=head_branch,
slug=slug, reason=f"{verb} merge failed",
)
await bot.delete_branch(
actor, owner=config.gitea_org, repo=(projects_mod.default_content_repo(config) or ""),
actor, owner=config.gitea_org, repo=meta_repo,
branch=head_branch, slug=slug,
action_kind="delete_post_merge_branch",
reason=f"{verb} merge failed",
+1 -1
View File
@@ -380,7 +380,7 @@ def _require_rfc(slug: str, viewer):
visibility gate is subtractive: a gated project's entries 404 to
non-members (§22.7)."""
row = db.conn().execute(
"SELECT slug, title, state, project_id FROM cached_rfcs WHERE slug = ?", (slug,),
"SELECT slug, title, state, (SELECT c.project_id FROM collections c WHERE c.id = cached_rfcs.collection_id) AS project_id FROM cached_rfcs WHERE slug = ?", (slug,),
).fetchone()
if row is None:
raise HTTPException(404, "RFC not found")
+311
View File
@@ -0,0 +1,311 @@
"""§22.8 S6 — request-to-join a scope + the cross-collection inbox.
A gated project or collection is invisible to non-members (§22.5), so joining is
by invite (an Owner grants directly — `api_memberships.py`) *or* by request: a
user who knows a scope exists asks to join it, naming a desired role. This module
is the request side:
* ``GET /api/scopes/{scope_type}/{scope_id}/join-target`` — what the join
form needs (the scope's name, the viewer's eligibility + whether they already
have a pending ask + their current role).
* ``POST /api/scopes/{scope_type}/{scope_id}/join-requests`` — submit the ask
(desired role + optional message); lands a row + one §15 notification per
Owner across the scope's subtree (the cross-collection inbox, §22.11).
* ``POST /api/scopes/{scope_type}/{scope_id}/join-requests/{id}/accept`` —
Owner: accept, which writes the `memberships` row via ``memberships.grant``
(the §22.8 "accepting writes the membership row"), then notifies the requester.
* ``POST /api/scopes/{scope_type}/{scope_id}/join-requests/{id}/decline`` —
Owner: decline; the request closes and the requester is notified.
Mirrors ``api_contributions.py`` (the per-RFC contribute-request flow) but at the
scope grain: the target is a ``(scope_type, scope_id)`` pair drawn from the
``memberships`` scope vocabulary (minus ``global`` — a deployment isn't a thing
one discovers and joins), and accept grants a scope role rather than minting an
RFC invitation.
The request POST deliberately does **not** require the scope be *readable*: the
whole point of request-to-join is to ask into a *gated* scope you were told about
but cannot see (§22.8). It is gated only on "you're signed in, granted, and not
already a member". Accept/decline are gated on Owner reach over the scope
(``auth.can_invite_at_project`` / ``auth.can_invite_at_collection``).
"""
from __future__ import annotations
import sqlite3
from typing import Any
from fastapi import APIRouter, HTTPException, Request
from pydantic import BaseModel, Field
from . import (
auth,
collections as collections_mod,
db,
memberships as memberships_mod,
notify,
)
_MESSAGE_MAX = 4000
class JoinRequestBody(BaseModel):
role: str
message: str | None = Field(default=None, max_length=_MESSAGE_MAX)
class DecideBody(BaseModel):
# On accept, the Owner may grant a role narrower than the one requested; a
# missing value grants exactly the requested role.
role: str | None = None
def _project_name(project_id: str) -> str | None:
row = db.conn().execute(
"SELECT name FROM projects WHERE id = ?", (project_id,)
).fetchone()
return row["name"] if row and row["name"] else None
def _resolve_scope(scope_type: str, scope_id: str) -> dict[str, Any]:
"""Resolve a `(scope_type, scope_id)` target to its display facts, or 404 if
it doesn't exist. Returns `{project_id, scope_name, project_name}`. The
`scope_type` itself must be one of the join-able scopes."""
if scope_type == "project":
row = db.conn().execute(
"SELECT id, name FROM projects WHERE id = ?", (scope_id,)
).fetchone()
if row is None:
raise HTTPException(404, "Not found")
name = row["name"] or scope_id
return {"project_id": scope_id, "scope_name": name, "project_name": name}
if scope_type == "collection":
col = collections_mod.get_collection(scope_id)
if col is None:
raise HTTPException(404, "Not found")
pid = col["project_id"]
return {
"project_id": pid,
"scope_name": col.get("name") or scope_id,
"project_name": _project_name(pid),
}
raise HTTPException(404, "Not found")
def _require_join_owner(viewer, scope_type: str, scope_id: str) -> None:
"""The accept/decline gate: an Owner whose reach covers the scope (§22.8 'the
scope's Owners across the subtree'). Reuses the S4 invite gates."""
ok = (
auth.can_invite_at_collection(viewer, scope_id)
if scope_type == "collection"
else auth.can_invite_at_project(viewer, scope_id)
)
if not ok:
raise HTTPException(403, "Only an Owner of this scope can act on join requests")
def _require_request(scope_type: str, scope_id: str, request_id: int):
row = db.conn().execute(
"""
SELECT id, scope_type, scope_id, requester_user_id, requested_role,
message, status
FROM join_requests
WHERE id = ? AND scope_type = ? AND scope_id = ?
""",
(request_id, scope_type, scope_id),
).fetchone()
if row is None:
raise HTTPException(404, "Join request not found")
return row
def make_router() -> APIRouter:
router = APIRouter()
# ---------------------------------------------------------------
# GET — what the join form needs to render + gate itself.
# ---------------------------------------------------------------
@router.get("/api/scopes/{scope_type}/{scope_id}/join-target")
async def join_target(scope_type: str, scope_id: str, request: Request) -> dict[str, Any]:
facts = _resolve_scope(scope_type, scope_id)
viewer = auth.current_user(request)
eligible = True
reason: str | None = None
already_requested = False
current_role = auth.effective_role_at_scope(viewer, scope_type, scope_id)
if viewer is None:
eligible, reason = False, "Sign in to request to join."
elif viewer.permission_state != "granted":
eligible, reason = False, "Your beta access request is in review."
elif current_role is not None:
eligible, reason = False, f"You already hold {('Owner' if current_role == 'owner' else 'RFC Contributor')} here."
else:
already_requested = bool(
db.conn().execute(
"""
SELECT 1 FROM join_requests
WHERE scope_type = ? AND scope_id = ? AND requester_user_id = ?
AND status = 'pending' LIMIT 1
""",
(scope_type, scope_id, viewer.user_id),
).fetchone()
)
return {
"scope_type": scope_type,
"scope_id": scope_id,
"name": facts["scope_name"],
"project_id": facts["project_id"],
"eligible": eligible and not already_requested,
"reason": reason,
"already_requested": already_requested,
"current_role": current_role,
}
# ---------------------------------------------------------------
# POST — submit a request to join.
# ---------------------------------------------------------------
@router.post("/api/scopes/{scope_type}/{scope_id}/join-requests")
async def create_join_request(
scope_type: str, scope_id: str, body: JoinRequestBody, request: Request
) -> dict[str, Any]:
viewer = auth.require_contributor(request)
facts = _resolve_scope(scope_type, scope_id)
role = (body.role or "").strip().lower()
if role not in memberships_mod.VALID_ROLES:
raise HTTPException(422, f"invalid role {body.role!r}")
# Already a member of the scope (at this or a broader grain)? Then there
# is nothing to request — a clear 409 rather than a useless self-request.
if auth.effective_role_at_scope(viewer, scope_type, scope_id) is not None:
raise HTTPException(409, "You already hold a role in this scope.")
message = (body.message or "").strip() or None
try:
cur = db.conn().execute(
"""
INSERT INTO join_requests
(scope_type, scope_id, requester_user_id, requested_role, message)
VALUES (?, ?, ?, ?, ?)
""",
(scope_type, scope_id, viewer.user_id, role, message),
)
except sqlite3.IntegrityError:
# The partial unique index — one open request per (scope, user).
raise HTTPException(409, "You already have a pending request to join this scope.")
request_id = cur.lastrowid
# One actionable notification per Owner across the subtree; stamp the
# first onto the row as the inbox-action handle (any Owner may act).
notif_ids = notify.fan_out_join_request(
scope_type=scope_type,
scope_id=scope_id,
scope_name=facts["scope_name"],
project_id=facts["project_id"],
project_name=facts["project_name"],
requester_user_id=viewer.user_id,
request_id=request_id,
requested_role=role,
message=message,
)
if notif_ids:
db.conn().execute(
"UPDATE join_requests SET notification_id = ? WHERE id = ?",
(notif_ids[0], request_id),
)
return {"id": request_id, "scope_type": scope_type, "scope_id": scope_id, "status": "pending"}
# ---------------------------------------------------------------
# POST — Owner accepts → write the membership row.
# ---------------------------------------------------------------
@router.post("/api/scopes/{scope_type}/{scope_id}/join-requests/{request_id}/accept")
async def accept_join_request(
scope_type: str, scope_id: str, request_id: int, body: DecideBody, request: Request
) -> dict[str, Any]:
viewer = auth.require_contributor(request)
facts = _resolve_scope(scope_type, scope_id)
_require_join_owner(viewer, scope_type, scope_id)
req = _require_request(scope_type, scope_id, request_id)
if req["status"] != "pending":
raise HTTPException(409, f"This request was already {req['status']}.")
# The Owner may narrow the requested role on accept; default to what was
# asked for. (Both are within the Owner's grant reach at this scope.)
granted_role = (body.role or req["requested_role"] or "").strip().lower()
if granted_role not in memberships_mod.VALID_ROLES:
raise HTTPException(422, f"invalid role {body.role!r}")
memberships_mod.grant(
scope_type=scope_type,
scope_id=scope_id,
user_id=req["requester_user_id"],
role=granted_role,
granted_by=viewer.user_id,
)
db.conn().execute(
"""
UPDATE join_requests
SET status = 'accepted', decided_at = datetime('now'),
decided_by_user_id = ?, granted_role = ?
WHERE id = ?
""",
(viewer.user_id, granted_role, request_id),
)
notify.notify_join_decided(
requester_user_id=req["requester_user_id"],
decider_user_id=viewer.user_id,
request_id=request_id,
scope_type=scope_type,
scope_id=scope_id,
scope_name=facts["scope_name"],
granted_role=granted_role,
accepted=True,
)
return {"ok": True, "status": "accepted", "granted_role": granted_role}
# ---------------------------------------------------------------
# POST — Owner declines.
# ---------------------------------------------------------------
@router.post("/api/scopes/{scope_type}/{scope_id}/join-requests/{request_id}/decline")
async def decline_join_request(
scope_type: str, scope_id: str, request_id: int, request: Request
) -> dict[str, Any]:
viewer = auth.require_contributor(request)
facts = _resolve_scope(scope_type, scope_id)
_require_join_owner(viewer, scope_type, scope_id)
req = _require_request(scope_type, scope_id, request_id)
if req["status"] != "pending":
raise HTTPException(409, f"This request was already {req['status']}.")
db.conn().execute(
"""
UPDATE join_requests
SET status = 'declined', decided_at = datetime('now'),
decided_by_user_id = ?
WHERE id = ?
""",
(viewer.user_id, request_id),
)
notify.notify_join_decided(
requester_user_id=req["requester_user_id"],
decider_user_id=viewer.user_id,
request_id=request_id,
scope_type=scope_type,
scope_id=scope_id,
scope_name=facts["scope_name"],
granted_role=None,
accepted=False,
)
return {"ok": True, "status": "declined"}
return router
+157
View File
@@ -0,0 +1,157 @@
"""§22 S4 (C.2) — the scope-role invitation surface.
An Owner grants `{owner, contributor}` at a scope their reach covers — the
project, or a single collection within it — to an existing account, looked up
by email. The grant writes a `memberships` row immediately and §15-notifies
the grantee (there is no accept round-trip; the C.2 scenarios name an existing
user and write the row directly). Endpoints:
GET /api/projects/:pid/members — list the project subtree's grants
POST /api/projects/:pid/members — grant at project scope, or
(with collection_id) at one collection
DELETE /api/projects/:pid/members/:user_id — revoke (optionally ?collection_id=)
The single POST keys on the optional `collection_id` so the invite UI's one
control (role picker + scope picker) maps to one endpoint:
* no `collection_id` → project-scope grant; gate `can_invite_at_project`.
* with `collection_id` → collection-scope grant; gate `can_invite_at_collection`.
There is deliberately no "grant at parent, exclude a child" parameter (C.2.5):
the only knobs are role ∈ {owner, contributor} and scope ∈ {project, one
collection}. Reach is bounded by the inviter's own Owner reach (C.2.3): a
collection Owner who is nothing more is refused the project-scope POST.
"""
from __future__ import annotations
from typing import Any
from fastapi import APIRouter, HTTPException, Request
from pydantic import BaseModel
from . import (
auth,
collections as collections_mod,
db,
memberships as memberships_mod,
notify,
)
class GrantBody(BaseModel):
email: str
role: str
collection_id: str | None = None
def _project_name(project_id: str) -> str | None:
row = db.conn().execute(
"SELECT name FROM projects WHERE id = ?", (project_id,)
).fetchone()
return row["name"] if row and row["name"] else None
def make_router() -> APIRouter:
router = APIRouter()
@router.get("/api/projects/{project_id}/members")
async def list_members(project_id: str, request: Request) -> dict[str, Any]:
user = auth.require_contributor(request)
auth.require_project_readable(user, project_id)
# The full subtree listing is a project-Owner view; a collection-only
# Owner manages membership through the collection-scoped POST/DELETE.
if not auth.can_invite_at_project(user, project_id):
raise HTTPException(403, "You may not manage membership in this project")
return {"items": memberships_mod.list_for_project(project_id)}
@router.post("/api/projects/{project_id}/members")
async def grant_member(
project_id: str, body: GrantBody, request: Request
) -> dict[str, Any]:
user = auth.require_contributor(request)
auth.require_project_readable(user, project_id)
role = (body.role or "").strip().lower()
if role not in memberships_mod.VALID_ROLES:
raise HTTPException(422, f"invalid role {body.role!r}")
cid = (body.collection_id or "").strip() or None
if cid is not None:
# Collection-scope grant — bounded by Owner reach over that collection.
col = collections_mod.get_collection(cid)
if col is None or col["project_id"] != project_id:
raise HTTPException(404, "Not found")
if not auth.can_invite_at_collection(user, cid):
raise HTTPException(403, "You may not manage membership in this collection")
scope_type, scope_id = "collection", cid
else:
# Project-scope grant — bounded by Owner reach over the project.
if not auth.can_invite_at_project(user, project_id):
raise HTTPException(403, "You may not manage membership in this project")
scope_type, scope_id = "project", project_id
grantee = memberships_mod.user_by_email(body.email)
if grantee is None:
raise HTTPException(
404,
"No account with that email — the invitee must sign in to the "
"deployment before they can be granted a role",
)
memberships_mod.grant(
scope_type=scope_type,
scope_id=scope_id,
user_id=grantee["id"],
role=role,
granted_by=user.user_id,
)
# §15 (C.2): name the project and role to the grantee.
col_name = None
if scope_type == "collection":
col = collections_mod.get_collection(scope_id)
col_name = (col.get("name") if col else None) or scope_id
notify.notify_scope_role_granted(
recipient_user_id=grantee["id"],
granter_user_id=user.user_id,
scope_type=scope_type,
scope_id=scope_id,
role=role,
project_id=project_id,
project_name=_project_name(project_id),
collection_name=col_name,
)
return {
"scope_type": scope_type,
"scope_id": scope_id,
"user_id": grantee["id"],
"role": role,
"pending": grantee["permission_state"] != "granted",
}
@router.delete("/api/projects/{project_id}/members/{user_id}")
async def revoke_member(
project_id: str, user_id: int, request: Request
) -> dict[str, Any]:
user = auth.require_contributor(request)
auth.require_project_readable(user, project_id)
cid = (request.query_params.get("collection_id") or "").strip() or None
if cid is not None:
col = collections_mod.get_collection(cid)
if col is None or col["project_id"] != project_id:
raise HTTPException(404, "Not found")
if not auth.can_invite_at_collection(user, cid):
raise HTTPException(403, "You may not manage membership in this collection")
removed = memberships_mod.revoke(
scope_type="collection", scope_id=cid, user_id=user_id
)
else:
if not auth.can_invite_at_project(user, project_id):
raise HTTPException(403, "You may not manage membership in this project")
removed = memberships_mod.revoke(
scope_type="project", scope_id=project_id, user_id=user_id
)
return {"removed": removed}
return router
+215
View File
@@ -0,0 +1,215 @@
"""§22.4a SLICE-4/5 — entry metadata edit endpoints.
`POST .../rfcs/<slug>/meta` writes schema-defined metadata to an entry's sidecar
with a direct commit (D7: direct commit for authorized roles), validated against
the collection's field schema at the write boundary (INV-4), lazy-migrating a
legacy entry to a clean body-only `.md` on first edit. The Owner-gated
`metadata.migrate_collection` operator endpoint also lives here (SLICE-4 carried
work); SLICE-5's bulk endpoint will join it.
"""
from __future__ import annotations
from typing import Any
from fastapi import APIRouter, HTTPException, Request
from pydantic import BaseModel
from . import (auth, cache, collections as collections_mod,
metadata as metadata_mod, metadata_schema,
projects as projects_mod)
from .bot import Bot
from .config import Config
from .gitea import Gitea, GiteaError
class MetaEditBody(BaseModel):
values: dict[str, Any]
class BulkMetaBody(BaseModel):
slugs: list[str]
op: str
field: str
value: Any = None
def _apply_op(entry: Any, op: str, field: str, value: Any) -> Any:
"""Return the new value for `field` after applying `op` to `entry`.
`set` → `value`; `add`/`remove` operate on the entry's current tags-list
value for `field` (the route restricts add/remove to tags-type fields).
"""
if op == "set":
return value
current = metadata_mod.metadata_dict(entry).get(field) or []
if not isinstance(current, list):
current = [current]
if op == "add":
return current if value in current else [*current, value]
if op == "remove":
return [x for x in current if x != value]
return value # unreachable; op validated by the route
def make_router(config: Config, gitea: Gitea, bot: Bot) -> APIRouter:
router = APIRouter()
def _content_repo(collection_id: str) -> tuple[str, str]:
# §22/G-15: the COLLECTION's project content_repo, not the deployment
# default — an entry in a non-default project writes its own repo.
# Falls back to the default repo for an unknown collection.
repo = (projects_mod.content_repo_for_collection(collection_id)
or (projects_mod.default_content_repo(config) or ""))
return config.gitea_org, repo
def _md_path(collection_id: str, slug: str) -> str:
sub = collections_mod.subfolder_of(collection_id) or ""
rfcs_dir = f"{sub}/rfcs" if sub else "rfcs"
return f"{rfcs_dir}/{slug}.md"
@router.post("/api/projects/{project_id}/collections/{collection_id}/rfcs/{slug}/meta")
async def edit_meta(
project_id: str, collection_id: str, slug: str,
body: MetaEditBody, request: Request,
) -> dict[str, Any]:
viewer = auth.current_user(request)
if collections_mod.project_of_collection(collection_id) != project_id:
raise HTTPException(404, "Collection not in project")
# INV-4: contributor+ on the collection (returns False for anonymous).
if not auth.can_contribute_in_collection(viewer, collection_id):
raise HTTPException(403, "Contributor access required to edit metadata")
col = collections_mod.get_collection(collection_id)
fields = (col or {}).get("fields") or {}
if not fields:
raise HTTPException(422, "Collection declares no editable fields")
if not body.values:
raise HTTPException(422, "Provide at least one field value")
unknown = [k for k in body.values if k not in fields]
if unknown:
raise HTTPException(422, f"Unknown field(s): {', '.join(sorted(unknown))}")
org, repo = _content_repo(collection_id)
md_path = _md_path(collection_id, slug)
st = await metadata_mod.read_entry_from_git(gitea, org, repo, md_path)
if st is None:
raise HTTPException(404, f"{md_path} not found")
# Validate the *raw* submitted values (INV-4): catch a type mismatch
# before `apply_values` coerces it — e.g. a scalar handed to a `tags`
# field would otherwise char-split into a valid-looking list.
problems = metadata_schema.validate(body.values, fields)
if problems:
raise HTTPException(422, {"problems": [p.as_dict() for p in problems]})
new_entry = metadata_mod.apply_values(st.entry, body.values)
files = metadata_mod.write_entry_files(md_path, new_entry, st)
try:
await bot.commit_entry_files(
viewer.as_actor(), org=org, repo=repo, files=files,
message=f"Edit metadata: {slug}", branch="main")
except GiteaError as e:
raise HTTPException(502, f"Gitea: {e.detail}")
await cache.refresh_meta_repo(config, gitea)
return {
"ok": True, "slug": slug,
"meta": metadata_mod.metadata_dict(new_entry),
}
@router.post("/api/projects/{project_id}/collections/{collection_id}/meta/bulk")
async def bulk_meta(
project_id: str, collection_id: str,
body: BulkMetaBody, request: Request,
) -> dict[str, Any]:
"""§22.4a PUC-2 (SLICE-5): apply one field op to many entries at once.
`set` works for any field; `add`/`remove` operate on a tags-type field.
Each passing entry's metadata is validated at the write boundary
(INV-4) and its sidecar staged; all stage into **one** commit (D7:
bulk = 1 commit, reusing the SLICE-4 sidecar write-through). Entries
that are missing or fail validation are reported in `rejected`; the
rest in `applied`. A no-op (value unchanged) is applied without writing.
"""
viewer = auth.current_user(request)
if collections_mod.project_of_collection(collection_id) != project_id:
raise HTTPException(404, "Collection not in project")
# INV-4: contributor+ on the collection (returns False for anonymous).
if not auth.can_contribute_in_collection(viewer, collection_id):
raise HTTPException(403, "Contributor access required to edit metadata")
col = collections_mod.get_collection(collection_id)
fields = (col or {}).get("fields") or {}
if not fields:
raise HTTPException(422, "Collection declares no editable fields")
if not body.slugs:
raise HTTPException(422, "Provide at least one entry")
if body.op not in ("set", "add", "remove"):
raise HTTPException(422, f"Unknown op: {body.op}")
if body.field not in fields:
raise HTTPException(422, f"Unknown field: {body.field}")
if body.op in ("add", "remove") and fields[body.field].get("type") != "tags":
raise HTTPException(422, f"op {body.op} requires a tags field")
org, repo = _content_repo(collection_id)
applied: list[str] = []
rejected: list[dict[str, str]] = []
all_ops: list[dict[str, Any]] = []
for slug in body.slugs:
md_path = _md_path(collection_id, slug)
st = await metadata_mod.read_entry_from_git(gitea, org, repo, md_path)
if st is None:
rejected.append({"slug": slug, "reason": "not found"})
continue
new_value = _apply_op(st.entry, body.op, body.field, body.value)
# Validate the *raw* new value before coercion (see edit_meta) so a
# scalar `set` onto a tags field is rejected, not char-split.
problems = metadata_schema.validate({body.field: new_value}, fields)
if problems:
rejected.append({"slug": slug,
"reason": "; ".join(p.message for p in problems)})
continue
applied.append(slug)
new_entry = metadata_mod.apply_values(st.entry, {body.field: new_value})
if metadata_mod.metadata_dict(new_entry) != metadata_mod.metadata_dict(st.entry):
all_ops.extend(metadata_mod.write_entry_files(md_path, new_entry, st))
committed = False
if all_ops:
n = len(applied)
msg = f"Bulk {body.op} {body.field}: {n} entr{'y' if n == 1 else 'ies'}"
try:
await bot.commit_entry_files(
viewer.as_actor(), org=org, repo=repo, files=all_ops,
message=msg, branch="main")
except GiteaError as e:
raise HTTPException(502, f"Gitea: {e.detail}")
committed = True
await cache.refresh_meta_repo(config, gitea)
return {"ok": True, "applied": applied,
"rejected": rejected, "committed": committed}
@router.post("/api/projects/{project_id}/collections/{collection_id}/migrate")
async def migrate(
project_id: str, collection_id: str, request: Request
) -> dict[str, Any]:
"""§22.4a PUC-5: migrate a collection's legacy-frontmatter entries to
clean body-only `.md` + sidecars, one commit per collection. Owner-gated
operator action. Safe to ship now that every entry write path is
sidecar-aware (SLICE-4 carried work). Idempotent."""
viewer = auth.current_user(request)
if collections_mod.project_of_collection(collection_id) != project_id:
raise HTTPException(404, "Collection not in project")
if not auth.is_collection_superuser(viewer, collection_id):
raise HTTPException(403, "Owner access required to migrate a collection")
org, repo = _content_repo(collection_id)
subfolder = collections_mod.subfolder_of(collection_id) or ""
try:
result = await metadata_mod.migrate_collection(
gitea, org=org, repo=repo, subfolder=subfolder,
actor=viewer.as_actor())
except GiteaError as e:
raise HTTPException(502, f"Gitea: {e.detail}")
if result["committed"]:
await cache.refresh_meta_repo(config, gitea)
return result
return router
+2 -2
View File
@@ -213,7 +213,7 @@ def make_router(config: Config) -> APIRouter:
@router.post("/api/rfcs/{slug}/watch")
async def set_watch(slug: str, body: WatchBody, request: Request) -> dict[str, Any]:
viewer = auth.require_user(request)
rfc = db.conn().execute("SELECT project_id FROM cached_rfcs WHERE slug = ?", (slug,)).fetchone()
rfc = db.conn().execute("SELECT (SELECT c.project_id FROM collections c WHERE c.id = cached_rfcs.collection_id) AS project_id FROM cached_rfcs WHERE slug = ?", (slug,)).fetchone()
if rfc is None:
raise HTTPException(404, "RFC not found")
# §22.5 visibility gate (subtractive): gated → 404 to non-members.
@@ -222,7 +222,7 @@ def make_router(config: Config) -> APIRouter:
"""
INSERT INTO watches (user_id, rfc_slug, state, set_by, set_at, last_participation_at)
VALUES (?, ?, ?, 'explicit', datetime('now'), datetime('now'))
ON CONFLICT(project_id, user_id, rfc_slug) DO UPDATE SET
ON CONFLICT(collection_id, user_id, rfc_slug) DO UPDATE SET
state = excluded.state,
set_by = 'explicit',
set_at = excluded.set_at
+27 -16
View File
@@ -23,7 +23,7 @@ from typing import Any
from fastapi import APIRouter, HTTPException, Request
from pydantic import BaseModel, Field
from . import auth, cache, chat as chat_layer, db, entry as entry_mod, funder, models_resolver, projects as projects_mod, rfc_links
from . import auth, cache, chat as chat_layer, db, entry as entry_mod, funder, metadata as metadata_mod, models_resolver, projects as projects_mod, rfc_links
from .bot import Bot
from .config import Config
from .gitea import Gitea, GiteaError
@@ -153,7 +153,7 @@ def make_router(
"""
INSERT INTO branch_visibility (rfc_slug, branch_name, read_public, contribute_mode)
VALUES (?, ?, 1, 'just-me')
ON CONFLICT(project_id, rfc_slug, branch_name) DO UPDATE SET read_public = 1
ON CONFLICT(collection_id, rfc_slug, branch_name) DO UPDATE SET read_public = 1
""",
(slug, branch),
)
@@ -189,7 +189,7 @@ def make_router(
"""
INSERT INTO proposed_use_cases (scope, rfc_slug, pr_number, use_case)
VALUES ('pr', ?, ?, ?)
ON CONFLICT(project_id, scope, pr_number) DO UPDATE SET use_case = excluded.use_case
ON CONFLICT(collection_id, scope, pr_number) DO UPDATE SET use_case = excluded.use_case
""",
(slug, pr["number"], use_case),
)
@@ -398,7 +398,7 @@ def make_router(
INSERT INTO pr_seen
(user_id, rfc_slug, pr_number, last_seen_commit_sha, last_seen_message_id, seen_at)
VALUES (?, ?, ?, ?, ?, datetime('now'))
ON CONFLICT(project_id, user_id, rfc_slug, pr_number) DO UPDATE SET
ON CONFLICT(collection_id, user_id, rfc_slug, pr_number) DO UPDATE SET
last_seen_commit_sha = excluded.last_seen_commit_sha,
last_seen_message_id = excluded.last_seen_message_id,
seen_at = excluded.seen_at
@@ -662,7 +662,7 @@ def make_router(
# ------------------------------------------------------------------
def _require_rfc(slug: str, viewer):
row = db.conn().execute("SELECT * FROM cached_rfcs WHERE slug = ?", (slug,)).fetchone()
row = db.conn().execute("SELECT *, (SELECT c.project_id FROM collections c WHERE c.id = cached_rfcs.collection_id) AS project_id FROM cached_rfcs WHERE slug = ?", (slug,)).fetchone()
if row is None:
raise HTTPException(404, "RFC not found")
# §22.5 visibility gate (subtractive, §22.7) — even §11.3 "PRs always
@@ -690,14 +690,20 @@ def make_router(
return not rfc["repo"]
def _owner_repo(rfc) -> tuple[str, str]:
# §22/G-15: a meta-resident entry's repo is its COLLECTION's project
# content_repo, not the deployment default (entry_location falls back to
# the default for a legacy/unknown collection).
if _is_meta_resident(rfc):
return config.gitea_org, (projects_mod.default_content_repo(config) or "")
org, repo, _ = projects_mod.entry_location(config, rfc["collection_id"], rfc["slug"])
return org, repo
owner, repo = rfc["repo"].split("/", 1)
return owner, repo
def _file_path_for(rfc) -> str:
# §22/G-15: the collection's `<subfolder>/rfcs/<slug>.md`.
if _is_meta_resident(rfc):
return f"rfcs/{rfc['slug']}.md"
_, _, path = projects_mod.entry_location(config, rfc["collection_id"], rfc["slug"])
return path
return RFC_FILE_PATH
def _extract_body(rfc, file_contents: str) -> str:
@@ -786,7 +792,7 @@ def _can_merge(rfc, viewer) -> bool:
"""§6.1 admin/owner or §22.6 project_admin OR §6.3 RFC owners/arbiters."""
if viewer is None:
return False
if auth.is_project_superuser(viewer, rfc["project_id"]):
if auth.is_collection_superuser(viewer, rfc["collection_id"]):
return True
owners = json.loads(rfc["owners_json"] or "[]")
arbiters = json.loads(rfc["arbiters_json"] or "[]")
@@ -1009,20 +1015,25 @@ async def _replay_changes(
def _extract_body_for_replay(is_super_draft: bool, content: str) -> str:
# §22.4a SLICE-4: a meta-resident entry may be legacy (frontmatter+body) or
# migrated (body-only). strip_frontmatter handles both without raising.
if not is_super_draft:
return content
try:
return entry_mod.parse(content).body
except Exception:
return content
return metadata_mod.strip_frontmatter(content)
def _wrap_body_for_replay(is_super_draft: bool, prior_content: str, new_body: str) -> str:
# §22.4a SLICE-4: identity for body-only (migrated) files — never re-grow
# frontmatter; preserve a legacy file's frontmatter until its next metadata
# edit migrates it.
nb = new_body if new_body.endswith("\n") else new_body + "\n"
if not is_super_draft:
return new_body
entry = entry_mod.parse(prior_content)
entry.body = new_body if new_body.endswith("\n") else new_body + "\n"
return entry_mod.serialize(entry)
return nb
if entry_mod.FRONTMATTER_RE.match(prior_content):
entry = entry_mod.parse(prior_content)
entry.body = nb
return entry_mod.serialize(entry)
return nb
def _resolution_branch_name(original_branch: str) -> str:
+379 -49
View File
@@ -16,6 +16,7 @@ from typing import Any
import httpx
from fastapi import HTTPException, Request
from . import collections as collections_mod
from . import db
from .bot import Actor
from .config import Config
@@ -144,6 +145,19 @@ def provision_user(config: Config, profile: dict[str, Any]) -> SessionUser:
c = db.conn()
existing = c.execute("SELECT * FROM users WHERE gitea_id = ?", (gitea_id,)).fetchone()
# No row for this gitea_id yet. A prior OTC sign-in (`provision_or_link_user`)
# may have created an email-only row with this same email and `gitea_id`
# NULL; `idx_users_email` is a unique index, so a blind INSERT below would
# raise IntegrityError and 500 the OAuth callback. Reconcile by email: link
# the OAuth identity onto that existing row instead. Mirror of the OTC
# linker, which links an OAuth-era email row on first OTC sign-in.
linked = False
if existing is None and email:
existing = c.execute(
"SELECT * FROM users WHERE email = ? COLLATE NOCASE AND gitea_id IS NULL LIMIT 1",
(email,),
).fetchone()
linked = existing is not None
if existing is None:
role = "owner" if config.owner_gitea_login and login == config.owner_gitea_login else "contributor"
# v0.8.0: a fresh OAuth-provisioned user is also subject to
@@ -165,15 +179,26 @@ def provision_user(config: Config, profile: dict[str, Any]) -> SessionUser:
permission_state = "granted"
else:
user_id = existing["id"]
role = existing["role"]
# On a fresh link this OAuth sign-in is the row's first, so apply the
# §6.1 owner-zero bootstrap (matching the INSERT path). For a row already
# matched by gitea_id the role is settled — preserve it. `permission_state`
# is preserved either way: linking an OAuth identity onto an existing
# email row must not silently change the human's admission status.
if linked and config.owner_gitea_login and login == config.owner_gitea_login:
role = "owner"
else:
role = existing["role"]
permission_state = existing["permission_state"] or "granted"
# Setting `gitea_id` matters only on the link path (it was NULL); on the
# gitea_id-matched path it re-writes the same value. `role` is likewise a
# no-op there. Always refresh the mutable profile fields + last_seen.
c.execute(
"""
UPDATE users
SET gitea_login = ?, email = ?, display_name = ?, avatar_url = ?, last_seen_at = datetime('now')
SET gitea_id = ?, gitea_login = ?, email = ?, display_name = ?, avatar_url = ?, role = ?, last_seen_at = datetime('now')
WHERE id = ?
""",
(login, email, display, avatar, user_id),
(gitea_id, login, email, display, avatar, role, user_id),
)
return SessionUser(
@@ -336,31 +361,73 @@ def project_visibility(project_id: str) -> str:
return row["visibility"] or "gated"
def _is_default_project(project_id: str) -> bool:
"""True iff `project_id` owns the migration-seeded `default` collection — the
deployment's primary project (§22.13), whatever its configured id. Only there
do M2's role rows (which migrated to collection scope `default`) stand in for
project-level authority."""
return collections_mod.project_of_collection(collections_mod.DEFAULT_COLLECTION_ID) == project_id
def project_member_role(user: SessionUser | None, project_id: str) -> str | None:
"""The user's *explicit* §22.6 project_members role in this project, or
None. This is the stored row only it does not fold in the deployment tier
or the implicit-on-public baseline (those live in the helpers below)."""
"""The user's *project-grain* §22.6 role at this project, or None — the
most-permissive of a **global** grant (inherits down to every project) and a
**project**-scope grant. Mapped back to the legacy
`project_admin`/`project_contributor` strings the project-grain authz speaks.
Back-compat: on the deployment's *default* project only, M2's rows live at
collection scope `default` (§B.3 migration), so a `default` collection-scope
grant there is read as project-level too. A collection grant on any other
project is NOT project authority that is the four-layer collection resolver
(`effective_scope_role`). Does not fold in the deployment tier
(`is_project_superuser` adds it) or the implicit-on-public baseline."""
if user is None:
return None
clauses = ["scope_type = 'global'", "(scope_type = 'project' AND scope_id = ?)"]
params: list = [user.user_id, project_id]
if _is_default_project(project_id):
clauses.append("(scope_type = 'collection' AND scope_id = ?)")
params.append(collections_mod.DEFAULT_COLLECTION_ID)
row = db.conn().execute(
"SELECT role FROM project_members WHERE project_id = ? AND user_id = ?",
(project_id, user.user_id),
"SELECT role FROM memberships WHERE user_id = ? AND (" + " OR ".join(clauses) + ") "
"ORDER BY CASE role WHEN 'owner' THEN 0 ELSE 1 END LIMIT 1",
params,
).fetchone()
return row["role"] if row else None
if row is None:
return None
return "project_admin" if row["role"] == "owner" else "project_contributor"
def project_of_rfc(rfc_slug: str) -> str:
"""The project an RFC belongs to (`cached_rfcs.project_id`). Falls back to
the default project when the slug isn't cached or the column is unset — the
same N=1 default migration 026 backfills."""
"""The project an RFC belongs to, via its collection
(`cached_rfcs.collection_id` -> `collections.project_id`, §22 three-tier).
Falls back to the default project when the slug isn't cached — the same N=1
default migration 026 backfills."""
row = db.conn().execute(
"SELECT project_id FROM cached_rfcs WHERE slug = ?", (rfc_slug,)
"SELECT c.project_id AS project_id "
"FROM cached_rfcs r JOIN collections c ON c.id = r.collection_id "
"WHERE r.slug = ?",
(rfc_slug,),
).fetchone()
if row is None:
return DEFAULT_PROJECT_ID
return row["project_id"] or DEFAULT_PROJECT_ID
def collection_of_rfc(rfc_slug: str) -> str:
"""The collection an RFC belongs to (`cached_rfcs.collection_id`). Falls back
to the default collection when the slug isn't cached. Mirrors
`project_of_rfc`'s first-match semantics; a slug shared across collections is
a known routing ambiguity (the RFC-grain helpers take a bare slug) resolved
by the collection-qualified routes in later slices."""
row = db.conn().execute(
"SELECT collection_id FROM cached_rfcs WHERE slug = ?", (rfc_slug,)
).fetchone()
if row is None or not row["collection_id"]:
return collections_mod.DEFAULT_COLLECTION_ID
return row["collection_id"]
def is_project_superuser(user: SessionUser | None, project_id: str) -> bool:
"""Maximal authority within a project: a deployment owner/admin (superuser
in every project, §22.7) or an explicit `project_admin` (§22.6). Both
@@ -373,20 +440,30 @@ def is_project_superuser(user: SessionUser | None, project_id: str) -> bool:
def can_read_project(user: SessionUser | None, project_id: str) -> bool:
"""The §22.5 visibility gate. `public`/`unlisted` are readable by anyone
(anonymous included `unlisted` is link-only but the link still reads);
`gated` is readable only by a deployment owner/admin or a granted project
member of any role. Used as the subtractive read gate (a gated project's
entries 404 to non-members)."""
"""The §22.5 visibility gate at the project grain. `public`/`unlisted` are
readable by anyone (anonymous included `unlisted` is link-only but the link
still reads); `gated` is readable only by a deployment owner/admin or a
holder of any scope grant reaching the project a global grant, a project
grant, or membership at *any* collection within it (seeing a collection
implies seeing its project). Used as the subtractive read gate (a gated
project's entries 404 to non-members)."""
vis = project_visibility(project_id)
if vis in ("public", "unlisted"):
return True
# gated — members + superusers only, subject to the §6 admission floor.
# gated — scope-role holders + superusers only, subject to the §6 floor.
if user is None or user.permission_state != "granted":
return False
if user.role in _DEPLOYMENT_SUPERUSER_ROLES:
return True
return project_member_role(user, project_id) is not None
row = db.conn().execute(
"SELECT 1 FROM memberships m WHERE m.user_id = ? AND ("
" m.scope_type = 'global'"
" OR (m.scope_type = 'project' AND m.scope_id = ?)"
" OR (m.scope_type = 'collection' AND m.scope_id IN "
" (SELECT id FROM collections WHERE project_id = ?))) LIMIT 1",
(user.user_id, project_id, project_id),
).fetchone()
return row is not None
def require_project_readable(user: SessionUser | None, project_id: str) -> None:
@@ -448,6 +525,262 @@ def visible_project_ids(user: SessionUser | None) -> list[str]:
return [r["id"] for r in rows if can_read_project(user, r["id"])]
# ===========================================================================
# §22 three-tier — S3. The four-layer scope-role resolver (§B.2) and the
# collection-grain visibility gate.
#
# A grant attaches the unified role {owner, contributor} at a scope: global,
# project, or collection (§B.1). Grants inherit downward, are additive, and
# admit no negative override (§B.2). Effective authority over a *collection* is
# the most-permissive union of the layers reaching it:
#
# global (users.role owner/admin memberships scope_type='global')
# project (memberships scope_type='project' at the collection's project)
# collection (memberships scope_type='collection' at the collection)
#
# minus the §22.5 visibility gate and §6.2 write-mute (subtractive, as today).
# Per-entry authority (owners / arbiters / rfc_collaborators) is a distinct,
# finer layer the RFC-grain helpers union in beneath collection.
#
# OPERATOR DECISIONS (S3, session 0076):
# * scope-role-primary — a plain granted account (users.role 'contributor', no
# membership) is a granted *account*, not a write-everywhere global role.
# Write standing comes from an explicit scope grant; the lone exception is the
# grandfathered implicit-public baseline below.
# * grandfathered baseline — the migration-seeded `default` collection keeps the
# pre-three-tier implicit-on-public write baseline (a granted deployment
# contributor may propose while it is public), so the N=1 deployment (§22.13)
# loses no capability. Every *explicitly-created* collection requires an
# explicit scope grant to write.
# * hidden-from-public — a `gated` collection is invisible to the public (404,
# omitted from the directory) yet visible to any scope-role holder reaching it
# (collection/project/global). A collection's visibility may be set only as
# strict or stricter than its project's (the rank ordering below).
# ===========================================================================
# The global scope is a single tier per deployment; its grant rows use this
# sentinel scope_id (migration 030).
GLOBAL_SCOPE_ID = "*"
# §22.5 visibility strictness on the public-exposure axis: `public` is least
# strict, `gated` most strict. A collection may narrow its project's visibility
# but never widen it (`rank(collection) >= rank(project)`).
_VISIBILITY_RANK = {"public": 0, "unlisted": 1, "gated": 2}
def visibility_rank(visibility: str | None) -> int:
"""The strictness rank of a §22.5 visibility (higher = stricter). An unknown
value reads as the strictest (`gated`) the safe default."""
return _VISIBILITY_RANK.get(visibility or "", _VISIBILITY_RANK["gated"])
def effective_scope_role(user: SessionUser | None, collection_id: str) -> str | None:
"""The most-permissive unified role ({'owner','contributor'}) the user holds
over `collection_id`, folding §B.2's global → project → collection layers.
Returns None when no scope grant reaches the collection. 'owner' outranks
'contributor'; there is no negative override (a parent grant is never
subtracted by a child). Subject to the §6 admission floor."""
if user is None or user.permission_state != "granted":
return None
# Global tier — a deployment owner/admin is a global Owner (§B.1).
if user.role in _DEPLOYMENT_SUPERUSER_ROLES:
return "owner"
pid = collections_mod.project_of_collection(collection_id)
row = db.conn().execute(
"SELECT role FROM memberships "
"WHERE user_id = ? AND ("
" scope_type = 'global'"
" OR (scope_type = 'project' AND scope_id = ?)"
" OR (scope_type = 'collection' AND scope_id = ?)) "
"ORDER BY CASE role WHEN 'owner' THEN 0 ELSE 1 END LIMIT 1",
(user.user_id, pid, collection_id),
).fetchone()
return row["role"] if row else None
def _effective_project_role(user: SessionUser | None, project_id: str) -> str | None:
"""The most-permissive role the user holds *over a project* — folding the
global tier (deployment owner/admin, or a `scope_type='global'` grant) and a
`scope_type='project'` grant on this project. Unlike `effective_scope_role`
(which keys on a collection), this answers the project grain directly, for the
§22.8 request-to-join membership check. Subject to the §6 admission floor."""
if user is None or user.permission_state != "granted":
return None
if user.role in _DEPLOYMENT_SUPERUSER_ROLES:
return "owner"
row = db.conn().execute(
"SELECT role FROM memberships "
"WHERE user_id = ? AND ("
" scope_type = 'global'"
" OR (scope_type = 'project' AND scope_id = ?)) "
"ORDER BY CASE role WHEN 'owner' THEN 0 ELSE 1 END LIMIT 1",
(user.user_id, project_id),
).fetchone()
return row["role"] if row else None
def effective_role_at_scope(
user: SessionUser | None, scope_type: str, scope_id: str
) -> str | None:
"""The most-permissive scope role the user holds over a `(scope_type,
scope_id)` target the scope-grain twin of `effective_scope_role`. A
`collection` target folds global project collection (the existing
resolver); a `project` target folds global project. Returns None when no
grant reaches the scope. Drives the §22.8 "already a member?" gate."""
if scope_type == "collection":
return effective_scope_role(user, scope_id)
if scope_type == "project":
return _effective_project_role(user, scope_id)
return None
def collection_visibility(collection_id: str) -> str:
"""The collection's own §22.5 visibility. A missing row reads as 'gated'
an unknown collection is invisible rather than open."""
row = db.conn().execute(
"SELECT visibility FROM collections WHERE id = ?", (collection_id,)
).fetchone()
if row is None or not row["visibility"]:
return "gated"
return row["visibility"]
def effective_collection_visibility(collection_id: str) -> str:
"""The stricter of the collection's own visibility and its project's (§22.5
'both gates'). A collection is constrained to be its project in strictness,
but we max() defensively so a misconfigured looser collection can never widen
its project's gate."""
cvis = collection_visibility(collection_id)
pid = collections_mod.project_of_collection(collection_id)
pvis = project_visibility(pid) if pid else "gated"
return cvis if visibility_rank(cvis) >= visibility_rank(pvis) else pvis
def can_read_collection(user: SessionUser | None, collection_id: str) -> bool:
"""§22.5 read/existence gate at the collection grain. `public`/`unlisted`
read by anyone (anonymous included `unlisted` is link-only but the link
reads); `gated` ("hidden from public existence") reads only for a scope-role
holder over the collection (collection/project/global) or a deployment
owner/admin. The subtractive read gate a gated collection 404s a
non-holder, indistinguishable from absent."""
vis = effective_collection_visibility(collection_id)
if vis in ("public", "unlisted"):
return True
if user is None or user.permission_state != "granted":
return False
return effective_scope_role(user, collection_id) is not None
def require_collection_readable(user: SessionUser | None, collection_id: str) -> None:
"""Raise 404 when the collection is not readable by this viewer (§22.5: a
hidden/gated collection is invisible to non-holders the shape matches an
unknown collection)."""
if not can_read_collection(user, collection_id):
raise HTTPException(status_code=404, detail="Not found")
def is_collection_superuser(user: SessionUser | None, collection_id: str) -> bool:
"""Maximal authority over a collection: an effective scope role of 'owner'
reaching it (a collection Owner, a project Owner of its project, a global
Owner, or a deployment owner/admin). Subsumes the per-entry owners/arbiters
tier within the collection."""
return effective_scope_role(user, collection_id) == "owner"
def _has_collection_write_baseline(user: SessionUser | None, collection_id: str) -> bool:
"""The grandfathered implicit-on-public write baseline, narrowed to the
migration-seeded `default` collection (§22.13 N=1 case). A granted deployment
`contributor` keeps its pre-three-tier write standing on the default
collection while its effective visibility is public; every explicitly-created
collection requires an explicit scope grant (S3 operator decision)."""
if user is None or user.permission_state != "granted":
return False
if collection_id != collections_mod.DEFAULT_COLLECTION_ID:
return False
return user.role == "contributor" and effective_collection_visibility(collection_id) == "public"
def can_contribute_in_collection(user: SessionUser | None, collection_id: str) -> bool:
"""May the user contribute *new* content to the collection (propose an entry)
the collection-level contribute standing. The union of the scope-role grant
(owner/contributor reaching the collection) and the grandfathered default
baseline, subject to the visibility read gate."""
if user is None or user.permission_state != "granted":
return False
if not can_read_collection(user, collection_id):
return False
if effective_scope_role(user, collection_id) is not None:
return True
return _has_collection_write_baseline(user, collection_id)
def can_discuss_in_collection(user: SessionUser | None, collection_id: str) -> bool:
"""May the user participate in discussion in the collection — the
collection-level discuss standing. A superset of contribute for this pass
(the read-only viewer tier is deferred, §B.3), so it mirrors
`can_contribute_in_collection`."""
return can_contribute_in_collection(user, collection_id)
def can_create_collection(user: SessionUser | None, project_id: str) -> bool:
"""May the user create a new collection in this project (§B.1)? Creating a
collection is a *project-level* action: a deployment owner/admin, or any
holder of a project-scope or global-scope grant (Owner OR RFC Contributor
'anyone at the project level with permission to create a collection'). A
*collection*-scope grant cannot create sibling collections."""
if user is None or user.permission_state != "granted":
return False
if user.role in _DEPLOYMENT_SUPERUSER_ROLES:
return True
row = db.conn().execute(
"SELECT 1 FROM memberships "
"WHERE user_id = ? AND ("
" scope_type = 'global'"
" OR (scope_type = 'project' AND scope_id = ?)) LIMIT 1",
(user.user_id, project_id),
).fetchone()
return row is not None
def can_create_project(user: SessionUser | None) -> bool:
"""§22 S5 (§A.2 / §B.1): may the user create a new project? "+ New project"
is a **global-Owner** action a deployment owner/admin (a global Owner per
§B.1) or a holder of an explicit `scope_type='global'` Owner grant. Creating
a project is deployment-level, so it is not reachable by a project- or
collection-scope grant nor by a global RFC Contributor (that role creates
collections, not projects). Subject to the §6 admission floor."""
if user is None or user.permission_state != "granted":
return False
if user.role in _DEPLOYMENT_SUPERUSER_ROLES:
return True
row = db.conn().execute(
"SELECT 1 FROM memberships "
"WHERE user_id = ? AND scope_type = 'global' AND role = 'owner' LIMIT 1",
(user.user_id,),
).fetchone()
return row is not None
def can_invite_at_project(user: SessionUser | None, project_id: str) -> bool:
"""§22 S4 (C.2): may the user grant scope roles at this project (or at any
collection within it)? Managing membership is an *Owner* capability whose
reach covers the project a deployment owner/admin, a global Owner, or this
project's Owner. An RFC Contributor does not manage membership (C.2.4); a
collection Owner's reach is its own collection only (C.2.3), so it is not
offered project-scope invites. Identical to `is_project_superuser` the
invite gate IS "is an Owner over this project"."""
return is_project_superuser(user, project_id)
def can_invite_at_collection(user: SessionUser | None, collection_id: str) -> bool:
"""§22 S4 (C.2): may the user grant scope roles at this collection? An Owner
whose reach covers it the collection's Owner, its project's Owner, a global
Owner, or a deployment owner/admin (`is_collection_superuser`). This is the
narrowest invite reach; a collection Owner who is nothing more may invite
here but not at the project or globally (C.2.3)."""
return is_collection_superuser(user, collection_id)
# v0.16.0 (roadmap item #12): per-RFC membership helpers.
#
# These don't replace `require_contributor` — they layer on top of it for
@@ -544,16 +877,14 @@ def can_discuss_rfc(user: SessionUser | None, rfc_slug: str) -> bool:
return False
if user.permission_state != "granted":
return False
pid = project_of_rfc(rfc_slug)
# §22.5 visibility gate is subtractive (§22.7) — no capability in a project
# the viewer cannot even read.
if not can_read_project(user, pid):
cid = collection_of_rfc(rfc_slug)
# §22.5 visibility gate is subtractive (§22.7) — no capability in a
# collection the viewer cannot even read.
if not can_read_collection(user, cid):
return False
# §22.7 union, override grants first — these bypass per-RFC curation
# (project_viewer ⊇ discussant; project_admin / deployment superuser ⊇ all).
if is_project_superuser(user, pid):
return True
if project_member_role(user, pid) in ("project_viewer", "project_contributor"):
# §B.2 union, scope-role grants first — these bypass per-RFC curation (a
# collection/project/global Owner or RFC Contributor ⊇ discussant).
if effective_scope_role(user, cid) is not None:
return True
# per-RFC authority (union term).
owners = _rfc_owners_set(rfc_slug)
@@ -561,11 +892,11 @@ def can_discuss_rfc(user: SessionUser | None, rfc_slug: str) -> bool:
return True
if is_rfc_collaborator(user, rfc_slug, role_in_rfc=None):
return True
# implicit-public baseline (curation preserved): a granted deployment
# contributor on a public project may discuss only while the RFC is
# unclaimed. The first §13.1 claim engages the per-RFC gate, mirroring the
# grandfathered implicit-public baseline (curation preserved): on the default
# collection a granted deployment contributor may discuss only while the RFC
# is unclaimed. The first §13.1 claim engages the per-RFC gate, mirroring the
# pre-multi-project v0.16.0 contract.
if not owners and _has_write_baseline(user, pid):
if not owners and _has_collection_write_baseline(user, cid):
return True
return False
@@ -589,14 +920,12 @@ def can_contribute_to_rfc(user: SessionUser | None, rfc_slug: str) -> bool:
return False
if user.permission_state != "granted":
return False
pid = project_of_rfc(rfc_slug)
if not can_read_project(user, pid):
cid = collection_of_rfc(rfc_slug)
if not can_read_collection(user, cid):
return False
# §22.7 union, override grants first (project_contributor ⊇
# rfc_collaborators(contributor); project_admin / superuser ⊇ all).
if is_project_superuser(user, pid):
return True
if project_member_role(user, pid) == "project_contributor":
# §B.2 union, scope-role grants first (a collection/project/global RFC
# Contributor ⊇ rfc_collaborators(contributor); an Owner ⊇ all).
if effective_scope_role(user, cid) is not None:
return True
# per-RFC authority (union term). A 'discussant' row is NOT sufficient —
# PRs are the higher-privilege surface.
@@ -605,9 +934,10 @@ def can_contribute_to_rfc(user: SessionUser | None, rfc_slug: str) -> bool:
return True
if is_rfc_collaborator(user, rfc_slug, role_in_rfc="contributor"):
return True
# implicit-public baseline (curation preserved): until an owner exists, a
# granted deployment contributor on a public project may contribute.
if not owners and _has_write_baseline(user, pid):
# grandfathered implicit-public baseline (curation preserved): until an owner
# exists, a granted deployment contributor on the public default collection
# may contribute.
if not owners and _has_collection_write_baseline(user, cid):
return True
return False
@@ -620,13 +950,13 @@ def can_invite_to_rfc(user: SessionUser | None, rfc_slug: str) -> bool:
return False
if user.permission_state != "granted":
return False
pid = project_of_rfc(rfc_slug)
if not can_read_project(user, pid):
cid = collection_of_rfc(rfc_slug)
if not can_read_collection(user, cid):
return False
# Deployment owner/admin or project_admin (§22.6) may invite; otherwise
# only the RFC's frontmatter owner. Per-RFC collaborators and the
# implicit-public baseline do not get the invite-others power.
if is_project_superuser(user, pid):
# An Owner reaching the collection (collection/project/global Owner, or a
# deployment owner/admin) may invite; otherwise only the RFC's frontmatter
# owner. Per-RFC collaborators and the baseline do not get the invite power.
if is_collection_superuser(user, cid):
return True
return is_rfc_owner(user, rfc_slug)
+218 -88
View File
@@ -27,7 +27,7 @@ import json
import logging
from dataclasses import dataclass
from . import db, entry as entry_mod, notify
from . import db, entry as entry_mod, metadata as metadata_mod, notify
from .gitea import Gitea, GiteaError
log = logging.getLogger(__name__)
@@ -163,6 +163,117 @@ class Bot:
def __init__(self, gitea: Gitea):
self._gitea = gitea
# ----- Content repo: collection structure (§22 S2) -----
async def create_collection(
self,
actor: Actor,
*,
org: str,
content_repo: str,
collection_id: str,
manifest_yaml: str,
) -> dict:
"""§22 S2: commit `<collection_id>/.collection.yaml` to the content
repo's main. A structural admin action — committed straight to main (no
PR), like the registry config it feeds; the registry mirror then upserts
the collections row (§22.2 keeps the registry the source of truth). Logs
an audit row for the §6.5 trail."""
path = f"{collection_id}/.collection.yaml"
created = await self._gitea.create_file(
org,
content_repo,
path,
content=manifest_yaml,
message=_stamp_single(f"chore: create collection {collection_id}", actor),
branch="main",
author_name=actor.display_name,
author_email=actor.email or f"{actor.gitea_login}@users.noreply",
)
_log(
actor,
"create_collection",
bot_commit_sha=created.get("commit", {}).get("sha"),
details={"collection_id": collection_id, "repo": content_repo},
)
return created
async def create_project(
self,
actor: Actor,
*,
org: str,
registry_repo: str,
content_repo: str,
project_id: str,
projects_yaml_new: str,
projects_yaml_sha: str,
readme_text: str,
) -> dict:
"""§22 S5 (§A.2): stand up a new project. A global-Owner action wrapping
a bot write at two git sources:
1. **provision the content repo** create `org/content_repo` if it
doesn't exist, then seed a `README.md` on `main` so the branch
exists (the contents API initialises the repo with that commit; the
corpus mirror and the propose path both need a `main` to write to).
2. **register the project** commit the caller-composed
`projects.yaml` (the existing doc with the new project appended) to
the registry repo's `main`.
Like `create_collection`, this is a structural admin action committed
straight to main (no PR), like the registry config it feeds; the caller
then re-runs the registry mirror so the new `projects` + default
`collections` rows flow from the registry (§22.2 keeps the registry the
source of truth). Logs a `create_project` audit row for the §6.5 trail.
Returns the registry update_file result (carries the new commit sha)."""
ae = actor.email or f"{actor.gitea_login}@users.noreply"
existing = await self._gitea.get_repo(org, content_repo)
if existing is None:
await self._gitea.create_org_repo(
org, content_repo, description=f"Content repo for project {project_id}"
)
# Seed a README if absent, which also establishes `main` on a freshly
# created (auto_init=False) repo — the contents API initialises the repo
# with that commit. Keyed on the README rather than the branch so it is
# idempotent and behaves identically whether the repo has a bare `main`
# or no branch at all. Mirrors `ensure_rfc_repo_seed`'s empty-repo seed.
readme = await self._gitea.get_contents(org, content_repo, "README.md", ref="main")
if readme is None:
await self._gitea.create_file(
org,
content_repo,
"README.md",
content=readme_text,
message=_stamp_single(f"chore: initialise content repo for {project_id}", actor),
branch="main",
author_name=actor.display_name,
author_email=ae,
)
result = await self._gitea.update_file(
org,
registry_repo,
"projects.yaml",
content=projects_yaml_new,
sha=projects_yaml_sha,
message=_stamp_single(f"chore: create project {project_id}", actor),
branch="main",
author_name=actor.display_name,
author_email=ae,
)
commit_sha = (
result.get("commit", {}).get("sha")
or result.get("content", {}).get("sha")
or ""
)
_log(
actor,
"create_project",
bot_commit_sha=commit_sha,
details={"project_id": project_id, "content_repo": content_repo},
)
return result
# ----- Meta repo: idea PRs (§9.1 / §9.2) -----
async def open_idea_pr(
@@ -175,12 +286,16 @@ class Bot:
file_contents: str,
pr_title: str,
pr_description: str,
rfcs_dir: str = "rfcs",
) -> dict:
"""Per §9.1: open a meta-repo PR adding one file under rfcs/.
"""Per §9.1: open a meta-repo PR adding one file under `<rfcs_dir>/`.
One file per PR keeps idea submissions atomic and conflict-free.
The PR title and the file-add commit subject share §9.2's fixed
pattern; callers compose `pr_title` as `Propose: <Title>`.
pattern; callers compose `pr_title` as `Propose: <Title>`. §22 S2:
`rfcs_dir` carries the target collection's `<subfolder>/rfcs` so a
propose into a named collection writes under its subfolder; it
defaults to `rfcs` (the default collection / shipped behaviour).
"""
branch = f"propose/{slug}"
await self._gitea.create_branch(org, meta_repo, branch, from_branch="main")
@@ -189,7 +304,7 @@ class Bot:
created = await self._gitea.create_file(
org,
meta_repo,
f"rfcs/{slug}.md",
f"{rfcs_dir}/{slug}.md",
content=file_contents,
message=commit_message,
branch=branch,
@@ -289,6 +404,43 @@ class Bot:
pr_number=pr_number,
)
# ----- Entry sidecar writes (§22.4a SLICE-4) -----
async def commit_entry_files(
self, actor: Actor, *, org: str, repo: str,
files: list[dict], message: str, branch: str = "main",
) -> dict:
"""Commit a set of entry file ops (sidecar + body-only `.md`, from
`metadata.write_entry_files`) in one commit. Used by the direct-commit
metadata paths and, on a branch, by `open_entry_pr`."""
return await self._gitea.change_files(
org, repo, files=files,
message=_stamp_single(message, actor), branch=branch,
author_name=actor.display_name,
author_email=actor.email or f"{actor.gitea_login}@users.noreply",
)
async def open_entry_pr(
self, actor: Actor, *, org: str, repo: str, slug: str,
files: list[dict], pr_title: str, pr_description: str,
branch_prefix: str = "metadata",
) -> dict:
"""Create a branch, commit entry file ops there, and open a PR — the
sidecar-aware successor to `open_metadata_pr`'s single-file write."""
import secrets
branch = f"{branch_prefix}-{slug}-{secrets.token_hex(3)}"
await self._gitea.create_branch(org, repo, branch, from_branch="main")
await self.commit_entry_files(
actor, org=org, repo=repo, files=files,
message=pr_title, branch=branch)
_subject, pr_body = _stamp("", pr_description, actor)
pr = await self._gitea.create_pull(
org, repo, title=pr_title, body=pr_body, head=branch, base="main")
_log(actor, "open_entry_pr", rfc_slug=slug, branch_name=branch,
pr_number=pr["number"], details={"pr_title": pr_title})
return pr
# ----- Meta repo: metadata-pane PRs (§9.5) -----
async def open_metadata_pr(
@@ -298,17 +450,19 @@ class Bot:
org: str,
meta_repo: str,
slug: str,
file_path: str,
new_file_contents: str,
prior_sha: str,
pr_title: str,
pr_description: str,
) -> dict:
"""Per §9.5: a metadata-pane edit (title or tags) on a super-draft
opens a tiny meta-repo PR that touches only the frontmatter of
`rfcs/<slug>.md`. One commit, one PR, easy to triage. The branch
name uses the dash-separated `metadata-<slug>-<6hex>` shape same
routing-friendly form Slice 4 picked for edit branches per the
§19.2 path-routing candidate.
opens a tiny meta-repo PR that touches only the frontmatter of the
entry's `.md`. `file_path` is the collection-resolved path (§22/G-15:
`<subfolder>/rfcs/<slug>.md`), not assumed to be at the repo root. One
commit, one PR, easy to triage. The branch name uses the dash-separated
`metadata-<slug>-<6hex>` shape same routing-friendly form Slice 4
picked for edit branches per the §19.2 path-routing candidate.
"""
import secrets
@@ -319,7 +473,7 @@ class Bot:
result = await self._gitea.update_file(
org,
meta_repo,
f"rfcs/{slug}.md",
file_path,
content=new_file_contents,
sha=prior_sha,
message=commit_message,
@@ -704,35 +858,29 @@ class Bot:
org: str,
meta_repo: str,
slug: str,
new_file_contents: str,
prior_sha: str,
files: list[dict],
rfc_id: str | None,
owners: list[str],
) -> dict:
"""§13.3 (meta-only): open a PR against the meta repo that flips the
entry's frontmatter to `state: active` with the graduation stamps
and **optionally** the integer `id`, **keeping the body
unchanged** (§1 meta-only topology; no repo is created and no body
is stripped). When `rfc_id` is None the entry graduates without a
number (id stays null, slug is canonical per §2.3, §13.2). Branch
name uses the `graduate-<slug>-<6hex>` shape dash-separated like
the other meta-repo branches per the §19.2 path-routing candidate.
entry to `state: active` with the graduation stamps and **optionally**
the integer `id`, **keeping the body unchanged** (§1 meta-only
topology; no repo is created and no body is stripped). The graduation
metadata is written to the entry's sidecar (§22.4a) via `files`; a legacy
`.md` is lazy-migrated to body-only in the same commit. When `rfc_id` is
None the entry graduates without a number (id stays null, slug is
canonical per §2.3, §13.2). Branch name uses the `graduate-<slug>-<6hex>`
shape dash-separated like the other meta-repo branches per the §19.2
path-routing candidate.
"""
import secrets
branch = f"graduate-{slug}-{secrets.token_hex(3)}"
await self._gitea.create_branch(org, meta_repo, branch, from_branch="main")
ae = actor.email or f"{actor.gitea_login}@users.noreply"
commit_subject = f"Graduate {slug}{rfc_id}" if rfc_id else f"Graduate {slug} (no number)"
commit_message = _stamp_single(commit_subject, actor)
result = await self._gitea.update_file(
org, meta_repo, f"rfcs/{slug}.md",
content=new_file_contents,
sha=prior_sha,
message=commit_message,
branch=branch,
author_name=actor.display_name, author_email=ae,
)
result = await self.commit_entry_files(
actor, org=org, repo=meta_repo, files=files,
message=commit_subject, branch=branch)
commit_sha = (
result.get("commit", {}).get("sha")
or result.get("content", {}).get("sha")
@@ -821,35 +969,26 @@ class Bot:
org: str,
meta_repo: str,
slug: str,
new_file_contents: str,
prior_sha: str,
files: list[dict],
verb: str,
target_state: str,
) -> dict:
"""§13.7: open a PR flipping `rfcs/<slug>.md` to `state:
<target_state>` for retire (`verb='retire'`, target `retired`)
or un-retire (`verb='unretire'`, target the restored prior state).
Only the frontmatter `state` changes; the body and every other
field (including the integer `id`) are kept, so an un-retire
restores the entry exactly. Branch shape mirrors graduation's
`<verb>-<slug>-<6hex>`.
"""§13.7: open a PR flipping an entry to `state: <target_state>` — for
retire (`verb='retire'`, target `retired`) or un-retire
(`verb='unretire'`, target the restored prior state). The `state` change
is written to the entry's metadata sidecar (§22.4a), keeping the `.md`
body and every other field, so an un-retire restores the entry exactly.
`files` come from `metadata.write_entry_files`. Branch shape mirrors
graduation's `<verb>-<slug>-<6hex>`.
"""
import secrets
branch = f"{verb}-{slug}-{secrets.token_hex(3)}"
await self._gitea.create_branch(org, meta_repo, branch, from_branch="main")
ae = actor.email or f"{actor.gitea_login}@users.noreply"
verb_title = "Retire" if verb == "retire" else "Un-retire"
commit_subject = f"{verb_title} {slug}"
commit_message = _stamp_single(commit_subject, actor)
result = await self._gitea.update_file(
org, meta_repo, f"rfcs/{slug}.md",
content=new_file_contents,
sha=prior_sha,
message=commit_message,
branch=branch,
author_name=actor.display_name, author_email=ae,
)
result = await self.commit_entry_files(
actor, org=org, repo=meta_repo, files=files,
message=f"{verb_title} {slug}", branch=branch)
commit_sha = (
result.get("commit", {}).get("sha")
or result.get("content", {}).get("sha")
@@ -1019,30 +1158,23 @@ class Bot:
org: str,
meta_repo: str,
slug: str,
new_file_contents: str,
prior_sha: str,
files: list[dict],
) -> dict:
"""§13.1: open a PR adding the actor to the entry's `owners:` list.
Touches only the frontmatter of `rfcs/<slug>.md`. Branch shape is
`claim/<slug>` single attempt per super-draft per actor (Gitea
refuses duplicate branch creation, which is the right behavior:
if the claim is still open, point the contributor at the existing
PR rather than opening a second one).
Writes the updated `owners:` to the entry's metadata sidecar (§22.4a)
via `files`; a legacy `.md` is lazy-migrated to body-only in the same
commit. Branch shape is `claim/<slug>` single attempt per super-draft
per actor (Gitea refuses duplicate branch creation, which is the right
behavior: if the claim is still open, point the contributor at the
existing PR rather than opening a second one).
"""
branch = f"claim/{slug}"
await self._gitea.create_branch(org, meta_repo, branch, from_branch="main")
ae = actor.email or f"{actor.gitea_login}@users.noreply"
commit_subject = f"Claim ownership of {slug} for {actor.gitea_login}"
commit_message = _stamp_single(commit_subject, actor)
result = await self._gitea.update_file(
org, meta_repo, f"rfcs/{slug}.md",
content=new_file_contents,
sha=prior_sha,
message=commit_message,
branch=branch,
author_name=actor.display_name, author_email=ae,
)
result = await self.commit_entry_files(
actor, org=org, repo=meta_repo, files=files,
message=commit_subject, branch=branch)
commit_sha = (
result.get("commit", {}).get("sha")
or result.get("content", {}).get("sha")
@@ -1080,30 +1212,28 @@ class Bot:
slug: str,
reviewed_by: str,
reviewed_at: str,
file_path: str | None = None,
) -> None:
"""Clear §22.4c unreviewed on an active entry by rewriting its
frontmatter on main. Stamps the commit with the §6.5 On-behalf-of
trailer and writes an actions-log row, mirroring the graduation
stamp's bot-write shape."""
path = f"rfcs/{slug}.md"
result = await self._gitea.read_file(org, meta_repo, path, ref="main")
if result is None:
"""Clear §22.4c unreviewed on an active entry by writing its metadata
sidecar on main (§22.4a). Dual-reads the entry (so a migrated body-only
`.md` doesn't crash) and lazy-migrates a legacy `.md` to body-only in the
same commit. Stamps the §6.5 On-behalf-of trailer and writes an
actions-log row, mirroring the graduation stamp's bot-write shape.
`file_path` is the collection-resolved entry path (§22/G-15:
`<subfolder>/rfcs/<slug>.md`); it defaults to the repo-root
`rfcs/<slug>.md` for the legacy single-corpus / default-collection case."""
path = file_path or f"rfcs/{slug}.md"
st = await metadata_mod.read_entry_from_git(self._gitea, org, meta_repo, path)
if st is None:
raise GiteaError(404, f"{path} not found")
text, sha = result
e = entry_mod.parse(text)
e.unreviewed = False
e.reviewed_at = reviewed_at
e.reviewed_by = reviewed_by
commit_message = _stamp_single(f"Mark {slug} reviewed", actor)
result = await self._gitea.update_file(
org, meta_repo, path,
content=entry_mod.serialize(e),
sha=sha,
message=commit_message,
branch="main",
author_name=actor.display_name,
author_email=actor.email or f"{actor.gitea_login}@users.noreply",
)
e = metadata_mod.apply_values(st.entry, {
"unreviewed": False, "reviewed_at": reviewed_at, "reviewed_by": reviewed_by,
})
files = metadata_mod.write_entry_files(path, e, st)
result = await self.commit_entry_files(
actor, org=org, repo=meta_repo, files=files,
message=f"Mark {slug} reviewed", branch="main")
commit_sha = (
result.get("commit", {}).get("sha")
or result.get("content", {}).get("sha")
+180 -80
View File
@@ -27,7 +27,15 @@ import asyncio
import json
import logging
from . import db, entry as entry_mod, projects as projects_mod, registry as registry_mod
from . import (
collections as collections_mod,
db,
entry as entry_mod,
metadata as metadata_mod,
metadata_schema,
projects as projects_mod,
registry as registry_mod,
)
from .config import Config
from .gitea import Gitea, GiteaError
@@ -54,12 +62,45 @@ async def refresh_meta_repo(config: Config, gitea: Gitea) -> None:
async def _refresh_project_corpus(org: str, project_id: str, repo: str, gitea: Gitea) -> None:
# §22 S2: the corpus grain is the collection. Mirror every collection of the
# project from its `<subfolder>/rfcs/` directory, keying cached_rfcs by the
# collection id. The default collection (subfolder '') reads `rfcs/` — the
# shipped path, unchanged. include_unlisted: the mirror serves every
# collection's content regardless of enumeration visibility.
from . import collections as collections_mod
for col in collections_mod.list_collections(project_id, include_unlisted=True):
await _refresh_collection_corpus(
org, project_id, repo, col["id"], col["subfolder"] or "", gitea
)
async def _refresh_collection_corpus(
org: str, project_id: str, repo: str, collection_id: str, subfolder: str, gitea: Gitea
) -> None:
rfcs_dir = f"{subfolder}/rfcs" if subfolder else "rfcs"
try:
files = await gitea.list_dir(org, repo, "rfcs", ref="main")
files = await gitea.list_dir(org, repo, rfcs_dir, ref="main")
except GiteaError as e:
log.warning("refresh_meta_repo: project %s: cannot list rfcs/: %s", project_id, e)
log.warning("refresh_meta_repo: %s/%s: cannot list %s: %s",
project_id, collection_id, rfcs_dir, e)
return
# §22.4a SLICE-2: a collection may declare a metadata field schema. Fetch it
# once for the whole corpus pass; entries whose stored values fail it are
# flagged malformed advisory-only (INV-3) — the read never hard-fails. A
# collection with no schema validates nothing (INV-5, the default unchanged).
col = collections_mod.get_collection(collection_id)
fields_schema = (col or {}).get("fields") or None
# §22.4a SLICE-1: an entry's metadata may live in a `<slug>.meta.yaml`
# sidecar (the source of truth) with the `.md` kept as pure prose. Map the
# sidecars surfaced by this listing so each `.md` can dual-read its sibling.
sidecar_path_by_slug = {
metadata_mod.slug_of_sidecar(f["name"]): f["path"]
for f in files
if f.get("type") == "file" and metadata_mod.is_sidecar(f.get("name", ""))
}
seen_slugs: set[str] = set()
for f in files:
if f.get("type") != "file" or not f.get("name", "").endswith(".md"):
@@ -68,47 +109,81 @@ async def _refresh_project_corpus(org: str, project_id: str, repo: str, gitea: G
if not result:
continue
text, sha = result
stem = f["name"][:-len(".md")]
sidecar_text: str | None = None
sidecar_path = sidecar_path_by_slug.get(stem)
if sidecar_path:
sc_result = await gitea.read_file(org, repo, sidecar_path, ref="main")
sidecar_text = sc_result[0] if sc_result else None
try:
entry = entry_mod.parse(text)
entry, malformed = metadata_mod.read_entry(text, sidecar_text, fallback_slug=stem)
except Exception as parse_err:
log.warning("refresh_meta_repo: %s: skipping %s: %s", project_id, f["path"], parse_err)
log.warning("refresh_meta_repo: %s/%s: skipping %s: %s",
project_id, collection_id, f["path"], parse_err)
continue
if not entry.slug:
log.warning("refresh_meta_repo: %s: skipping %s: missing slug", project_id, f["path"])
log.warning("refresh_meta_repo: %s/%s: skipping %s: missing slug",
project_id, collection_id, f["path"])
continue
if malformed:
log.warning("refresh_meta_repo: %s/%s: %s has malformed metadata sidecar",
project_id, collection_id, f["path"])
# §22.4a SLICE-2: advisory schema validation (INV-3). A schema violation
# flags the entry malformed without blocking the read, OR-ed onto any
# sidecar-syntax malformation above.
if fields_schema:
problems = metadata_schema.validate(
metadata_mod.metadata_dict(entry), fields_schema
)
if problems:
malformed = True
log.warning("refresh_meta_repo: %s/%s: %s fails its field schema: %s",
project_id, collection_id, f["path"],
"; ".join(p.message for p in problems))
seen_slugs.add(entry.slug)
_upsert_cached_rfc(entry, body_sha=sha, project_id=project_id)
_upsert_cached_rfc(entry, body_sha=sha, collection_id=collection_id,
metadata_malformed=malformed)
# Entries removed from a project's rfcs/ — the spec keeps withdrawn entries
# Entries removed from a collection's rfcs/ — the spec keeps withdrawn entries
# as historical record (§3), so this fires only for out-of-band deletes;
# leave the row, scoped to this project, for reconciler attention.
# leave the row, scoped to this collection, for reconciler attention.
existing = {
row["slug"]
for row in db.conn().execute(
"SELECT slug FROM cached_rfcs WHERE project_id = ?", (project_id,)
"SELECT slug FROM cached_rfcs WHERE collection_id = ?", (collection_id,)
)
}
for missing in existing - seen_slugs:
log.info("refresh_meta_repo: %s/%s no longer in rfcs/ — leaving cache row", project_id, missing)
log.info("refresh_meta_repo: %s/%s/%s no longer present — leaving cache row",
project_id, collection_id, missing)
def _upsert_cached_rfc(entry: entry_mod.Entry, body_sha: str, project_id: str = "default") -> None:
def _upsert_cached_rfc(
entry: entry_mod.Entry,
body_sha: str,
collection_id: str = "default",
metadata_malformed: bool = False,
) -> None:
# §6.6: models_json stays NULL when the frontmatter key is absent
# (inherit operator universe) and '[]' for the explicit opt-out.
models_json = json.dumps(entry.models) if entry.models is not None else None
# §6.7: funder_login mirrors the optional `funder:` frontmatter
# field. NULL means absent — operator credentials are used.
funder_login = entry.funder or None
# §22.4a SLICE-3: persist the full per-entry metadata mapping (known keys +
# extra, never the body) so facet/filter can read any declared field. Stored
# via metadata_dict so the sidecar's forward-compat keys (INV-7) ride along.
meta_json = json.dumps(metadata_mod.metadata_dict(entry))
db.conn().execute(
"""
INSERT INTO cached_rfcs
(slug, title, state, rfc_id, repo, proposed_by, proposed_at,
graduated_at, graduated_by, owners_json, arbiters_json, tags_json,
models_json, funder_login, body, body_sha,
unreviewed, reviewed_at, reviewed_by, project_id,
last_entry_commit_at, updated_at)
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, datetime('now'), datetime('now'))
ON CONFLICT(project_id, slug) DO UPDATE SET
unreviewed, reviewed_at, reviewed_by, collection_id,
metadata_malformed, meta_json, last_entry_commit_at, updated_at)
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, datetime('now'), datetime('now'))
ON CONFLICT(collection_id, slug) DO UPDATE SET
title = excluded.title,
state = excluded.state,
rfc_id = excluded.rfc_id,
@@ -127,6 +202,8 @@ def _upsert_cached_rfc(entry: entry_mod.Entry, body_sha: str, project_id: str =
unreviewed = excluded.unreviewed,
reviewed_at = excluded.reviewed_at,
reviewed_by = excluded.reviewed_by,
metadata_malformed = excluded.metadata_malformed,
meta_json = excluded.meta_json,
last_entry_commit_at = datetime('now'),
updated_at = datetime('now')
""",
@@ -150,7 +227,9 @@ def _upsert_cached_rfc(entry: entry_mod.Entry, body_sha: str, project_id: str =
1 if entry.unreviewed else 0,
entry.reviewed_at,
entry.reviewed_by,
project_id,
collection_id,
1 if metadata_malformed else 0,
meta_json,
),
)
@@ -210,7 +289,7 @@ async def refresh_rfc_repo(config: Config, gitea: Gitea, slug: str) -> None:
"""
INSERT INTO cached_branches (rfc_slug, branch_name, head_sha, state, last_commit_at)
VALUES (?, ?, ?, 'open', ?)
ON CONFLICT(project_id, rfc_slug, branch_name) DO UPDATE SET
ON CONFLICT(collection_id, rfc_slug, branch_name) DO UPDATE SET
head_sha = excluded.head_sha,
state = CASE WHEN cached_branches.state = 'closed' THEN 'closed' ELSE 'open' END,
last_commit_at = excluded.last_commit_at
@@ -347,73 +426,85 @@ async def refresh_meta_branches(config: Config, gitea: Gitea) -> None:
of slashes per the §19.2 path-routing candidate.
"""
org = config.gitea_org
repo = projects_mod.default_content_repo(config)
if not repo:
log.warning("refresh_meta_branches: default project has no content_repo yet; skipping")
return
try:
branches = await gitea.list_branches(org, repo)
except GiteaError as e:
log.warning("refresh_meta_branches: %s", e)
# §22/G-15: scan EVERY project's content_repo (not just the default), so an
# entry in a non-default project gets its edit branches + synthesized `main`
# row cached and its branch dropdown / has-commits-ahead check work. Mirrors
# refresh_meta_pulls' per-project loop.
prows = db.conn().execute(
"SELECT id, content_repo FROM projects WHERE content_repo IS NOT NULL AND content_repo != ''"
).fetchall()
if not prows:
log.warning("refresh_meta_branches: no projects with a content_repo yet; skipping")
return
meta_main_sha = ""
meta_main_ts = None
edit_keys_seen: set[tuple[str, str]] = set()
for b in branches:
name = b.get("name") or ""
head_sha = (b.get("commit") or {}).get("id") or ""
last_commit_at = (b.get("commit") or {}).get("timestamp")
if name == "main":
meta_main_sha = head_sha
meta_main_ts = last_commit_at
for prow in prows:
repo = prow["content_repo"]
try:
branches = await gitea.list_branches(org, repo)
except GiteaError as e:
log.warning("refresh_meta_branches: %s (%s)", e, repo)
continue
slug = _slug_from_branch_name(name)
if not slug:
continue
rfc = db.conn().execute(
"SELECT state, repo FROM cached_rfcs WHERE slug = ?", (slug,)
).fetchone()
# Meta-only topology (§1): edit branches live on the meta repo for
# every meta-resident entry — super-drafts and active RFCs alike
# (active RFCs are graduated in place and keep editing here, §13).
# A legacy per-RFC repo (repo set) is the only thing excluded.
if not rfc or rfc["repo"] or rfc["state"] not in ("super-draft", "active"):
continue
edit_keys_seen.add((slug, name))
db.conn().execute(
"""
INSERT INTO cached_branches (rfc_slug, branch_name, head_sha, state, last_commit_at)
VALUES (?, ?, ?, 'open', ?)
ON CONFLICT(project_id, rfc_slug, branch_name) DO UPDATE SET
head_sha = excluded.head_sha,
state = CASE WHEN cached_branches.state = 'closed' THEN 'closed' ELSE 'open' END,
last_commit_at = excluded.last_commit_at
""",
(slug, name, head_sha, last_commit_at),
)
# Synthesize a per-slug `main` row for every super-draft entry, so the
# §10.1 has-commits-ahead check in api_prs.py works uniformly. The
# head_sha is the meta-repo main's tip — every super-draft edit branch
# diverges from this single point.
if meta_main_sha:
super_drafts = db.conn().execute(
"SELECT slug FROM cached_rfcs "
"WHERE repo IS NULL AND state IN ('super-draft', 'active')"
).fetchall()
for r in super_drafts:
meta_main_sha = ""
meta_main_ts = None
for b in branches:
name = b.get("name") or ""
head_sha = (b.get("commit") or {}).get("id") or ""
last_commit_at = (b.get("commit") or {}).get("timestamp")
if name == "main":
meta_main_sha = head_sha
meta_main_ts = last_commit_at
continue
slug = _slug_from_branch_name(name)
if not slug:
continue
rfc = db.conn().execute(
"SELECT state, repo FROM cached_rfcs WHERE slug = ?", (slug,)
).fetchone()
# Meta-only topology (§1): edit branches live on the content repo for
# every meta-resident entry — super-drafts and active RFCs alike
# (active RFCs are graduated in place and keep editing here, §13).
# A legacy per-RFC repo (repo set) is the only thing excluded.
if not rfc or rfc["repo"] or rfc["state"] not in ("super-draft", "active"):
continue
edit_keys_seen.add((slug, name))
db.conn().execute(
"""
INSERT INTO cached_branches (rfc_slug, branch_name, head_sha, state, last_commit_at)
VALUES (?, 'main', ?, 'open', ?)
ON CONFLICT(project_id, rfc_slug, branch_name) DO UPDATE SET
VALUES (?, ?, ?, 'open', ?)
ON CONFLICT(collection_id, rfc_slug, branch_name) DO UPDATE SET
head_sha = excluded.head_sha,
state = CASE WHEN cached_branches.state = 'closed' THEN 'closed' ELSE 'open' END,
last_commit_at = excluded.last_commit_at
""",
(r["slug"], meta_main_sha, meta_main_ts),
(slug, name, head_sha, last_commit_at),
)
# Synthesize a per-slug `main` row for this project's super-draft/active
# entries, so the §10.1 has-commits-ahead check works uniformly. The
# head_sha is this content repo's main tip — every edit branch in the
# project diverges from that single point.
if meta_main_sha:
super_drafts = db.conn().execute(
"SELECT r.slug AS slug FROM cached_rfcs r "
"JOIN collections c ON c.id = r.collection_id "
"WHERE r.repo IS NULL AND r.state IN ('super-draft', 'active') "
" AND c.project_id = ?",
(prow["id"],),
).fetchall()
for r in super_drafts:
db.conn().execute(
"""
INSERT INTO cached_branches (rfc_slug, branch_name, head_sha, state, last_commit_at)
VALUES (?, 'main', ?, 'open', ?)
ON CONFLICT(collection_id, rfc_slug, branch_name) DO UPDATE SET
head_sha = excluded.head_sha,
last_commit_at = excluded.last_commit_at
""",
(r["slug"], meta_main_sha, meta_main_ts),
)
# Mark previously-known edit branches that disappeared as deleted per
# §11.5 / §12. Keep the row so chat history survives the branch's
# deletion in Gitea.
@@ -468,20 +559,28 @@ async def refresh_meta_pulls(config: Config, gitea: Gitea) -> None:
login as last resort.
"""
org = config.gitea_org
repo = projects_mod.default_content_repo(config)
if not repo:
log.warning("refresh_meta_pulls: default project has no content_repo yet; skipping")
bot_login = config.gitea_bot_user
rows = db.conn().execute(
"SELECT id, content_repo FROM projects WHERE content_repo IS NOT NULL AND content_repo != ''"
).fetchall()
if not rows:
log.warning("refresh_meta_pulls: no projects with a content_repo yet; skipping")
return
for prow in rows:
await _refresh_project_pulls(org, prow["id"], prow["content_repo"], gitea, bot_login)
async def _refresh_project_pulls(
org: str, project_id: str, repo: str, gitea: Gitea, bot_login: str
) -> None:
repo_full = f"{org}/{repo}"
try:
open_pulls = await gitea.list_pulls(org, repo, state="open")
closed_pulls = await gitea.list_pulls(org, repo, state="closed")
except GiteaError as e:
log.warning("refresh_meta_pulls: %s", e)
log.warning("refresh_meta_pulls: project %s: %s", project_id, e)
return
bot_login = config.gitea_bot_user
for pull in open_pulls + closed_pulls:
head_branch = pull.get("head", {}).get("ref", "")
# A merged-and-deleted PR's branch is no longer reported by Gitea
@@ -536,8 +635,8 @@ async def refresh_meta_pulls(config: Config, gitea: Gitea) -> None:
INSERT INTO cached_prs
(rfc_slug, pr_kind, repo, pr_number, title, description, state,
opened_by, opened_at, merged_at, closed_at,
head_branch, base_branch, head_sha, merge_commit_sha)
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
head_branch, base_branch, head_sha, merge_commit_sha, project_id)
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
ON CONFLICT(repo, pr_number) DO UPDATE SET
title = excluded.title,
description = excluded.description,
@@ -564,6 +663,7 @@ async def refresh_meta_pulls(config: Config, gitea: Gitea) -> None:
(pull.get("base") or {}).get("ref") or "main",
(pull.get("head") or {}).get("sha"),
merge_commit_sha,
project_id,
),
)
+146
View File
@@ -0,0 +1,146 @@
"""§22 collection grain — resolution helpers beneath the project tier.
In S1 each project has exactly one collection (the default). These helpers
recover the collection for a project and read the per-corpus fields (`type`,
`initial_state`) that moved down from `projects` in migration 029. Project-grain
authz (auth.py) recovers a row's project by joining `collections` on
`collection_id`.
"""
from __future__ import annotations
import json
from . import db
DEFAULT_COLLECTION_ID = "default"
# §22.4a item (2): the displayed noun for an entry is a type-driven label, a
# framework concept (like role names), not deployment content. The chrome reads
# this from the API rather than hardcoding "RFC", so a `specification` collection
# says "Spec" and a `bdd` collection says "Feature" with no per-deployment config.
ENTRY_NOUN = {
"document": "RFC",
"specification": "Spec",
"bdd": "Feature",
}
_DEFAULT_ENTRY_NOUN = "RFC"
def entry_noun(collection_type: str) -> str:
"""The §22.4a entry noun for a collection type. Unknown types fall back to
the generic 'RFC' so a future type is never label-less."""
return ENTRY_NOUN.get(collection_type, _DEFAULT_ENTRY_NOUN)
def _enabled_models_from_config(config_json: str | None) -> list[str] | None:
"""§22.12 per-collection enabled_models from a `config_json` blob, or None
when unset (the collection inherits its project's universe)."""
if not config_json:
return None
try:
cfg = json.loads(config_json)
except (json.JSONDecodeError, TypeError):
return None
em = cfg.get("enabled_models") if isinstance(cfg, dict) else None
return [str(m) for m in em] if isinstance(em, list) else None
def _fields_from_config(config_json: str | None) -> dict | None:
"""§22.4a SLICE-2 per-collection metadata field schema from a `config_json`
blob, or None when the collection declares no `fields:`. The stored value is
already normalized by `metadata_schema.parse_fields` at ingest, so it's
served verbatim."""
if not config_json:
return None
try:
cfg = json.loads(config_json)
except (json.JSONDecodeError, TypeError):
return None
fields = cfg.get("fields") if isinstance(cfg, dict) else None
return fields if isinstance(fields, dict) and fields else None
def default_collection_id(project_id: str) -> str:
"""The id of a project's default (S1: sole) collection. Falls back to the
literal 'default' when the project has no collection row yet."""
row = db.conn().execute(
"SELECT id FROM collections WHERE project_id = ? ORDER BY created_at, id LIMIT 1",
(project_id,),
).fetchone()
return row["id"] if row else DEFAULT_COLLECTION_ID
def project_of_collection(collection_id: str) -> str | None:
"""The project a collection belongs to, or None if unknown."""
row = db.conn().execute(
"SELECT project_id FROM collections WHERE id = ?", (collection_id,)
).fetchone()
return row["project_id"] if row else None
def collection_initial_state(collection_id: str) -> str:
"""§22.4b landing state for new entries in a collection. 'super-draft'
default for an unknown/unset row (today's safe flow)."""
row = db.conn().execute(
"SELECT initial_state FROM collections WHERE id = ?", (collection_id,)
).fetchone()
if row is None or not row["initial_state"]:
return "super-draft"
return row["initial_state"]
def collection_type(collection_id: str) -> str:
"""The collection's immutable §22.4a type. 'document' default for unknown."""
row = db.conn().execute(
"SELECT type FROM collections WHERE id = ?", (collection_id,)
).fetchone()
return row["type"] if row and row["type"] else "document"
def subfolder_of(collection_id: str) -> str:
"""The content-repo subfolder a collection lives under (§22.3). Empty string
for the default collection (entries at the repo root `rfcs/`)."""
row = db.conn().execute(
"SELECT subfolder FROM collections WHERE id = ?", (collection_id,)
).fetchone()
return (row["subfolder"] if row else "") or ""
def get_collection(collection_id: str) -> dict | None:
"""The full collection row as a dict, or None if unknown. `enabled_models`
(§22.12) is unpacked from `config_json` as a list, or None when unset."""
row = db.conn().execute(
"SELECT id, project_id, type, subfolder, initial_state, visibility, name, "
"config_json FROM collections WHERE id = ?",
(collection_id,),
).fetchone()
if row is None:
return None
out = dict(row)
config_json = out.pop("config_json", None)
out["enabled_models"] = _enabled_models_from_config(config_json)
# §22.4a SLICE-2: serve the collection's metadata field schema (None when
# the collection declares no `fields:` — INV-5, the default `document`
# collection is unaffected).
out["fields"] = _fields_from_config(config_json)
out["entry_noun"] = entry_noun(out["type"])
return out
def list_collections(project_id: str, include_unlisted: bool = False) -> list[dict]:
"""Collections in a project, the default first then by name (§22.5). `unlisted`
is omitted from enumeration unless include_unlisted (a direct-id read or the
corpus mirror, which serves every collection)."""
rows = db.conn().execute(
"SELECT id, project_id, type, subfolder, initial_state, visibility, name "
"FROM collections WHERE project_id = ? ORDER BY (id != 'default'), name, id",
(project_id,),
).fetchall()
out: list[dict] = []
for r in rows:
if not include_unlisted and r["visibility"] == "unlisted":
continue
item = dict(r)
item["entry_noun"] = entry_noun(item["type"])
out.append(item)
return out
+44 -2
View File
@@ -58,6 +58,20 @@ class Entry:
reviewed_at: str | None = None
reviewed_by: str | None = None
body: str = ""
# §22.4a (configurable collection metadata, SLICE-1): frontmatter / sidecar
# keys outside the known set above are preserved here verbatim so they ride
# along untouched through a parse→serialize round-trip and the
# frontmatter→sidecar migration (INV-7). Includes future collection-`fields:`
# schema values, which the engine does not interpret.
extra: dict[str, Any] = field(default_factory=dict)
# Frontmatter keys the Entry models explicitly; everything else is `extra`.
KNOWN_KEYS = {
"slug", "title", "state", "id", "repo", "proposed_by", "proposed_at",
"graduated_at", "graduated_by", "owners", "arbiters", "tags", "models",
"funder", "unreviewed", "reviewed_at", "reviewed_by",
}
def parse(text: str) -> Entry:
@@ -66,6 +80,16 @@ def parse(text: str) -> Entry:
raise ValueError("Entry file missing frontmatter")
fm = yaml.safe_load(match.group(1)) or {}
body = match.group(2).lstrip("\n")
return from_frontmatter(fm, body)
def from_frontmatter(fm: dict[str, Any], body: str = "") -> Entry:
"""Build an Entry from an already-parsed metadata mapping + body.
Shared by `parse()` (legacy `.md` frontmatter) and the SLICE-1 dual-read
sidecar path (`metadata.read_entry`), so both produce identical records
(INV-6). `fm` keys outside `KNOWN_KEYS` are preserved on `Entry.extra`.
"""
raw_models = fm.get("models", _ABSENT)
if raw_models is _ABSENT or raw_models is None:
models: list[str] | None = None
@@ -74,6 +98,7 @@ def parse(text: str) -> Entry:
raw_funder = fm.get("funder")
funder = str(raw_funder).strip() if raw_funder else None
unreviewed = bool(fm.get("unreviewed") or False)
extra = {k: v for k, v in fm.items() if k not in KNOWN_KEYS}
return Entry(
slug=str(fm.get("slug") or ""),
title=str(fm.get("title") or ""),
@@ -93,11 +118,18 @@ def parse(text: str) -> Entry:
reviewed_at=fm.get("reviewed_at") or None,
reviewed_by=fm.get("reviewed_by") or None,
body=body,
extra=extra,
)
def serialize(entry: Entry) -> str:
"""Emit canonical entry file text — frontmatter then body."""
def to_frontmatter_dict(entry: Entry) -> dict[str, Any]:
"""The canonical ordered metadata mapping for an entry.
Shared by `serialize()` (which wraps it in `---` fences over the body) and
the SLICE-1 sidecar writer (`metadata.sidecar_yaml`, which emits the same
mapping as a standalone `<slug>.meta.yaml`). Known keys first in canonical
order, then `extra` (INV-7).
"""
fm: dict[str, Any] = {
"slug": entry.slug,
"title": entry.title,
@@ -129,6 +161,16 @@ def serialize(entry: Entry) -> str:
fm["reviewed_at"] = entry.reviewed_at
if entry.reviewed_by:
fm["reviewed_by"] = entry.reviewed_by
# INV-7: forward-compat / unknown keys ride along after the known ones.
for k, v in entry.extra.items():
if k not in fm:
fm[k] = v
return fm
def serialize(entry: Entry) -> str:
"""Emit canonical entry file text — frontmatter then body."""
fm = to_frontmatter_dict(entry)
yaml_text = yaml.safe_dump(fm, sort_keys=False, default_flow_style=False).rstrip()
body = entry.body.lstrip("\n")
if body:
+110
View File
@@ -0,0 +1,110 @@
"""§22.4a SLICE-3 — faceted catalog filtering + counts (read).
Pure functions over already-mirrored entries; no I/O, no DB. An "entry" here is
a plain dict carrying at least:
- "state": the lifecycle state column,
- "metadata_malformed": bool,
- "meta": the per-entry metadata mapping (from cached_rfcs.meta_json).
Facetable fields (§5.1, plan decision 1): a collection's declared `enum` and
`tags` fields, in declaration order, plus the built-in `state` facet appended
last but only when the collection declares a schema (INV-5: a no-`fields:`
collection has no facets at all, so the frontend keeps its legacy chips). `text`
fields are not faceted in v1 (they get a detail control in SLICE-4).
Counts use drill-down semantics (plan decision 2): the count for a value of
field F is taken over entries matching every OTHER field's selection (and the
malformed toggle), not F's own — so within-field values stay switchable (OR
within a field, AND across fields). The returned items list applies ALL
selections.
"""
from __future__ import annotations
from typing import Any
# enum + tags are facetable; text is rendered as a detail control (SLICE-4).
FACETABLE_TYPES = {"enum", "tags"}
def facet_fields(fields: dict[str, dict] | None) -> list[tuple[str, str]]:
"""Ordered `[(name, type), ...]` facetable from the schema, `state` last.
Empty when the collection declares no schema (INV-5)."""
if not fields:
return []
out = [
(name, spec.get("type"))
for name, spec in fields.items()
if spec.get("type") in FACETABLE_TYPES
]
out.append(("state", "enum"))
return out
def allowed_filter_keys(fields: dict[str, dict] | None) -> set[str]:
"""Query-param keys the collection-scoped list accepts (plan decision 6)."""
keys = {name for name, _ in facet_fields(fields)}
keys.update({"unreviewed", "malformed"})
return keys
def _values_for(entry: dict[str, Any], name: str, ftype: str) -> list[str]:
"""The facet value(s) an entry contributes for field `name` (str-cast)."""
if name == "state":
v = entry.get("state")
return [str(v)] if v else []
meta = entry.get("meta") or {}
v = meta.get(name)
if v is None:
return []
if ftype == "tags":
return [str(x) for x in v] if isinstance(v, list) else []
return [str(v)]
def _matches(entry: dict[str, Any], name: str, ftype: str, selected: set[str]) -> bool:
if not selected:
return True
return bool(set(_values_for(entry, name, ftype)) & selected) # OR within field
def filter_and_count(
entries: list[dict[str, Any]],
fields: dict[str, dict] | None,
selections: dict[str, set[str]],
only_malformed: bool = False,
) -> tuple[list[dict[str, Any]], dict[str, dict[str, int]]]:
"""Filter `entries` by `selections` and compute drill-down facet counts.
`selections` maps a facet field name the set of selected values (OR within
the field; AND across fields). `only_malformed` narrows items and counts to
entries flagged malformed (INV-3). Returns `(items, facets)` where
`facets = {field: {value: count}}`. With no schema `([all passing], {})`.
"""
facetable = facet_fields(fields)
def passes_malformed(e: dict[str, Any]) -> bool:
return (not only_malformed) or bool(e.get("metadata_malformed"))
items = [
e for e in entries
if passes_malformed(e)
and all(_matches(e, n, t, selections.get(n, set())) for n, t in facetable)
]
facets: dict[str, dict[str, int]] = {}
for name, ftype in facetable:
counts: dict[str, int] = {}
for e in entries:
if not passes_malformed(e):
continue
if not all(
_matches(e, on, ot, selections.get(on, set()))
for on, ot in facetable
if on != name
):
continue
for val in _values_for(e, name, ftype):
counts[val] = counts.get(val, 0) + 1
facets[name] = counts
return items, facets
+1 -1
View File
@@ -220,7 +220,7 @@ def add_consent(user_id: int, slug: str) -> None:
db.conn().execute(
"""
INSERT INTO funder_consents (user_id, rfc_slug) VALUES (?, ?)
ON CONFLICT(project_id, user_id, rfc_slug) DO NOTHING
ON CONFLICT(collection_id, user_id, rfc_slug) DO NOTHING
""",
(user_id, slug),
)
+34
View File
@@ -193,6 +193,40 @@ class Gitea:
resp = await self._request("PUT", f"/repos/{owner}/{repo}/contents/{path}", json=body)
return resp.json()
async def change_files(
self,
owner: str,
repo: str,
*,
files: list[dict[str, Any]],
message: str,
branch: str,
author_name: str | None = None,
author_email: str | None = None,
) -> dict:
"""Create/update/delete several files in ONE commit (Gitea ChangeFiles).
Each `files` entry is `{"operation": "create"|"update"|"delete",
"path": str, "content": str (for create/update), "sha": str (required
for update/delete)}`. Plaintext `content` is base64-encoded here.
Backs the §22.4a frontmattersidecar migration's "one commit per
collection" (`metadata_migrate`).
"""
out_files: list[dict[str, Any]] = []
for f in files:
item: dict[str, Any] = {"operation": f["operation"], "path": f["path"]}
if "content" in f and f["content"] is not None:
item["content"] = base64.b64encode(f["content"].encode("utf-8")).decode("ascii")
if f.get("sha"):
item["sha"] = f["sha"]
out_files.append(item)
body: dict[str, Any] = {"message": message, "branch": branch, "files": out_files}
if author_name and author_email:
body["author"] = {"name": author_name, "email": author_email}
body["committer"] = {"name": author_name, "email": author_email}
resp = await self._request("POST", f"/repos/{owner}/{repo}/contents", json=body)
return resp.json()
# ----- Pull requests -----
async def list_pulls(self, owner: str, repo: str, state: str = "open") -> list[dict]:
+5 -4
View File
@@ -293,20 +293,21 @@ async def _delete_branch_via_bot(
(we leave the branch row in place a subsequent reconciler sweep
will reconcile or the operator can intervene)."""
rfc = db.conn().execute(
"SELECT state, repo FROM cached_rfcs WHERE slug = ?", (slug,)
"SELECT state, repo, collection_id FROM cached_rfcs WHERE slug = ?", (slug,)
).fetchone()
if rfc is None:
log.warning("hygiene: cannot delete %s/%s — slug missing from cache", slug, branch)
return False
if not rfc["repo"]:
repo = projects_mod.default_content_repo(config)
# §22/G-15: the edit/graduation branch lives on the entry's COLLECTION's
# project content_repo, not the deployment default.
owner, repo, _ = projects_mod.entry_location(config, rfc["collection_id"], slug)
if not repo:
log.warning(
"hygiene: default project has no content_repo; skipping branch delete for %s/%s",
"hygiene: no content_repo resolved; skipping branch delete for %s/%s",
slug, branch,
)
return False
owner = config.gitea_org
elif "/" in rfc["repo"]:
owner, repo = rfc["repo"].split("/", 1)
else:
+86
View File
@@ -66,6 +66,13 @@ class OtcVerifyBody(BaseModel):
trust_device: bool = False
class TestLoginBody(BaseModel):
# The single configured test identity to sign in as. Must equal
# E2E_TEST_AUTH_EMAIL (case-insensitive) or the request is refused —
# see `/auth/test/login`.
email: str = Field(min_length=3, max_length=320)
class PasscodeSetBody(BaseModel):
passcode: str = Field(min_length=1, max_length=64)
@@ -101,7 +108,26 @@ async def lifespan(app: FastAPI):
config = load_config()
db.run_migrations(config)
db.init(config)
# v0.52.0: shout if the deployed-env E2E test-auth shortcut is live.
# It mints owner sessions for one configured identity (see
# `/auth/test/login`); it must only ever be on for a pre-prod (PPE)
# host. A loud startup line means an accidental prod enablement is
# visible in the logs rather than silent.
if os.environ.get("E2E_TEST_AUTH_SECRET", "").strip() and os.environ.get(
"E2E_TEST_AUTH_EMAIL", ""
).strip():
log.warning(
"E2E TEST-AUTH IS ENABLED: POST /auth/test/login will mint an owner "
"session for %s. This must NEVER be set on production.",
os.environ["E2E_TEST_AUTH_EMAIL"].strip(),
)
gitea = Gitea(config)
# §22 framework heal: reconcile a divergent default-project collection id
# (migration 029's ≥2-projects seed names it after the project, e.g. 'ohm',
# but the mirror expects 'default') BEFORE the mirror runs, so the mirror
# merges onto the canonical 'default' collection instead of duplicating it.
# Idempotent no-op on fresh / single-project / already-aligned deployments.
projects.reconcile_default_collection_id(config)
# §22.2: mirror the registry before anything reads projects/content_repo.
# First boot has no last-good rows, so a missing/invalid registry is fatal
# (loud-fail per separation-of-concerns); the reconciler sweep keeps it
@@ -112,6 +138,11 @@ async def lifespan(app: FastAPI):
raise RuntimeError(
f"registry mirror failed at startup ({config.registry_repo_full}/projects.yaml): {e}"
) from e
# §22.13 step 1: re-stamp the M1 bootstrap 'default' project id to the
# deployment's configured id (DEFAULT_PROJECT_ID) once the registry row
# exists, so the original corpus lands at a meaningful /p/<id>/ and
# 'default' is never a public URL. Idempotent no-op once done.
projects.restamp_default_project(config)
if projects.default_content_repo(config) is None:
raise RuntimeError(
f"registry does not describe the default project "
@@ -375,6 +406,61 @@ def _oauth_router(config) -> APIRouter:
"needs_profile": needs_profile,
}
# ---------------------------------------------------------------
# v0.52.0: deployed-environment E2E test-auth shortcut.
#
# Running the Playwright E2E suite against a *deployed* environment
# (PPE) is the §9 pre-prod gate. But the deployed env has neither of
# the two scaffolds the Tier-1 docker stack relies on for auth: a
# Mailpit sink to read the OTC code from, and direct SQLite access to
# inject a granted-owner row. This endpoint replaces both with a
# single gated gesture: it mints an authenticated OWNER session for
# one pre-configured throwaway identity.
#
# It is FAIL-CLOSED and must never function in production:
# * 404 unless BOTH `E2E_TEST_AUTH_SECRET` and `E2E_TEST_AUTH_EMAIL`
# are set — a prod deployment that sets neither cannot be coaxed
# into minting a session, and the route is invisible.
# * The caller must present the shared secret in `X-Test-Auth-Secret`
# (constant-time compare); a wrong/absent secret 404s (the route
# does not advertise itself to an unauthenticated caller).
# * Only the one configured email may be minted; any other address
# is refused (403). So an enabled PPE exposes exactly one
# throwaway owner identity, with the secret as the trust boundary.
#
# The hard secrets rule (§6.3) holds: the secret is a Secret Manager
# ref injected as env on the VM (never a literal in the repo), and the
# E2E runner presents it from SM at runtime (never echoed).
@router.post("/auth/test/login")
async def test_login(body: TestLoginBody, request: Request):
secret = os.environ.get("E2E_TEST_AUTH_SECRET", "").strip()
configured_email = os.environ.get("E2E_TEST_AUTH_EMAIL", "").strip()
# Feature off (the default, incl. production): route is invisible.
if not secret or not configured_email:
raise HTTPException(404, "Not Found")
presented = request.headers.get("x-test-auth-secret", "")
if not secrets.compare_digest(presented, secret):
# Don't reveal that the route exists to a caller without the
# secret — mirror the "off" shape exactly.
raise HTTPException(404, "Not Found")
if body.email.strip().lower() != configured_email.lower():
raise HTTPException(403, "email not permitted")
# Provision-or-link the row, then force it to a granted owner so
# the metadata write paths (SLICE-4/5) accept it — the deployed
# equivalent of the Tier-1 docker-compose backend-seed owner row.
user = otc.provision_or_link_user(body.email)
db.conn().execute(
"UPDATE users SET role = 'owner', permission_state = 'granted', "
"last_seen_at = datetime('now') WHERE id = ?",
(user.user_id,),
)
db.conn().commit()
user.role = "owner"
user.permission_state = "granted"
auth.store_session(request, user)
return {"ok": True}
# ---------------------------------------------------------------
# v0.10.0: user-set passcodes after OTC (§6.2, roadmap item #8).
#
+141
View File
@@ -0,0 +1,141 @@
"""§22 S4 (C.2) — scope-role grant operations over the `memberships` table.
The membership *gates* (who may invite, who holds which role) live in
`auth.py`; this module holds the *mutations* the invitation surface drives
granting, the "broader scope supersedes narrower" cleanup, listing, and
revocation mirroring how `invites.py` owns the create/claim/list of per-user
invite tokens while the gate (`auth.can_invite_to_rfc`) lives in `auth.py`.
The model (Part B / S3): a `memberships` row is `(scope_type {global,
project, collection}, scope_id, user_id, role {owner, contributor})`, unique
per `(scope_type, scope_id, user_id)`. A grant is a direct write of that row
(the C.2 scenarios write the row immediately and §15-notify an existing
account there is no accept round-trip; inviting a not-yet-account email is
out of S4 scope and handled by the admin-create-invite path).
The "broader scope supersedes narrower" rule (C.2.6): granting a role at a
broader scope removes this user's narrower rows that the new grant *subsumes*
a narrower row whose role is no more permissive than the new one. A narrower
row that is *more* permissive is kept (no negative override: a child Owner
grant survives a parent Contributor grant, and the §B.2 resolver still unions
most-permissively).
"""
from __future__ import annotations
from typing import Any
from . import collections as collections_mod
from . import db
# Higher rank = more permissive. Used by the supersede rule: a narrower grant
# is pruned only when its rank ≤ the new broader grant's rank.
_ROLE_RANK = {"contributor": 1, "owner": 2}
VALID_SCOPE_TYPES = ("global", "project", "collection")
VALID_ROLES = ("owner", "contributor")
GLOBAL_SCOPE_ID = "*"
def user_by_email(email: str) -> dict[str, Any] | None:
"""The `users` row (id, display_name, email, permission_state) for an
email, case-insensitively, or None. The grantee must already be an account
S4 grants a scope role to an existing user, it does not provision one."""
row = db.conn().execute(
"SELECT id, display_name, email, permission_state, role "
"FROM users WHERE lower(email) = lower(?) "
"ORDER BY id LIMIT 1",
(email.strip(),),
).fetchone()
return dict(row) if row else None
def grant(
*,
scope_type: str,
scope_id: str,
user_id: int,
role: str,
granted_by: int | None,
) -> None:
"""Write (or update) the membership row, then apply the C.2.6
broader-scope-supersedes cleanup. Idempotent on `(scope_type, scope_id,
user_id)` re-granting at the same scope updates the role and the grantor.
The grant is recorded regardless of the grantee's deployment
`permission_state`: a `pending` account's row is written (C.2.7), but the
§6 admission floor in `auth.effective_scope_role` keeps it conferring no
write until the account is granted at the deployment."""
db.conn().execute(
"INSERT INTO memberships (scope_type, scope_id, user_id, role, granted_by) "
"VALUES (?, ?, ?, ?, ?) "
"ON CONFLICT (scope_type, scope_id, user_id) "
"DO UPDATE SET role = excluded.role, granted_by = excluded.granted_by, "
"granted_at = datetime('now')",
(scope_type, scope_id, user_id, role, granted_by),
)
_prune_subsumed(scope_type=scope_type, scope_id=scope_id, user_id=user_id, role=role)
def _prune_subsumed(*, scope_type: str, scope_id: str, user_id: int, role: str) -> None:
"""Remove this user's narrower rows that the just-written broader grant
subsumes (same-or-lower role rank within the broader scope's subtree). A
collection grant subsumes nothing narrower (the per-entry tier is separate);
a project grant subsumes its collections; a global grant subsumes every
project and collection."""
rank = _ROLE_RANK[role]
keep_ranks = [r for r, v in _ROLE_RANK.items() if v <= rank]
if not keep_ranks:
return
placeholders = ",".join("?" for _ in keep_ranks)
if scope_type == "project":
# Narrower = collection-scope rows for collections in this project.
db.conn().execute(
f"DELETE FROM memberships "
f"WHERE user_id = ? AND scope_type = 'collection' "
f" AND role IN ({placeholders}) "
f" AND scope_id IN (SELECT id FROM collections WHERE project_id = ?)",
(user_id, *keep_ranks, scope_id),
)
elif scope_type == "global":
# Narrower = every project- and collection-scope row for this user.
db.conn().execute(
f"DELETE FROM memberships "
f"WHERE user_id = ? AND scope_type IN ('project', 'collection') "
f" AND role IN ({placeholders})",
(user_id, *keep_ranks),
)
def revoke(*, scope_type: str, scope_id: str, user_id: int) -> bool:
"""Remove a membership row at exactly this scope. Returns True if a row was
removed. Revocation is scope-exact: it does not cascade to broader or
narrower grants (each is its own administrative act)."""
cur = db.conn().execute(
"DELETE FROM memberships WHERE scope_type = ? AND scope_id = ? AND user_id = ?",
(scope_type, scope_id, user_id),
)
return cur.rowcount > 0
def list_for_project(project_id: str) -> list[dict[str, Any]]:
"""Every project-scope grant on this project plus every collection-scope
grant on its collections, joined to the grantee's display fields — the data
behind the project-Owner membership panel. Ordered project grants first,
then by collection, then by role (Owner before Contributor)."""
rows = db.conn().execute(
"""
SELECT m.scope_type, m.scope_id, m.user_id, m.role, m.granted_at,
u.display_name, u.email, u.permission_state,
c.name AS collection_name
FROM memberships m
JOIN users u ON u.id = m.user_id
LEFT JOIN collections c ON c.id = m.scope_id AND m.scope_type = 'collection'
WHERE (m.scope_type = 'project' AND m.scope_id = ?)
OR (m.scope_type = 'collection'
AND m.scope_id IN (SELECT id FROM collections WHERE project_id = ?))
ORDER BY (m.scope_type != 'project'), m.scope_id,
CASE m.role WHEN 'owner' THEN 0 ELSE 1 END, u.display_name
""",
(project_id, project_id),
).fetchall()
return [dict(r) for r in rows]
+311
View File
@@ -0,0 +1,311 @@
"""§22.4a configurable collection metadata — sidecar storage + dual-read.
SLICE-1 of docs/design/2026-06-06-configurable-collection-metadata.md.
Entry metadata is collection-configured and stored in a per-entry sidecar,
`<slug>.meta.yaml`, with the `.md` body kept as pure prose (INV-2). This module
is the storage/compat layer:
- the **dual-read** parser (`read_entry`) read the sidecar if present, else
legacy top-of-document frontmatter, with identical resulting records
(INV-6);
- sidecar (de)serialization that preserves unknown / forward-compat keys
(INV-7), reusing `entry`'s canonical field semantics;
- lenient parsing that never hard-fails a read a malformed sidecar surfaces
a flag, not an exception (INV-3).
The collection `fields:` schema and per-field validation are SLICE-2; faceted
filtering and the edit UIs are later slices. This module interprets no field
values it only moves metadata between git and in-memory `Entry` records.
"""
from __future__ import annotations
from dataclasses import dataclass
from typing import Any
import yaml
from . import entry as entry_mod
from .entry import Entry
SIDECAR_SUFFIX = ".meta.yaml"
# ----- filename helpers -----
def sidecar_name(slug: str) -> str:
"""The sidecar filename for an entry whose markdown is `<slug>.md`."""
return f"{slug}{SIDECAR_SUFFIX}"
def is_sidecar(name: str) -> bool:
return name.endswith(SIDECAR_SUFFIX)
def slug_of_sidecar(name: str) -> str:
"""The entry stem for a `<slug>.meta.yaml` filename."""
return name[: -len(SIDECAR_SUFFIX)] if is_sidecar(name) else name
def sidecar_path_for(md_path: str) -> str:
"""The sidecar path sibling to a `<dir>/<slug>.md` entry file."""
assert md_path.endswith(".md"), md_path
return md_path[: -len(".md")] + SIDECAR_SUFFIX
# ----- metadata <-> sidecar -----
def metadata_dict(entry: Entry) -> dict[str, Any]:
"""The full metadata mapping for an entry (known fields + `extra`)."""
return entry_mod.to_frontmatter_dict(entry)
def sidecar_yaml(entry: Entry) -> str:
"""Render an entry's metadata as standalone `<slug>.meta.yaml` text."""
return yaml.safe_dump(
metadata_dict(entry), sort_keys=False, default_flow_style=False
)
def parse_sidecar(text: str) -> tuple[dict[str, Any], bool]:
"""Parse sidecar YAML leniently → `(values, malformed)`.
`malformed` is True when the text is not a YAML mapping (a list, a scalar,
or a YAML syntax error). An empty / whitespace-only sidecar is an empty
mapping, not malformed. Never raises (INV-3).
"""
try:
raw = yaml.safe_load(text)
except yaml.YAMLError:
return {}, True
if raw is None:
return {}, False
if not isinstance(raw, dict):
return {}, True
return raw, False
# ----- frontmatter stripping (INV-2) -----
def strip_frontmatter(md_text: str) -> str:
"""Return the prose body of a `.md`, dropping a leading `---…---` block.
A migrated entry's `.md` is body-only and passes through unchanged. A
not-yet-migrated `.md` still carrying frontmatter yields just its body, so
the dual-read body is the same either way.
"""
match = entry_mod.FRONTMATTER_RE.match(md_text)
if not match:
return md_text
return match.group(2).lstrip("\n")
# ----- dual-read (INV-6) -----
def read_entry(
md_text: str, sidecar_text: str | None, *, fallback_slug: str | None = None
) -> tuple[Entry, bool]:
"""Read an entry from its `.md` and optional sidecar → `(Entry, malformed)`.
- **Sidecar present and well-formed (non-empty):** metadata comes from the
sidecar; the body is the `.md` stripped of any leading frontmatter. The
sidecar is the source of truth (INV-1) and wins over stale `.md`
frontmatter.
- **Sidecar present but malformed:** the entry still loads from the legacy
`.md` frontmatter (if any) and is flagged `malformed` (INV-3).
- **Sidecar present but empty:** it has nothing to override with, so fall
back to the `.md` frontmatter (not flagged).
- **No sidecar:** the legacy path parse the `.md` frontmatter (INV-6).
`fallback_slug` (typically the filename stem) backstops the entry's slug
whenever the metadata source lacks one so a degenerate sidecar never
yields a slug-less record the caller has to silently drop (INV-3).
"""
def _with_slug(entry: Entry) -> Entry:
if not entry.slug and fallback_slug:
entry.slug = fallback_slug
return entry
if sidecar_text is None:
return _with_slug(entry_mod.parse(md_text)), False
values, malformed = parse_sidecar(sidecar_text)
if malformed or not values:
# Malformed or empty sidecar: load from the legacy .md so the entry
# still loads; flag only when the sidecar was actually malformed.
try:
entry = entry_mod.parse(md_text)
except ValueError:
entry = entry_mod.from_frontmatter({}, strip_frontmatter(md_text))
return _with_slug(entry), malformed
body = strip_frontmatter(md_text)
return _with_slug(entry_mod.from_frontmatter(values, body)), False
# ----- value editing (SLICE-4) -----
def apply_values(entry: Entry, values: dict[str, Any]) -> Entry:
"""Return a new Entry with `values` merged over the entry's metadata.
Known keys (`tags`, `state`, `reviewed_by`, ) land on their typed fields;
unknown keys land on `extra` (INV-7). The body is carried through unchanged
this mutates metadata only. Unspecified keys are preserved.
"""
merged = metadata_dict(entry)
merged.update(values)
return entry_mod.from_frontmatter(merged, entry.body)
# ----- git-aware read/write (SLICE-4) -----
@dataclass
class EntryGitState:
"""An entry's on-disk state across its `.md` and optional sidecar.
Captured by `read_entry_from_git` and consumed by `write_entry_files` to
decide create-vs-update for the sidecar and whether the `.md` still needs
its frontmatter stripped (lazy migration).
"""
entry: Entry
md_text: str
md_sha: str
sidecar_text: str | None
sidecar_sha: str | None
malformed: bool
async def read_entry_from_git(
gitea: Any, org: str, repo: str, md_path: str, *, ref: str = "main"
) -> "EntryGitState | None":
"""Dual-read an entry from git → `EntryGitState`, or None if the `.md` is
missing. Reads the `.md` and its sibling sidecar (if any); never raises on
bad metadata (INV-3)."""
md = await gitea.read_file(org, repo, md_path, ref=ref)
if md is None:
return None
md_text, md_sha = md
sc_path = sidecar_path_for(md_path)
sc = await gitea.read_file(org, repo, sc_path, ref=ref)
sidecar_text, sidecar_sha = (sc[0], sc[1]) if sc else (None, None)
stem = md_path.rsplit("/", 1)[-1][: -len(".md")]
entry, malformed = read_entry(md_text, sidecar_text, fallback_slug=stem)
return EntryGitState(
entry=entry, md_text=md_text, md_sha=md_sha,
sidecar_text=sidecar_text, sidecar_sha=sidecar_sha, malformed=malformed,
)
def _md_has_frontmatter(md_text: str) -> bool:
return entry_mod.FRONTMATTER_RE.match(md_text) is not None
def write_entry_files(
md_path: str, entry: Entry, state: "EntryGitState"
) -> list[dict[str, Any]]:
"""Produce `change_files` ops that persist `entry`'s metadata to its sidecar
and keep the `.md` as pure prose (INV-1/INV-2).
- Sidecar: `create` when none existed, else `update` at its prior sha.
- `.md`: rewritten body-only **only when it still carries frontmatter**
(lazy migration, INV-6); an already-clean body is left untouched.
"""
sc_path = sidecar_path_for(md_path)
ops: list[dict[str, Any]] = []
sc_op: dict[str, Any] = {
"operation": "update" if state.sidecar_sha else "create",
"path": sc_path,
"content": sidecar_yaml(entry),
}
if state.sidecar_sha:
sc_op["sha"] = state.sidecar_sha
ops.append(sc_op)
if _md_has_frontmatter(state.md_text):
body = strip_frontmatter(state.md_text)
new_md = body if (body == "" or body.endswith("\n")) else body + "\n"
ops.append({
"operation": "update", "path": md_path,
"content": new_md, "sha": state.md_sha,
})
return ops
# ----- frontmatter -> sidecar migration (PUC-5) -----
async def migrate_collection(
gitea: Any,
*,
org: str,
repo: str,
subfolder: str = "",
actor: Any = None,
branch: str = "main",
) -> dict[str, Any]:
"""Migrate a collection's legacy-frontmatter entries to sidecars.
Walks `<subfolder>/rfcs`; for each `<slug>.md` that has legacy frontmatter
and **no** `<slug>.meta.yaml` sibling yet, it stages two file changes
create the sidecar (the entry's metadata, unknown keys preserved, INV-7)
and rewrite the `.md` to body-only (INV-2) and commits all of them in a
single ChangeFiles commit (§6.5: one commit per collection).
Idempotent: an entry that already has a sidecar is skipped; a second run
with nothing left to migrate makes no commit. Returns
`{"migrated": [...], "skipped": [...], "committed": bool}`.
"""
rfcs_dir = f"{subfolder}/rfcs" if subfolder else "rfcs"
listing = await gitea.list_dir(org, repo, rfcs_dir, ref=branch)
names = {f.get("name") for f in listing if f.get("type") == "file"}
ops: list[dict[str, Any]] = []
migrated: list[str] = []
skipped: list[str] = []
for f in listing:
if f.get("type") != "file" or not f.get("name", "").endswith(".md"):
continue
slug = f["name"][:-len(".md")]
if sidecar_name(slug) in names:
skipped.append(slug) # already migrated
continue
result = await gitea.read_file(org, repo, f["path"], ref=branch)
if not result:
continue
text, sha = result
try:
e = entry_mod.parse(text)
except ValueError:
# No frontmatter to lift (e.g. an already-clean body without a
# sidecar) — nothing to migrate; leave it untouched.
skipped.append(slug)
continue
body = strip_frontmatter(text)
new_md = body if (body == "" or body.endswith("\n")) else body + "\n"
ops.append({
"operation": "create",
"path": f"{rfcs_dir}/{sidecar_name(slug)}",
"content": sidecar_yaml(e),
})
ops.append({
"operation": "update",
"path": f["path"],
"content": new_md,
"sha": sha,
})
migrated.append(slug)
committed = False
if ops:
n = len(migrated)
message = f"Migrate {n} entr{'y' if n == 1 else 'ies'} to metadata sidecars (§22.4a)"
kwargs: dict[str, Any] = {}
if actor is not None:
kwargs = {
"author_name": actor.display_name,
"author_email": actor.email or f"{actor.gitea_login}@users.noreply",
}
await gitea.change_files(
org, repo, files=ops, message=message, branch=branch, **kwargs
)
committed = True
return {"migrated": migrated, "skipped": skipped, "committed": committed}
+147
View File
@@ -0,0 +1,147 @@
"""§22.4a configurable collection metadata — field schema + central validation.
SLICE-2 of docs/design/2026-06-06-configurable-collection-metadata.md.
A collection declares a small **field schema** in its `.collection.yaml`
(`fields:` block) so its entries can carry structured metadata priority, tags,
and any custom fields the deployment defines. This module is the **one place**
that knows a collection's field shapes (modeled on `registry.py`):
- `parse_fields` normalize + validate the declared schema, leniently: a bad
block or a bad field def is skipped with a warning, never raised, so a typo
in one field can't nuke the collection mirror (INV-3 spirit). The normalized
schema is a plain, JSON-serializable mapping that rides in
`collections.config_json` (no DB migration) and is served verbatim by the
collection API.
- `validate` check an entry's stored values against the schema, returning a
list of advisory `Problem`s. Empty list = clean. Used **advisory at read**
(the corpus mirror flags a non-empty result as `metadata_malformed`, INV-3)
and is the enforcement point at the **write** boundary (the metadata-edit
endpoints land in SLICE-4/5).
Field types (v1): `enum` (single scalar, controlled by a required `values:`
list), `tags` (a list; free-form unless `values:` given), `text` (a free
string). `ref` / `multi-enum` are future (design §2, Q2). Unknown types are
ignored with a warning. Keys an entry carries that the schema does **not**
declare ride along untouched and are never flagged (INV-7).
"""
from __future__ import annotations
import logging
from dataclasses import dataclass
from typing import Any
log = logging.getLogger(__name__)
# §6.3 v1 field types. `ref` (typed cross-entry link) and `multi-enum` are
# deferred (design §2, Q2) — declared with an unknown type they're skipped.
VALID_FIELD_TYPES = {"enum", "tags", "text"}
@dataclass
class Problem:
"""One advisory schema-validation problem against a declared field.
`code` is a stable machine token (`not-in-values`, `wrong-type`); `message`
is human-facing (surfaced at the write boundary in SLICE-4/5)."""
field: str
code: str
message: str
def as_dict(self) -> dict[str, str]:
return {"field": self.field, "code": self.code, "message": self.message}
# ----- schema parsing (lenient) -----
def parse_fields(raw: Any) -> dict[str, dict]:
"""Normalize a `.collection.yaml` `fields:` block → `{name: {type, ...}}`.
Pure (no I/O). Lenient (INV-3): a non-mapping block yields `{}`; an
individual field def that is not a mapping, has an unknown/missing `type`, or
is an `enum` without a non-empty `values:` list is **skipped with a warning**
never raised. Order is preserved (facet display order, SLICE-3). The
result is plain dicts so it serializes straight into `config_json` and the
collection API.
"""
if not isinstance(raw, dict):
if raw is not None:
log.warning("metadata_schema: fields block is not a mapping (%s); ignoring",
type(raw).__name__)
return {}
out: dict[str, dict] = {}
for name, spec in raw.items():
if not isinstance(spec, dict):
log.warning("metadata_schema: field %r def is not a mapping; skipping", name)
continue
ftype = str(spec.get("type") or "").strip()
if ftype not in VALID_FIELD_TYPES:
log.warning("metadata_schema: field %r has unknown type %r; skipping",
name, ftype)
continue
values = spec.get("values")
norm_values: list[str] | None = None
if values is not None:
if not isinstance(values, list):
log.warning("metadata_schema: field %r values is not a list; ignoring",
name)
else:
norm_values = [str(v) for v in values]
if ftype == "enum" and not norm_values:
log.warning("metadata_schema: enum field %r needs a non-empty values "
"list; skipping", name)
continue
field_def: dict[str, Any] = {"type": ftype}
if norm_values is not None:
field_def["values"] = norm_values
label = spec.get("label")
if label:
field_def["label"] = str(label)
out[str(name)] = field_def
return out
# ----- value validation (advisory) -----
def _is_scalar(v: Any) -> bool:
return isinstance(v, (str, int, float, bool))
def validate(values: dict[str, Any], fields: dict[str, dict]) -> list[Problem]:
"""Check an entry's metadata `values` against a collection's field schema.
Returns advisory `Problem`s (empty = clean). Only **declared** fields are
checked; a field the entry omits is fine (no required fields in v1), and a
key the schema doesn't declare rides along untouched (INV-7). With an empty
schema, everything is clean (INV-5). Never raises (INV-3).
"""
problems: list[Problem] = []
for name, spec in fields.items():
if name not in values:
continue
value = values[name]
if value is None:
continue
ftype = spec.get("type")
allowed = spec.get("values")
if ftype == "enum":
if not _is_scalar(value):
problems.append(Problem(name, "wrong-type",
f"{name!r} must be a single value, got {type(value).__name__}"))
elif allowed is not None and str(value) not in allowed:
problems.append(Problem(name, "not-in-values",
f"{name!r} value {value!r} is not one of {allowed}"))
elif ftype == "tags":
if not isinstance(value, list):
problems.append(Problem(name, "wrong-type",
f"{name!r} must be a list, got {type(value).__name__}"))
elif allowed is not None:
for member in value:
if str(member) not in allowed:
problems.append(Problem(name, "not-in-values",
f"{name!r} value {member!r} is not one of {allowed}"))
elif ftype == "text":
if not _is_scalar(value):
problems.append(Problem(name, "wrong-type",
f"{name!r} must be a string, got {type(value).__name__}"))
return problems
+62 -6
View File
@@ -38,22 +38,78 @@ from . import db, funder
from .providers import BaseProvider
def _models_from_config(config_json: str | None) -> list[str] | None:
"""The `enabled_models` list inside a project/collection `config_json`,
or None when the key is absent (meaning "no narrowing at this tier")."""
if not config_json:
return None
try:
cfg = json.loads(config_json)
except (json.JSONDecodeError, TypeError):
return None
em = cfg.get("enabled_models") if isinstance(cfg, dict) else None
return [str(m) for m in em] if isinstance(em, list) else None
def _narrow(universe: list[str], allowed: list[str] | None) -> list[str]:
"""Intersect `universe` with `allowed`, preserving universe order. `allowed`
None means no narrowing at this tier; an empty list narrows to empty (an
opt-out), exactly like the §6.6 per-entry `models: []`."""
if allowed is None:
return universe
allow = set(allowed)
return [k for k in universe if k in allow]
def _scope_narrowed_universe(
collection_id: str | None, operator_keys: list[str]
) -> list[str]:
"""§22.12 — narrow the operator (deployment) universe by the entry's
project then its collection `enabled_models`. Each tier may only narrow;
a missing config at a tier is a no-op. The collection cannot widen its
project because narrowing composes from the operator ceiling downward."""
if collection_id is None:
return list(operator_keys)
conn = db.conn()
crow = conn.execute(
"SELECT project_id, config_json FROM collections WHERE id = ?",
(collection_id,),
).fetchone()
universe = list(operator_keys)
if crow is None:
return universe
prow = conn.execute(
"SELECT config_json FROM projects WHERE id = ?", (crow["project_id"],)
).fetchone()
universe = _narrow(universe, _models_from_config(prow["config_json"] if prow else None))
universe = _narrow(universe, _models_from_config(crow["config_json"]))
return universe
def resolve_models_for_rfc(
slug: str, providers: dict[str, BaseProvider]
) -> list[str]:
"""Return the per-RFC resolved model keys per §6.6, extended by §6.7.
"""Return the per-RFC resolved model keys per §6.6, extended by §6.7 and
§22.12.
The first entry is the RFC's default model. An empty list means
AI is unavailable on this RFC and callers refuse the AI surface.
"""
# §6.7: the funder universe (if any) replaces the operator universe
# as the base set the §6.6 frontmatter intersects against.
funder_universe = funder.resolve_funder_universe(slug, providers)
base_universe = funder_universe if funder_universe is not None else list(providers.keys())
row = db.conn().execute(
"SELECT models_json FROM cached_rfcs WHERE slug = ?",
"SELECT collection_id, models_json FROM cached_rfcs WHERE slug = ?",
(slug,),
).fetchone()
collection_id = row["collection_id"] if row is not None else None
# §22.12: first narrow the operator universe by the entry's project +
# collection enabled_models (the deployment → project → collection chain).
scope_universe = _scope_narrowed_universe(collection_id, list(providers.keys()))
# §6.7: a consenting funder universe (if any) replaces the operator universe
# as the base set — still bounded by the §22.12 scope narrowing above.
funder_universe = funder.resolve_funder_universe(slug, providers)
if funder_universe is not None:
base_universe = _narrow(list(funder_universe), scope_universe)
else:
base_universe = scope_universe
if row is None or row["models_json"] is None:
return list(base_universe)
try:
+196
View File
@@ -324,6 +324,48 @@ def fan_out_contribution_request(
return notif_ids
def notify_scope_role_granted(
*,
recipient_user_id: int,
granter_user_id: int | None,
scope_type: str,
scope_id: str,
role: str,
project_id: str | None,
project_name: str | None,
collection_name: str | None,
) -> int | None:
"""§22 S4 (C.2): a scope Owner granted `recipient` a role at a scope.
Personal-direct the recipient is the named subject so it rides the
`email_personal_direct` gate like the other owner-facing personal events.
The scope facts ride in the payload so the inbox row (and email body) names
the project and role without a second fetch. Actor is the granter (§15.9);
a system/administrative grant with no granter renders as "the app".
Returns the notification id, or None when the grantee would be notifying
themselves (a self-grant no notification)."""
if granter_user_id is not None and recipient_user_id == granter_user_id:
return None
details = {
"scope_type": scope_type,
"scope_id": scope_id,
"role": role,
"project_id": project_id or "",
"project_name": project_name or "",
"collection_name": collection_name or "",
}
return _emit_one(
recipient_user_id=recipient_user_id,
event_kind="scope_role_granted",
category=CATEGORY_PERSONAL,
actor_user_id=granter_user_id,
rfc_slug=None,
branch_name=None,
pr_number=None,
details=details,
)
def notify_contribution_decided(
*,
rfc_slug: str,
@@ -350,6 +392,95 @@ def notify_contribution_decided(
)
def fan_out_join_request(
*,
scope_type: str,
scope_id: str,
scope_name: str | None,
project_id: str | None,
project_name: str | None,
requester_user_id: int,
request_id: int,
requested_role: str,
message: str | None,
) -> list[int]:
"""§22.8: a user asked to join a scope. Land one actionable notification per
Owner across the scope's subtree (the cross-collection inbox, §22.11) and
return their ids (the caller stamps the first onto the request row as the
inbox-action handle any of them can act on it).
Personal-direct: each Owner is a named subject able to act, so it rides the
`email_personal_direct` gate like the other owner-facing personal events. The
requested role + message ride in the payload so the inbox row shows the full
ask inline. Actor is the requester per §15.9.
"""
requester = db.conn().execute(
"SELECT display_name FROM users WHERE id = ?", (requester_user_id,)
).fetchone()
display = (requester["display_name"] if requester else None) or "Someone"
details = {
"request_id": request_id,
"scope_type": scope_type,
"scope_id": scope_id,
"scope_name": scope_name or scope_id,
"project_id": project_id or "",
"project_name": project_name or "",
"requested_role": requested_role,
"requester_user_id": requester_user_id,
"requester_display": display,
"message": message or "",
}
notif_ids: list[int] = []
for recipient_id in _scope_owner_user_ids(scope_type, scope_id):
if recipient_id == requester_user_id:
continue
notif_ids.append(
_emit_one(
recipient_user_id=recipient_id,
event_kind="join_request_on_scope",
category=CATEGORY_PERSONAL,
actor_user_id=requester_user_id,
rfc_slug=None,
branch_name=None,
pr_number=None,
details=details,
)
)
return notif_ids
def notify_join_decided(
*,
requester_user_id: int,
decider_user_id: int,
request_id: int,
scope_type: str,
scope_id: str,
scope_name: str | None,
granted_role: str | None,
accepted: bool,
) -> None:
"""§22.8: tell the requester an Owner accepted (writing their `memberships`
row) or declined their request to join. The scope + granted role ride in the
payload so the inbox row names where they were let in without a second fetch."""
_emit_one(
recipient_user_id=requester_user_id,
event_kind=("join_request_accepted" if accepted else "join_request_declined"),
category=CATEGORY_PERSONAL,
actor_user_id=decider_user_id,
rfc_slug=None,
branch_name=None,
pr_number=None,
details={
"request_id": request_id,
"scope_type": scope_type,
"scope_id": scope_id,
"scope_name": scope_name or scope_id,
"granted_role": granted_role or "",
},
)
def fan_out_chat_message(
*,
actor_user_id: int,
@@ -625,6 +756,41 @@ def _admin_user_ids() -> set[int]:
}
def _scope_owner_user_ids(scope_type: str, scope_id: str) -> set[int]:
"""The Owners who administer a scope *across the subtree* (§22.8 / §22.11) —
the recipients of a request-to-join, aggregated upward so the request reaches
everyone who could grant it. For a `collection`: its collection-scope Owners,
its project's Owners, the global Owners, and deployment owners/admins. For a
`project`: its project-scope Owners plus global Owners and deployment
owners/admins. (Mirrors the upward fold in `auth.can_invite_at_*`.)"""
# Deployment owners/admins are global Owners by §B.1; explicit
# scope_type='global' Owner grants join them.
ids: set[int] = set(_admin_user_ids())
for r in db.conn().execute(
"SELECT user_id AS id FROM memberships WHERE scope_type = 'global' AND role = 'owner'"
):
ids.add(r["id"])
def _owners_at(stype: str, sid: str) -> None:
for r in db.conn().execute(
"SELECT user_id AS id FROM memberships "
"WHERE scope_type = ? AND scope_id = ? AND role = 'owner'",
(stype, sid),
):
ids.add(r["id"])
if scope_type == "collection":
_owners_at("collection", scope_id)
prow = db.conn().execute(
"SELECT project_id FROM collections WHERE id = ?", (scope_id,)
).fetchone()
if prow and prow["project_id"]:
_owners_at("project", prow["project_id"])
elif scope_type == "project":
_owners_at("project", scope_id)
return ids
def _proposer_user_id(rfc_slug: str) -> set[int]:
row = db.conn().execute(
"""
@@ -859,6 +1025,36 @@ def render_summary(event_kind: str, actor_display: str | None, rfc_title: str |
return f"{actor} accepted your request to contribute to {title} — check your email to accept the invitation."
if event_kind == "contribution_request_declined":
return f"{actor} declined your request to contribute to {title}."
if event_kind == "scope_role_granted":
# §22 S4 (C.2): names the role and the scope (the project, and the
# collection when collection-scoped) per "a §15 notification naming the
# project and role".
role_label = "Owner" if extras.get("role") == "owner" else "RFC Contributor"
project_label = extras.get("project_name") or extras.get("project_id") or "a project"
scope_type = extras.get("scope_type")
if scope_type == "collection":
col_label = extras.get("collection_name") or extras.get("scope_id") or "a collection"
return f"{actor} granted you {role_label} on collection {project_label}/{col_label}."
if scope_type == "global":
return f"{actor} granted you {role_label} across the whole deployment."
return f"{actor} granted you {role_label} on project {project_label}."
if event_kind == "join_request_on_scope":
# §22.8: owner-facing, actionable. Names who wants in, where, and as
# what; the inbox row renders Accept/Decline beneath this line.
role_label = "Owner" if extras.get("requested_role") == "owner" else "RFC Contributor"
scope_type = extras.get("scope_type")
scope_label = extras.get("scope_name") or extras.get("scope_id") or "a scope"
where = (
f"collection {scope_label}" if scope_type == "collection" else f"project {scope_label}"
)
return f"{actor} asked to join {where} as {role_label}."
if event_kind == "join_request_accepted":
role_label = "Owner" if extras.get("granted_role") == "owner" else "RFC Contributor"
scope_label = extras.get("scope_name") or extras.get("scope_id") or "the scope"
return f"{actor} accepted your request to join {scope_label} — you're in as {role_label}."
if event_kind == "join_request_declined":
scope_label = extras.get("scope_name") or extras.get("scope_id") or "the scope"
return f"{actor} declined your request to join {scope_label}."
if event_kind == "new_beta_request":
# v0.9.0: framework-scoped, not RFC-scoped. The actor (the
# requester) and the captured full name + email read as
+207 -7
View File
@@ -7,9 +7,13 @@ the registry mirror (`registry.refresh_registry`) is authoritative.
"""
from __future__ import annotations
import logging
from . import db
from .config import Config
log = logging.getLogger(__name__)
DEFAULT_PROJECT_ID = "default"
@@ -20,6 +24,163 @@ def resolved_default_id(config: Config) -> str:
return config.default_project_id.strip() or DEFAULT_PROJECT_ID
def restamp_default_project(config: Config) -> None:
"""§22.13 step 1 — one-time rename of the M1 bootstrap project id
(DEFAULT_PROJECT_ID = 'default') to the deployment's configured default id
(the DEFAULT_PROJECT_ID env var, e.g. 'ohm'), so the deployment's original
corpus lands at a meaningful `/p/<id>/` and `default` is never a public URL.
Renames `project_id` across every project-scoped table (discovered by
column, so it stays correct as the schema grows), then drops the stale
bootstrap `projects` row (its data has moved to the configured row, which
the registry mirror already created). Idempotent and a no-op when the
configured id is still 'default' or no bootstrap rows remain. Runs at
startup after the registry mirror, with FK enforcement off for the rename
(the composite FKs are kept consistent because parent and child rows are
renamed together) and a foreign_key_check backstop before commit.
"""
target = resolved_default_id(config)
if target == DEFAULT_PROJECT_ID:
return
conn = db.conn()
# §22 three-tier: the entry-corpus tables key on collection_id now; detect a
# lingering bootstrap project by the project-grain `collections.project_id`
# (the PRAGMA scan below still renames every project_id column dynamically).
has_rows = conn.execute(
"SELECT 1 FROM collections WHERE project_id = ? LIMIT 1", (DEFAULT_PROJECT_ID,)
).fetchone()
stale_proj = conn.execute(
"SELECT 1 FROM projects WHERE id = ? LIMIT 1", (DEFAULT_PROJECT_ID,)
).fetchone()
if not has_rows and not stale_proj:
return
if conn.execute("SELECT 1 FROM projects WHERE id = ? LIMIT 1", (target,)).fetchone() is None:
log.warning("restamp: target project %r not in registry yet; skipping", target)
return
tables = [r["name"] for r in conn.execute("SELECT name FROM sqlite_master WHERE type='table'")]
pid_tables = [
t for t in tables
if any(c["name"] == "project_id" for c in conn.execute(f"PRAGMA table_info({t})"))
]
conn.execute("PRAGMA foreign_keys = OFF")
try:
conn.execute("BEGIN")
for t in pid_tables:
conn.execute(
f"UPDATE {t} SET project_id = ? WHERE project_id = ?",
(target, DEFAULT_PROJECT_ID),
)
# The bootstrap row's data has moved to the configured (registry) row.
conn.execute("DELETE FROM projects WHERE id = ?", (DEFAULT_PROJECT_ID,))
violations = conn.execute("PRAGMA foreign_key_check").fetchall()
if violations:
conn.execute("ROLLBACK")
raise RuntimeError(
f"restamp left foreign-key violations: {[tuple(v) for v in violations]}"
)
conn.execute("COMMIT")
except Exception:
try:
conn.execute("ROLLBACK")
except Exception:
pass
raise
finally:
conn.execute("PRAGMA foreign_keys = ON")
log.info("restamp: renamed bootstrap project %r -> %r across %d tables",
DEFAULT_PROJECT_ID, target, len(pid_tables))
def reconcile_default_collection_id(config: Config) -> None:
"""Heal the §22 migration-029 vs registry-mirror divergence for the default
project's collection id on a multi-project deployment.
Migration 029 seeds each project's default collection id as the literal
'default' only when the DB holds a single project at migration time; with
2 projects it falls back to the *project id* (avoiding a PK collision
029 can't read DEFAULT_PROJECT_ID, there is no env in SQL). But the registry
mirror (`registry._default_collection_id`) expects the deployment's default
project to own the collection id 'default'. On an upgrade whose DB already
held 2 projects when 029 ran, the default project's collection is therefore
named after the project (e.g. 'ohm'), and the next mirror would INSERT a
second, empty 'default' collection instead of merging duplicating the
default corpus and orphaning the entries (which point at 'ohm').
This is the collection-grain twin of `restamp_default_project`. Run at
startup BEFORE the registry mirror so the canonical 'default' collection
already exists when the mirror upserts (merge, not duplicate). Renames the
divergent collection's id to 'default' and cascades `collection_id` across
every collection-keyed table, with FK enforcement off for the atomic rename
and a `foreign_key_check` backstop before commit. Idempotent; a no-op on
fresh / single-project / already-aligned deployments.
"""
from .collections import DEFAULT_COLLECTION_ID
target = resolved_default_id(config)
if target == DEFAULT_COLLECTION_ID: # default project already owns 'default'
return
conn = db.conn()
# The 029 ≥2-projects seed names the default project's collection after the
# project itself; the canonical id the mirror expects is 'default'.
divergent = conn.execute(
"SELECT 1 FROM collections WHERE id = ? AND project_id = ? LIMIT 1",
(target, target),
).fetchone()
if not divergent:
return
if conn.execute(
"SELECT 1 FROM collections WHERE id = ? LIMIT 1", (DEFAULT_COLLECTION_ID,)
).fetchone():
# A 'default' collection already exists (e.g. a prior buggy mirror left a
# duplicate). Don't auto-merge data — that needs care; leave both for
# operator cleanup and log loudly.
log.warning(
"reconcile: default project %r owns both a %r and a 'default' "
"collection; skipping auto-rename (manual merge required)",
target, target,
)
return
cid_tables = [
t["name"]
for t in conn.execute("SELECT name FROM sqlite_master WHERE type='table'")
if any(c["name"] == "collection_id"
for c in conn.execute(f"PRAGMA table_info({t['name']})"))
]
conn.execute("PRAGMA foreign_keys = OFF")
try:
conn.execute("BEGIN")
conn.execute(
"UPDATE collections SET id = ? WHERE id = ?",
(DEFAULT_COLLECTION_ID, target),
)
for t in cid_tables:
conn.execute(
f"UPDATE {t} SET collection_id = ? WHERE collection_id = ?",
(DEFAULT_COLLECTION_ID, target),
)
violations = conn.execute("PRAGMA foreign_key_check").fetchall()
if violations:
conn.execute("ROLLBACK")
raise RuntimeError(
f"reconcile left foreign-key violations: {[tuple(v) for v in violations]}"
)
conn.execute("COMMIT")
except Exception:
try:
conn.execute("ROLLBACK")
except Exception:
pass
raise
finally:
conn.execute("PRAGMA foreign_keys = ON")
log.info(
"reconcile: renamed default-project collection %r -> 'default' across %d tables",
target, len(cid_tables),
)
def default_content_repo(config: Config) -> str | None:
"""The content repo the single-corpus mirror reads, from the default
project's row (filled by the registry mirror). Replaces the retired
@@ -31,12 +192,51 @@ def default_content_repo(config: Config) -> str | None:
return row["content_repo"] if row and row["content_repo"] else None
def project_initial_state(project_id: str) -> str:
"""§22.4b landing state for new entries in a project. Defaults to
'super-draft' for an unknown/unset row (the safe, today's-flow default)."""
def content_repo(project_id: str) -> str | None:
"""The content repo for a specific project (§22.3). None if unknown/unset.
The per-project successor to `default_content_repo` for the write path."""
row = db.conn().execute(
"SELECT initial_state FROM projects WHERE id = ?", (project_id,)
"SELECT content_repo FROM projects WHERE id = ?", (project_id,)
).fetchone()
if row is None or not row["initial_state"]:
return "super-draft"
return row["initial_state"]
return row["content_repo"] if row and row["content_repo"] else None
def content_repo_for_collection(collection_id: str) -> str | None:
"""The content repo a collection's entries live in (§22 three-tier write
path, G-15): collection project content_repo. None if the collection or
its project is unknown/unset. The per-collection successor to
`default_content_repo` for the WRITE path an entry in a non-default
project must read/write that project's repo, not the deployment default."""
from . import collections as collections_mod
pid = collections_mod.project_of_collection(collection_id)
return content_repo(pid) if pid else None
def entry_location(config: Config, collection_id: str, slug: str) -> tuple[str, str, str]:
"""The git location `(gitea_org, content_repo, md_path)` of an entry, resolved
from its collection (§22 three-tier, G-15).
Repo: the collection's project content_repo, falling back to the deployment
default project's repo when the collection (or its project) is unknown — so a
legacy/single-corpus entry still resolves to a usable location rather than an
empty repo. Path: `<subfolder>/rfcs/<slug>.md`, or `rfcs/<slug>.md` at the
repo root for a default (subfolder-less) collection.
This is the single resolver the branch/edit/body/metadata/graduation write
paths share, replacing the hardcoded `default_content_repo` + `rfcs/<slug>.md`.
"""
from . import collections as collections_mod
repo = content_repo_for_collection(collection_id) or (default_content_repo(config) or "")
sub = collections_mod.subfolder_of(collection_id)
rfcs_dir = f"{sub}/rfcs" if sub else "rfcs"
return config.gitea_org, repo, f"{rfcs_dir}/{slug}.md"
def project_initial_state(project_id: str) -> str:
"""§22.4b landing state for new entries in a project's default collection
(the per-corpus field moved down to the collection in migration 029).
Defaults to 'super-draft' for an unknown/unset row (today's-flow default)."""
from . import collections as collections_mod
return collections_mod.collection_initial_state(
collections_mod.default_collection_id(project_id)
)
+22 -3
View File
@@ -19,11 +19,27 @@ lockouts). Both layers run together.
"""
from __future__ import annotations
import os
import threading
import time
from collections import defaultdict, deque
def _max_events(env_name: str, default: int) -> int:
"""Per-limiter budget, overridable via env (e.g. a test/PPE stack that
drives the auth endpoints repeatedly from one IP). Production leaves these
unset and gets the secure defaults below. A non-positive / unparseable
value falls back to the default."""
raw = os.environ.get(env_name, "").strip()
if not raw:
return default
try:
n = int(raw)
except ValueError:
return default
return n if n > 0 else default
class SlidingWindowLimiter:
"""Allow at most `max_events` per `window_seconds` per key.
@@ -66,12 +82,15 @@ class SlidingWindowLimiter:
# * verify: 10 attempts / 5 min / IP across the auth verify surfaces.
# * otc request: 5 sends / 5 min / IP (Turnstile is the primary gate;
# this is defense in depth against a solved-challenge replay loop).
verify_limiter = SlidingWindowLimiter(max_events=10, window_seconds=300)
otc_request_limiter = SlidingWindowLimiter(max_events=5, window_seconds=300)
verify_limiter = SlidingWindowLimiter(
max_events=_max_events("RATELIMIT_VERIFY_MAX", 10), window_seconds=300)
otc_request_limiter = SlidingWindowLimiter(
max_events=_max_events("RATELIMIT_OTC_REQUEST_MAX", 5), window_seconds=300)
# /auth/passcode/check is an anonymous has-passcode oracle (audit 0026 L3).
# It's a legitimate Login-flow affordance, so the budget is generous —
# enough for a human typing emails, tight enough to stop bulk scraping.
check_limiter = SlidingWindowLimiter(max_events=30, window_seconds=300)
check_limiter = SlidingWindowLimiter(
max_events=_max_events("RATELIMIT_CHECK_MAX", 30), window_seconds=300)
def _reset_all_for_tests() -> None:
+193 -20
View File
@@ -18,6 +18,7 @@ from dataclasses import dataclass, field
import yaml
from . import db
from . import metadata_schema
from .config import Config
from .gitea import Gitea
@@ -55,6 +56,18 @@ class ProjectEntry:
config: dict = field(default_factory=dict) # theme, enabled_models
@dataclass
class CollectionEntry:
"""A named collection declared by a `.collection.yaml` manifest inside a
project's content repo (S2). `visibility=None` means "inherit the project's
visibility"."""
type: str
visibility: str | None
initial_state: str
name: str | None
config: dict = field(default_factory=dict) # §22.12 enabled_models
@dataclass
class RegistryDoc:
deployment_name: str
@@ -114,44 +127,117 @@ def parse_registry(text: str) -> RegistryDoc:
)
def apply_registry(doc: RegistryDoc, registry_sha: str) -> None:
"""Upsert the parsed registry into projects + deployment. Idempotent.
def parse_collection_manifest(text: str) -> CollectionEntry:
"""Parse + validate a `.collection.yaml`. Pure (no I/O). Raises RegistryError.
§22.4a: `type` is immutable a change against an existing row is rejected
(skip + log), never applied. Projects absent from the registry are left in
place (archival is out of scope for M3; they simply stop refreshing).
`type` is required and immutable (§22.4a, enforced at upsert). `visibility`
is optional omitted means inherit the project's. `initial_state` defaults
per type (§22.4b)."""
raw = yaml.safe_load(text) or {}
if not isinstance(raw, dict):
raise RegistryError("collection manifest must be a mapping")
ctype = str(raw.get("type") or "").strip()
if ctype not in VALID_TYPES:
raise RegistryError(f"collection has invalid type {ctype!r}")
vis = raw.get("visibility")
if vis is not None:
vis = str(vis).strip()
if vis not in VALID_VISIBILITY:
raise RegistryError(f"collection has invalid visibility {vis!r}")
initial_state = str(
raw.get("initial_state") or _TYPE_DEFAULT_INITIAL_STATE[ctype]
).strip()
if initial_state not in VALID_INITIAL_STATE:
raise RegistryError(f"collection has invalid initial_state {initial_state!r}")
name = raw.get("name")
name = str(name).strip() if name else None
# §22.12: an optional per-collection enabled_models list that narrows the
# project's universe. Absent → no narrowing (inherit). Present (incl. empty)
# → narrowing applies; [] opts the collection out of AI.
cfg: dict = {}
if raw.get("enabled_models") is not None:
em = raw["enabled_models"]
if not isinstance(em, list):
raise RegistryError("collection enabled_models must be a list")
cfg["enabled_models"] = [str(m) for m in em]
# §22.4a SLICE-2: the collection's metadata field schema. Parsed leniently —
# a bad field def is skipped with a warning, never fatal (INV-3), so a typo
# in one field can't drop the whole collection from the mirror. Stored only
# when at least one valid field survives.
fields = metadata_schema.parse_fields(raw.get("fields"))
if fields:
cfg["fields"] = fields
return CollectionEntry(ctype, vis, initial_state, name, cfg)
def _default_collection_id(project_id: str, default_id: str) -> str:
"""The id of a project's default collection. The deployment's primary
project (== `default_id`, the §22.13 resolved default) gets the stable
literal `'default'` matching migration 029's seed so the upsert *merges*
onto the migration-seeded row rather than duplicating it (critical on a
fresh deploy where the bootstrap `default` project is later restamped to the
configured id). Any additional project keys its default collection by its own
id, keeping the collection PK globally unique (pre-S5 multi-project)."""
return "default" if project_id == default_id else project_id
def apply_registry(doc: RegistryDoc, registry_sha: str, default_id: str) -> None:
"""Upsert the parsed registry into projects + their default collections +
the deployment singleton. Idempotent.
§22 three-tier (S1): a project carries the grouping-tier fields (name,
content_repo, visibility, config); the per-corpus fields (`type`,
`initial_state`) live on the project's default collection. §22.4a: `type` is
immutable a change against an existing collection is rejected (skip the
type change + log), never applied. Projects absent from the registry are
left in place (archival is out of scope for M3; they stop refreshing).
"""
with db.tx() as conn:
for e in doc.projects:
cid = _default_collection_id(e.id, default_id)
existing = conn.execute(
"SELECT type FROM projects WHERE id = ?", (e.id,)
"SELECT type FROM collections WHERE id = ?", (cid,)
).fetchone()
if existing is not None and existing["type"] != e.type:
type_locked = existing is not None and existing["type"] != e.type
if type_locked:
log.error(
"registry: refusing immutable type change on project %s (%s -> %s)",
e.id, existing["type"], e.type,
"registry: refusing immutable type change on collection %s (%s -> %s)",
cid, existing["type"], e.type,
)
continue
# The project (grouping tier) always refreshes.
conn.execute(
"""
INSERT INTO projects
(id, name, type, content_repo, visibility, initial_state,
config_json, registry_sha, updated_at)
VALUES (?, ?, ?, ?, ?, ?, ?, ?, datetime('now'))
(id, name, content_repo, visibility, config_json, registry_sha, updated_at)
VALUES (?, ?, ?, ?, ?, ?, datetime('now'))
ON CONFLICT(id) DO UPDATE SET
name = excluded.name,
type = excluded.type,
content_repo = excluded.content_repo,
visibility = excluded.visibility,
initial_state = excluded.initial_state,
config_json = excluded.config_json,
registry_sha = excluded.registry_sha,
updated_at = datetime('now')
""",
(
e.id, e.name, e.type, e.content_repo, e.visibility,
e.initial_state, json.dumps(e.config), registry_sha,
),
(e.id, e.name, e.content_repo, e.visibility, json.dumps(e.config), registry_sha),
)
# The default collection (corpus tier). On an immutable-type
# conflict, keep the stored type but still refresh the rest.
effective_type = existing["type"] if type_locked else e.type
conn.execute(
"""
INSERT INTO collections
(id, project_id, type, subfolder, initial_state, visibility, name, registry_sha, updated_at)
VALUES (?, ?, ?, '', ?, ?, ?, ?, datetime('now'))
ON CONFLICT(id) DO UPDATE SET
project_id = excluded.project_id,
type = excluded.type,
initial_state = excluded.initial_state,
visibility = excluded.visibility,
name = excluded.name,
registry_sha = excluded.registry_sha,
updated_at = datetime('now')
""",
(cid, e.id, effective_type, e.initial_state, e.visibility, e.name, registry_sha),
)
conn.execute(
"""
@@ -163,6 +249,90 @@ def apply_registry(doc: RegistryDoc, registry_sha: str) -> None:
)
def _strictest_visibility(a: str, b: str) -> str:
"""The stricter of two §22.5 visibilities on the public-exposure axis
(`public` < `unlisted` < `gated`). Used to enforce that a collection is set
only as strict or stricter than its project (S3 operator decision)."""
rank = {"public": 0, "unlisted": 1, "gated": 2}
return a if rank.get(a, 2) >= rank.get(b, 2) else b
def _upsert_named_collection(
proj: ProjectEntry, subdir: str, ce: CollectionEntry, sha: str
) -> None:
"""Upsert one named collection (S2). Type is immutable (§22.4a): a type
change against an existing row is refused (logged, not applied). A None
manifest visibility inherits the project's visibility; a manifest that tries
to be *looser* than its project is clamped to the project's (S3 strictness:
a collection may narrow but never widen its project's visibility)."""
requested = ce.visibility or proj.visibility
visibility = _strictest_visibility(requested, proj.visibility)
if visibility != requested:
log.warning(
"registry: collection %s visibility %r looser than project %s %r"
"clamped to %r (S3 strictness)",
subdir, requested, proj.id, proj.visibility, visibility,
)
with db.tx() as conn:
existing = conn.execute(
"SELECT type FROM collections WHERE id = ?", (subdir,)
).fetchone()
if existing is not None and existing["type"] != ce.type:
log.error(
"registry: refusing immutable type change on collection %s (%s -> %s)",
subdir, existing["type"], ce.type,
)
return
conn.execute(
"""
INSERT INTO collections
(id, project_id, type, subfolder, initial_state, visibility, name, config_json, registry_sha, updated_at)
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, datetime('now'))
ON CONFLICT(id) DO UPDATE SET
project_id = excluded.project_id,
initial_state = excluded.initial_state,
visibility = excluded.visibility,
name = excluded.name,
config_json = excluded.config_json,
registry_sha = excluded.registry_sha,
updated_at = datetime('now')
""",
(subdir, proj.id, ce.type, subdir, ce.initial_state, visibility, ce.name,
json.dumps(ce.config), sha),
)
async def _mirror_named_collections(config: Config, gitea: Gitea, doc: RegistryDoc, sha: str) -> None:
"""§22 S2: named collections are declared by `.collection.yaml` manifests
inside each project's content repo (the default collection comes from
projects.yaml). Walk each content repo root; a subdir carrying a manifest
becomes a collection keyed by the subdir name. Tolerant: a transport or
parse failure on one project/collection logs and is skipped, never aborts
the wider mirror (keep last-good)."""
for proj in doc.projects:
try:
items = await gitea.list_dir(config.gitea_org, proj.content_repo, "", ref="main")
except Exception as e: # noqa: BLE001 — GiteaError/transport: tolerate
log.warning("registry: cannot list %s root: %s", proj.content_repo, e)
continue
for it in items:
if it.get("type") != "dir":
continue
subdir = it["name"]
manifest = await gitea.get_contents(
config.gitea_org, proj.content_repo, f"{subdir}/.collection.yaml", ref="main"
)
if not manifest or manifest.get("type") != "file":
continue
mtext = base64.b64decode(manifest["content"]).decode("utf-8")
try:
ce = parse_collection_manifest(mtext)
except RegistryError as e:
log.error("registry: bad manifest %s/%s: %s", proj.content_repo, subdir, e)
continue
_upsert_named_collection(proj, subdir, ce, sha)
async def refresh_registry(config: Config, gitea: Gitea) -> None:
"""Mirror REGISTRY_REPO/projects.yaml into projects + deployment.
@@ -181,5 +351,8 @@ async def refresh_registry(config: Config, gitea: Gitea) -> None:
# includes it on the contents response); fall back to the blob sha.
sha = item.get("last_commit_sha") or item.get("sha") or ""
doc = parse_registry(text)
apply_registry(doc, sha)
from . import projects as projects_mod
apply_registry(doc, sha, projects_mod.resolved_default_id(config))
# §22 S2: discover + upsert named collections from each content repo.
await _mirror_named_collections(config, gitea, doc, sha)
log.info("registry: mirrored %d project(s) at %s", len(doc.projects), sha)
+12 -5
View File
@@ -80,10 +80,17 @@ def make_router(config: Config, gitea: Gitea) -> APIRouter:
payload = {}
repo_full = (payload.get("repository") or {}).get("full_name") or ""
registry_full = f"{config.gitea_org}/{config.registry_repo}"
content_repo = projects_mod.default_content_repo(config)
if not content_repo:
log.warning("webhook: default project content_repo is unknown; corpus refresh skipped")
content_full = f"{config.gitea_org}/{content_repo}" if content_repo else None
# §22/G-15: a corpus push can land on ANY project's content_repo, not
# just the default — recognise the full set so a non-default project's
# push triggers the (multi-project) corpus/branch/PR refresh.
content_fulls = {
f"{config.gitea_org}/{r['content_repo']}"
for r in db.conn().execute(
"SELECT content_repo FROM projects "
"WHERE content_repo IS NOT NULL AND content_repo != ''")
}
if not content_fulls:
log.warning("webhook: no project content_repo is known; corpus refresh skipped")
try:
if repo_full == registry_full:
# §22.2: a registry-repo push re-mirrors the projects table.
@@ -94,7 +101,7 @@ def make_router(config: Config, gitea: Gitea) -> APIRouter:
await registry_mod.refresh_registry(config, gitea)
except registry_mod.RegistryError:
log.exception("registry webhook: invalid projects.yaml; keeping last-good")
elif content_full and (repo_full == content_full or not repo_full):
elif content_fulls and (repo_full in content_fulls or not repo_full):
await cache.refresh_meta_repo(config, gitea)
await cache.refresh_meta_branches(config, gitea)
await cache.refresh_meta_pulls(config, gitea)
+474
View File
@@ -0,0 +1,474 @@
-- migrate:no-foreign-keys
--
-- §22 three-tier refactor — S1. Insert a *collection* grain beneath project.
--
-- (1) a `collections` table beneath `projects`;
-- (2) move the per-corpus fields (type, initial_state) down from `projects`
-- (projects keeps id, name, content_repo, visibility, config_json, …);
-- (3) one default collection per project (id='default' for the standard
-- single-project deployment, subfolder = repo root), inheriting the
-- project's type / initial_state / visibility;
-- (4) re-key the 13 entry-corpus tables (project_id, slug) -> (collection_id,
-- slug) via the migration-028 rebuild pattern, mapping each row to its
-- project's default collection by JOIN;
-- (5) generalise project_members -> memberships(scope_type ∈ {project,
-- collection}, scope_id, …), collapsing the role enum to {owner,
-- contributor} (§B.3).
--
-- SQLite can't ALTER a PK/UNIQUE in place, so each keyed table is rebuilt by the
-- official create-copy-drop-rename procedure. FK enforcement is OFF for the file
-- (the `migrate:no-foreign-keys` marker tells the runner to toggle it and run
-- foreign_key_check after). cached_rfcs is rebuilt FIRST so the child tables can
-- re-point their composite FK at its new (collection_id, slug) key.
--
-- The tables 026 tagged with project_id but 028 did NOT key (threads, changes,
-- notifications, actions, pr_resolution_branches, cached_prs) keep project_id —
-- they carry a project-grain tag, untouched in S1. See
-- docs/design/2026-06-05-three-tier-projects-collections.md §A.6 / Part E.
-- ── §22.13 repair: re-stamp stale satellite project_id before rekeying ──────
-- The §22.13 default→ohm re-stamp (v0.39.0, `projects.restamp_default_project`)
-- updated `cached_rfcs.project_id` but NOT the entry-satellite tables, leaving
-- rows with a stale `project_id` (e.g. 'default') that the per-project collection
-- backfill below cannot map — the subquery returns NULL and the NOT NULL rebuild
-- fails (`cached_branches__new.collection_id`). Before rebuilding, re-derive each
-- satellite's `project_id` from its entry (`cached_rfcs`, joined by slug — slugs
-- are unique per collection and, pre-rebuild, globally), and drop rows whose
-- entry no longer exists (stale cache; the `cached_*` tables are rebuildable from
-- gitea). On a clean/fresh deployment every satellite is empty or already
-- consistent, so this whole block is a no-op. (Discovered on the OHM data:
-- ~1.3k `cached_branches` rows stranded at project_id='default'.)
-- First drop stale rows that DUPLICATE an already-correctly-stamped row (the same
-- branch cached under both the stale and the real project_id) — re-stamping them
-- would collide on the (project_id, rfc_slug, branch_name) key. The correctly-
-- stamped copy is kept (it carries the current head_sha / visibility). Only the
-- branch-keyed tables can hold such a pair; the others key on (rfc_slug,user_id)
-- /(scope,pr_number) and have no stale data here, so they need no dedup.
DELETE FROM cached_branches WHERE project_id NOT IN (SELECT id FROM projects)
AND EXISTS (SELECT 1 FROM cached_branches o WHERE o.rfc_slug = cached_branches.rfc_slug AND o.branch_name = cached_branches.branch_name AND o.project_id IN (SELECT id FROM projects));
DELETE FROM branch_visibility WHERE project_id NOT IN (SELECT id FROM projects)
AND EXISTS (SELECT 1 FROM branch_visibility o WHERE o.rfc_slug = branch_visibility.rfc_slug AND o.branch_name = branch_visibility.branch_name AND o.project_id IN (SELECT id FROM projects));
UPDATE rfc_invitations SET project_id = (SELECT r.project_id FROM cached_rfcs r WHERE r.slug = rfc_invitations.rfc_slug) WHERE rfc_slug IN (SELECT slug FROM cached_rfcs);
DELETE FROM rfc_invitations WHERE rfc_slug NOT IN (SELECT slug FROM cached_rfcs);
UPDATE cached_branches SET project_id = (SELECT r.project_id FROM cached_rfcs r WHERE r.slug = cached_branches.rfc_slug) WHERE rfc_slug IN (SELECT slug FROM cached_rfcs);
DELETE FROM cached_branches WHERE rfc_slug NOT IN (SELECT slug FROM cached_rfcs);
UPDATE branch_visibility SET project_id = (SELECT r.project_id FROM cached_rfcs r WHERE r.slug = branch_visibility.rfc_slug) WHERE rfc_slug IN (SELECT slug FROM cached_rfcs);
DELETE FROM branch_visibility WHERE rfc_slug NOT IN (SELECT slug FROM cached_rfcs);
UPDATE branch_contribute_grants SET project_id = (SELECT r.project_id FROM cached_rfcs r WHERE r.slug = branch_contribute_grants.rfc_slug) WHERE rfc_slug IN (SELECT slug FROM cached_rfcs);
DELETE FROM branch_contribute_grants WHERE rfc_slug NOT IN (SELECT slug FROM cached_rfcs);
UPDATE stars SET project_id = (SELECT r.project_id FROM cached_rfcs r WHERE r.slug = stars.rfc_slug) WHERE rfc_slug IN (SELECT slug FROM cached_rfcs);
DELETE FROM stars WHERE rfc_slug NOT IN (SELECT slug FROM cached_rfcs);
UPDATE watches SET project_id = (SELECT r.project_id FROM cached_rfcs r WHERE r.slug = watches.rfc_slug) WHERE rfc_slug IN (SELECT slug FROM cached_rfcs);
DELETE FROM watches WHERE rfc_slug NOT IN (SELECT slug FROM cached_rfcs);
UPDATE pr_seen SET project_id = (SELECT r.project_id FROM cached_rfcs r WHERE r.slug = pr_seen.rfc_slug) WHERE rfc_slug IN (SELECT slug FROM cached_rfcs);
DELETE FROM pr_seen WHERE rfc_slug NOT IN (SELECT slug FROM cached_rfcs);
UPDATE branch_chat_seen SET project_id = (SELECT r.project_id FROM cached_rfcs r WHERE r.slug = branch_chat_seen.rfc_slug) WHERE rfc_slug IN (SELECT slug FROM cached_rfcs);
DELETE FROM branch_chat_seen WHERE rfc_slug NOT IN (SELECT slug FROM cached_rfcs);
UPDATE funder_consents SET project_id = (SELECT r.project_id FROM cached_rfcs r WHERE r.slug = funder_consents.rfc_slug) WHERE rfc_slug IN (SELECT slug FROM cached_rfcs);
DELETE FROM funder_consents WHERE rfc_slug NOT IN (SELECT slug FROM cached_rfcs);
UPDATE rfc_collaborators SET project_id = (SELECT r.project_id FROM cached_rfcs r WHERE r.slug = rfc_collaborators.rfc_slug) WHERE rfc_slug IN (SELECT slug FROM cached_rfcs);
DELETE FROM rfc_collaborators WHERE rfc_slug NOT IN (SELECT slug FROM cached_rfcs);
UPDATE contribution_requests SET project_id = (SELECT r.project_id FROM cached_rfcs r WHERE r.slug = contribution_requests.rfc_slug) WHERE rfc_slug IN (SELECT slug FROM cached_rfcs);
DELETE FROM contribution_requests WHERE rfc_slug NOT IN (SELECT slug FROM cached_rfcs);
UPDATE proposed_use_cases SET project_id = (SELECT r.project_id FROM cached_rfcs r WHERE r.slug = proposed_use_cases.rfc_slug) WHERE rfc_slug IN (SELECT slug FROM cached_rfcs);
DELETE FROM proposed_use_cases WHERE rfc_slug NOT IN (SELECT slug FROM cached_rfcs);
-- ── collections: the new typed-corpus grain beneath projects ───────────────
CREATE TABLE collections (
id TEXT NOT NULL,
project_id TEXT NOT NULL REFERENCES projects(id) ON DELETE CASCADE,
type TEXT NOT NULL DEFAULT 'document'
CHECK (type IN ('document', 'specification', 'bdd')),
subfolder TEXT NOT NULL DEFAULT '',
initial_state TEXT NOT NULL DEFAULT 'super-draft'
CHECK (initial_state IN ('super-draft', 'active')),
visibility TEXT NOT NULL DEFAULT 'gated'
CHECK (visibility IN ('gated', 'public', 'unlisted')),
name TEXT,
registry_sha TEXT,
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
PRIMARY KEY (id)
);
CREATE INDEX idx_collections_project ON collections(project_id);
-- One default collection per project. id='default' for the standard
-- single-project deployment (a stable literal across deploy histories); the
-- project_id is used as a unique fallback id only if a non-standard
-- multi-project deployment migrates (pre-S5; avoids a PK collision).
INSERT INTO collections (id, project_id, type, subfolder, initial_state, visibility, name)
SELECT
CASE WHEN (SELECT COUNT(*) FROM projects) <= 1 THEN 'default' ELSE p.id END,
p.id, p.type, '', p.initial_state, p.visibility, p.name
FROM projects p;
-- ── projects: rebuild to DROP the per-corpus fields (type, initial_state) ───
CREATE TABLE projects__new (
id TEXT PRIMARY KEY,
name TEXT NOT NULL,
content_repo TEXT,
visibility TEXT NOT NULL DEFAULT 'gated'
CHECK (visibility IN ('gated', 'public', 'unlisted')),
config_json TEXT,
registry_sha TEXT,
created_at TEXT NOT NULL DEFAULT (datetime('now')),
updated_at TEXT NOT NULL DEFAULT (datetime('now'))
);
INSERT INTO projects__new (id, name, content_repo, visibility, config_json, registry_sha, created_at, updated_at)
SELECT id, name, content_repo, visibility, config_json, registry_sha, created_at, updated_at FROM projects;
DROP TABLE projects;
ALTER TABLE projects__new RENAME TO projects;
-- ── cached_rfcs: PRIMARY KEY (project_id, slug) -> (collection_id, slug) ────
CREATE TABLE cached_rfcs__new (
slug TEXT NOT NULL,
title TEXT NOT NULL,
state TEXT NOT NULL CHECK (state IN ('super-draft', 'active', 'withdrawn', 'retired')),
rfc_id TEXT,
repo TEXT,
proposed_by TEXT,
proposed_at TEXT,
graduated_at TEXT,
graduated_by TEXT,
owners_json TEXT NOT NULL DEFAULT '[]',
arbiters_json TEXT NOT NULL DEFAULT '[]',
tags_json TEXT NOT NULL DEFAULT '[]',
body TEXT,
body_sha TEXT,
last_main_commit_at TEXT,
last_entry_commit_at TEXT,
updated_at TEXT NOT NULL DEFAULT (datetime('now')),
models_json TEXT,
funder_login TEXT,
proposed_use_case TEXT,
collection_id TEXT NOT NULL DEFAULT 'default' REFERENCES collections(id),
unreviewed INTEGER NOT NULL DEFAULT 0,
reviewed_at TEXT,
reviewed_by TEXT,
PRIMARY KEY (collection_id, slug)
);
INSERT INTO cached_rfcs__new
(slug, title, state, rfc_id, repo, proposed_by, proposed_at, graduated_at,
graduated_by, owners_json, arbiters_json, tags_json, body, body_sha,
last_main_commit_at, last_entry_commit_at, updated_at, models_json,
funder_login, proposed_use_case, collection_id, unreviewed, reviewed_at, reviewed_by)
SELECT
r.slug, r.title, r.state, r.rfc_id, r.repo, r.proposed_by, r.proposed_at, r.graduated_at,
r.graduated_by, r.owners_json, r.arbiters_json, r.tags_json, r.body, r.body_sha,
r.last_main_commit_at, r.last_entry_commit_at, r.updated_at, r.models_json,
r.funder_login, r.proposed_use_case,
(SELECT c.id FROM collections c WHERE c.project_id = r.project_id LIMIT 1),
r.unreviewed, r.reviewed_at, r.reviewed_by
FROM cached_rfcs r;
DROP TABLE cached_rfcs;
ALTER TABLE cached_rfcs__new RENAME TO cached_rfcs;
CREATE INDEX idx_cached_rfcs_state ON cached_rfcs (state);
CREATE INDEX idx_cached_rfcs_last_active ON cached_rfcs (
COALESCE(last_main_commit_at, last_entry_commit_at) DESC
);
CREATE INDEX idx_cached_rfcs_collection ON cached_rfcs(collection_id);
-- ── rfc_invitations: single-col FK -> composite (collection_id, rfc_slug) ───
CREATE TABLE rfc_invitations__new (
id INTEGER PRIMARY KEY AUTOINCREMENT,
rfc_slug TEXT NOT NULL,
inviter_user_id INTEGER REFERENCES users(id) ON DELETE SET NULL,
invitee_email TEXT NOT NULL,
role_in_rfc TEXT NOT NULL CHECK (role_in_rfc IN ('contributor', 'discussant')),
status TEXT NOT NULL DEFAULT 'pending'
CHECK (status IN ('pending', 'accepted', 'revoked', 'expired')),
token TEXT NOT NULL,
expires_at TEXT NOT NULL,
created_at TEXT NOT NULL DEFAULT (datetime('now')),
accepted_at TEXT,
accepted_by_user_id INTEGER REFERENCES users(id) ON DELETE SET NULL,
collection_id TEXT NOT NULL DEFAULT 'default',
FOREIGN KEY (collection_id, rfc_slug) REFERENCES cached_rfcs(collection_id, slug) ON DELETE CASCADE
);
INSERT INTO rfc_invitations__new
(id, rfc_slug, inviter_user_id, invitee_email, role_in_rfc, status, token,
expires_at, created_at, accepted_at, accepted_by_user_id, collection_id)
SELECT
i.id, i.rfc_slug, i.inviter_user_id, i.invitee_email, i.role_in_rfc, i.status, i.token,
i.expires_at, i.created_at, i.accepted_at, i.accepted_by_user_id,
(SELECT c.id FROM collections c WHERE c.project_id = i.project_id LIMIT 1)
FROM rfc_invitations i;
DROP TABLE rfc_invitations;
ALTER TABLE rfc_invitations__new RENAME TO rfc_invitations;
CREATE UNIQUE INDEX idx_rfc_invitations_token ON rfc_invitations (token);
CREATE INDEX idx_rfc_invitations_rfc_status ON rfc_invitations (rfc_slug, status);
CREATE INDEX idx_rfc_invitations_email_status ON rfc_invitations (invitee_email, status);
-- ── cached_branches: UNIQUE (project_id, rfc_slug, branch_name) -> collection
CREATE TABLE cached_branches__new (
id INTEGER PRIMARY KEY AUTOINCREMENT,
rfc_slug TEXT NOT NULL,
branch_name TEXT NOT NULL,
head_sha TEXT,
state TEXT NOT NULL DEFAULT 'open' CHECK (state IN ('open', 'closed', 'deleted')),
pinned INTEGER NOT NULL DEFAULT 0,
created_at TEXT NOT NULL DEFAULT (datetime('now')),
last_commit_at TEXT,
closed_at TEXT,
collection_id TEXT NOT NULL DEFAULT 'default',
UNIQUE (collection_id, rfc_slug, branch_name)
);
INSERT INTO cached_branches__new
(id, rfc_slug, branch_name, head_sha, state, pinned, created_at, last_commit_at, closed_at, collection_id)
SELECT
b.id, b.rfc_slug, b.branch_name, b.head_sha, b.state, b.pinned, b.created_at, b.last_commit_at, b.closed_at,
(SELECT c.id FROM collections c WHERE c.project_id = b.project_id LIMIT 1)
FROM cached_branches b;
DROP TABLE cached_branches;
ALTER TABLE cached_branches__new RENAME TO cached_branches;
CREATE INDEX idx_cached_branches_rfc ON cached_branches (rfc_slug, state);
-- ── branch_visibility: UNIQUE (project_id, rfc_slug, branch_name) -> collection
CREATE TABLE branch_visibility__new (
id INTEGER PRIMARY KEY AUTOINCREMENT,
rfc_slug TEXT NOT NULL,
branch_name TEXT NOT NULL,
read_public INTEGER NOT NULL DEFAULT 1,
contribute_mode TEXT NOT NULL DEFAULT 'just-me' CHECK (contribute_mode IN ('just-me', 'specific', 'any-contributor')),
collection_id TEXT NOT NULL DEFAULT 'default',
UNIQUE (collection_id, rfc_slug, branch_name)
);
INSERT INTO branch_visibility__new
(id, rfc_slug, branch_name, read_public, contribute_mode, collection_id)
SELECT
v.id, v.rfc_slug, v.branch_name, v.read_public, v.contribute_mode,
(SELECT c.id FROM collections c WHERE c.project_id = v.project_id LIMIT 1)
FROM branch_visibility v;
DROP TABLE branch_visibility;
ALTER TABLE branch_visibility__new RENAME TO branch_visibility;
-- ── branch_contribute_grants: UNIQUE (..., grantee) -> +collection_id ───────
CREATE TABLE branch_contribute_grants__new (
id INTEGER PRIMARY KEY AUTOINCREMENT,
rfc_slug TEXT NOT NULL,
branch_name TEXT NOT NULL,
grantee_user_id INTEGER NOT NULL REFERENCES users(id) ON DELETE CASCADE,
granted_by INTEGER NOT NULL REFERENCES users(id) ON DELETE SET NULL,
granted_at TEXT NOT NULL DEFAULT (datetime('now')),
collection_id TEXT NOT NULL DEFAULT 'default',
UNIQUE (collection_id, rfc_slug, branch_name, grantee_user_id)
);
INSERT INTO branch_contribute_grants__new
(id, rfc_slug, branch_name, grantee_user_id, granted_by, granted_at, collection_id)
SELECT
g.id, g.rfc_slug, g.branch_name, g.grantee_user_id, g.granted_by, g.granted_at,
(SELECT c.id FROM collections c WHERE c.project_id = g.project_id LIMIT 1)
FROM branch_contribute_grants g;
DROP TABLE branch_contribute_grants;
ALTER TABLE branch_contribute_grants__new RENAME TO branch_contribute_grants;
CREATE INDEX idx_grants_lookup ON branch_contribute_grants (rfc_slug, branch_name);
CREATE INDEX idx_grants_grantee ON branch_contribute_grants (grantee_user_id);
-- ── stars: UNIQUE (project_id, user_id, rfc_slug) -> collection_id ──────────
CREATE TABLE stars__new (
id INTEGER PRIMARY KEY AUTOINCREMENT,
user_id INTEGER NOT NULL REFERENCES users(id) ON DELETE CASCADE,
rfc_slug TEXT NOT NULL,
starred_at TEXT NOT NULL DEFAULT (datetime('now')),
collection_id TEXT NOT NULL DEFAULT 'default',
UNIQUE (collection_id, user_id, rfc_slug)
);
INSERT INTO stars__new (id, user_id, rfc_slug, starred_at, collection_id)
SELECT s.id, s.user_id, s.rfc_slug, s.starred_at,
(SELECT c.id FROM collections c WHERE c.project_id = s.project_id LIMIT 1)
FROM stars s;
DROP TABLE stars;
ALTER TABLE stars__new RENAME TO stars;
CREATE INDEX idx_stars_user ON stars (user_id);
CREATE INDEX idx_stars_rfc ON stars (rfc_slug);
-- ── watches: UNIQUE (project_id, user_id, rfc_slug) -> collection_id ────────
CREATE TABLE watches__new (
id INTEGER PRIMARY KEY AUTOINCREMENT,
user_id INTEGER NOT NULL REFERENCES users(id) ON DELETE CASCADE,
rfc_slug TEXT NOT NULL,
state TEXT NOT NULL CHECK (state IN ('watching', 'following', 'muted')),
set_by TEXT NOT NULL CHECK (set_by IN ('auto', 'explicit')),
set_at TEXT NOT NULL DEFAULT (datetime('now')),
last_participation_at TEXT,
collection_id TEXT NOT NULL DEFAULT 'default',
UNIQUE (collection_id, user_id, rfc_slug)
);
INSERT INTO watches__new
(id, user_id, rfc_slug, state, set_by, set_at, last_participation_at, collection_id)
SELECT
w.id, w.user_id, w.rfc_slug, w.state, w.set_by, w.set_at, w.last_participation_at,
(SELECT c.id FROM collections c WHERE c.project_id = w.project_id LIMIT 1)
FROM watches w;
DROP TABLE watches;
ALTER TABLE watches__new RENAME TO watches;
CREATE INDEX idx_watches_user ON watches (user_id);
CREATE INDEX idx_watches_rfc ON watches (rfc_slug);
CREATE INDEX idx_watches_decay ON watches (state, last_participation_at);
-- ── pr_seen: UNIQUE (project_id, user_id, rfc_slug, pr_number) -> collection ─
CREATE TABLE pr_seen__new (
id INTEGER PRIMARY KEY AUTOINCREMENT,
user_id INTEGER NOT NULL REFERENCES users(id) ON DELETE CASCADE,
rfc_slug TEXT NOT NULL,
pr_number INTEGER NOT NULL,
last_seen_commit_sha TEXT,
last_seen_message_id INTEGER REFERENCES thread_messages(id) ON DELETE SET NULL,
seen_at TEXT NOT NULL DEFAULT (datetime('now')),
collection_id TEXT NOT NULL DEFAULT 'default',
UNIQUE (collection_id, user_id, rfc_slug, pr_number)
);
INSERT INTO pr_seen__new
(id, user_id, rfc_slug, pr_number, last_seen_commit_sha, last_seen_message_id, seen_at, collection_id)
SELECT
p.id, p.user_id, p.rfc_slug, p.pr_number, p.last_seen_commit_sha, p.last_seen_message_id, p.seen_at,
(SELECT c.id FROM collections c WHERE c.project_id = p.project_id LIMIT 1)
FROM pr_seen p;
DROP TABLE pr_seen;
ALTER TABLE pr_seen__new RENAME TO pr_seen;
-- ── branch_chat_seen: UNIQUE (project_id, user_id, rfc_slug, branch) -> coll ─
CREATE TABLE branch_chat_seen__new (
id INTEGER PRIMARY KEY AUTOINCREMENT,
user_id INTEGER NOT NULL REFERENCES users(id) ON DELETE CASCADE,
rfc_slug TEXT NOT NULL,
branch_name TEXT NOT NULL,
last_seen_message_id INTEGER REFERENCES thread_messages(id) ON DELETE SET NULL,
seen_at TEXT NOT NULL DEFAULT (datetime('now')),
collection_id TEXT NOT NULL DEFAULT 'default',
UNIQUE (collection_id, user_id, rfc_slug, branch_name)
);
INSERT INTO branch_chat_seen__new
(id, user_id, rfc_slug, branch_name, last_seen_message_id, seen_at, collection_id)
SELECT
s.id, s.user_id, s.rfc_slug, s.branch_name, s.last_seen_message_id, s.seen_at,
(SELECT c.id FROM collections c WHERE c.project_id = s.project_id LIMIT 1)
FROM branch_chat_seen s;
DROP TABLE branch_chat_seen;
ALTER TABLE branch_chat_seen__new RENAME TO branch_chat_seen;
-- ── funder_consents: PRIMARY KEY (project_id, user_id, rfc_slug) -> collection
CREATE TABLE funder_consents__new (
user_id INTEGER NOT NULL,
rfc_slug TEXT NOT NULL,
created_at TEXT NOT NULL DEFAULT (datetime('now')),
collection_id TEXT NOT NULL DEFAULT 'default',
PRIMARY KEY (collection_id, user_id, rfc_slug),
FOREIGN KEY (user_id) REFERENCES users(id) ON DELETE CASCADE
);
INSERT INTO funder_consents__new (user_id, rfc_slug, created_at, collection_id)
SELECT f.user_id, f.rfc_slug, f.created_at,
(SELECT c.id FROM collections c WHERE c.project_id = f.project_id LIMIT 1)
FROM funder_consents f;
DROP TABLE funder_consents;
ALTER TABLE funder_consents__new RENAME TO funder_consents;
CREATE INDEX idx_funder_consents_slug ON funder_consents (rfc_slug);
-- ── rfc_collaborators: UNIQUE idx + composite FK -> collection_id ───────────
CREATE TABLE rfc_collaborators__new (
id INTEGER PRIMARY KEY AUTOINCREMENT,
rfc_slug TEXT NOT NULL,
user_id INTEGER NOT NULL REFERENCES users(id) ON DELETE CASCADE,
role_in_rfc TEXT NOT NULL CHECK (role_in_rfc IN ('contributor', 'discussant')),
invitation_id INTEGER REFERENCES rfc_invitations(id) ON DELETE SET NULL,
created_at TEXT NOT NULL DEFAULT (datetime('now')),
collection_id TEXT NOT NULL DEFAULT 'default',
FOREIGN KEY (collection_id, rfc_slug) REFERENCES cached_rfcs(collection_id, slug) ON DELETE CASCADE
);
INSERT INTO rfc_collaborators__new
(id, rfc_slug, user_id, role_in_rfc, invitation_id, created_at, collection_id)
SELECT
rc.id, rc.rfc_slug, rc.user_id, rc.role_in_rfc, rc.invitation_id, rc.created_at,
(SELECT c.id FROM collections c WHERE c.project_id = rc.project_id LIMIT 1)
FROM rfc_collaborators rc;
DROP TABLE rfc_collaborators;
ALTER TABLE rfc_collaborators__new RENAME TO rfc_collaborators;
CREATE UNIQUE INDEX idx_rfc_collaborators_unique ON rfc_collaborators (collection_id, rfc_slug, user_id);
CREATE INDEX idx_rfc_collaborators_user ON rfc_collaborators (user_id);
-- ── contribution_requests: UNIQUE idx (pending) + composite FK -> collection ─
CREATE TABLE contribution_requests__new (
id INTEGER PRIMARY KEY AUTOINCREMENT,
rfc_slug TEXT NOT NULL,
requester_user_id INTEGER NOT NULL REFERENCES users(id) ON DELETE CASCADE,
matched_term TEXT NOT NULL,
who_i_am TEXT NOT NULL,
why TEXT NOT NULL,
use_case TEXT,
status TEXT NOT NULL DEFAULT 'pending'
CHECK (status IN ('pending', 'accepted', 'declined')),
created_at TEXT NOT NULL DEFAULT (datetime('now')),
decided_at TEXT,
decided_by_user_id INTEGER REFERENCES users(id) ON DELETE SET NULL,
invitation_id INTEGER REFERENCES rfc_invitations(id) ON DELETE SET NULL,
notification_id INTEGER REFERENCES notifications(id) ON DELETE SET NULL,
collection_id TEXT NOT NULL DEFAULT 'default',
FOREIGN KEY (collection_id, rfc_slug) REFERENCES cached_rfcs(collection_id, slug) ON DELETE CASCADE
);
INSERT INTO contribution_requests__new
(id, rfc_slug, requester_user_id, matched_term, who_i_am, why, use_case, status,
created_at, decided_at, decided_by_user_id, invitation_id, notification_id, collection_id)
SELECT
cr.id, cr.rfc_slug, cr.requester_user_id, cr.matched_term, cr.who_i_am, cr.why, cr.use_case, cr.status,
cr.created_at, cr.decided_at, cr.decided_by_user_id, cr.invitation_id, cr.notification_id,
(SELECT c.id FROM collections c WHERE c.project_id = cr.project_id LIMIT 1)
FROM contribution_requests cr;
DROP TABLE contribution_requests;
ALTER TABLE contribution_requests__new RENAME TO contribution_requests;
CREATE INDEX idx_contribution_requests_rfc ON contribution_requests(rfc_slug, status);
CREATE INDEX idx_contribution_requests_requester ON contribution_requests(requester_user_id, status);
CREATE UNIQUE INDEX idx_contribution_requests_one_open
ON contribution_requests(collection_id, rfc_slug, requester_user_id)
WHERE status = 'pending';
-- ── proposed_use_cases: UNIQUE (project_id, scope, pr_number) -> collection ──
CREATE TABLE proposed_use_cases__new (
id INTEGER PRIMARY KEY AUTOINCREMENT,
scope TEXT NOT NULL CHECK (scope IN ('rfc', 'pr')),
rfc_slug TEXT NOT NULL,
pr_number INTEGER NOT NULL,
use_case TEXT NOT NULL,
created_at TEXT NOT NULL DEFAULT (datetime('now')),
collection_id TEXT NOT NULL DEFAULT 'default',
UNIQUE (collection_id, scope, pr_number)
);
INSERT INTO proposed_use_cases__new
(id, scope, rfc_slug, pr_number, use_case, created_at, collection_id)
SELECT
u.id, u.scope, u.rfc_slug, u.pr_number, u.use_case, u.created_at,
(SELECT c.id FROM collections c WHERE c.project_id = u.project_id LIMIT 1)
FROM proposed_use_cases u;
DROP TABLE proposed_use_cases;
ALTER TABLE proposed_use_cases__new RENAME TO proposed_use_cases;
CREATE INDEX idx_proposed_use_cases_lookup ON proposed_use_cases (scope, pr_number);
CREATE INDEX idx_proposed_use_cases_slug ON proposed_use_cases (scope, rfc_slug);
-- ── project_members -> memberships(scope_type, scope_id, …); roles collapsed ─
CREATE TABLE memberships (
id INTEGER PRIMARY KEY AUTOINCREMENT,
scope_type TEXT NOT NULL CHECK (scope_type IN ('project', 'collection')),
scope_id TEXT NOT NULL,
user_id INTEGER NOT NULL REFERENCES users(id) ON DELETE CASCADE,
role TEXT NOT NULL CHECK (role IN ('owner', 'contributor')),
granted_by INTEGER REFERENCES users(id) ON DELETE SET NULL,
granted_at TEXT NOT NULL DEFAULT (datetime('now')),
UNIQUE (scope_type, scope_id, user_id)
);
CREATE INDEX idx_memberships_user ON memberships(user_id);
CREATE INDEX idx_memberships_scope ON memberships(scope_type, scope_id);
-- M2 project_members rows attached at what is now the *collection*; collapse the
-- role enum (project_admin -> owner, project_contributor -> contributor;
-- project_viewer dropped this pass, §B.3) and migrate onto the default
-- collection of each project.
INSERT INTO memberships (scope_type, scope_id, user_id, role, granted_by, granted_at)
SELECT 'collection',
(SELECT c.id FROM collections c WHERE c.project_id = pm.project_id LIMIT 1),
pm.user_id,
CASE pm.role WHEN 'project_admin' THEN 'owner'
WHEN 'project_contributor' THEN 'contributor'
ELSE 'contributor' END,
pm.granted_by, pm.granted_at
FROM project_members pm
WHERE pm.role IN ('project_admin', 'project_contributor');
DROP TABLE project_members;
+34
View File
@@ -0,0 +1,34 @@
-- migrate:no-foreign-keys
--
-- §22 three-tier — S3. Admit a *global*-scope grant to the memberships table.
--
-- §B.2's resolver folds four layers (global → project → collection → per-entry).
-- Migration 029 created `memberships` with scope_type ∈ {project, collection}
-- only; the global tier was left to S3. A global grant is how a "global RFC
-- Contributor" (a contributor who may propose in every collection of every
-- project, distinct from a deployment owner/admin) is represented — see
-- docs/design/2026-06-05-three-tier-projects-collections.md §B.2/§B.3 and the
-- C.1 "cleo" scenario.
--
-- SQLite can't ALTER a CHECK constraint in place, so the table is rebuilt by the
-- create-copy-drop-rename procedure (the 028/029 pattern). The global scope uses
-- a stable sentinel scope_id of '*' (one global tier per deployment); the
-- UNIQUE(scope_type, scope_id, user_id) then admits exactly one global grant per
-- user, mirroring the project/collection rows.
CREATE TABLE memberships__new (
id INTEGER PRIMARY KEY AUTOINCREMENT,
scope_type TEXT NOT NULL CHECK (scope_type IN ('global', 'project', 'collection')),
scope_id TEXT NOT NULL,
user_id INTEGER NOT NULL REFERENCES users(id) ON DELETE CASCADE,
role TEXT NOT NULL CHECK (role IN ('owner', 'contributor')),
granted_by INTEGER REFERENCES users(id) ON DELETE SET NULL,
granted_at TEXT NOT NULL DEFAULT (datetime('now')),
UNIQUE (scope_type, scope_id, user_id)
);
INSERT INTO memberships__new (id, scope_type, scope_id, user_id, role, granted_by, granted_at)
SELECT id, scope_type, scope_id, user_id, role, granted_by, granted_at FROM memberships;
DROP TABLE memberships;
ALTER TABLE memberships__new RENAME TO memberships;
CREATE INDEX idx_memberships_user ON memberships(user_id);
CREATE INDEX idx_memberships_scope ON memberships(scope_type, scope_id);
@@ -0,0 +1,9 @@
-- §22.12 S6 — per-collection model universe.
--
-- A collection's `.collection.yaml` may carry an `enabled_models` list that
-- NARROWS its project's universe (which narrows the deployment ENABLED_MODELS).
-- Mirrored into a `config_json` blob on the collection row, paralleling
-- `projects.config_json` (which already holds the project's enabled_models +
-- theme). Additive only — no rebuild. NULL means "no per-collection narrowing;
-- inherit the project's universe."
ALTER TABLE collections ADD COLUMN config_json TEXT;
+58
View File
@@ -0,0 +1,58 @@
-- §22.8 S6 — request-to-join a scope + the cross-collection inbox.
--
-- A gated project or collection is invisible to non-members (§22.5), so a user
-- who knows a scope exists can ask to join it: they name a desired role and the
-- request is recorded here, then fanned out to that scope's Owners *across the
-- subtree* (a collection request reaches the collection's Owners, its project's
-- Owners, and global Owners — the cross-collection inbox, §22.11). An Owner
-- accepts (which writes the `memberships` row via memberships.grant) or declines;
-- the requester is §15-notified of the decision either way.
--
-- This mirrors `contribution_requests` (migration 024) but at the scope grain
-- instead of the per-RFC grain: the target is a `(scope_type, scope_id)` pair
-- (matching the `memberships` scope vocabulary, minus 'global' — joining is for a
-- project or collection a user discovers, not the deployment), and accept grants
-- a scope role rather than minting an RFC invitation.
--
-- The request row is the persistent record; the inbox notification is the
-- owner-facing actionable surface keyed back to it via `notification_id`.
CREATE TABLE IF NOT EXISTS join_requests (
id INTEGER PRIMARY KEY AUTOINCREMENT,
-- The target scope. 'global' is intentionally excluded: the deployment is
-- not a thing one "discovers and joins" (§22.8 names a project/collection).
scope_type TEXT NOT NULL
CHECK (scope_type IN ('project', 'collection')),
scope_id TEXT NOT NULL,
requester_user_id INTEGER NOT NULL
REFERENCES users(id) ON DELETE CASCADE,
-- The role the requester is asking for ({owner, contributor}, the §22.6
-- unified vocabulary). The accepting Owner may grant this or a narrower role.
requested_role TEXT NOT NULL
CHECK (requested_role IN ('owner', 'contributor')),
-- Optional free text — "who I am / why I want in". Bounded by the API layer.
message TEXT,
status TEXT NOT NULL DEFAULT 'pending'
CHECK (status IN ('pending', 'accepted', 'declined')),
created_at TEXT NOT NULL DEFAULT (datetime('now')),
decided_at TEXT,
decided_by_user_id INTEGER REFERENCES users(id) ON DELETE SET NULL,
-- The role actually granted on accept (may differ from requested_role if the
-- Owner narrowed it); NULL until accepted.
granted_role TEXT CHECK (granted_role IN ('owner', 'contributor')),
-- The owner-facing notification row that carries the Accept/Decline action.
notification_id INTEGER REFERENCES notifications(id) ON DELETE SET NULL
);
CREATE INDEX IF NOT EXISTS idx_join_requests_scope
ON join_requests(scope_type, scope_id, status);
CREATE INDEX IF NOT EXISTS idx_join_requests_requester
ON join_requests(requester_user_id, status);
-- At most one open (pending) request per (scope, requester): a second ask while
-- one is still pending is a 409, not a duplicate row. A decided request
-- (accepted/declined) does not block a fresh ask later.
CREATE UNIQUE INDEX IF NOT EXISTS idx_join_requests_one_open
ON join_requests(scope_type, scope_id, requester_user_id)
WHERE status = 'pending';
@@ -0,0 +1,8 @@
-- §22.4a SLICE-1 — configurable collection metadata: malformed-sidecar flag.
--
-- The corpus mirror reads an entry's metadata from its `<slug>.meta.yaml`
-- sidecar (dual-read: sidecar-else-legacy-frontmatter). A sidecar that does not
-- parse as a YAML mapping never hard-fails the read (INV-3) — the entry still
-- loads (from the legacy `.md` frontmatter if present) and this derived flag
-- marks it so the catalog can surface it. Additive only — no rebuild. 0 = ok.
ALTER TABLE cached_rfcs ADD COLUMN metadata_malformed INTEGER NOT NULL DEFAULT 0;
@@ -0,0 +1,12 @@
-- §22.4a SLICE-3 — configurable collection metadata: cache per-entry values.
--
-- Faceted filtering (§5.1) needs each entry's metadata values (priority, custom
-- enum/tags fields) to compute facet counts and honour filter params. Today the
-- mirror keeps only `tags_json` + the lifecycle columns and drops `Entry.extra`,
-- so a declared field's values are unrecoverable. This column persists the full
-- per-entry metadata mapping (`metadata.metadata_dict(entry)`, known keys +
-- extra, never the body) as JSON, so `app/facets.py` can read any declared
-- field uniformly. Additive + nullable — no rebuild. NULL = not yet re-ingested
-- (the reconciler/webhook fills it on the next sweep) → that entry contributes
-- no facet values until then. SLICE-4/5 edit panels read the same column.
ALTER TABLE cached_rfcs ADD COLUMN meta_json TEXT;
+14 -7
View File
@@ -9,11 +9,18 @@ from test_propose_vertical import ( # noqa: F401
def _add_project(pid, name, vis, typ="document"):
# §22 three-tier: a project (grouping tier) + its default collection (the
# per-corpus type/initial_state moved down in migration 029). The default
# collection keys by the project id so it is globally unique in tests.
from app import db
db.conn().execute(
"INSERT OR REPLACE INTO projects (id, name, type, content_repo, visibility, initial_state) "
"VALUES (?, ?, ?, ?, ?, 'super-draft')",
(pid, name, typ, pid, vis),
"INSERT OR REPLACE INTO projects (id, name, content_repo, visibility) VALUES (?, ?, ?, ?)",
(pid, name, pid, vis),
)
db.conn().execute(
"INSERT OR REPLACE INTO collections (id, project_id, type, subfolder, initial_state, visibility, name) "
"VALUES (?, ?, ?, '', 'super-draft', ?, ?)",
(pid, pid, typ, vis, name),
)
@@ -93,7 +100,7 @@ def test_rfc_root_url_redirects_308_to_project_scoped(app_with_fake_gitea):
with TestClient(app) as client:
r = client.get("/rfc/human", follow_redirects=False)
assert r.status_code == 308
assert r.headers["location"] == "/p/default/e/human"
assert r.headers["location"] == "/p/default/c/default/e/human"
def test_rfc_pr_url_redirects_308_to_project_scoped(app_with_fake_gitea):
@@ -102,7 +109,7 @@ def test_rfc_pr_url_redirects_308_to_project_scoped(app_with_fake_gitea):
with TestClient(app) as client:
r = client.get("/rfc/human/pr/7", follow_redirects=False)
assert r.status_code == 308
assert r.headers["location"] == "/p/default/e/human/pr/7"
assert r.headers["location"] == "/p/default/c/default/e/human/pr/7"
def test_proposals_root_url_redirects_308_to_project_scoped(app_with_fake_gitea):
@@ -110,7 +117,7 @@ def test_proposals_root_url_redirects_308_to_project_scoped(app_with_fake_gitea)
with TestClient(app) as client:
r = client.get("/proposals/42", follow_redirects=False)
assert r.status_code == 308
assert r.headers["location"] == "/p/default/proposals/42"
assert r.headers["location"] == "/p/default/c/default/proposals/42"
def test_gated_project_visible_and_readable_to_member(app_with_fake_gitea):
@@ -120,7 +127,7 @@ def test_gated_project_visible_and_readable_to_member(app_with_fake_gitea):
_add_project("teamx", "Team X", "gated")
provision_user_row(user_id=5, login="mia", role="contributor")
db.conn().execute(
"INSERT INTO project_members (project_id, user_id, role) VALUES ('teamx', 5, 'project_viewer')"
"INSERT INTO memberships (scope_type, scope_id, user_id, role) VALUES ('collection', 'teamx', 5, 'contributor')"
)
sign_in_as(client, user_id=5, gitea_login="mia", display_name="Mia", role="contributor")
# member sees the gated project in the deployment directory
+62
View File
@@ -0,0 +1,62 @@
"""SLICE-4 — bot multi-file commit + PR primitives."""
from __future__ import annotations
import asyncio
from app import gitea as gitea_mod, metadata
from app.bot import Actor, Bot
from app.config import load_config
from test_propose_vertical import ( # noqa: F401
app_with_fake_gitea,
provision_user_row,
tmp_env,
)
LEGACY = "---\nslug: alpha\ntitle: Alpha\nstate: active\ntags:\n- one\n---\n\nBody.\n"
def _actor():
return Actor(user_id=1, gitea_login="ben.stull", display_name="Ben", email="ben@x.io")
def test_commit_entry_files_direct_to_main(app_with_fake_gitea):
_app, fake = app_with_fake_gitea
gitea = gitea_mod.Gitea(load_config())
bot = Bot(gitea)
fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.md")] = {
"content": LEGACY, "sha": "s1"}
st = asyncio.run(metadata.read_entry_from_git(gitea, "wiggleverse", "meta", "rfcs/alpha.md"))
e2 = metadata.apply_values(st.entry, {"tags": ["two"]})
ops = metadata.write_entry_files("rfcs/alpha.md", e2, st)
asyncio.run(bot.commit_entry_files(
_actor(), org="wiggleverse", repo="meta", files=ops,
message="Edit metadata", branch="main"))
sc = fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.meta.yaml")]["content"]
assert "two" in sc
# .md is now body-only
assert "---" not in fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.md")]["content"]
def test_open_entry_pr_commits_on_branch_and_opens_pr(app_with_fake_gitea):
from fastapi.testclient import TestClient
app, fake = app_with_fake_gitea
gitea = gitea_mod.Gitea(load_config())
bot = Bot(gitea)
fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.md")] = {
"content": LEGACY, "sha": "s1"}
with TestClient(app): # lifespan inits the DB
provision_user_row(user_id=1, login="ben.stull", role="owner")
st = asyncio.run(metadata.read_entry_from_git(gitea, "wiggleverse", "meta", "rfcs/alpha.md"))
e2 = metadata.apply_values(st.entry, {"tags": ["two"]})
ops = metadata.write_entry_files("rfcs/alpha.md", e2, st)
pr = asyncio.run(bot.open_entry_pr(
_actor(), org="wiggleverse", repo="meta", slug="alpha", files=ops,
pr_title="Metadata: Alpha", pr_description="edit", branch_prefix="metadata"))
assert pr["number"] >= 1
head = pr["head"]["ref"]
assert head.startswith("metadata-alpha-")
# committed on the branch, main untouched
assert ("wiggleverse", "meta", head, "rfcs/alpha.meta.yaml") in fake.files
assert ("wiggleverse", "meta", "main", "rfcs/alpha.meta.yaml") not in fake.files
@@ -0,0 +1,83 @@
"""§22 S2 — create-collection vertical: a deployment owner/admin POSTs, the bot
commits a `.collection.yaml`, and the registry mirror upserts the collections
row (registry stays the source of truth)."""
from __future__ import annotations
from fastapi.testclient import TestClient
from app import db
from test_propose_vertical import ( # noqa: F401
app_with_fake_gitea, tmp_env, provision_user_row, sign_in_as,
)
def test_create_collection_commits_manifest_and_mirrors(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
provision_user_row(user_id=1, login="ben", role="owner")
sign_in_as(client, user_id=1, gitea_login="ben", display_name="Ben",
role="owner", email="ben@test")
r = client.post("/api/projects/default/collections",
json={"collection_id": "features", "type": "bdd", "name": "Features"})
assert r.status_code == 200, r.text
assert r.json()["type"] == "bdd"
# The bot committed the manifest to the content repo's main.
f = fake.files.get(("wiggleverse", "meta", "main", "features/.collection.yaml"))
assert f is not None
assert "type: bdd" in f["content"]
# The registry refresh mirrored it into a collections row.
row = db.conn().execute(
"SELECT type, project_id, subfolder FROM collections WHERE id='features'"
).fetchone()
assert (row["type"], row["project_id"], row["subfolder"]) == ("bdd", "default", "features")
# It is now navigable via the directory + scoped serve.
items = client.get("/api/projects/default/collections").json()["items"]
assert any(c["id"] == "features" for c in items)
assert client.get("/api/projects/default/collections/features/rfcs").status_code == 200
def test_create_collection_requires_admin(app_with_fake_gitea):
app, _ = app_with_fake_gitea
with TestClient(app) as client:
provision_user_row(user_id=2, login="alice", role="contributor")
sign_in_as(client, user_id=2, gitea_login="alice", display_name="Alice",
role="contributor", email="alice@test")
r = client.post("/api/projects/default/collections",
json={"collection_id": "x", "type": "bdd"})
assert r.status_code in (401, 403)
def test_create_collection_anonymous_rejected(app_with_fake_gitea):
app, _ = app_with_fake_gitea
with TestClient(app) as client:
r = client.post("/api/projects/default/collections",
json={"collection_id": "x", "type": "bdd"})
assert r.status_code in (401, 403)
def test_create_collection_rejects_duplicate(app_with_fake_gitea):
app, _ = app_with_fake_gitea
with TestClient(app) as client:
provision_user_row(user_id=1, login="ben", role="owner")
sign_in_as(client, user_id=1, gitea_login="ben", display_name="Ben",
role="owner", email="ben@test")
ok = client.post("/api/projects/default/collections",
json={"collection_id": "features", "type": "bdd"})
assert ok.status_code == 200, ok.text
dup = client.post("/api/projects/default/collections",
json={"collection_id": "features", "type": "bdd"})
assert dup.status_code == 409
def test_create_collection_rejects_reserved_default_id(app_with_fake_gitea):
app, _ = app_with_fake_gitea
with TestClient(app) as client:
provision_user_row(user_id=1, login="ben", role="owner")
sign_in_as(client, user_id=1, gitea_login="ben", display_name="Ben",
role="owner", email="ben@test")
r = client.post("/api/projects/default/collections",
json={"collection_id": "default", "type": "bdd"})
assert r.status_code == 422
+85
View File
@@ -0,0 +1,85 @@
"""§22 S2 — collection read helpers: list_collections / get_collection /
subfolder_of."""
from __future__ import annotations
import json
import tempfile
from pathlib import Path
from app import collections as collections_mod, db
from app.config import Config
def _db() -> Config:
cfg = Config(
gitea_url="x", gitea_bot_user="x", gitea_bot_token="x", gitea_org="x",
registry_repo="registry", oauth_client_id="x",
oauth_client_secret="x", app_url="x", secret_key="x",
database_path=Path(tempfile.mkdtemp(prefix="colhelp-")) / "t.db",
owner_gitea_login="x", webhook_secret="x",
)
db.run_migrations(cfg)
if db._CONN is not None:
db._CONN.close()
db._CONN = None
db.init(cfg)
return cfg
def _seed(project_id="ohm"):
db.conn().execute(
"INSERT OR REPLACE INTO projects (id, name, content_repo, visibility, updated_at) "
"VALUES (?, 'Ohm', 'ohm-rfc', 'public', datetime('now'))", (project_id,))
for cid, sub, vis, name in [
("default", "", "public", "Model"),
("features", "features", "public", "Features"),
("secret", "secret", "unlisted", "Secret"),
]:
db.conn().execute(
"INSERT OR REPLACE INTO collections (id, project_id, type, subfolder, initial_state, "
"visibility, name, created_at, updated_at) VALUES (?,?, 'document', ?, "
"'super-draft', ?, ?, datetime('now'), datetime('now'))",
(cid, project_id, sub, vis, name))
def test_list_collections_excludes_unlisted():
_db()
_seed()
ids = [c["id"] for c in collections_mod.list_collections("ohm", include_unlisted=False)]
assert ids == ["default", "features"] # default first, then by name; 'secret' omitted
def test_list_collections_include_unlisted():
_db()
_seed()
ids = {c["id"] for c in collections_mod.list_collections("ohm", include_unlisted=True)}
assert ids == {"default", "features", "secret"}
def test_get_collection_and_subfolder():
_db()
_seed()
assert collections_mod.get_collection("features")["name"] == "Features"
assert collections_mod.subfolder_of("features") == "features"
assert collections_mod.subfolder_of("default") == ""
assert collections_mod.get_collection("nope") is None
# ---- §22.4a SLICE-2: field schema served on the collection ----
def test_get_collection_fields_none_when_unset():
# INV-5: a collection with no `fields:` exposes fields=None (the default
# `document` collection sees zero change).
_db()
_seed()
assert collections_mod.get_collection("features")["fields"] is None
def test_get_collection_exposes_field_schema():
_db()
_seed()
schema = {"priority": {"type": "enum", "values": ["P0", "P1"]}}
db.conn().execute(
"UPDATE collections SET config_json = ? WHERE id = 'features'",
(json.dumps({"fields": schema}),))
assert collections_mod.get_collection("features")["fields"] == schema
+176
View File
@@ -0,0 +1,176 @@
"""§22 S2 — the registry mirror reads `.collection.yaml` manifests inside each
project's content repo and upserts a named collection per manifest. The default
collection still flows from projects.yaml (test_registry.py)."""
from __future__ import annotations
import asyncio
import base64
import tempfile
from pathlib import Path
import pytest
from app import db, registry
from app.config import Config
def _db() -> Config:
cfg = Config(
gitea_url="x", gitea_bot_user="x", gitea_bot_token="x", gitea_org="wiggleverse",
registry_repo="registry", oauth_client_id="x",
oauth_client_secret="x", app_url="x", secret_key="x",
database_path=Path(tempfile.mkdtemp(prefix="colreg-")) / "t.db",
owner_gitea_login="x", webhook_secret="x",
)
db.run_migrations(cfg)
if db._CONN is not None:
db._CONN.close()
db._CONN = None
db.init(cfg)
return cfg
# --- pure parser --------------------------------------------------------------
def test_parse_collection_manifest_minimal():
doc = registry.parse_collection_manifest("type: bdd\n")
assert doc.type == "bdd"
# §22.4b: bdd defaults to 'active'; visibility inherits (None == inherit).
assert doc.initial_state == "active"
assert doc.visibility is None
assert doc.name is None
def test_parse_collection_manifest_full():
doc = registry.parse_collection_manifest(
"type: document\nvisibility: public\ninitial_state: active\nname: Model\n"
)
assert (doc.type, doc.visibility, doc.initial_state, doc.name) == (
"document", "public", "active", "Model",
)
def test_parse_collection_manifest_rejects_bad_type():
with pytest.raises(registry.RegistryError):
registry.parse_collection_manifest("type: nonsense\n")
def test_parse_collection_manifest_rejects_bad_visibility():
with pytest.raises(registry.RegistryError):
registry.parse_collection_manifest("type: bdd\nvisibility: nope\n")
# ---- §22.4a SLICE-2: a `fields:` block flows into the collection config ----
def test_parse_collection_manifest_reads_field_schema():
doc = registry.parse_collection_manifest(
"type: bdd\n"
"fields:\n"
" priority:\n"
" type: enum\n"
" values: [P0, P1, P2]\n"
" tags:\n"
" type: tags\n"
)
assert doc.config["fields"] == {
"priority": {"type": "enum", "values": ["P0", "P1", "P2"]},
"tags": {"type": "tags"},
}
def test_parse_collection_manifest_lenient_on_bad_field():
# A bad field def is skipped (INV-3), the manifest still parses, and a
# manifest with no surviving fields carries no `fields` config key at all.
doc = registry.parse_collection_manifest(
"type: bdd\n"
"fields:\n"
" broken:\n"
" type: ref\n"
)
assert "fields" not in doc.config
# --- mirror discovery ---------------------------------------------------------
class _FakeGitea:
"""Minimal Gitea stub: projects.yaml in the registry repo + a content repo
whose root holds a `features/` subdir carrying a `.collection.yaml`."""
def __init__(self, projects_yaml: str, repo_tree: dict[str, dict[str, str]]):
self._projects_yaml = projects_yaml
self._repo_tree = repo_tree # {repo: {path: text}}
async def get_contents(self, org, repo, path, ref="main"):
if path == "projects.yaml":
return {"type": "file",
"content": base64.b64encode(self._projects_yaml.encode()).decode(),
"sha": "regsha-test"}
text = self._repo_tree.get(repo, {}).get(path)
if text is None:
return None
return {"type": "file",
"content": base64.b64encode(text.encode()).decode(), "sha": "c0ffee"}
async def list_dir(self, org, repo, path, ref="main"):
# Root listing: surface each top-level segment as a 'dir' entry.
prefix = (path.rstrip("/") + "/") if path else ""
dirs = set()
for p in self._repo_tree.get(repo, {}):
if not p.startswith(prefix):
continue
rest = p[len(prefix):]
if "/" in rest:
dirs.add(rest.split("/", 1)[0])
return [{"type": "dir", "name": n, "path": prefix + n} for n in sorted(dirs)]
_PROJECTS = (
"deployment:\n name: Ohm\n tagline: t\n"
"projects:\n - id: ohm\n name: Ohm\n type: document\n"
" content_repo: ohm-rfc\n visibility: public\n"
)
def test_refresh_registry_mirrors_named_collection():
cfg = _db()
gitea = _FakeGitea(
projects_yaml=_PROJECTS,
repo_tree={"ohm-rfc": {"features/.collection.yaml": "type: bdd\nname: Features\n"}},
)
asyncio.run(registry.refresh_registry(cfg, gitea))
row = db.conn().execute(
"SELECT type, subfolder, name, project_id, visibility FROM collections WHERE id='features'"
).fetchone()
assert row is not None
assert (row["type"], row["subfolder"], row["project_id"]) == ("bdd", "features", "ohm")
assert row["name"] == "Features"
# visibility inherits the project's (public) when the manifest omits it.
assert row["visibility"] == "public"
def test_refresh_registry_leaves_default_collection_intact():
cfg = _db()
gitea = _FakeGitea(
projects_yaml=_PROJECTS,
repo_tree={"ohm-rfc": {"features/.collection.yaml": "type: bdd\n"}},
)
asyncio.run(registry.refresh_registry(cfg, gitea))
# The default collection (from projects.yaml) and the named one coexist.
ids = {r["id"] for r in db.conn().execute("SELECT id FROM collections")}
assert {"default", "features"} <= ids
def test_refresh_registry_immutable_type_on_named_collection():
cfg = _db()
gitea = _FakeGitea(
projects_yaml=_PROJECTS,
repo_tree={"ohm-rfc": {"features/.collection.yaml": "type: bdd\n"}},
)
asyncio.run(registry.refresh_registry(cfg, gitea))
# A later manifest that flips the type is refused (§22.4a immutable type).
gitea._repo_tree["ohm-rfc"]["features/.collection.yaml"] = "type: document\n"
asyncio.run(registry.refresh_registry(cfg, gitea))
t = db.conn().execute("SELECT type FROM collections WHERE id='features'").fetchone()["type"]
assert t == "bdd"
@@ -0,0 +1,164 @@
"""§22 S2 — collection-grained corpus mirror + collection-scoped serve/propose.
The mirror test drives cache.refresh_meta_repo against an in-memory content repo
holding entries under both the default `rfcs/` and a named collection's
`features/rfcs/`, and asserts cached_rfcs is keyed by the right collection_id."""
from __future__ import annotations
import asyncio
import tempfile
from pathlib import Path
from app import cache, db
from app.config import Config
def _db() -> Config:
cfg = Config(
gitea_url="x", gitea_bot_user="x", gitea_bot_token="x", gitea_org="wiggleverse",
registry_repo="registry", oauth_client_id="x",
oauth_client_secret="x", app_url="x", secret_key="x",
database_path=Path(tempfile.mkdtemp(prefix="colserve-")) / "t.db",
owner_gitea_login="x", webhook_secret="x",
)
db.run_migrations(cfg)
if db._CONN is not None:
db._CONN.close()
db._CONN = None
db.init(cfg)
return cfg
class _CorpusGitea:
"""A content repo modelled as a flat {path: text} map, listing files under a
directory prefix and reading them back."""
def __init__(self, tree: dict[str, str]):
self._tree = tree
async def list_dir(self, org, repo, path, ref="main"):
out = []
prefix = (path.rstrip("/") + "/") if path else ""
for p in self._tree:
if p.startswith(prefix) and "/" not in p[len(prefix):]:
out.append({"type": "file", "name": p.split("/")[-1], "path": p})
return out
async def read_file(self, org, repo, path, ref="main"):
t = self._tree.get(path)
return (t, "sha-" + path) if t is not None else None
def _entry_md(slug, title):
return f"---\nslug: {slug}\ntitle: {title}\nstate: active\n---\nbody\n"
def _seed_project_with_two_collections():
db.conn().execute(
"INSERT OR REPLACE INTO projects (id, name, content_repo, visibility, updated_at) "
"VALUES ('ohm','Ohm','ohm-rfc','public', datetime('now'))")
for cid, sub in [("default", ""), ("features", "features")]:
db.conn().execute(
"INSERT OR REPLACE INTO collections (id, project_id, type, subfolder, initial_state, "
"visibility, created_at, updated_at) VALUES (?, 'ohm','document',?, "
"'super-draft','public', datetime('now'), datetime('now'))", (cid, sub))
def test_mirror_keys_entries_by_collection():
cfg = _db()
_seed_project_with_two_collections()
gitea = _CorpusGitea({
"rfcs/a.md": _entry_md("a", "Default A"),
"features/rfcs/b.md": _entry_md("b", "Feature B"),
})
asyncio.run(cache.refresh_meta_repo(cfg, gitea))
got = {(r["collection_id"], r["slug"]) for r in
db.conn().execute("SELECT collection_id, slug FROM cached_rfcs")}
assert got == {("default", "a"), ("features", "b")}
# --- collection-scoped serve + propose (full app) -----------------------------
from fastapi.testclient import TestClient # noqa: E402
from test_propose_vertical import ( # noqa: E402,F401
app_with_fake_gitea, tmp_env, provision_user_row, sign_in_as,
)
def _add_features_collection(content_repo="meta"):
"""Add a named 'features' collection (subfolder 'features') under the seeded
default project, plus a single entry under features/rfcs/ in the db cache."""
db.conn().execute(
"INSERT OR REPLACE INTO collections (id, project_id, type, subfolder, "
"initial_state, visibility, name, created_at, updated_at) VALUES "
"('features','default','document','features','super-draft','public','Features', "
"datetime('now'), datetime('now'))")
def test_scoped_list_returns_only_that_collection(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_add_features_collection()
# Seed one entry under each collection's rfcs dir + mirror them in.
fake.files[("wiggleverse", "meta", "main", "rfcs/a.md")] = {
"content": _entry_md("a", "Default A"), "sha": "sa"}
fake.files[("wiggleverse", "meta", "main", "features/rfcs/b.md")] = {
"content": _entry_md("b", "Feature B"), "sha": "sb"}
from app import cache as cache_mod, gitea as gitea_mod
from app.config import load_config
cfg = load_config()
asyncio.run(cache_mod.refresh_meta_repo(cfg, gitea_mod.Gitea(cfg)))
r = client.get("/api/projects/default/collections/features/rfcs")
assert r.status_code == 200, r.text
assert [i["slug"] for i in r.json()["items"]] == ["b"]
# The default collection still serves only its own entry.
r2 = client.get("/api/projects/default/collections/default/rfcs")
assert [i["slug"] for i in r2.json()["items"]] == ["a"]
def test_scoped_propose_writes_into_collection_subfolder(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_add_features_collection()
provision_user_row(user_id=3, login="alice", role="contributor")
# §22 S3: an explicitly-created collection requires an explicit scope
# grant to write (the grandfathered baseline covers only `default`).
db.conn().execute(
"INSERT OR REPLACE INTO memberships (scope_type, scope_id, user_id, role) "
"VALUES ('collection', 'features', 3, 'contributor')")
sign_in_as(client, user_id=3, gitea_login="alice", display_name="Alice",
role="contributor", email="alice@test")
r = client.post(
"/api/projects/default/collections/features/rfcs/propose",
json={"title": "New B", "slug": "newb", "pitch": "x", "tags": []})
assert r.status_code == 200, r.text
# The bot wrote the entry under features/rfcs/, not rfcs/.
keys = {(k[1], k[3]) for k in fake.files
if k[1] == "meta" and k[3].endswith("newb.md")}
assert ("meta", "features/rfcs/newb.md") in keys
def test_scoped_routes_404_for_collection_outside_project(app_with_fake_gitea):
app, _ = app_with_fake_gitea
with TestClient(app) as client:
r = client.get("/api/projects/default/collections/nope/rfcs")
assert r.status_code == 404
def test_s2_anonymous_empty_public_collection(app_with_fake_gitea):
"""C3.6 (@S2): a public collection with no entries; an anonymous visitor
lands on its catalog an empty catalog (200, no items), and the propose
action is not available to them (the propose route rejects anonymous)."""
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_add_features_collection() # public, no entries
# Anonymous (no session cookie) reads the empty catalog — 200, [].
r = client.get("/api/projects/default/collections/features/rfcs")
assert r.status_code == 200, r.text
assert r.json()["items"] == []
# No propose action for an anonymous visitor.
r2 = client.post(
"/api/projects/default/collections/features/rfcs/propose",
json={"title": "X", "slug": "x", "pitch": "p", "tags": []})
assert r2.status_code == 401
@@ -0,0 +1,202 @@
"""§22 S5 — create-project vertical: a global Owner POSTs `/api/projects`, the
bot provisions a Gitea content repo + commits the project to `projects.yaml`, and
the registry mirror upserts the `projects` + default `collections` rows (registry
stays the source of truth). Plus the deployment-directory empty-state signals
(`viewer.can_create_project`, `default_project_readable`) that drive C3.1/C3.2.
"""
from __future__ import annotations
from fastapi.testclient import TestClient
from app import db
from test_propose_vertical import ( # noqa: F401
app_with_fake_gitea, tmp_env, provision_user_row, sign_in_as,
)
def test_create_project_provisions_repo_commits_registry_and_mirrors(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
provision_user_row(user_id=1, login="ben", role="owner")
sign_in_as(client, user_id=1, gitea_login="ben", display_name="Ben",
role="owner", email="ben@test")
r = client.post("/api/projects",
json={"project_id": "acme", "name": "Acme", "type": "bdd",
"visibility": "public"})
assert r.status_code == 200, r.text
body = r.json()
assert body["id"] == "acme"
assert body["type"] == "bdd"
assert body["visibility"] == "public"
# The bot provisioned the content repo (default name <id>-content) and
# seeded a README so `main` exists.
assert ("wiggleverse", "acme-content") in fake.repos
readme = fake.files.get(("wiggleverse", "acme-content", "main", "README.md"))
assert readme is not None and "acme" in readme["content"]
# The bot committed the new project into projects.yaml.
reg = fake.files.get(("wiggleverse", "registry", "main", "projects.yaml"))
assert reg is not None
assert "id: acme" in reg["content"]
assert "content_repo: acme-content" in reg["content"]
# The registry refresh mirrored a projects row + its default collection.
prow = db.conn().execute(
"SELECT name, content_repo, visibility FROM projects WHERE id='acme'"
).fetchone()
assert (prow["name"], prow["content_repo"], prow["visibility"]) == (
"Acme", "acme-content", "public")
crow = db.conn().execute(
"SELECT type, project_id, subfolder FROM collections WHERE id='acme'"
).fetchone()
assert (crow["type"], crow["project_id"], crow["subfolder"]) == ("bdd", "acme", "")
# It is now visible in the deployment directory and readable.
ids = {p["id"] for p in client.get("/api/deployment").json()["projects"]}
assert "acme" in ids
assert client.get("/api/projects/acme").status_code == 200
# An audit row records the structural action with the global Owner actor.
act = db.conn().execute(
"SELECT actor_user_id FROM actions WHERE action_kind='create_project'"
).fetchone()
assert act is not None and act["actor_user_id"] == 1
def test_create_project_custom_content_repo_name(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
provision_user_row(user_id=1, login="ben", role="owner")
sign_in_as(client, user_id=1, gitea_login="ben", display_name="Ben",
role="owner", email="ben@test")
r = client.post("/api/projects",
json={"project_id": "beta", "name": "Beta", "type": "document",
"content_repo": "beta-corpus"})
assert r.status_code == 200, r.text
assert ("wiggleverse", "beta-corpus") in fake.repos
prow = db.conn().execute(
"SELECT content_repo FROM projects WHERE id='beta'"
).fetchone()
assert prow["content_repo"] == "beta-corpus"
def test_create_project_requires_global_owner(app_with_fake_gitea):
# A plain deployment contributor is not a global Owner (C: + New project is a
# global-Owner action), even though they may create collections.
app, _ = app_with_fake_gitea
with TestClient(app) as client:
provision_user_row(user_id=2, login="alice", role="contributor")
sign_in_as(client, user_id=2, gitea_login="alice", display_name="Alice",
role="contributor", email="alice@test")
r = client.post("/api/projects",
json={"project_id": "x", "name": "X", "type": "bdd"})
assert r.status_code == 403
def test_create_project_global_owner_grant_permitted(app_with_fake_gitea):
# An explicit global-scope Owner grant (not a deployment owner/admin) may
# create projects.
app, _ = app_with_fake_gitea
with TestClient(app) as client:
provision_user_row(user_id=3, login="gina", role="contributor")
db.conn().execute(
"INSERT INTO memberships (scope_type, scope_id, user_id, role) "
"VALUES ('global', '*', 3, 'owner')"
)
sign_in_as(client, user_id=3, gitea_login="gina", display_name="Gina",
role="contributor", email="gina@test")
r = client.post("/api/projects",
json={"project_id": "gproj", "name": "G", "type": "document"})
assert r.status_code == 200, r.text
def test_create_project_anonymous_rejected(app_with_fake_gitea):
app, _ = app_with_fake_gitea
with TestClient(app) as client:
r = client.post("/api/projects",
json={"project_id": "x", "name": "X", "type": "bdd"})
assert r.status_code in (401, 403)
def test_create_project_rejects_duplicate(app_with_fake_gitea):
app, _ = app_with_fake_gitea
with TestClient(app) as client:
provision_user_row(user_id=1, login="ben", role="owner")
sign_in_as(client, user_id=1, gitea_login="ben", display_name="Ben",
role="owner", email="ben@test")
ok = client.post("/api/projects",
json={"project_id": "acme", "name": "Acme", "type": "bdd"})
assert ok.status_code == 200, ok.text
dup = client.post("/api/projects",
json={"project_id": "acme", "name": "Acme 2", "type": "bdd"})
assert dup.status_code == 409
def test_create_project_rejects_reserved_default_id(app_with_fake_gitea):
app, _ = app_with_fake_gitea
with TestClient(app) as client:
provision_user_row(user_id=1, login="ben", role="owner")
sign_in_as(client, user_id=1, gitea_login="ben", display_name="Ben",
role="owner", email="ben@test")
r = client.post("/api/projects",
json={"project_id": "default", "name": "X", "type": "bdd"})
assert r.status_code == 422
def test_create_project_rejects_bad_type(app_with_fake_gitea):
app, _ = app_with_fake_gitea
with TestClient(app) as client:
provision_user_row(user_id=1, login="ben", role="owner")
sign_in_as(client, user_id=1, gitea_login="ben", display_name="Ben",
role="owner", email="ben@test")
r = client.post("/api/projects",
json={"project_id": "x", "name": "X", "type": "nonsense"})
assert r.status_code == 422
# --- C3.1 / C3.2: deployment-directory empty-state signals ------------------
def test_deployment_owner_sees_create_project_capability(app_with_fake_gitea):
# C3.1: a global Owner is offered the create-project action.
app, _ = app_with_fake_gitea
with TestClient(app) as client:
provision_user_row(user_id=1, login="ben", role="owner")
sign_in_as(client, user_id=1, gitea_login="ben", display_name="Ben", role="owner")
body = client.get("/api/deployment").json()
assert body["viewer"]["can_create_project"] is True
def test_deployment_non_owner_no_create_capability(app_with_fake_gitea):
# C3.2: a granted account with no roles is not offered create-project.
app, _ = app_with_fake_gitea
with TestClient(app) as client:
provision_user_row(user_id=2, login="vee", role="contributor")
sign_in_as(client, user_id=2, gitea_login="vee", display_name="Vee", role="contributor")
body = client.get("/api/deployment").json()
assert body["viewer"]["can_create_project"] is False
def test_deployment_default_readable_when_default_is_public(app_with_fake_gitea):
# The seeded default project is public → readable → the N=1 redirect target
# is valid (land-in-corpus preserved).
app, _ = app_with_fake_gitea
with TestClient(app) as client:
body = client.get("/api/deployment").json()
assert body["default_project_readable"] is True
def test_deployment_default_not_readable_when_only_gated(app_with_fake_gitea):
# C3.2: the only project is gated; a granted non-member sees no visible
# projects AND default_project_readable False → the frontend renders the
# empty directory (no 404 bounce).
app, _ = app_with_fake_gitea
with TestClient(app) as client:
db.conn().execute("UPDATE projects SET visibility='gated' WHERE id='default'")
db.conn().execute("UPDATE collections SET visibility='gated' WHERE id='default'")
provision_user_row(user_id=2, login="vee", role="contributor")
sign_in_as(client, user_id=2, gitea_login="vee", display_name="Vee", role="contributor")
body = client.get("/api/deployment").json()
assert body["projects"] == []
assert body["default_project_readable"] is False
@@ -0,0 +1,165 @@
"""End-to-end integration tests for the deployed-environment E2E
test-auth path (`POST /auth/test/login`).
This endpoint is a **deliberately gated auth shortcut** for running the
Playwright E2E suite against a *deployed* environment (PPE) that has no
Mailpit OTC sink and no direct SQLite access to inject an owner row
the two scaffolds the Tier-1 stack relies on. It mints an authenticated
**owner** session for a single, pre-configured test identity, but ONLY
when the deployment has explicitly opted in by setting BOTH
`E2E_TEST_AUTH_SECRET` and `E2E_TEST_AUTH_EMAIL`. It is fail-closed:
* Off by default with neither (or only one) env var set, the route
is invisible (404), so a production deployment that never sets them
cannot be coaxed into minting a session.
* Even when enabled, it requires the caller to present the shared
secret in the `X-Test-Auth-Secret` header (constant-time compare),
and it will only mint a session for the one configured email any
other address is refused (403). So the blast radius of an enabled
PPE is a single throwaway owner identity, and the secret is the
trust boundary.
The tests below pin every branch of that gate.
"""
from __future__ import annotations
from test_propose_vertical import ( # noqa: F401
FakeGitea,
app_with_fake_gitea,
tmp_env,
)
SECRET = "ppe-e2e-shared-secret-value"
EMAIL = "e2e-owner@example.test"
def test_test_login_is_404_when_disabled(app_with_fake_gitea):
"""Neither env var set (the default, incl. production) → the route
does not exist."""
from fastapi.testclient import TestClient
app, _fake = app_with_fake_gitea
with TestClient(app) as client:
r = client.post(
"/auth/test/login",
json={"email": EMAIL},
headers={"X-Test-Auth-Secret": SECRET},
)
assert r.status_code == 404, r.text
def test_test_login_is_404_when_only_email_is_set(app_with_fake_gitea, monkeypatch):
"""Half-configured (email but no secret) must NOT open the route —
a framework auth shortcut gated only by a known email would be far
too weak."""
from fastapi.testclient import TestClient
monkeypatch.setenv("E2E_TEST_AUTH_EMAIL", EMAIL)
monkeypatch.delenv("E2E_TEST_AUTH_SECRET", raising=False)
app, _fake = app_with_fake_gitea
with TestClient(app) as client:
r = client.post(
"/auth/test/login",
json={"email": EMAIL},
headers={"X-Test-Auth-Secret": SECRET},
)
assert r.status_code == 404, r.text
def test_test_login_refuses_wrong_secret(app_with_fake_gitea, monkeypatch):
"""Enabled, but a bad/absent secret → 404 (don't advertise the
route's existence to an unauthenticated caller)."""
from fastapi.testclient import TestClient
monkeypatch.setenv("E2E_TEST_AUTH_SECRET", SECRET)
monkeypatch.setenv("E2E_TEST_AUTH_EMAIL", EMAIL)
app, _fake = app_with_fake_gitea
with TestClient(app) as client:
# Wrong secret.
r = client.post(
"/auth/test/login",
json={"email": EMAIL},
headers={"X-Test-Auth-Secret": "not-the-secret"},
)
assert r.status_code == 404, r.text
# Absent secret.
r = client.post("/auth/test/login", json={"email": EMAIL})
assert r.status_code == 404, r.text
def test_test_login_refuses_unconfigured_email(app_with_fake_gitea, monkeypatch):
"""Right secret but an email other than the single configured
identity 403. Even a secret-bearer can only mint the one test
owner."""
from fastapi.testclient import TestClient
monkeypatch.setenv("E2E_TEST_AUTH_SECRET", SECRET)
monkeypatch.setenv("E2E_TEST_AUTH_EMAIL", EMAIL)
app, _fake = app_with_fake_gitea
with TestClient(app) as client:
r = client.post(
"/auth/test/login",
json={"email": "someone-else@example.test"},
headers={"X-Test-Auth-Secret": SECRET},
)
assert r.status_code == 403, r.text
def test_test_login_mints_owner_session(app_with_fake_gitea, monkeypatch):
"""The happy path: right secret + configured email → an authenticated
session whose user is a GRANTED OWNER (so the metadata write paths
SLICE-4 edit, SLICE-5 bulk accept it), persisted on a fresh
`users` row."""
from fastapi.testclient import TestClient
monkeypatch.setenv("E2E_TEST_AUTH_SECRET", SECRET)
monkeypatch.setenv("E2E_TEST_AUTH_EMAIL", EMAIL)
app, _fake = app_with_fake_gitea
with TestClient(app) as client:
r = client.post(
"/auth/test/login",
json={"email": EMAIL},
headers={"X-Test-Auth-Secret": SECRET},
)
assert r.status_code == 200, r.text
# The session cookie now surfaces an authenticated owner.
me = client.get("/api/auth/me").json()
assert me["authenticated"] is True
assert me["user"]["email"] == EMAIL
assert me["user"]["role"] == "owner"
assert me["user"]["permission_state"] == "granted"
# Idempotent: a second login reuses the same row (still owner).
r = client.post(
"/auth/test/login",
json={"email": EMAIL},
headers={"X-Test-Auth-Secret": SECRET},
)
assert r.status_code == 200, r.text
from app import db
rows = db.conn().execute(
"SELECT role, permission_state FROM users WHERE email = ? COLLATE NOCASE",
(EMAIL,),
).fetchall()
assert len(rows) == 1
assert rows[0]["role"] == "owner"
assert rows[0]["permission_state"] == "granted"
def test_test_login_is_case_insensitive_on_email(app_with_fake_gitea, monkeypatch):
"""The configured-email check matches case-insensitively, mirroring
how the rest of the auth stack treats email."""
from fastapi.testclient import TestClient
monkeypatch.setenv("E2E_TEST_AUTH_SECRET", SECRET)
monkeypatch.setenv("E2E_TEST_AUTH_EMAIL", EMAIL)
app, _fake = app_with_fake_gitea
with TestClient(app) as client:
r = client.post(
"/auth/test/login",
json={"email": EMAIL.upper()},
headers={"X-Test-Auth-Secret": SECRET},
)
assert r.status_code == 200, r.text
+87
View File
@@ -0,0 +1,87 @@
"""§22.4a SLICE-3 — pure facet field-set + filter/count (PUC-3).
Per docs/design/2026-06-06-configurable-collection-metadata.md §5.1, §6.4.
"""
from app import facets
PRIORITY = {"priority": {"type": "enum", "values": ["P0", "P1", "P2"]}}
SCHEMA = {
"priority": {"type": "enum", "values": ["P0", "P1", "P2"]},
"tags": {"type": "tags"},
"owner": {"type": "text"},
}
def _e(slug, state="active", malformed=False, **meta):
return {"slug": slug, "state": state, "metadata_malformed": malformed, "meta": meta}
def test_facet_fields_orders_declared_then_state_skips_text():
# enum + tags in declaration order, text skipped, state appended last.
assert facets.facet_fields(SCHEMA) == [
("priority", "enum"), ("tags", "tags"), ("state", "enum")]
def test_no_schema_yields_no_facets():
# INV-5: a collection with no fields has no facets (frontend keeps chips).
assert facets.facet_fields(None) == []
assert facets.facet_fields({}) == []
def test_filter_and_count_basic_counts():
entries = [
_e("a", priority="P0", tags=["checkout"]),
_e("b", priority="P0", tags=["cart"]),
_e("c", priority="P1", tags=["checkout", "cart"]),
]
items, fac = facets.filter_and_count(entries, SCHEMA, {})
assert {i["slug"] for i in items} == {"a", "b", "c"}
assert fac["priority"] == {"P0": 2, "P1": 1}
assert fac["tags"] == {"checkout": 2, "cart": 2}
assert fac["state"] == {"active": 3}
def test_filter_compose_or_within_and_across():
entries = [
_e("a", priority="P0", tags=["checkout"]), # P0 + checkout
_e("b", priority="P0", tags=["cart"]), # P0, no checkout
_e("c", priority="P1", tags=["checkout"]), # checkout, not P0
]
# priority=P0 AND tags=checkout → only "a".
items, _ = facets.filter_and_count(
entries, SCHEMA, {"priority": {"P0"}, "tags": {"checkout"}})
assert {i["slug"] for i in items} == {"a"}
def test_drilldown_counts_exclude_own_field_selection():
entries = [
_e("a", priority="P0"),
_e("b", priority="P1"),
_e("c", priority="P1"),
]
# With P0 selected, the priority facet still counts P1 over the set that
# ignores priority's own selection — so P1 stays switchable.
_, fac = facets.filter_and_count(entries, SCHEMA, {"priority": {"P0"}})
assert fac["priority"] == {"P0": 1, "P1": 2}
def test_malformed_toggle_narrows_items_and_counts():
entries = [
_e("a", state="active", malformed=True, priority="P9"),
_e("b", state="active", malformed=False, priority="P0"),
]
items, fac = facets.filter_and_count(entries, SCHEMA, {}, only_malformed=True)
assert {i["slug"] for i in items} == {"a"}
assert fac["priority"] == {"P9": 1}
def test_missing_value_contributes_no_facet_value():
entries = [_e("a", priority="P0"), _e("b")] # b has no priority
_, fac = facets.filter_and_count(entries, SCHEMA, {})
assert fac["priority"] == {"P0": 1}
def test_allowed_filter_keys():
assert facets.allowed_filter_keys(PRIORITY) == {"priority", "state",
"unreviewed", "malformed"}
+134
View File
@@ -0,0 +1,134 @@
"""§22.4a SLICE-3 integration — faceted list endpoint (PUC-3, §6.4).
Through the real API: schema-declared facets, filter params (OR within / AND
across), drill-down counts, malformed toggle, unknown-field 400, and meta_json
persistence. Reuses the fake-Gitea harness from test_metadata_cache.
"""
from __future__ import annotations
import asyncio
import json
import yaml
from fastapi.testclient import TestClient
from app import cache, db, gitea as gitea_mod
from app.config import load_config
from test_propose_vertical import ( # noqa: F401 (fixtures)
app_with_fake_gitea,
tmp_env,
)
# The fake-Gitea harness seeds a single project whose id is the literal
# 'default' (see test_s1_collection_grain_vertical), served at
# /api/projects/default/rfcs.
PID = "default"
def _refresh():
cfg = load_config()
asyncio.run(cache.refresh_meta_repo(cfg, gitea_mod.Gitea(cfg)))
def _set_default_fields_schema(schema):
db.conn().execute(
"UPDATE collections SET config_json = ? WHERE id = 'default'",
(json.dumps({"fields": schema}),))
def _seed(fake, slug, *, state="active", **front):
fm = {"slug": slug, "title": slug.title(), "state": state, **front}
body = yaml.safe_dump(fm, sort_keys=False).strip()
fake.files[("wiggleverse", "meta", "main", f"rfcs/{slug}.md")] = {
"content": f"---\n{body}\n---\n\nBody.\n", "sha": slug}
def test_facets_and_counts_returned(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_set_default_fields_schema({
"priority": {"type": "enum", "values": ["P0", "P1", "P2"]},
"tags": {"type": "tags"},
})
_seed(fake, "a", priority="P0", tags=["checkout"])
_seed(fake, "b", priority="P0", tags=["cart"])
_seed(fake, "c", priority="P1", tags=["checkout", "cart"])
_refresh()
res = client.get(f"/api/projects/{PID}/rfcs")
assert res.status_code == 200
body = res.json()
assert {i["slug"] for i in body["items"]} == {"a", "b", "c"}
assert body["facets"]["priority"] == {"P0": 2, "P1": 1}
assert body["facets"]["tags"] == {"checkout": 2, "cart": 2}
assert body["facets"]["state"] == {"active": 3}
def test_filter_params_compose(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_set_default_fields_schema({
"priority": {"type": "enum", "values": ["P0", "P1"]},
"tags": {"type": "tags"},
})
_seed(fake, "a", priority="P0", tags=["checkout"])
_seed(fake, "b", priority="P0", tags=["cart"])
_seed(fake, "c", priority="P1", tags=["checkout"])
_refresh()
res = client.get(
f"/api/projects/{PID}/rfcs",
params={"priority": "P0", "tags": "checkout"})
assert {i["slug"] for i in res.json()["items"]} == {"a"}
def test_unknown_filter_field_400(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_set_default_fields_schema(
{"priority": {"type": "enum", "values": ["P0"]}})
_seed(fake, "a", priority="P0")
_refresh()
res = client.get(f"/api/projects/{PID}/rfcs",
params={"nonsense": "x"})
assert res.status_code == 400
def test_malformed_toggle(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_set_default_fields_schema(
{"priority": {"type": "enum", "values": ["P0", "P1"]}})
_seed(fake, "good", priority="P0")
_seed(fake, "bad", priority="P9") # not in values → malformed (INV-3)
_refresh()
res = client.get(f"/api/projects/{PID}/rfcs",
params={"malformed": "true"})
slugs = {i["slug"] for i in res.json()["items"]}
assert slugs == {"bad"}
def test_no_schema_no_facets(app_with_fake_gitea):
# INV-5: the default document collection (no fields) returns empty facets.
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_seed(fake, "plain", tags=["whatever"])
_refresh()
body = client.get(f"/api/projects/{PID}/rfcs").json()
assert body["facets"] == {}
def test_meta_json_persisted_at_ingest(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app):
_seed(fake, "withmeta", priority="P0", tags=["x"])
_refresh()
row = db.conn().execute(
"SELECT meta_json FROM cached_rfcs WHERE slug = 'withmeta'"
).fetchone()
meta = json.loads(row["meta_json"])
assert meta["priority"] == "P0"
assert meta["tags"] == ["x"]
@@ -0,0 +1,192 @@
"""G-15 — the branch/body subsystem is three-tier (project/collection) aware.
Before G-15 the branch-body GET resolved every meta-resident entry to the
default project's content repo at `rfcs/<slug>.md`, so an entry in a named
collection (subfolder) or a non-default project's repo rendered a BLANK
canonical body and its edit/PR/body-write paths hit the wrong file. These tests
seed entries outside the default collection and assert the collection-scoped
body-read routes (and the now-collection-aware slug-only routes) read the
correct repo + subfolder.
"""
from __future__ import annotations
import asyncio
from app import cache as cache_mod, db, gitea as gitea_mod
from app.config import load_config
from fastapi.testclient import TestClient # noqa: E402
from test_propose_vertical import ( # noqa: E402,F401
app_with_fake_gitea, tmp_env, provision_user_row, sign_in_as,
)
def _entry_md(slug, title, state="active"):
return f"---\nslug: {slug}\ntitle: {title}\nstate: {state}\n---\nthe canonical body\n"
def _add_features_collection():
"""A named 'features' collection (subfolder 'features') under the seeded
default project same content repo ('meta'), distinct subfolder."""
db.conn().execute(
"INSERT OR REPLACE INTO collections (id, project_id, type, subfolder, "
"initial_state, visibility, name, created_at, updated_at) VALUES "
"('features','default','document','features','super-draft','public','Features', "
"datetime('now'), datetime('now'))")
def _add_distinct_project():
"""A second project with its OWN content repo + a default (root) collection
the live OHM dogfood shape (rfc-app project at rfc-app-content)."""
db.conn().execute(
"INSERT OR REPLACE INTO projects (id, name, content_repo, visibility, updated_at) "
"VALUES ('rfc-app','RFC App','rfc-app-content','public', datetime('now'))")
db.conn().execute(
"INSERT OR REPLACE INTO collections (id, project_id, type, subfolder, "
"initial_state, visibility, name, created_at, updated_at) VALUES "
"('rfc-app','rfc-app','document','','super-draft','public','RFC App', "
"datetime('now'), datetime('now'))")
def _mirror():
cfg = load_config()
asyncio.run(cache_mod.refresh_meta_repo(cfg, gitea_mod.Gitea(cfg)))
# --- named collection (subfolder) under the default project -------------------
def test_branch_view_named_collection_reads_subfolder(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_add_features_collection()
fake.files[("wiggleverse", "meta", "main", "features/rfcs/feat.md")] = {
"content": _entry_md("feat", "Feature Entry"), "sha": "sf"}
_mirror()
# cached_rfcs is keyed by the named collection.
assert db.conn().execute(
"SELECT collection_id FROM cached_rfcs WHERE slug='feat'"
).fetchone()["collection_id"] == "features"
# Collection-scoped branch view renders the body (was blank pre-G-15).
r = client.get(
"/api/projects/default/collections/features/rfcs/feat/branches/main")
assert r.status_code == 200, r.text
assert r.json()["body"] == "the canonical body\n"
# The slug-only legacy route is now collection-aware too: it resolves
# the entry's own subfolder via the cached row, so it also renders.
r2 = client.get("/api/rfcs/feat/branches/main")
assert r2.status_code == 200, r2.text
assert r2.json()["body"] == "the canonical body\n"
def test_main_view_named_collection_scoped_route(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_add_features_collection()
fake.files[("wiggleverse", "meta", "main", "features/rfcs/feat.md")] = {
"content": _entry_md("feat", "Feature Entry"), "sha": "sf"}
_mirror()
r = client.get("/api/projects/default/collections/features/rfcs/feat/main")
assert r.status_code == 200, r.text
assert r.json()["title"] == "Feature Entry"
# --- non-default project with a DISTINCT content repo (the OHM dogfood) -------
def test_branch_view_distinct_project_repo(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_add_distinct_project()
fake.files[("wiggleverse", "rfc-app-content", "main", "rfcs/scoped.md")] = {
"content": _entry_md("scoped", "Scoped Admin IA"), "sha": "s1"}
_mirror()
assert db.conn().execute(
"SELECT collection_id FROM cached_rfcs WHERE slug='scoped'"
).fetchone()["collection_id"] == "rfc-app"
# Scoped read resolves the entry's OWN content repo (rfc-app-content).
r = client.get(
"/api/projects/rfc-app/collections/rfc-app/rfcs/scoped/branches/main")
assert r.status_code == 200, r.text
assert r.json()["body"] == "the canonical body\n"
# And the slug-only route resolves the right repo via the cached row.
r2 = client.get("/api/rfcs/scoped/branches/main")
assert r2.status_code == 200, r2.text
assert r2.json()["body"] == "the canonical body\n"
# --- guards + regression ------------------------------------------------------
def test_scoped_branch_route_404_for_collection_outside_project(app_with_fake_gitea):
app, _ = app_with_fake_gitea
with TestClient(app) as client:
r = client.get(
"/api/projects/default/collections/nope/rfcs/x/branches/main")
assert r.status_code == 404
def _seed_super_draft_in_collection(fake, *, slug, collection_id, subfolder, owners):
import json as _json
import yaml
md_path = f"{subfolder}/rfcs/{slug}.md" if subfolder else f"rfcs/{slug}.md"
fm = {"slug": slug, "title": slug.title(), "state": "super-draft", "id": None,
"repo": None, "proposed_by": owners[0], "proposed_at": "2026-05-23",
"graduated_at": None, "graduated_by": None,
"owners": owners, "arbiters": owners[:1], "tags": []}
body = "the body\n"
text = f"---\n{yaml.safe_dump(fm, sort_keys=False).rstrip()}\n---\n\n{body}"
sha = fake._next_sha()
fake.files[("wiggleverse", "meta", "main", md_path)] = {"content": text, "sha": sha}
db.conn().execute(
"INSERT OR REPLACE INTO cached_rfcs (slug, title, state, rfc_id, repo, "
"proposed_by, proposed_at, owners_json, arbiters_json, tags_json, body, "
"body_sha, collection_id, last_main_commit_at, last_entry_commit_at) "
"VALUES (?,?, 'super-draft', NULL, NULL, ?, '2026-05-23', ?, ?, '[]', ?, ?, ?, "
"datetime('now'), datetime('now'))",
(slug, slug.title(), owners[0], _json.dumps(owners), _json.dumps(owners[:1]),
body, sha, collection_id))
def test_graduate_in_named_collection_writes_to_subfolder(app_with_fake_gitea):
"""G-15 write path: graduating a super-draft that lives in a named
collection flips the entry in that collection's `<subfolder>/rfcs/<slug>.md`,
not the default `rfcs/<slug>.md`."""
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_add_features_collection()
provision_user_row(user_id=1, login="ben", role="owner")
_seed_super_draft_in_collection(
fake, slug="gradme", collection_id="features", subfolder="features",
owners=["ben"])
sign_in_as(client, user_id=1, gitea_login="ben", display_name="Ben",
role="owner", email="ben@x")
r = client.post("/api/rfcs/gradme/graduate?_sync=1",
json={"rfc_id": "RFC-0007", "owners": ["ben"]})
assert r.status_code == 200, r.text
# The flip landed in the collection's subfolder, not the repo root.
sc = fake.files.get(
("wiggleverse", "meta", "main", "features/rfcs/gradme.meta.yaml"))
assert sc is not None, "sidecar not written under the collection subfolder"
import yaml as _yaml
assert _yaml.safe_load(sc["content"])["state"] == "active"
# Nothing was written to the default repo-root path.
assert ("wiggleverse", "meta", "main", "rfcs/gradme.meta.yaml") not in fake.files
def test_default_collection_entry_still_renders(app_with_fake_gitea):
"""Regression: the default-collection path (repo root `rfcs/<slug>.md` in
the default content repo) is unchanged by the G-15 resolution."""
app, fake = app_with_fake_gitea
with TestClient(app) as client:
fake.files[("wiggleverse", "meta", "main", "rfcs/base.md")] = {
"content": _entry_md("base", "Baseline"), "sha": "sb"}
_mirror()
r = client.get("/api/rfcs/base/branches/main")
assert r.status_code == 200, r.text
assert r.json()["body"] == "the canonical body\n"
r2 = client.get("/api/projects/default/collections/default/rfcs/base/branches/main")
assert r2.status_code == 200, r2.text
assert r2.json()["body"] == "the canonical body\n"
+108
View File
@@ -0,0 +1,108 @@
"""G-15 — the §22 three-tier write-path resolver.
`projects.content_repo_for_collection` and `projects.entry_location` resolve an
entry's git location (org, content_repo, md_path) from its *collection*
(collection project content_repo, plus the collection subfolder) instead of
the deployment default. This is the keystone the branch/edit/body/graduation
write paths share so an entry outside the default project's default collection
reads/writes the correct file.
"""
from __future__ import annotations
import tempfile
from pathlib import Path
from app import collections as collections_mod, db, projects as projects_mod
from app.config import Config
def _db() -> Config:
cfg = Config(
gitea_url="x", gitea_bot_user="x", gitea_bot_token="x", gitea_org="wiggleverse",
registry_repo="registry", oauth_client_id="x",
oauth_client_secret="x", app_url="x", secret_key="x",
database_path=Path(tempfile.mkdtemp(prefix="g15loc-")) / "t.db",
owner_gitea_login="x", webhook_secret="x",
)
db.run_migrations(cfg)
if db._CONN is not None:
db._CONN.close()
db._CONN = None
db.init(cfg)
return cfg
def _seed():
# Default project (its content_repo is the deployment default) + a second
# project with a DISTINCT content_repo, each with a default + a named
# (subfolder) collection.
db.conn().execute(
"INSERT OR REPLACE INTO projects (id, name, content_repo, visibility, updated_at) "
"VALUES ('default','Default','meta','public', datetime('now'))")
db.conn().execute(
"INSERT OR REPLACE INTO projects (id, name, content_repo, visibility, updated_at) "
"VALUES ('rfc-app','RFC App','rfc-app-content','public', datetime('now'))")
rows = [
("default", "default", ""),
("features", "default", "features"),
("rfc-app", "rfc-app", ""),
("specs", "rfc-app", "specs"),
]
for cid, pid, sub in rows:
db.conn().execute(
"INSERT OR REPLACE INTO collections (id, project_id, type, subfolder, "
"initial_state, visibility, created_at, updated_at) VALUES "
"(?,?, 'document', ?, 'super-draft','public', datetime('now'), datetime('now'))",
(cid, pid, sub))
def test_content_repo_for_collection_resolves_per_project():
_db()
_seed()
# Default project's collections → the default content repo.
assert projects_mod.content_repo_for_collection("default") == "meta"
assert projects_mod.content_repo_for_collection("features") == "meta"
# The second project's collections → its own content repo.
assert projects_mod.content_repo_for_collection("rfc-app") == "rfc-app-content"
assert projects_mod.content_repo_for_collection("specs") == "rfc-app-content"
def test_content_repo_for_collection_unknown_is_none():
_db()
_seed()
assert projects_mod.content_repo_for_collection("nope") is None
def test_entry_location_default_collection_repo_root():
cfg = _db()
_seed()
org, repo, path = projects_mod.entry_location(cfg, "default", "alpha")
assert (org, repo, path) == ("wiggleverse", "meta", "rfcs/alpha.md")
def test_entry_location_named_collection_uses_subfolder():
cfg = _db()
_seed()
org, repo, path = projects_mod.entry_location(cfg, "features", "beta")
assert (org, repo, path) == ("wiggleverse", "meta", "features/rfcs/beta.md")
def test_entry_location_other_project_distinct_repo():
cfg = _db()
_seed()
# Named collection in a non-default project: distinct repo AND subfolder.
org, repo, path = projects_mod.entry_location(cfg, "specs", "gamma")
assert (org, repo, path) == ("wiggleverse", "rfc-app-content", "specs/rfcs/gamma.md")
# Default collection of the non-default project: distinct repo, repo root.
org, repo, path = projects_mod.entry_location(cfg, "rfc-app", "delta")
assert (org, repo, path) == ("wiggleverse", "rfc-app-content", "rfcs/delta.md")
def test_entry_location_unknown_collection_falls_back_to_default_repo():
cfg = _db()
_seed()
# An entry whose collection_id is missing/unknown must still resolve to a
# usable location (the deployment default repo, repo root) rather than an
# empty repo — the legacy single-corpus behaviour.
org, repo, path = projects_mod.entry_location(cfg, "nope", "epsilon")
assert (org, repo, path) == ("wiggleverse", "meta", "rfcs/epsilon.md")
+45 -10
View File
@@ -48,6 +48,18 @@ PITCH = (
)
def _entry_from_git(fake, slug, branch="main"):
"""§22.4a SLICE-4: read an entry's combined metadata+body from git via the
dual-read parser graduation/claim now write metadata to the sidecar and
keep the body in the `.md`, so an Entry is reconstructed from both."""
from app import metadata
md = fake.files[("wiggleverse", "meta", branch, f"rfcs/{slug}.md")]["content"]
sc = fake.files.get(
("wiggleverse", "meta", branch, f"rfcs/{slug}.meta.yaml"), {}).get("content")
e, _ = metadata.read_entry(md, sc, fallback_slug=slug)
return e
def seed_owned_super_draft(fake: FakeGitea, *, slug: str, title: str, pitch: str,
owners: list[str], arbiters: list[str] | None = None,
proposed_by: str = "alice", tags: list[str] | None = None) -> None:
@@ -190,9 +202,8 @@ def test_graduate_happy_path_flips_in_place_keeping_body(app_with_fake_gitea):
k[1].startswith("rfc-0042") for k in fake.repos
), f"a per-RFC repo was created: {fake.repos}"
# Meta entry on main: state flipped, body KEPT, repo null.
meta_text = fake.files[("wiggleverse", "meta", "main", "rfcs/ohm.md")]["content"]
graduated = entry_mod.parse(meta_text)
# Meta entry on main: state flipped (sidecar), body KEPT (.md), repo null.
graduated = _entry_from_git(fake, "ohm")
assert graduated.state == "active"
assert graduated.id == "RFC-0042"
assert graduated.repo is None
@@ -224,6 +235,34 @@ def test_graduate_happy_path_flips_in_place_keeping_body(app_with_fake_gitea):
assert gone not in kinds, f"retired audit row present: {gone}"
def test_graduate_preserves_unknown_frontmatter_keys(app_with_fake_gitea):
"""§22.4a INV-7: a forward-compat / unknown frontmatter key on the
super-draft entry must ride through the graduation rebuild, not be dropped."""
from fastapi.testclient import TestClient
from app import entry as entry_mod
app, fake = app_with_fake_gitea
with TestClient(app) as client:
provision_user_row(user_id=1, login="ben", role="owner")
seed_owned_super_draft(fake, slug="ohm", title="OHM", pitch=PITCH,
owners=["ben"], arbiters=["ben"])
# Inject an unknown key into the seeded entry's frontmatter.
key = ("wiggleverse", "meta", "main", "rfcs/ohm.md")
e = entry_mod.parse(fake.files[key]["content"])
e.extra["priority"] = "P1"
fake.files[key]["content"] = entry_mod.serialize(e)
sign_in_as(client, user_id=1, gitea_login="ben",
display_name="Ben", role="owner", email="ben@test")
r = client.post("/api/rfcs/ohm/graduate?_sync=1",
json={"rfc_id": "RFC-0042", "owners": ["ben"]})
assert r.status_code == 200, r.text
graduated = _entry_from_git(fake, "ohm")
assert graduated.state == "active"
assert graduated.extra.get("priority") == "P1"
def test_graduate_coexists_with_open_body_edit_pr(app_with_fake_gitea):
"""§9.8 (meta-only): an open meta-repo body-edit PR no longer blocks
graduation the body is kept, so they coexist. /check stays
@@ -571,8 +610,7 @@ def test_graduate_without_number_flips_to_active_null_id_by_slug(app_with_fake_g
assert d["rfc_id"] is None
# Meta entry: active, id null, body kept, graduation stamped.
meta_text = fake.files[("wiggleverse", "meta", "main", "rfcs/ohm.md")]["content"]
graduated = entry_mod.parse(meta_text)
graduated = _entry_from_git(fake, "ohm")
assert graduated.state == "active"
assert graduated.id is None
assert graduated.graduated_by == "ben"
@@ -621,9 +659,7 @@ def test_graduate_with_number_unchanged_when_id_absent_field(app_with_fake_gitea
r = client.post("/api/rfcs/ohm/graduate?_sync=1", json={"owners": ["ben"]})
assert r.status_code == 200, r.text
assert r.json()["rfc_id"] is None
graduated = entry_mod.parse(
fake.files[("wiggleverse", "meta", "main", "rfcs/ohm.md")]["content"]
)
graduated = _entry_from_git(fake, "ohm")
assert graduated.state == "active"
assert graduated.id is None
@@ -649,8 +685,7 @@ def test_claim_opens_meta_pr(app_with_fake_gitea):
d = r.json()
assert d["branch_name"] == "claim/ohm"
text = fake.files[("wiggleverse", "meta", "claim/ohm", "rfcs/ohm.md")]["content"]
ent = entry_mod.parse(text)
ent = _entry_from_git(fake, "ohm", branch="claim/ohm")
assert "alice" in ent.owners
row = db.conn().execute(
+1 -1
View File
@@ -33,7 +33,7 @@ def test_active_initial_state_lands_active_unreviewed(app_with_fake_gitea):
from app import db, entry as entry_mod
app, fake = app_with_fake_gitea
with TestClient(app) as client:
db.conn().execute("UPDATE projects SET initial_state='active' WHERE id='default'")
db.conn().execute("UPDATE collections SET initial_state='active' WHERE project_id='default'")
provision_user_row(user_id=1, login="ben", role="owner")
sign_in_as(client, user_id=1, gitea_login="ben", display_name="Ben", role="owner")
assert _propose(client).status_code == 200
@@ -0,0 +1,339 @@
"""§22.8 S6 — request-to-join a scope + the cross-collection inbox.
A user who knows a (gated) scope exists asks to join it, naming a desired role;
the request is recorded and fanned out to that scope's Owners *across the
subtree* (the cross-collection inbox, §22.11). An Owner accepts which writes
the `memberships` row via memberships.grant or declines, and the requester is
§15-notified either way.
Built by analogy to test_contributions_vertical.py (the per-RFC contribute flow)
and test_s4_invitations_vertical.py (the scope/membership world-builders).
World: project "ohm" owns collections "model" (document, gated) and "features"
(bdd, gated). eve is project Owner; dan is collection Owner of features only;
zoe is a global Owner; ada is a deployment admin. ben is a plain granted account
(no scope role) the would-be joiner.
"""
from __future__ import annotations
from fastapi.testclient import TestClient
from app import db
from test_propose_vertical import ( # noqa: F401 — fixtures land via import
app_with_fake_gitea,
provision_user_row,
sign_in_as,
tmp_env,
)
# ---------------------------------------------------------------------------
# World-builders (mirror the S4 vertical)
# ---------------------------------------------------------------------------
def _project(pid: str, visibility: str = "gated", content_repo: str = "meta") -> None:
db.conn().execute(
"INSERT OR REPLACE INTO projects (id, name, content_repo, visibility, updated_at) "
"VALUES (?, ?, ?, ?, datetime('now'))",
(pid, pid.capitalize(), content_repo, visibility),
)
def _collection(cid: str, project_id: str, *, ctype: str = "document",
visibility: str = "gated") -> None:
db.conn().execute(
"INSERT OR REPLACE INTO collections "
"(id, project_id, type, subfolder, initial_state, visibility, name, created_at, updated_at) "
"VALUES (?, ?, ?, ?, 'super-draft', ?, ?, datetime('now'), datetime('now'))",
(cid, project_id, ctype, cid, visibility, cid.capitalize()),
)
def _grant(scope_type: str, scope_id: str, user_id: int, role: str) -> None:
db.conn().execute(
"INSERT OR REPLACE INTO memberships (scope_type, scope_id, user_id, role) "
"VALUES (?, ?, ?, ?)",
(scope_type, scope_id, user_id, role),
)
def _membership(user_id: int):
rows = db.conn().execute(
"SELECT scope_type, scope_id, role FROM memberships WHERE user_id = ?",
(user_id,),
).fetchall()
return {(r["scope_type"], r["scope_id"], r["role"]) for r in rows}
def _join_requests(scope_type: str, scope_id: str):
rows = db.conn().execute(
"SELECT id, requester_user_id, requested_role, status, granted_role "
"FROM join_requests WHERE scope_type = ? AND scope_id = ?",
(scope_type, scope_id),
).fetchall()
return [dict(r) for r in rows]
def _join_notif_recipients(event_kind: str = "join_request_on_scope") -> set[int]:
return {
r["recipient_user_id"]
for r in db.conn().execute(
"SELECT recipient_user_id FROM notifications WHERE event_kind = ?",
(event_kind,),
)
}
def _seed_world() -> None:
_project("ohm", "gated")
_collection("model", "ohm", ctype="document")
_collection("features", "ohm", ctype="bdd")
provision_user_row(user_id=2, login="ben", role="contributor") # the joiner
provision_user_row(user_id=4, login="dan", role="contributor") # collection Owner (features)
provision_user_row(user_id=5, login="eve", role="contributor") # project Owner
provision_user_row(user_id=6, login="zoe", role="contributor") # global Owner
provision_user_row(user_id=7, login="ada", role="admin") # deployment admin
_grant("project", "ohm", 5, "owner")
_grant("collection", "features", 4, "owner")
_grant("global", "*", 6, "owner")
def _login(client, uid: int, login: str, role: str = "contributor") -> None:
sign_in_as(client, user_id=uid, gitea_login=login, display_name=login.capitalize(), role=role)
# ---------------------------------------------------------------------------
# Request → cross-collection fan-out
# ---------------------------------------------------------------------------
def test_request_to_join_collection_fans_out_to_subtree_owners(app_with_fake_gitea):
"""A request to join a collection lands a row and notifies every Owner whose
reach covers it the collection's Owner, the project's Owner, a global
Owner, and the deployment admin (the cross-collection inbox) never the
requester."""
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
_login(client, 2, "ben")
r = client.post(
"/api/scopes/collection/features/join-requests",
json={"role": "contributor", "message": "I work on BDD corpora."},
)
assert r.status_code == 200, r.text
assert r.json()["status"] == "pending"
reqs = _join_requests("collection", "features")
assert len(reqs) == 1
assert reqs[0]["requester_user_id"] == 2
assert reqs[0]["requested_role"] == "contributor"
assert reqs[0]["status"] == "pending"
# Owners across the subtree are notified; ben (requester) is not.
recips = _join_notif_recipients()
assert {4, 5, 6, 7}.issubset(recips) # dan, eve, zoe, ada
assert 2 not in recips
def test_request_to_join_project_reaches_project_and_global_owners(app_with_fake_gitea):
"""A project-scope request reaches the project's Owners + global Owners +
admin, but NOT a collection-only Owner (their reach doesn't cover the
project)."""
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
_login(client, 2, "ben")
r = client.post(
"/api/scopes/project/ohm/join-requests",
json={"role": "owner"},
)
assert r.status_code == 200, r.text
recips = _join_notif_recipients()
assert {5, 6, 7}.issubset(recips) # eve (project), zoe (global), ada (admin)
assert 4 not in recips # dan is only a collection Owner
# ---------------------------------------------------------------------------
# Accept → writes membership + notifies
# ---------------------------------------------------------------------------
def test_owner_accept_writes_membership_and_notifies(app_with_fake_gitea):
"""The collection Owner accepts; a `memberships` row is written at the
requested scope/role and the requester gets a join_request_accepted inbox
row."""
from app import auth
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
_login(client, 2, "ben")
client.post(
"/api/scopes/collection/features/join-requests",
json={"role": "contributor"},
)
req_id = _join_requests("collection", "features")[0]["id"]
# dan (collection Owner of features) accepts.
_login(client, 4, "dan")
r = client.post(
f"/api/scopes/collection/features/join-requests/{req_id}/accept",
json={},
)
assert r.status_code == 200, r.text
assert r.json()["granted_role"] == "contributor"
# ben now holds the collection role and can contribute there.
assert ("collection", "features", "contributor") in _membership(2)
ben = auth.SessionUser(
user_id=2, gitea_id=2, gitea_login="ben", display_name="Ben",
email="ben@test", avatar_url="", role="contributor", permission_state="granted",
)
assert auth.can_contribute_in_collection(ben, "features") is True
assert auth.can_contribute_in_collection(ben, "model") is False
# the row is closed; the requester is notified.
assert _join_requests("collection", "features")[0]["status"] == "accepted"
_login(client, 2, "ben")
inbox = client.get("/api/notifications").json()["items"]
accepted = [n for n in inbox if n["event_kind"] == "join_request_accepted"]
assert accepted, inbox
assert "Features" in accepted[0]["summary"]
def test_owner_may_narrow_role_on_accept(app_with_fake_gitea):
"""A request for Owner may be accepted as RFC Contributor — the Owner narrows
the grant; the membership row carries the granted (not requested) role."""
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
_login(client, 2, "ben")
client.post("/api/scopes/collection/features/join-requests", json={"role": "owner"})
req_id = _join_requests("collection", "features")[0]["id"]
_login(client, 5, "eve") # project Owner — reach covers the collection
r = client.post(
f"/api/scopes/collection/features/join-requests/{req_id}/accept",
json={"role": "contributor"},
)
assert r.status_code == 200, r.text
assert ("collection", "features", "contributor") in _membership(2)
assert _join_requests("collection", "features")[0]["granted_role"] == "contributor"
def test_owner_decline_notifies_and_grants_nothing(app_with_fake_gitea):
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
_login(client, 2, "ben")
client.post("/api/scopes/collection/features/join-requests", json={"role": "contributor"})
req_id = _join_requests("collection", "features")[0]["id"]
_login(client, 4, "dan")
r = client.post(f"/api/scopes/collection/features/join-requests/{req_id}/decline")
assert r.status_code == 200, r.text
assert _membership(2) == set()
assert _join_requests("collection", "features")[0]["status"] == "declined"
_login(client, 2, "ben")
inbox = client.get("/api/notifications").json()["items"]
assert any(n["event_kind"] == "join_request_declined" for n in inbox)
# ---------------------------------------------------------------------------
# Gates & guards
# ---------------------------------------------------------------------------
def test_duplicate_pending_request_is_conflict(app_with_fake_gitea):
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
_login(client, 2, "ben")
r1 = client.post("/api/scopes/collection/features/join-requests", json={"role": "contributor"})
assert r1.status_code == 200, r1.text
r2 = client.post("/api/scopes/collection/features/join-requests", json={"role": "contributor"})
assert r2.status_code == 409, r2.text
def test_existing_member_cannot_request(app_with_fake_gitea):
"""dan already owns the collection — there is nothing to request (409)."""
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
_login(client, 4, "dan")
r = client.post("/api/scopes/collection/features/join-requests", json={"role": "contributor"})
assert r.status_code == 409, r.text
def test_non_owner_cannot_accept(app_with_fake_gitea):
"""A plain requester (or any non-Owner) is refused the accept action."""
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
_login(client, 2, "ben")
client.post("/api/scopes/collection/features/join-requests", json={"role": "contributor"})
req_id = _join_requests("collection", "features")[0]["id"]
# provision a second plain account that tries to accept
provision_user_row(user_id=12, login="mal", role="contributor")
_login(client, 12, "mal")
r = client.post(
f"/api/scopes/collection/features/join-requests/{req_id}/accept", json={}
)
assert r.status_code == 403, r.text
assert _membership(2) == set()
def test_collection_owner_cannot_act_on_sibling_collection(app_with_fake_gitea):
"""dan owns 'features' only; a request to join 'model' is not his to act on."""
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
_login(client, 2, "ben")
client.post("/api/scopes/collection/model/join-requests", json={"role": "contributor"})
req_id = _join_requests("collection", "model")[0]["id"]
_login(client, 4, "dan")
r = client.post(
f"/api/scopes/collection/model/join-requests/{req_id}/accept", json={}
)
assert r.status_code == 403, r.text
def test_unknown_scope_404(app_with_fake_gitea):
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
_login(client, 2, "ben")
assert client.post(
"/api/scopes/collection/nope/join-requests", json={"role": "contributor"}
).status_code == 404
assert client.post(
"/api/scopes/project/nope/join-requests", json={"role": "contributor"}
).status_code == 404
# 'global' is not a join-able scope_type.
assert client.post(
"/api/scopes/global/*/join-requests", json={"role": "contributor"}
).status_code == 404
def test_join_target_reports_eligibility(app_with_fake_gitea):
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
# ben: eligible (granted, no role).
_login(client, 2, "ben")
t = client.get("/api/scopes/collection/features/join-target").json()
assert t["eligible"] is True
assert t["name"] == "Features"
assert t["current_role"] is None
# after requesting, already_requested flips and eligible drops.
client.post("/api/scopes/collection/features/join-requests", json={"role": "contributor"})
t2 = client.get("/api/scopes/collection/features/join-target").json()
assert t2["already_requested"] is True
assert t2["eligible"] is False
# dan: already a member → ineligible with current_role.
_login(client, 4, "dan")
t3 = client.get("/api/scopes/collection/features/join-target").json()
assert t3["eligible"] is False
assert t3["current_role"] == "owner"
+8 -5
View File
@@ -47,12 +47,15 @@ def test_mark_reviewed_clears_flag(app_with_fake_gitea):
assert row["unreviewed"] == 0
assert row["reviewed_by"] == "ben"
assert row["reviewed_at"] # provenance stamped
# git-side: the entry file on main was rewritten with the cleared flag.
from app import entry as entry_mod
# git-side (§22.4a SLICE-4): the cleared flag now lands in the metadata
# sidecar and the `.md` is lazy-migrated to a clean body-only file (INV-2).
import yaml
sidecar = fake.files[("wiggleverse", "meta", "main", "rfcs/feat.meta.yaml")]["content"]
sc = yaml.safe_load(sidecar)
assert not sc.get("unreviewed") # cleared (omitted when False)
assert sc.get("reviewed_by") == "ben"
written = fake.files[("wiggleverse", "meta", "main", "rfcs/feat.md")]["content"]
e = entry_mod.parse(written)
assert e.unreviewed is False
assert e.reviewed_by == "ben"
assert "---" not in written # body-only, no frontmatter
def test_mark_reviewed_forbidden_for_non_superuser(app_with_fake_gitea):
+223
View File
@@ -0,0 +1,223 @@
"""SLICE-1 unit tests — sidecar metadata: dual-read, unknown-key preservation,
frontmatter stripping, malformed detection.
Pure functions only (no DB / no Gitea). Per
docs/design/2026-06-06-configurable-collection-metadata.md §7.2 (SLICE-1) and
INV-6 (dual-read), INV-7 (unknown keys ride along), INV-2 (clean body).
"""
from __future__ import annotations
import yaml
from app import entry as entry_mod
from app import metadata
LEGACY_MD = """---
slug: view-metrics
title: View today's metrics
state: active
owners:
- ben.stull
tags:
- dashboard
- analytics
priority: P1
owner: hasan
---
This is the prose body.
Second paragraph.
"""
# ---- INV-7: unknown keys ride along ----
def test_parse_preserves_unknown_keys_in_extra():
e = entry_mod.parse(LEGACY_MD)
assert e.extra == {"priority": "P1", "owner": "hasan"}
def test_serialize_round_trip_preserves_unknown_keys():
e = entry_mod.parse(LEGACY_MD)
text = entry_mod.serialize(e)
e2 = entry_mod.parse(text)
assert e2.extra == {"priority": "P1", "owner": "hasan"}
assert e2.tags == ["dashboard", "analytics"]
assert e2.owners == ["ben.stull"]
def test_known_keys_never_leak_into_extra():
e = entry_mod.parse(LEGACY_MD)
for known in ("slug", "title", "state", "owners", "tags"):
assert known not in e.extra
# ---- metadata_dict / sidecar_yaml ----
def test_metadata_dict_merges_known_and_extra():
e = entry_mod.parse(LEGACY_MD)
d = metadata.metadata_dict(e)
assert d["slug"] == "view-metrics"
assert d["title"] == "View today's metrics"
assert d["state"] == "active"
assert d["tags"] == ["dashboard", "analytics"]
# forward-compat keys present
assert d["priority"] == "P1"
assert d["owner"] == "hasan"
def test_sidecar_yaml_is_parseable_and_has_no_frontmatter_fences():
e = entry_mod.parse(LEGACY_MD)
sc = metadata.sidecar_yaml(e)
assert "---" not in sc.splitlines()[0]
loaded = yaml.safe_load(sc)
assert loaded["slug"] == "view-metrics"
assert loaded["priority"] == "P1"
# ---- strip_frontmatter (INV-2) ----
def test_strip_frontmatter_removes_leading_block():
body = metadata.strip_frontmatter(LEGACY_MD)
assert body.startswith("This is the prose body.")
assert "slug:" not in body
assert "priority:" not in body
def test_strip_frontmatter_passthrough_when_no_frontmatter():
plain = "Just a body.\n\nNo frontmatter here.\n"
assert metadata.strip_frontmatter(plain).strip() == plain.strip()
# ---- parse_sidecar (malformed detection, INV-3) ----
def test_parse_sidecar_good():
values, malformed = metadata.parse_sidecar("slug: a\ntitle: A\npriority: P0\n")
assert malformed is False
assert values == {"slug": "a", "title": "A", "priority": "P0"}
def test_parse_sidecar_non_mapping_is_malformed():
values, malformed = metadata.parse_sidecar("- just\n- a\n- list\n")
assert malformed is True
assert values == {}
def test_parse_sidecar_invalid_yaml_is_malformed():
values, malformed = metadata.parse_sidecar("slug: : : not yaml\n bad: [unclosed\n")
assert malformed is True
assert values == {}
def test_parse_sidecar_empty_is_empty_not_malformed():
values, malformed = metadata.parse_sidecar("")
assert malformed is False
assert values == {}
# ---- read_entry dual-read equivalence (INV-6) ----
def test_dual_read_sidecar_matches_legacy():
legacy_entry, legacy_bad = metadata.read_entry(LEGACY_MD, None)
# The migrated form: body-only .md + a sidecar holding the metadata.
migrated_md = metadata.strip_frontmatter(LEGACY_MD)
sidecar_text = metadata.sidecar_yaml(legacy_entry)
sidecar_entry, sidecar_bad = metadata.read_entry(migrated_md, sidecar_text)
assert legacy_bad is False
assert sidecar_bad is False
# Identical resulting records (INV-6).
assert sidecar_entry.slug == legacy_entry.slug
assert sidecar_entry.title == legacy_entry.title
assert sidecar_entry.state == legacy_entry.state
assert sidecar_entry.owners == legacy_entry.owners
assert sidecar_entry.tags == legacy_entry.tags
assert sidecar_entry.extra == legacy_entry.extra
assert sidecar_entry.body.strip() == legacy_entry.body.strip()
def test_read_entry_sidecar_takes_precedence_over_md_frontmatter():
# A not-yet-migrated .md still carrying frontmatter, plus a sidecar that
# disagrees: the sidecar wins for metadata; the body comes from the .md.
md_with_fm = "---\nslug: old\ntitle: Old Title\nstate: super-draft\n---\n\nBody.\n"
sidecar = "slug: new\ntitle: New Title\nstate: active\n"
e, malformed = metadata.read_entry(md_with_fm, sidecar)
assert malformed is False
assert e.title == "New Title"
assert e.state == "active"
assert e.body.strip() == "Body."
def test_read_entry_malformed_sidecar_still_loads_entry():
# INV-3: a malformed sidecar never hard-fails the read. The entry loads
# (from the .md frontmatter if present) and is flagged malformed.
md = "---\nslug: x\ntitle: X\nstate: active\n---\n\nBody.\n"
e, malformed = metadata.read_entry(md, "- not a mapping\n")
assert malformed is True
assert e.slug == "x"
assert e.title == "X"
assert e.body.strip() == "Body."
# ---- dual-read robustness: degenerate sidecars never drop the entry ----
def test_empty_sidecar_falls_back_to_md_frontmatter():
# An empty sidecar has no metadata to override with — keep the .md's.
md = "---\nslug: keep\ntitle: Keep Me\nstate: active\n---\n\nBody.\n"
e, malformed = metadata.read_entry(md, "", fallback_slug="keep")
assert malformed is False
assert e.slug == "keep"
assert e.title == "Keep Me"
def test_malformed_sidecar_on_body_only_md_loads_with_fallback_slug():
# The .md is already body-only (migrated) and the sidecar is corrupt:
# the entry must still load (INV-3), taking its slug from the filename stem.
e, malformed = metadata.read_entry("Just a body.\n", "- a\n- list\n", fallback_slug="foo")
assert malformed is True
assert e.slug == "foo"
def test_slugless_sidecar_uses_fallback_slug():
md = "Body only.\n"
sidecar = "title: No Slug Here\nstate: active\n"
e, malformed = metadata.read_entry(md, sidecar, fallback_slug="bar")
assert malformed is False
assert e.slug == "bar"
assert e.title == "No Slug Here"
# ---- sidecar filename helpers ----
def test_sidecar_filename_helpers():
assert metadata.sidecar_name("view-metrics") == "view-metrics.meta.yaml"
assert metadata.is_sidecar("view-metrics.meta.yaml") is True
assert metadata.is_sidecar("view-metrics.md") is False
assert metadata.slug_of_sidecar("view-metrics.meta.yaml") == "view-metrics"
# ---- SLICE-4: sidecar_path_for + apply_values ----
def test_sidecar_path_for_derives_sibling():
assert metadata.sidecar_path_for("rfcs/alpha.md") == "rfcs/alpha.meta.yaml"
assert metadata.sidecar_path_for("x/y/beta.md") == "x/y/beta.meta.yaml"
def test_apply_values_updates_known_and_extra_fields():
e = entry_mod.parse(LEGACY_MD) # has tags + extra priority/owner
e2 = metadata.apply_values(e, {"tags": ["x"], "priority": "P0", "owner": "sam"})
assert e2.tags == ["x"]
assert e2.extra["priority"] == "P0"
assert e2.extra["owner"] == "sam"
# body preserved unchanged
assert e2.body == e.body
def test_apply_values_preserves_unspecified_keys():
e = entry_mod.parse(LEGACY_MD)
e2 = metadata.apply_values(e, {"priority": "P0"})
assert e2.tags == e.tags # untouched
assert e2.extra["owner"] == "hasan" # untouched
@@ -0,0 +1,232 @@
"""SLICE-5 — bulk metadata edit endpoint (PUC-2, §6.4/§6.5).
Through the real API: contributor+ gating (INV-4), set/add/remove ops,
validation at the write boundary, one commit for N sidecars (D7), and
partial-rejection reporting. Reuses the fake-Gitea harness.
"""
from __future__ import annotations
import asyncio
import json
import yaml
from fastapi.testclient import TestClient
from app import cache, db, gitea as gitea_mod
from app.config import load_config
from test_propose_vertical import ( # noqa: F401 (fixtures)
app_with_fake_gitea,
provision_user_row,
sign_in_as,
tmp_env,
)
PID = "default"
CID = "default"
BASE = f"/api/projects/{PID}/collections/{CID}"
def _refresh():
cfg = load_config()
asyncio.run(cache.refresh_meta_repo(cfg, gitea_mod.Gitea(cfg)))
def _set_fields(schema):
db.conn().execute(
"UPDATE collections SET config_json = ? WHERE id = 'default'",
(json.dumps({"fields": schema}),))
def _seed_legacy(fake, slug, *, state="active", **front):
fm = {"slug": slug, "title": slug.title(), "state": state, **front}
body = yaml.safe_dump(fm, sort_keys=False).strip()
fake.files[("wiggleverse", "meta", "main", f"rfcs/{slug}.md")] = {
"content": f"---\n{body}\n---\n\nBody.\n", "sha": slug}
def _login_owner(client):
provision_user_row(user_id=1, login="ben", role="owner")
sign_in_as(client, user_id=1, gitea_login="ben", display_name="Ben", role="owner")
def _has_sidecar(fake, slug):
return ("wiggleverse", "meta", "main", f"rfcs/{slug}.meta.yaml") in fake.files
def _sidecar(fake, slug):
return yaml.safe_load(
fake.files[("wiggleverse", "meta", "main", f"rfcs/{slug}.meta.yaml")]["content"])
# ---- Task 1: happy path, one commit ----
def test_bulk_set_applies_to_all_and_one_commit(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_set_fields({"priority": {"type": "enum", "values": ["P0", "P1", "P2"]}})
_seed_legacy(fake, "a", priority="P2")
_seed_legacy(fake, "b", priority="P1")
_refresh()
_login_owner(client)
commits_before = fake.change_files_calls
r = client.post(f"{BASE}/meta/bulk",
json={"slugs": ["a", "b"], "op": "set",
"field": "priority", "value": "P0"})
assert r.status_code == 200, r.text
body = r.json()
assert set(body["applied"]) == {"a", "b"}
assert body["rejected"] == []
assert body["committed"] is True
# exactly one ChangeFiles commit covered both entries (D7)
assert fake.change_files_calls - commits_before == 1
assert _sidecar(fake, "a")["priority"] == "P0"
assert _sidecar(fake, "b")["priority"] == "P0"
# ---- Task 2: add/remove tags ----
def test_bulk_add_tag(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_set_fields({"tags": {"type": "tags"}})
_seed_legacy(fake, "a", tags=["x"])
_seed_legacy(fake, "b", tags=["x", "y"])
_refresh()
_login_owner(client)
r = client.post(f"{BASE}/meta/bulk",
json={"slugs": ["a", "b"], "op": "add",
"field": "tags", "value": "y"})
assert r.status_code == 200, r.text
assert set(r.json()["applied"]) == {"a", "b"}
# "a" gained y; "b" already had y (no-op write skipped → no sidecar written)
assert _sidecar(fake, "a")["tags"] == ["x", "y"]
assert not _has_sidecar(fake, "b")
def test_bulk_remove_tag(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_set_fields({"tags": {"type": "tags"}})
_seed_legacy(fake, "a", tags=["x", "y"])
_refresh()
_login_owner(client)
r = client.post(f"{BASE}/meta/bulk",
json={"slugs": ["a"], "op": "remove",
"field": "tags", "value": "x"})
assert r.status_code == 200, r.text
assert r.json()["applied"] == ["a"]
assert _sidecar(fake, "a")["tags"] == ["y"]
def test_bulk_set_scalar_on_tags_rejected(app_with_fake_gitea):
# A scalar `set` onto a tags field must reject, not char-split into a list.
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_set_fields({"tags": {"type": "tags"}})
_seed_legacy(fake, "a", tags=["x"])
_refresh()
_login_owner(client)
r = client.post(f"{BASE}/meta/bulk",
json={"slugs": ["a"], "op": "set",
"field": "tags", "value": "checkout"})
assert r.status_code == 200, r.text
assert r.json()["applied"] == []
assert len(r.json()["rejected"]) == 1
assert r.json()["committed"] is False
assert not _has_sidecar(fake, "a")
def test_bulk_set_list_on_tags_ok(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_set_fields({"tags": {"type": "tags"}})
_seed_legacy(fake, "a", tags=["x"])
_refresh()
_login_owner(client)
r = client.post(f"{BASE}/meta/bulk",
json={"slugs": ["a"], "op": "set",
"field": "tags", "value": ["x", "y"]})
assert r.status_code == 200, r.text
assert r.json()["applied"] == ["a"]
assert _sidecar(fake, "a")["tags"] == ["x", "y"]
def test_bulk_add_remove_requires_tags_field(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_set_fields({"priority": {"type": "enum", "values": ["P0"]}})
_seed_legacy(fake, "a", priority="P0")
_refresh()
_login_owner(client)
r = client.post(f"{BASE}/meta/bulk",
json={"slugs": ["a"], "op": "add",
"field": "priority", "value": "z"})
assert r.status_code == 422, r.text
# ---- Task 3: partial rejection, authz, validation guards ----
def test_bulk_partial_reject_missing_entry(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_set_fields({"priority": {"type": "enum", "values": ["P0", "P1"]}})
_seed_legacy(fake, "a", priority="P1")
_refresh()
_login_owner(client)
r = client.post(f"{BASE}/meta/bulk",
json={"slugs": ["a", "ghost"], "op": "set",
"field": "priority", "value": "P0"})
assert r.status_code == 200, r.text
body = r.json()
assert body["applied"] == ["a"]
assert body["rejected"] == [{"slug": "ghost", "reason": "not found"}]
assert body["committed"] is True
assert _sidecar(fake, "a")["priority"] == "P0"
def test_bulk_invalid_value_rejects_all_no_commit(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_set_fields({"priority": {"type": "enum", "values": ["P0", "P1"]}})
_seed_legacy(fake, "a", priority="P1")
_refresh()
_login_owner(client)
r = client.post(f"{BASE}/meta/bulk",
json={"slugs": ["a"], "op": "set",
"field": "priority", "value": "ZZZ"})
assert r.status_code == 200, r.text
assert r.json()["applied"] == []
assert len(r.json()["rejected"]) == 1
assert r.json()["committed"] is False
assert not _has_sidecar(fake, "a")
def test_bulk_forbidden_for_anonymous(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_set_fields({"priority": {"type": "enum", "values": ["P0"]}})
_seed_legacy(fake, "a", priority="P0")
_refresh()
r = client.post(f"{BASE}/meta/bulk",
json={"slugs": ["a"], "op": "set",
"field": "priority", "value": "P0"})
assert r.status_code == 403, r.text
def test_bulk_unknown_field_op_and_empty(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_set_fields({"priority": {"type": "enum", "values": ["P0"]}})
_seed_legacy(fake, "a", priority="P0")
_refresh()
_login_owner(client)
assert client.post(f"{BASE}/meta/bulk",
json={"slugs": ["a"], "op": "set",
"field": "nope", "value": "P0"}).status_code == 422
assert client.post(f"{BASE}/meta/bulk",
json={"slugs": ["a"], "op": "frobnicate",
"field": "priority", "value": "P0"}).status_code == 422
assert client.post(f"{BASE}/meta/bulk",
json={"slugs": [], "op": "set",
"field": "priority", "value": "P0"}).status_code == 422
+177
View File
@@ -0,0 +1,177 @@
"""SLICE-1 integration — the corpus mirror reads sidecars (dual-read) and
derives the malformed flag (PUC-6, INV-3/INV-6).
Per docs/design/2026-06-06-configurable-collection-metadata.md §6.2-6.3.
"""
from __future__ import annotations
import asyncio
import json
from fastapi.testclient import TestClient
from app import cache, db, gitea as gitea_mod
from app.config import load_config
from test_propose_vertical import ( # noqa: F401 (fixtures)
app_with_fake_gitea,
tmp_env,
)
def _refresh():
cfg = load_config()
asyncio.run(cache.refresh_meta_repo(cfg, gitea_mod.Gitea(cfg)))
def _row(slug):
return db.conn().execute(
"SELECT * FROM cached_rfcs WHERE slug = ?", (slug,)
).fetchone()
def test_mirror_reads_metadata_from_sidecar(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app):
# A migrated entry: body-only .md + a sidecar holding the metadata.
fake.files[("wiggleverse", "meta", "main", "rfcs/sidecar-one.md")] = {
"content": "Just the prose body.\n", "sha": "s1"}
fake.files[("wiggleverse", "meta", "main", "rfcs/sidecar-one.meta.yaml")] = {
"content": "slug: sidecar-one\ntitle: From Sidecar\nstate: active\ntags:\n- alpha\n",
"sha": "m1"}
_refresh()
row = _row("sidecar-one")
assert row is not None
assert row["title"] == "From Sidecar"
assert row["state"] == "active"
assert row["body"].strip() == "Just the prose body."
assert row["metadata_malformed"] == 0
def test_malformed_sidecar_flags_but_still_loads(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app):
# .md still has frontmatter; sidecar is malformed (a list, not a map).
fake.files[("wiggleverse", "meta", "main", "rfcs/bad-meta.md")] = {
"content": "---\nslug: bad-meta\ntitle: Legacy Title\nstate: active\n---\n\nBody.\n",
"sha": "b1"}
fake.files[("wiggleverse", "meta", "main", "rfcs/bad-meta.meta.yaml")] = {
"content": "- not\n- a\n- mapping\n", "sha": "b2"}
_refresh()
row = _row("bad-meta")
assert row is not None # INV-3: still loads
assert row["metadata_malformed"] == 1
# Falls back to the legacy .md frontmatter for the metadata.
assert row["title"] == "Legacy Title"
def test_malformed_flag_surfaces_in_catalog_api(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
fake.files[("wiggleverse", "meta", "main", "rfcs/flagged.md")] = {
"content": "---\nslug: flagged\ntitle: Flagged\nstate: active\n---\n\nB.\n",
"sha": "f1"}
fake.files[("wiggleverse", "meta", "main", "rfcs/flagged.meta.yaml")] = {
"content": "just a scalar\n", "sha": "f2"}
fake.files[("wiggleverse", "meta", "main", "rfcs/clean.md")] = {
"content": "---\nslug: clean\ntitle: Clean\nstate: active\n---\n\nB.\n",
"sha": "c1"}
_refresh()
items = {i["slug"]: i for i in client.get("/api/rfcs").json()["items"]}
assert items["flagged"]["metadata_malformed"] is True
assert items["clean"]["metadata_malformed"] is False
# And on the detail view.
assert client.get("/api/rfcs/flagged").json()["metadata_malformed"] is True
def test_malformed_sidecar_on_migrated_entry_still_loads_flagged(app_with_fake_gitea):
# INV-3 regression: a migrated (body-only .md) entry whose sidecar is
# corrupt must NOT vanish from the catalog — it loads (slug from the
# filename stem) and is flagged malformed.
app, fake = app_with_fake_gitea
with TestClient(app):
fake.files[("wiggleverse", "meta", "main", "rfcs/orphaned.md")] = {
"content": "Just the body, no frontmatter.\n", "sha": "o1"}
fake.files[("wiggleverse", "meta", "main", "rfcs/orphaned.meta.yaml")] = {
"content": "- corrupt\n- list\n", "sha": "o2"}
_refresh()
row = _row("orphaned")
assert row is not None # did not vanish
assert row["metadata_malformed"] == 1
assert row["body"].strip() == "Just the body, no frontmatter."
def _set_default_fields_schema(schema):
# apply_registry leaves the default collection's config_json untouched, so a
# schema set here survives a corpus refresh (refresh_meta_repo only).
db.conn().execute(
"UPDATE collections SET config_json = ? WHERE id = 'default'",
(json.dumps({"fields": schema}),))
# ---- §22.4a SLICE-2: advisory schema validation at ingest (INV-3) ----
def test_schema_violation_flags_malformed(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app):
_set_default_fields_schema(
{"priority": {"type": "enum", "values": ["P0", "P1", "P2"]}})
fake.files[("wiggleverse", "meta", "main", "rfcs/bad-prio.md")] = {
"content": "---\nslug: bad-prio\ntitle: Bad\nstate: active\npriority: P9\n---\n\nB.\n",
"sha": "bp1"}
fake.files[("wiggleverse", "meta", "main", "rfcs/good-prio.md")] = {
"content": "---\nslug: good-prio\ntitle: Good\nstate: active\npriority: P0\n---\n\nB.\n",
"sha": "gp1"}
_refresh()
assert _row("bad-prio")["metadata_malformed"] == 1 # INV-3: flagged
assert _row("bad-prio")["title"] == "Bad" # still loads
assert _row("good-prio")["metadata_malformed"] == 0
def test_schema_ignores_undeclared_keys(app_with_fake_gitea):
# INV-7: keys the schema doesn't declare ride along and never flag malformed.
app, fake = app_with_fake_gitea
with TestClient(app):
_set_default_fields_schema(
{"priority": {"type": "enum", "values": ["P0", "P1"]}})
fake.files[("wiggleverse", "meta", "main", "rfcs/extra-key.md")] = {
"content": "---\nslug: extra-key\ntitle: Extra\nstate: active\nowner: hasan\n---\n\nB.\n",
"sha": "ek1"}
_refresh()
assert _row("extra-key")["metadata_malformed"] == 0
def test_no_schema_never_flags(app_with_fake_gitea):
# INV-5: a collection with no field schema validates nothing, even when an
# entry carries values that would fail a schema if one existed.
app, fake = app_with_fake_gitea
with TestClient(app):
fake.files[("wiggleverse", "meta", "main", "rfcs/anything.md")] = {
"content": "---\nslug: anything\ntitle: Any\nstate: active\npriority: whatever\n---\n\nB.\n",
"sha": "an1"}
_refresh()
assert _row("anything")["metadata_malformed"] == 0
def test_legacy_collection_without_sidecars_unchanged(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app):
fake.files[("wiggleverse", "meta", "main", "rfcs/legacy.md")] = {
"content": "---\nslug: legacy\ntitle: Legacy\nstate: super-draft\n---\n\nPitch.\n",
"sha": "l1"}
_refresh()
row = _row("legacy")
assert row is not None
assert row["title"] == "Legacy"
assert row["state"] == "super-draft"
assert row["body"].strip() == "Pitch."
assert row["metadata_malformed"] == 0
@@ -0,0 +1,190 @@
"""SLICE-4 — single-entry metadata edit endpoint (PUC-1, §6.4).
Through the real API: contributor+ gating (INV-4), schema validation at the
write boundary, direct commit to the sidecar with lazy migration, re-ingest,
and the GET RFC `meta` + `can_edit_meta` exposure. Reuses the fake-Gitea harness.
"""
from __future__ import annotations
import asyncio
import json
import yaml
from fastapi.testclient import TestClient
from app import cache, db, gitea as gitea_mod
from app.config import load_config
from test_propose_vertical import ( # noqa: F401 (fixtures)
app_with_fake_gitea,
provision_user_row,
sign_in_as,
tmp_env,
)
PID = "default"
CID = "default"
META = f"/api/projects/{PID}/collections/{CID}"
def _refresh():
cfg = load_config()
asyncio.run(cache.refresh_meta_repo(cfg, gitea_mod.Gitea(cfg)))
def _set_fields(schema):
db.conn().execute(
"UPDATE collections SET config_json = ? WHERE id = 'default'",
(json.dumps({"fields": schema}),))
def _seed_legacy(fake, slug, *, state="active", **front):
fm = {"slug": slug, "title": slug.title(), "state": state, **front}
body = yaml.safe_dump(fm, sort_keys=False).strip()
fake.files[("wiggleverse", "meta", "main", f"rfcs/{slug}.md")] = {
"content": f"---\n{body}\n---\n\nBody.\n", "sha": slug}
def _seed_migrated(fake, slug, sidecar):
fake.files[("wiggleverse", "meta", "main", f"rfcs/{slug}.md")] = {
"content": "Body.\n", "sha": f"{slug}-md"}
fake.files[("wiggleverse", "meta", "main", f"rfcs/{slug}.meta.yaml")] = {
"content": sidecar, "sha": f"{slug}-sc"}
def _login_owner(client):
provision_user_row(user_id=1, login="ben", role="owner")
sign_in_as(client, user_id=1, gitea_login="ben", display_name="Ben", role="owner")
def test_edit_meta_sets_value_commits_sidecar_and_lazy_migrates(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_set_fields({"priority": {"type": "enum", "values": ["P0", "P1", "P2"]},
"tags": {"type": "tags"}})
_seed_legacy(fake, "a", priority="P1", tags=["x"])
_refresh()
_login_owner(client)
r = client.post(f"{META}/rfcs/a/meta", json={"values": {"priority": "P0"}})
assert r.status_code == 200, r.text
assert r.json()["meta"]["priority"] == "P0"
# sidecar written, .md lazy-migrated to body-only
sc = yaml.safe_load(
fake.files[("wiggleverse", "meta", "main", "rfcs/a.meta.yaml")]["content"])
assert sc["priority"] == "P0"
assert "---" not in fake.files[("wiggleverse", "meta", "main", "rfcs/a.md")]["content"]
# cache reflects the new value
row = db.conn().execute(
"SELECT meta_json FROM cached_rfcs WHERE slug='a'").fetchone()
assert json.loads(row["meta_json"])["priority"] == "P0"
def test_edit_meta_rejects_value_outside_enum(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_set_fields({"priority": {"type": "enum", "values": ["P0", "P1"]}})
_seed_legacy(fake, "a", priority="P1")
_refresh()
_login_owner(client)
r = client.post(f"{META}/rfcs/a/meta", json={"values": {"priority": "ZZZ"}})
assert r.status_code == 422, r.text
# nothing committed
assert ("wiggleverse", "meta", "main", "rfcs/a.meta.yaml") not in fake.files
def test_edit_meta_rejects_unknown_field(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_set_fields({"priority": {"type": "enum", "values": ["P0"]}})
_seed_legacy(fake, "a", priority="P0")
_refresh()
_login_owner(client)
r = client.post(f"{META}/rfcs/a/meta", json={"values": {"nope": "x"}})
assert r.status_code == 422, r.text
def test_edit_meta_forbidden_for_anonymous(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_set_fields({"priority": {"type": "enum", "values": ["P0"]}})
_seed_legacy(fake, "a", priority="P0")
_refresh()
r = client.post(f"{META}/rfcs/a/meta", json={"values": {"priority": "P0"}})
assert r.status_code == 403, r.text
def test_edit_meta_on_already_migrated_entry(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_set_fields({"priority": {"type": "enum", "values": ["P0", "P1"]}})
_seed_migrated(fake, "a", "slug: a\ntitle: A\nstate: active\npriority: P1\n")
_refresh()
_login_owner(client)
r = client.post(f"{META}/rfcs/a/meta", json={"values": {"priority": "P0"}})
assert r.status_code == 200, r.text
sc = yaml.safe_load(
fake.files[("wiggleverse", "meta", "main", "rfcs/a.meta.yaml")]["content"])
assert sc["priority"] == "P0"
# .md untouched (still body-only)
assert fake.files[("wiggleverse", "meta", "main", "rfcs/a.md")]["content"] == "Body.\n"
def test_get_rfc_exposes_meta_and_can_edit(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_set_fields({"priority": {"type": "enum", "values": ["P0", "P1"]}})
_seed_legacy(fake, "a", priority="P1")
_refresh()
# anonymous: meta present, can_edit_meta False
r = client.get(f"{META}/rfcs/a")
assert r.status_code == 200, r.text
assert r.json()["meta"]["priority"] == "P1"
assert r.json()["can_edit_meta"] is False
# owner: can_edit_meta True
_login_owner(client)
r = client.get(f"{META}/rfcs/a")
assert r.json()["can_edit_meta"] is True
# ---- Owner-gated collection migrate endpoint (PUC-5) ----
def test_migrate_collection_endpoint_owner(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_seed_legacy(fake, "a", priority="P1", tags=["x"])
_seed_legacy(fake, "b", priority="P0")
_refresh()
_login_owner(client)
r = client.post(f"{META}/migrate")
assert r.status_code == 200, r.text
assert r.json()["committed"] is True
assert set(r.json()["migrated"]) == {"a", "b"}
# both entries now body-only + sidecar
for slug in ("a", "b"):
assert "---" not in fake.files[("wiggleverse", "meta", "main", f"rfcs/{slug}.md")]["content"]
assert ("wiggleverse", "meta", "main", f"rfcs/{slug}.meta.yaml") in fake.files
def test_migrate_collection_forbidden_for_contributor(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_seed_legacy(fake, "a", priority="P1")
_refresh()
provision_user_row(user_id=2, login="carol", role="contributor")
sign_in_as(client, user_id=2, gitea_login="carol",
display_name="Carol", role="contributor")
r = client.post(f"{META}/migrate")
assert r.status_code == 403, r.text
def test_migrate_collection_idempotent(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_seed_legacy(fake, "a", priority="P1")
_refresh()
_login_owner(client)
assert client.post(f"{META}/migrate").json()["committed"] is True
# second run: nothing left to migrate
r2 = client.post(f"{META}/migrate")
assert r2.status_code == 200, r2.text
assert r2.json()["committed"] is False
+108
View File
@@ -0,0 +1,108 @@
"""SLICE-4 — git-aware sidecar read/write helpers.
Uses the FakeGitea from the propose-vertical fixtures (no network).
"""
from __future__ import annotations
import asyncio
import yaml
from app import gitea as gitea_mod, metadata
from app.config import load_config
from test_propose_vertical import app_with_fake_gitea, tmp_env # noqa: F401
LEGACY = """---
slug: alpha
title: Alpha
state: active
owners:
- ben.stull
tags:
- one
priority: P1
---
Alpha body.
"""
MIGRATED_MD = "Alpha body.\n"
MIGRATED_SIDECAR = """slug: alpha
title: Alpha
state: active
owners:
- ben.stull
tags:
- one
priority: P1
"""
def _gitea():
return gitea_mod.Gitea(load_config())
def test_read_entry_from_git_legacy(app_with_fake_gitea):
_app, fake = app_with_fake_gitea
fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.md")] = {
"content": LEGACY, "sha": "s1"}
st = asyncio.run(metadata.read_entry_from_git(
_gitea(), "wiggleverse", "meta", "rfcs/alpha.md"))
assert st is not None
assert st.entry.slug == "alpha"
assert st.entry.extra["priority"] == "P1"
assert st.sidecar_sha is None # no sidecar yet
assert st.malformed is False
def test_read_entry_from_git_migrated(app_with_fake_gitea):
_app, fake = app_with_fake_gitea
fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.md")] = {
"content": MIGRATED_MD, "sha": "s1"}
fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.meta.yaml")] = {
"content": MIGRATED_SIDECAR, "sha": "s2"}
st = asyncio.run(metadata.read_entry_from_git(
_gitea(), "wiggleverse", "meta", "rfcs/alpha.md"))
assert st.entry.extra["priority"] == "P1" # from sidecar
assert st.entry.body == "Alpha body.\n"
assert st.sidecar_sha == "s2"
def test_read_entry_from_git_missing(app_with_fake_gitea):
st = asyncio.run(metadata.read_entry_from_git(
_gitea(), "wiggleverse", "meta", "rfcs/nope.md"))
assert st is None
def test_write_entry_files_lazy_migrates_legacy(app_with_fake_gitea):
_app, fake = app_with_fake_gitea
fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.md")] = {
"content": LEGACY, "sha": "s1"}
st = asyncio.run(metadata.read_entry_from_git(
_gitea(), "wiggleverse", "meta", "rfcs/alpha.md"))
e2 = metadata.apply_values(st.entry, {"priority": "P0"})
ops = metadata.write_entry_files("rfcs/alpha.md", e2, st)
paths = {o["path"]: o for o in ops}
# sidecar created, .md rewritten body-only
assert "rfcs/alpha.meta.yaml" in paths
assert paths["rfcs/alpha.meta.yaml"]["operation"] == "create"
assert paths["rfcs/alpha.md"]["operation"] == "update"
assert "---" not in paths["rfcs/alpha.md"]["content"] # INV-2 clean body
assert yaml.safe_load(paths["rfcs/alpha.meta.yaml"]["content"])["priority"] == "P0"
def test_write_entry_files_already_migrated_touches_sidecar_only(app_with_fake_gitea):
_app, fake = app_with_fake_gitea
fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.md")] = {
"content": MIGRATED_MD, "sha": "s1"}
fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.meta.yaml")] = {
"content": MIGRATED_SIDECAR, "sha": "s2"}
st = asyncio.run(metadata.read_entry_from_git(
_gitea(), "wiggleverse", "meta", "rfcs/alpha.md"))
e2 = metadata.apply_values(st.entry, {"priority": "P0"})
ops = metadata.write_entry_files("rfcs/alpha.md", e2, st)
paths = {o["path"]: o for o in ops}
assert set(paths) == {"rfcs/alpha.meta.yaml"} # .md untouched
assert paths["rfcs/alpha.meta.yaml"]["operation"] == "update"
assert paths["rfcs/alpha.meta.yaml"]["sha"] == "s2"
+135
View File
@@ -0,0 +1,135 @@
"""SLICE-1 integration — frontmatter→sidecar migration tool (PUC-5).
Per docs/design/2026-06-06-configurable-collection-metadata.md §6.5 / §7.2:
a tool walks a collection; for each entry with legacy frontmatter it writes
`<slug>.meta.yaml` and rewrites `<slug>.md` to the body only one commit per
collection, idempotent, preserving unknown keys (INV-7).
"""
from __future__ import annotations
import asyncio
import yaml
from app import gitea as gitea_mod, metadata
from app.config import load_config
from test_propose_vertical import ( # noqa: F401 (fixtures)
app_with_fake_gitea,
tmp_env,
)
ALPHA_MD = """---
slug: alpha
title: Alpha
state: active
owners:
- ben.stull
tags:
- one
priority: P1
---
Alpha body prose.
"""
BETA_MD = """---
slug: beta
title: Beta
state: super-draft
---
Beta body prose.
"""
def _seed_entries(fake):
fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.md")] = {
"content": ALPHA_MD, "sha": "a0001"}
fake.files[("wiggleverse", "meta", "main", "rfcs/beta.md")] = {
"content": BETA_MD, "sha": "b0001"}
def _run_migration(subfolder=""):
cfg = load_config()
gitea = gitea_mod.Gitea(cfg)
return asyncio.run(
metadata.migrate_collection(
gitea, org="wiggleverse", repo="meta", subfolder=subfolder
)
)
def test_migration_writes_sidecars_and_strips_bodies(app_with_fake_gitea):
_app, fake = app_with_fake_gitea
_seed_entries(fake)
summary = _run_migration()
assert sorted(summary["migrated"]) == ["alpha", "beta"]
# Sidecars now exist.
alpha_sc = fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.meta.yaml")]["content"]
beta_sc = fake.files[("wiggleverse", "meta", "main", "rfcs/beta.meta.yaml")]["content"]
alpha_vals = yaml.safe_load(alpha_sc)
assert alpha_vals["slug"] == "alpha"
assert alpha_vals["title"] == "Alpha"
assert alpha_vals["tags"] == ["one"]
# INV-7: the unknown key rides along into the sidecar.
assert alpha_vals["priority"] == "P1"
# .md bodies are stripped of frontmatter (INV-2).
alpha_md = fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.md")]["content"]
assert "---" not in alpha_md
assert "priority:" not in alpha_md
assert alpha_md.strip() == "Alpha body prose."
assert beta_sc # beta got a sidecar too
def test_migration_is_idempotent(app_with_fake_gitea):
_app, fake = app_with_fake_gitea
_seed_entries(fake)
first = _run_migration()
assert sorted(first["migrated"]) == ["alpha", "beta"]
assert first["committed"] is True
commits_after_first = fake._commit_counter
alpha_md_after_first = fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.md")]["content"]
second = _run_migration()
assert second["migrated"] == []
assert second["committed"] is False
# No new commit; files untouched.
assert fake._commit_counter == commits_after_first
assert fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.md")]["content"] == alpha_md_after_first
def test_migration_one_commit_for_whole_collection(app_with_fake_gitea):
_app, fake = app_with_fake_gitea
_seed_entries(fake)
before = fake._commit_counter
_run_migration()
# Two entries migrated in exactly one commit (ChangeFiles batch).
assert fake._commit_counter == before + 1
def test_migration_dual_read_equivalence_after_migrate(app_with_fake_gitea):
"""An entry reads identically before and after migration (INV-6)."""
_app, fake = app_with_fake_gitea
_seed_entries(fake)
before, _ = metadata.read_entry(ALPHA_MD, None)
_run_migration()
md = fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.md")]["content"]
sc = fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.meta.yaml")]["content"]
after, malformed = metadata.read_entry(md, sc)
assert malformed is False
assert after.slug == before.slug
assert after.title == before.title
assert after.state == before.state
assert after.tags == before.tags
assert after.extra == before.extra
assert after.body.strip() == before.body.strip()
+149
View File
@@ -0,0 +1,149 @@
"""SLICE-2 unit tests — collection field schema + central validation.
Pure functions only (no DB / no Gitea). Per
docs/design/2026-06-06-configurable-collection-metadata.md §7.2 (SLICE-2):
`metadata_schema.parse_fields` (lenient schema parsing) and
`metadata_schema.validate` (advisory at read / enforcement point at write).
Honors INV-3 (never hard-fails), INV-5 (no-fields unchanged), INV-7 (undeclared
keys ride along).
"""
from __future__ import annotations
from app import metadata_schema as ms
# ---- parse_fields: normalization + leniency ----
def test_parse_fields_each_type():
raw = {
"priority": {"type": "enum", "values": ["P0", "P1", "P2"], "label": "Priority"},
"tags": {"type": "tags"},
"owner": {"type": "text"},
}
fields = ms.parse_fields(raw)
assert list(fields) == ["priority", "tags", "owner"] # order preserved
assert fields["priority"] == {
"type": "enum",
"values": ["P0", "P1", "P2"],
"label": "Priority",
}
assert fields["tags"] == {"type": "tags"}
assert fields["owner"] == {"type": "text"}
def test_parse_fields_controlled_tags_keeps_values():
fields = ms.parse_fields({"area": {"type": "tags", "values": ["a", "b"]}})
assert fields["area"] == {"type": "tags", "values": ["a", "b"]}
def test_parse_fields_missing_block_is_empty():
assert ms.parse_fields(None) == {}
assert ms.parse_fields({}) == {}
def test_parse_fields_non_mapping_block_skipped():
assert ms.parse_fields(["not", "a", "mapping"]) == {}
assert ms.parse_fields("nope") == {}
def test_parse_fields_enum_without_values_skipped():
# enum requires a non-empty values list — skipped, not fatal.
assert ms.parse_fields({"p": {"type": "enum"}}) == {}
assert ms.parse_fields({"p": {"type": "enum", "values": []}}) == {}
def test_parse_fields_unknown_type_skipped():
fields = ms.parse_fields(
{"good": {"type": "text"}, "bad": {"type": "ref"}, "huh": {"type": "frob"}}
)
assert list(fields) == ["good"]
def test_parse_fields_non_mapping_def_skipped():
fields = ms.parse_fields({"good": {"type": "text"}, "bad": "scalar"})
assert list(fields) == ["good"]
def test_parse_fields_values_coerced_to_str_list():
fields = ms.parse_fields({"p": {"type": "enum", "values": [0, 1, 2]}})
assert fields["p"]["values"] == ["0", "1", "2"]
# ---- validate: advisory problem reporting ----
SCHEMA = ms.parse_fields(
{
"priority": {"type": "enum", "values": ["P0", "P1", "P2"]},
"tags": {"type": "tags"},
"area": {"type": "tags", "values": ["checkout", "cart"]},
"owner": {"type": "text"},
}
)
def test_validate_happy():
values = {
"priority": "P0",
"tags": ["anything", "free"],
"area": ["checkout"],
"owner": "ben",
}
assert ms.validate(values, SCHEMA) == []
def test_validate_absent_fields_ok():
# A declared field that the entry omits is fine (no required fields in v1).
assert ms.validate({}, SCHEMA) == []
def test_validate_enum_bad_value():
problems = ms.validate({"priority": "P9"}, SCHEMA)
assert [p.field for p in problems] == ["priority"]
assert problems[0].code == "not-in-values"
def test_validate_enum_wrong_type():
problems = ms.validate({"priority": ["P0"]}, SCHEMA)
assert problems[0].field == "priority"
assert problems[0].code == "wrong-type"
def test_validate_tags_free_form_ok():
assert ms.validate({"tags": ["x", "y", "z"]}, SCHEMA) == []
def test_validate_tags_wrong_type():
problems = ms.validate({"tags": "notalist"}, SCHEMA)
assert problems[0].field == "tags"
assert problems[0].code == "wrong-type"
def test_validate_controlled_tags_bad_member():
problems = ms.validate({"area": ["checkout", "nope"]}, SCHEMA)
assert problems[0].field == "area"
assert problems[0].code == "not-in-values"
def test_validate_text_wrong_type():
problems = ms.validate({"owner": ["a", "b"]}, SCHEMA)
assert problems[0].field == "owner"
assert problems[0].code == "wrong-type"
def test_validate_undeclared_keys_ignored():
# INV-7: keys outside the schema ride along untouched, never flagged.
assert ms.validate({"random": "value", "slug": "x", "title": "y"}, SCHEMA) == []
def test_validate_empty_schema_no_problems():
# INV-5: a collection with no fields validates everything as clean.
assert ms.validate({"priority": "anything", "x": 1}, {}) == []
def test_problem_as_dict():
p = ms.Problem(field="priority", code="not-in-values", message="bad")
assert p.as_dict() == {
"field": "priority",
"code": "not-in-values",
"message": "bad",
}
+145
View File
@@ -0,0 +1,145 @@
"""SLICE-4 — write paths are sidecar-aware: a migrated (body-only `.md` +
sidecar) entry never crashes `entry.parse` nor re-grows frontmatter."""
from __future__ import annotations
import asyncio
from app import gitea as gitea_mod, metadata
from app.bot import Actor, Bot
from app.config import load_config
from test_propose_vertical import ( # noqa: F401
app_with_fake_gitea,
provision_user_row,
tmp_env,
)
BODY_ONLY = "Alpha prose body.\n"
SIDECAR = ("slug: alpha\ntitle: Alpha\nstate: active\n"
"owners:\n- ben.stull\ntags:\n- one\npriority: P1\n")
def _seed_migrated(fake, repo="meta"):
fake.files[("wiggleverse", repo, "main", "rfcs/alpha.md")] = {
"content": BODY_ONLY, "sha": "m1"}
fake.files[("wiggleverse", repo, "main", "rfcs/alpha.meta.yaml")] = {
"content": SIDECAR, "sha": "m2"}
def _actor():
return Actor(user_id=1, gitea_login="ben.stull", display_name="Ben", email="ben@x.io")
def test_mark_entry_reviewed_on_migrated_entry(app_with_fake_gitea):
from fastapi.testclient import TestClient
app, fake = app_with_fake_gitea
_seed_migrated(fake)
gitea = gitea_mod.Gitea(load_config())
bot = Bot(gitea)
with TestClient(app):
provision_user_row(user_id=1, login="ben.stull", role="owner")
# Must not raise (legacy code parsed body-only .md → ValueError).
asyncio.run(bot.mark_entry_reviewed(
_actor(), org="wiggleverse", meta_repo="meta", slug="alpha",
reviewed_by="ben.stull", reviewed_at="2026-06-07"))
md = fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.md")]["content"]
sc = fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.meta.yaml")]["content"]
assert "---" not in md # body stays clean (no re-grown FM)
assert "reviewed_by: ben.stull" in sc # review stamp landed in the sidecar
# ---- Task 3.2: body extract/wrap helpers (api_branches) ----
def test_extract_wrap_body_on_body_only_md():
from app import api_branches
rfc = {"state": "super-draft", "repo": None, "slug": "alpha", "collection_id": "default"}
body = api_branches._extract_body_pure(rfc, BODY_ONLY, "main", is_meta=True)
assert body == BODY_ONLY
wrapped = api_branches._wrap_body_pure(rfc, BODY_ONLY, "new body\n", "main", is_meta=True)
assert wrapped == "new body\n" # stays clean — no re-grown frontmatter
def test_wrap_body_preserves_legacy_frontmatter():
from app import api_branches
legacy = "---\nslug: alpha\ntitle: Alpha\nstate: active\n---\n\nold body\n"
rfc = {"state": "super-draft", "repo": None, "slug": "alpha", "collection_id": "default"}
wrapped = api_branches._wrap_body_pure(rfc, legacy, "new body\n", "main", is_meta=True)
assert wrapped.startswith("---") # legacy frontmatter preserved
assert "new body" in wrapped
# ---- Task 3.3: PR-replay body wrappers (api_prs) ----
def test_replay_wrappers_on_body_only():
from app import api_prs
assert api_prs._extract_body_for_replay(True, BODY_ONLY) == BODY_ONLY
out = api_prs._wrap_body_for_replay(True, BODY_ONLY, "new\n")
assert out == "new\n" # clean, no re-grown frontmatter
legacy = "---\nslug: a\ntitle: A\nstate: active\n---\n\nold\n"
out2 = api_prs._wrap_body_for_replay(True, legacy, "new\n")
assert out2.startswith("---") # legacy preserved
# ---- Task 3.4: retire a fully-migrated (body-only + sidecar) entry ----
def test_retire_already_migrated_entry_does_not_crash(app_with_fake_gitea):
from fastapi.testclient import TestClient
from app import cache, db
from app.config import load_config
app, fake = app_with_fake_gitea
_seed_migrated(fake)
with TestClient(app) as client:
provision_user_row(user_id=1, login="ben", role="owner")
from test_propose_vertical import sign_in_as
sign_in_as(client, user_id=1, gitea_login="ben", display_name="Ben", role="owner")
# Ingest the migrated entry so the catalog/cache knows it.
asyncio.run(cache.refresh_meta_repo(load_config(), gitea_mod.Gitea(load_config())))
assert db.conn().execute(
"SELECT state FROM cached_rfcs WHERE slug='alpha'").fetchone()["state"] == "active"
r = client.post("/api/rfcs/alpha/retire")
assert r.status_code == 200, r.text
assert r.json()["state"] == "retired"
import yaml as _yaml
sc = _yaml.safe_load(
fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.meta.yaml")]["content"])
assert sc["state"] == "retired"
md = fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.md")]["content"]
assert "---" not in md # body stayed clean
# ---- Task 3.5: graduate a fully-migrated super-draft entry ----
SUPER_SIDECAR = ("slug: alpha\ntitle: Alpha\nstate: super-draft\n"
"owners:\n- ben\ntags:\n- one\npriority: P1\n")
def test_graduate_already_migrated_super_draft(app_with_fake_gitea):
from fastapi.testclient import TestClient
from app import cache, db
from app.config import load_config
app, fake = app_with_fake_gitea
fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.md")] = {
"content": BODY_ONLY, "sha": "m1"}
fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.meta.yaml")] = {
"content": SUPER_SIDECAR, "sha": "m2"}
with TestClient(app) as client:
from test_propose_vertical import sign_in_as
provision_user_row(user_id=1, login="ben", role="owner")
sign_in_as(client, user_id=1, gitea_login="ben", display_name="Ben", role="owner")
asyncio.run(cache.refresh_meta_repo(load_config(), gitea_mod.Gitea(load_config())))
assert db.conn().execute(
"SELECT state FROM cached_rfcs WHERE slug='alpha'").fetchone()["state"] == "super-draft"
r = client.post("/api/rfcs/alpha/graduate?_sync=1",
json={"rfc_id": "RFC-0007", "owners": ["ben"]})
assert r.status_code == 200, r.text
import yaml as _yaml
sc = _yaml.safe_load(
fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.meta.yaml")]["content"])
assert sc["state"] == "active"
assert sc["id"] == "RFC-0007"
assert sc.get("priority") == "P1" # INV-7 carried through graduation
md = fake.files[("wiggleverse", "meta", "main", "rfcs/alpha.md")]["content"]
assert "---" not in md # body stayed clean
+8 -3
View File
@@ -21,12 +21,17 @@ def _fresh_config() -> Config:
def test_027_adds_project_type_and_initial_state():
# 027 added type/initial_state to `projects`; migration 029 (three-tier)
# moved those per-corpus fields *down* onto `collections`. After the full
# migration chain they live on the collection, not the project.
cfg = _fresh_config()
db.run_migrations(cfg)
conn = db.connect(cfg.database_path)
cols = {r["name"]: r for r in conn.execute("PRAGMA table_info(projects)")}
assert "type" in cols and cols["type"]["dflt_value"] == "'document'"
assert "initial_state" in cols and cols["initial_state"]["dflt_value"] == "'super-draft'"
proj_cols = {r["name"] for r in conn.execute("PRAGMA table_info(projects)")}
assert "type" not in proj_cols and "initial_state" not in proj_cols
coll_cols = {r["name"]: r for r in conn.execute("PRAGMA table_info(collections)")}
assert "type" in coll_cols and coll_cols["type"]["dflt_value"] == "'document'"
assert "initial_state" in coll_cols and coll_cols["initial_state"]["dflt_value"] == "'super-draft'"
conn.close()
@@ -1,95 +0,0 @@
"""§22.13 / migration 028 — the slug-keyed PK/UNIQUE rebuild that activates
project #2. Proves two projects can hold the same slug, that (project_id, slug)
is still unique within a project, that the rebuilt FK is composite + enforced,
and that the no-foreign-keys migration runner left no dangling references."""
from __future__ import annotations
import sqlite3
import tempfile
from pathlib import Path
import pytest
from app import db
class _Cfg:
def __init__(self, path):
self.database_path = path
def _fresh_db():
d = tempfile.mkdtemp()
path = Path(d) / "t.db"
db.run_migrations(_Cfg(str(path)))
return db.connect(str(path))
def _seed_two_projects(conn):
for pid in ("default", "ecomm"):
conn.execute(
"INSERT OR IGNORE INTO projects (id, name, type, content_repo, visibility, initial_state) "
"VALUES (?, ?, 'document', ?, 'public', 'super-draft')",
(pid, pid.title(), pid + "-content"),
)
def test_same_slug_coexists_across_projects():
conn = _fresh_db()
_seed_two_projects(conn)
for pid in ("default", "ecomm"):
conn.execute(
"INSERT INTO cached_rfcs (slug, title, state, project_id) "
"VALUES ('intro', 'Intro', 'active', ?)",
(pid,),
)
rows = conn.execute(
"SELECT project_id FROM cached_rfcs WHERE slug = 'intro' ORDER BY project_id"
).fetchall()
assert [r["project_id"] for r in rows] == ["default", "ecomm"]
def test_slug_still_unique_within_a_project():
conn = _fresh_db()
_seed_two_projects(conn)
conn.execute(
"INSERT INTO cached_rfcs (slug, title, state, project_id) "
"VALUES ('intro', 'Intro', 'active', 'default')"
)
with pytest.raises(sqlite3.IntegrityError):
conn.execute(
"INSERT INTO cached_rfcs (slug, title, state, project_id) "
"VALUES ('intro', 'Dup', 'active', 'default')"
)
def test_rfc_collaborators_composite_fk_enforced():
conn = _fresh_db()
_seed_two_projects(conn)
conn.execute("INSERT INTO users (id, gitea_login, display_name, role) VALUES (1, 'a', 'A', 'contributor')")
conn.execute(
"INSERT INTO cached_rfcs (slug, title, state, project_id) "
"VALUES ('intro', 'Intro', 'active', 'ecomm')"
)
# Matching (project_id, slug) — FK holds.
conn.execute(
"INSERT INTO rfc_collaborators (rfc_slug, user_id, role_in_rfc, project_id) "
"VALUES ('intro', 1, 'contributor', 'ecomm')"
)
# Same slug but a project with no such entry — composite FK must reject.
with pytest.raises(sqlite3.IntegrityError):
conn.execute(
"INSERT INTO rfc_collaborators (rfc_slug, user_id, role_in_rfc, project_id) "
"VALUES ('intro', 1, 'contributor', 'default')"
)
def test_stars_unique_now_scoped_by_project():
conn = _fresh_db()
_seed_two_projects(conn)
conn.execute("INSERT INTO users (id, gitea_login, display_name, role) VALUES (1, 'a', 'A', 'contributor')")
# Same (user, slug) under two projects coexist; a duplicate within one rejects.
conn.execute("INSERT INTO stars (user_id, rfc_slug, project_id) VALUES (1, 'intro', 'default')")
conn.execute("INSERT INTO stars (user_id, rfc_slug, project_id) VALUES (1, 'intro', 'ecomm')")
with pytest.raises(sqlite3.IntegrityError):
conn.execute("INSERT INTO stars (user_id, rfc_slug, project_id) VALUES (1, 'intro', 'default')")
@@ -0,0 +1,199 @@
"""Migration 029 — the collection grain beneath projects (§22 three-tier S1).
Proves: a `collections` table exists with one default collection per project
(id='default', subfolder=repo root); the per-corpus fields (type, initial_state)
moved off `projects`; the 13 entry-corpus tables re-key (project_id,slug) ->
(collection_id,slug) with the composite PK/FK enforced; and project_members
generalises into memberships(scope_type, ) with the role enum collapsed.
Template: test_migration_028_project_scoped_keys.py.
"""
from __future__ import annotations
import sqlite3
import tempfile
from pathlib import Path
import pytest
from app import db
class _Cfg:
def __init__(self, path):
self.database_path = path
def _fresh_db():
d = tempfile.mkdtemp()
path = Path(d) / "t.db"
db.run_migrations(_Cfg(str(path)))
return db.connect(str(path))
def test_collections_table_exists_with_default_per_project():
conn = _fresh_db()
cols = {r["name"] for r in conn.execute("PRAGMA table_info(collections)")}
assert {"id", "project_id", "type", "subfolder",
"initial_state", "visibility", "name", "registry_sha"} <= cols
# one default collection seeded for the bootstrap 'default' project (026)
row = conn.execute(
"SELECT id, project_id, subfolder FROM collections WHERE project_id='default'"
).fetchone()
assert row is not None
assert row["id"] == "default"
assert row["subfolder"] == "" # repo root
def test_per_corpus_fields_moved_off_projects():
conn = _fresh_db()
proj_cols = {r["name"] for r in conn.execute("PRAGMA table_info(projects)")}
assert "type" not in proj_cols
assert "initial_state" not in proj_cols
# projects keeps the grouping-tier fields
assert {"id", "name", "content_repo", "visibility"} <= proj_cols
def test_entry_tables_rekeyed_to_collection_id():
conn = _fresh_db()
for t in ("cached_rfcs", "cached_branches", "stars", "watches",
"rfc_collaborators", "contribution_requests", "proposed_use_cases",
"branch_visibility", "branch_contribute_grants", "pr_seen",
"branch_chat_seen", "funder_consents", "rfc_invitations"):
cols = {r["name"] for r in conn.execute(f"PRAGMA table_info({t})")}
assert "collection_id" in cols, f"{t} missing collection_id"
assert "project_id" not in cols, f"{t} still has project_id"
def test_cached_rfcs_pk_is_collection_slug():
conn = _fresh_db()
# a second collection under the default project
conn.execute(
"INSERT INTO collections (id, project_id, type, subfolder, initial_state, visibility, name) "
"VALUES ('c2','default','document','specs','active','public','Specs')"
)
conn.execute("INSERT INTO cached_rfcs (slug, title, state, collection_id) VALUES ('intro','A','active','default')")
conn.execute("INSERT INTO cached_rfcs (slug, title, state, collection_id) VALUES ('intro','B','active','c2')")
n = conn.execute("SELECT COUNT(*) c FROM cached_rfcs WHERE slug='intro'").fetchone()["c"]
assert n == 2
with pytest.raises(sqlite3.IntegrityError):
conn.execute("INSERT INTO cached_rfcs (slug, title, state, collection_id) VALUES ('intro','dup','active','default')")
def test_cached_rfcs_collection_fk_enforced():
conn = _fresh_db()
conn.execute("PRAGMA foreign_keys=ON")
with pytest.raises(sqlite3.IntegrityError):
conn.execute("INSERT INTO cached_rfcs (slug, title, state, collection_id) VALUES ('x','X','active','nope')")
def test_collaborator_fk_is_composite_on_collection():
conn = _fresh_db()
conn.execute(
"INSERT INTO collections (id, project_id, type, subfolder, initial_state, visibility, name) "
"VALUES ('c2','default','document','specs','active','public','Specs')"
)
conn.execute("INSERT INTO cached_rfcs (slug, title, state, collection_id) VALUES ('intro','A','active','c2')")
conn.execute("INSERT INTO users (id, gitea_login, display_name, role) VALUES (1,'a','A','contributor')")
conn.execute("PRAGMA foreign_keys=ON")
conn.execute(
"INSERT INTO rfc_collaborators (rfc_slug, user_id, role_in_rfc, collection_id) "
"VALUES ('intro',1,'contributor','c2')"
)
with pytest.raises(sqlite3.IntegrityError):
# same slug, a collection with no such entry — composite FK rejects
conn.execute(
"INSERT INTO rfc_collaborators (rfc_slug, user_id, role_in_rfc, collection_id) "
"VALUES ('intro',1,'contributor','default')"
)
def test_stars_unique_now_scoped_by_collection():
conn = _fresh_db()
conn.execute(
"INSERT INTO collections (id, project_id, type, subfolder, initial_state, visibility, name) "
"VALUES ('c2','default','document','specs','active','public','Specs')"
)
conn.execute("INSERT INTO users (id, gitea_login, display_name, role) VALUES (1,'a','A','contributor')")
conn.execute("INSERT INTO cached_rfcs (slug, title, state, collection_id) VALUES ('intro','A','active','default')")
conn.execute("INSERT INTO cached_rfcs (slug, title, state, collection_id) VALUES ('intro','B','active','c2')")
conn.execute("INSERT INTO stars (user_id, rfc_slug, collection_id) VALUES (1,'intro','default')")
conn.execute("INSERT INTO stars (user_id, rfc_slug, collection_id) VALUES (1,'intro','c2')")
with pytest.raises(sqlite3.IntegrityError):
conn.execute("INSERT INTO stars (user_id, rfc_slug, collection_id) VALUES (1,'intro','default')")
def test_memberships_table_replaces_project_members():
conn = _fresh_db()
cols = {r["name"] for r in conn.execute("PRAGMA table_info(memberships)")}
assert {"scope_type", "scope_id", "user_id", "role", "granted_by", "granted_at"} <= cols
# project_members is gone
assert conn.execute(
"SELECT name FROM sqlite_master WHERE type='table' AND name='project_members'"
).fetchone() is None
conn.execute("INSERT INTO users (id, gitea_login, display_name, role) VALUES (9,'x','X','contributor')")
conn.execute("INSERT INTO memberships (scope_type, scope_id, user_id, role) VALUES ('project','default',9,'owner')")
# scope_type and role are CHECK-constrained
with pytest.raises(sqlite3.IntegrityError):
conn.execute("INSERT INTO memberships (scope_type, scope_id, user_id, role) VALUES ('bogus','default',9,'owner')")
with pytest.raises(sqlite3.IntegrityError):
conn.execute("INSERT INTO memberships (scope_type, scope_id, user_id, role) VALUES ('project','default',9,'viewer')")
# ── regression: §22.13 satellite re-stamp repair (the OHM-data deploy fault) ──
# Reproduces the shape that crashed the v0.46.0 deploy: the default→ohm re-stamp
# updated cached_rfcs but left cached_branches at the stale project_id='default',
# with (a) a stale row duplicating a freshly-stamped one, (b) a stale row with no
# fresh counterpart, and (c) a stale row whose RFC no longer exists. 029 must
# repair all three rather than hit NOT NULL / UNIQUE on the rebuild.
def _apply_through(path, ceiling):
conn = sqlite3.connect(path, isolation_level=None)
conn.row_factory = sqlite3.Row
conn.execute("CREATE TABLE IF NOT EXISTS schema_migrations (version TEXT PRIMARY KEY, applied_at TEXT NOT NULL DEFAULT (datetime('now')))")
done = {r["version"] for r in conn.execute("SELECT version FROM schema_migrations")}
for p in sorted(db.MIGRATIONS_DIR.glob("*.sql")):
v = p.stem
if v in done or v > ceiling:
continue
sql = p.read_text()
if "-- migrate:no-foreign-keys" in sql:
conn.execute("PRAGMA foreign_keys = OFF")
conn.executescript("BEGIN; " + sql + "; COMMIT;")
conn.execute("PRAGMA foreign_keys = ON")
else:
conn.executescript("BEGIN; " + sql + "; COMMIT;")
conn.execute("INSERT INTO schema_migrations (version) VALUES (?)", (v,))
return conn
def test_029_repairs_stale_duplicate_and_orphan_satellite_rows():
d = tempfile.mkdtemp()
path = str(Path(d) / "t.db")
conn = _apply_through(path, "028_project_scoped_keys")
# simulate the §22.13 re-stamp having renamed the default project + its RFCs
# to 'ohm', but NOT the satellite tables (the actual prod fault).
conn.execute("UPDATE projects SET id='ohm' WHERE id='default'")
conn.execute("INSERT INTO cached_rfcs (slug, title, state, project_id) VALUES ('human','Human','active','ohm')")
conn.execute("INSERT INTO cached_branches (rfc_slug, branch_name, project_id) VALUES ('human','main','ohm')") # fresh/correct
conn.execute("INSERT INTO cached_branches (rfc_slug, branch_name, project_id) VALUES ('human','main','default')") # stale DUP of the fresh one
conn.execute("INSERT INTO cached_branches (rfc_slug, branch_name, project_id) VALUES ('human','edit-1','default')")# stale, unique -> re-stamp+keep
conn.execute("INSERT INTO cached_branches (rfc_slug, branch_name, project_id) VALUES ('ghost','main','default')") # no live RFC -> drop
conn.close()
# apply 029+ (the patched migration). Must NOT raise.
db.run_migrations(_Cfg(path))
conn = db.connect(path)
rows = conn.execute(
"SELECT rfc_slug, branch_name, collection_id FROM cached_branches"
).fetchall()
got = {(r["rfc_slug"], r["branch_name"]) for r in rows}
# every surviving row mapped to a collection (the single-project 'default' one)
assert all(r["collection_id"] is not None for r in rows)
assert {r["collection_id"] for r in rows} == {"default"}
# the duplicate collapsed to exactly one human/main
assert len([r for r in rows if (r["rfc_slug"], r["branch_name"]) == ("human", "main")]) == 1
# the unique stale row survived (re-stamped)
assert ("human", "edit-1") in got
# the no-RFC stale row was dropped
assert ("ghost", "main") not in got
@@ -0,0 +1,91 @@
"""Migration 030 — the global-scope grant (§22 three-tier S3).
Proves: `memberships.scope_type` now admits 'global' alongside 'project' and
'collection' (§B.2's four-layer resolver), existing rows survive the rebuild,
and the UNIQUE(scope_type, scope_id, user_id) shape is preserved.
Template: test_migration_029_collections.py.
"""
from __future__ import annotations
import sqlite3
import tempfile
from pathlib import Path
import pytest
from app import db
class _Cfg:
def __init__(self, path):
self.database_path = path
def _fresh_db():
d = tempfile.mkdtemp()
path = Path(d) / "t.db"
db.run_migrations(_Cfg(str(path)))
return db.connect(str(path))
def _add_user(conn, uid, login):
conn.execute(
"INSERT INTO users (id, gitea_id, gitea_login, display_name, role) "
"VALUES (?, ?, ?, ?, 'contributor')",
(uid, uid, login, login.capitalize()),
)
def test_global_scope_type_is_accepted():
conn = _fresh_db()
_add_user(conn, 1, "cleo")
# global grant — the new tier — is accepted.
conn.execute(
"INSERT INTO memberships (scope_type, scope_id, user_id, role) "
"VALUES ('global', '*', 1, 'contributor')"
)
row = conn.execute(
"SELECT scope_type, scope_id, role FROM memberships WHERE user_id = 1"
).fetchone()
assert row["scope_type"] == "global"
assert row["scope_id"] == "*"
assert row["role"] == "contributor"
def test_project_and_collection_scopes_still_accepted():
conn = _fresh_db()
_add_user(conn, 1, "ben")
conn.execute(
"INSERT INTO memberships (scope_type, scope_id, user_id, role) "
"VALUES ('project', 'default', 1, 'owner')"
)
conn.execute(
"INSERT INTO memberships (scope_type, scope_id, user_id, role) "
"VALUES ('collection', 'default', 1, 'contributor')"
)
n = conn.execute("SELECT COUNT(*) AS n FROM memberships WHERE user_id = 1").fetchone()["n"]
assert n == 2
def test_unknown_scope_type_still_rejected():
conn = _fresh_db()
_add_user(conn, 1, "x")
with pytest.raises(sqlite3.IntegrityError):
conn.execute(
"INSERT INTO memberships (scope_type, scope_id, user_id, role) "
"VALUES ('deployment', '*', 1, 'owner')"
)
def test_one_global_grant_per_user():
conn = _fresh_db()
_add_user(conn, 1, "cleo")
conn.execute(
"INSERT INTO memberships (scope_type, scope_id, user_id, role) "
"VALUES ('global', '*', 1, 'contributor')"
)
with pytest.raises(sqlite3.IntegrityError):
conn.execute(
"INSERT INTO memberships (scope_type, scope_id, user_id, role) "
"VALUES ('global', '*', 1, 'owner')"
)
@@ -0,0 +1,95 @@
"""Migration 032 — the join_requests table (§22.8 S6).
Proves: the table exists with its CHECK constraints (scope_type
{project,collection}; role/status enums), the one-open-per-(scope,user) partial
unique index holds, and a decided request frees a fresh ask.
Template: test_migration_030_global_scope.py.
"""
from __future__ import annotations
import sqlite3
import tempfile
from pathlib import Path
import pytest
from app import db
class _Cfg:
def __init__(self, path):
self.database_path = path
def _fresh_db():
d = tempfile.mkdtemp()
path = Path(d) / "t.db"
db.run_migrations(_Cfg(str(path)))
return db.connect(str(path))
def _add_user(conn, uid, login):
conn.execute(
"INSERT INTO users (id, gitea_id, gitea_login, display_name, role) "
"VALUES (?, ?, ?, ?, 'contributor')",
(uid, uid, login, login.capitalize()),
)
def _request(conn, scope_type="collection", scope_id="features", uid=1, role="contributor"):
conn.execute(
"INSERT INTO join_requests (scope_type, scope_id, requester_user_id, requested_role) "
"VALUES (?, ?, ?, ?)",
(scope_type, scope_id, uid, role),
)
def test_join_request_row_round_trips():
conn = _fresh_db()
_add_user(conn, 1, "ben")
_request(conn)
row = conn.execute("SELECT * FROM join_requests WHERE requester_user_id = 1").fetchone()
assert row["scope_type"] == "collection"
assert row["requested_role"] == "contributor"
assert row["status"] == "pending"
assert row["granted_role"] is None
def test_global_scope_type_is_rejected():
conn = _fresh_db()
_add_user(conn, 1, "ben")
with pytest.raises(sqlite3.IntegrityError):
_request(conn, scope_type="global", scope_id="*")
def test_bad_role_and_status_rejected():
conn = _fresh_db()
_add_user(conn, 1, "ben")
with pytest.raises(sqlite3.IntegrityError):
_request(conn, role="viewer")
with pytest.raises(sqlite3.IntegrityError):
conn.execute(
"INSERT INTO join_requests (scope_type, scope_id, requester_user_id, requested_role, status) "
"VALUES ('project', 'ohm', 1, 'owner', 'maybe')"
)
def test_one_open_request_per_scope_user():
conn = _fresh_db()
_add_user(conn, 1, "ben")
_request(conn)
with pytest.raises(sqlite3.IntegrityError):
_request(conn)
def test_decided_request_frees_a_fresh_ask():
conn = _fresh_db()
_add_user(conn, 1, "ben")
_request(conn)
conn.execute("UPDATE join_requests SET status = 'declined' WHERE requester_user_id = 1")
# a second open ask is now allowed
_request(conn)
n = conn.execute(
"SELECT COUNT(*) AS n FROM join_requests WHERE requester_user_id = 1"
).fetchone()["n"]
assert n == 2
@@ -64,39 +64,55 @@ def _set_visibility(project_id: str, visibility: str) -> None:
)
def _add_member(project_id: str, user_id: int, role: str) -> None:
from app import db
# §22 three-tier (§B.3): M2's three project roles collapse to {owner,
# contributor} in the unified `memberships` table at the project's default
# collection. The read-only `viewer` tier is deferred (folded into contributor
# for this pass), so the legacy role names map: admin→owner, contributor and
# viewer→contributor.
_ROLE_MAP = {
"project_admin": "owner",
"project_contributor": "contributor",
"project_viewer": "contributor",
}
def _add_member(project_id: str, user_id: int, role: str) -> None:
from app import collections as collections_mod, db
cid = collections_mod.default_collection_id(project_id)
db.conn().execute(
"INSERT OR REPLACE INTO project_members (project_id, user_id, role) VALUES (?, ?, ?)",
(project_id, user_id, role),
"INSERT OR REPLACE INTO memberships (scope_type, scope_id, user_id, role) "
"VALUES ('collection', ?, ?, ?)",
(cid, user_id, _ROLE_MAP[role]),
)
def _remove_member(project_id: str, user_id: int) -> None:
from app import db
from app import collections as collections_mod, db
cid = collections_mod.default_collection_id(project_id)
db.conn().execute(
"DELETE FROM project_members WHERE project_id = ? AND user_id = ?",
(project_id, user_id),
"DELETE FROM memberships WHERE scope_type = 'collection' AND scope_id = ? AND user_id = ?",
(cid, user_id),
)
def _seed_rfc(slug: str, *, state: str = "active", owners=None, project_id: str = "default") -> None:
"""A minimal cached_rfcs row — enough for the authz gates (state, owners,
project_id). project_id defaults to 'default' via migration 026 but we set
it explicitly for clarity."""
collection grain). The entry lands in the project's default collection
(id == project_id for the single 'default' project under test)."""
import json
from app import db
from app import collections as collections_mod, db
cid = collections_mod.default_collection_id(project_id)
db.conn().execute(
"""
INSERT OR REPLACE INTO cached_rfcs
(slug, title, state, owners_json, arbiters_json, tags_json, project_id)
(slug, title, state, owners_json, arbiters_json, tags_json, collection_id)
VALUES (?, ?, ?, ?, '[]', '[]', ?)
""",
(slug, slug.capitalize(), state, json.dumps(owners or []), project_id),
(slug, slug.capitalize(), state, json.dumps(owners or []), cid),
)
@@ -154,18 +170,16 @@ def test_resolver_gated_project_requires_membership(app_with_fake_gitea):
assert auth.can_read_project(owner, "default") is True
assert auth.is_project_superuser(owner, "default") is True
# project_viewer → read + discuss, but not contribute.
_add_member("default", 1, "project_viewer")
# §22 three-tier (§B.3): the read-only viewer tier is deferred — the
# smallest grant is `contributor`, which grants read + discuss +
# contribute across the subtree.
_add_member("default", 1, "project_contributor")
assert auth.can_read_project(contributor, "default") is True
assert auth.can_discuss_in_project(contributor, "default") is True
assert auth.can_contribute_in_project(contributor, "default") is False
# project_contributor → contribute.
_add_member("default", 1, "project_contributor")
assert auth.can_contribute_in_project(contributor, "default") is True
assert auth.is_project_superuser(contributor, "default") is False
# project_admin → superuser within the project.
# project_admin → owner → superuser within the project.
_add_member("default", 1, "project_admin")
assert auth.is_project_superuser(contributor, "default") is True
@@ -270,7 +284,7 @@ def test_gated_propose_requires_project_contributor(app_with_fake_gitea):
assert client.post("/api/rfcs/propose", json=body).status_code != 403
def test_gated_viewer_can_discuss_contributor_can_contribute(app_with_fake_gitea):
def test_gated_member_can_discuss_and_contribute(app_with_fake_gitea):
from app import auth
app, _ = app_with_fake_gitea
@@ -285,14 +299,11 @@ def test_gated_viewer_can_discuss_contributor_can_contribute(app_with_fake_gitea
sign_in_as(client, user_id=2, gitea_login="bob", display_name="Bob", role="contributor")
assert client.post("/api/rfcs/spec/discussion/threads", json={"message": "q"}).status_code == 404
# project_viewer: can discuss (200) but cannot contribute (resolver).
_add_member("default", 2, "project_viewer")
# §22 three-tier (§B.3): a `contributor` member can both discuss and
# contribute (the viewer-only read tier is deferred this pass).
_add_member("default", 2, "project_contributor")
r = client.post("/api/rfcs/spec/discussion/threads", json={"message": "q"})
assert r.status_code == 200, r.text
assert auth.can_contribute_to_rfc(bob, "spec") is False
# project_contributor: can contribute.
_add_member("default", 2, "project_contributor")
assert auth.can_contribute_to_rfc(bob, "spec") is True
@@ -16,14 +16,18 @@ from test_propose_vertical import ( # noqa: F401
tmp_env,
)
# The 19 tables migration 026 threads project_id onto (docs/design/
# multi-project-spec.md §5 amendment list).
SLUG_TABLES = [
"cached_rfcs", "cached_branches", "cached_prs", "branch_visibility",
"branch_contribute_grants", "stars", "threads", "changes", "pr_seen",
"branch_chat_seen", "watches", "notifications", "actions",
"pr_resolution_branches", "funder_consents", "rfc_invitations",
"rfc_collaborators", "proposed_use_cases", "contribution_requests",
# §22 three-tier (migration 029): the entry-corpus grain is the collection, so
# the 13 tables migration 028 keyed by project_id re-key to collection_id. The
# remaining tables 026 tagged keep their denormalised project_id (project grain).
COLLECTION_TABLES = [
"cached_rfcs", "cached_branches", "branch_visibility",
"branch_contribute_grants", "stars", "pr_seen", "branch_chat_seen",
"watches", "funder_consents", "rfc_invitations", "rfc_collaborators",
"proposed_use_cases", "contribution_requests",
]
PROJECT_TAG_TABLES = [
"cached_prs", "threads", "changes", "notifications", "actions",
"pr_resolution_branches",
]
@@ -45,25 +49,28 @@ def test_default_project_seeded_and_backfilled(app_with_fake_gitea):
assert row["content_repo"] == "meta"
def test_project_id_on_every_slug_table(app_with_fake_gitea):
def test_grain_columns_on_every_slug_table(app_with_fake_gitea):
from app import db
app, _ = app_with_fake_gitea
with TestClient(app):
for table in SLUG_TABLES:
cols = {r["name"]: r for r in db.conn().execute(
f"PRAGMA table_info({table})"
)}
assert "project_id" in cols, f"{table} missing project_id"
col = cols["project_id"]
# NOT NULL with the constant 'default' backfill default.
assert col["notnull"] == 1, f"{table}.project_id should be NOT NULL"
assert col["dflt_value"] == "'default'", f"{table}.project_id default"
# The entry-corpus tables key on collection_id (NOT NULL, 'default').
for table in COLLECTION_TABLES:
cols = {r["name"]: r for r in db.conn().execute(f"PRAGMA table_info({table})")}
assert "collection_id" in cols, f"{table} missing collection_id"
assert "project_id" not in cols, f"{table} should no longer have project_id"
col = cols["collection_id"]
assert col["notnull"] == 1, f"{table}.collection_id should be NOT NULL"
assert col["dflt_value"] == "'default'", f"{table}.collection_id default"
# The project-tag tables keep their denormalised project_id.
for table in PROJECT_TAG_TABLES:
cols = {r["name"] for r in db.conn().execute(f"PRAGMA table_info({table})")}
assert "project_id" in cols, f"{table} missing project_id tag"
def test_existing_row_backfills_to_default(app_with_fake_gitea):
"""A row inserted the old way (no project_id) lands in the default
project the trick that keeps every pre-multi-project INSERT working."""
"""A row inserted the old way (no collection grain) lands in the default
collection the trick that keeps every pre-three-tier INSERT working."""
from app import db
app, _ = app_with_fake_gitea
@@ -73,35 +80,36 @@ def test_existing_row_backfills_to_default(app_with_fake_gitea):
("human", "Human", "active"),
)
got = db.conn().execute(
"SELECT project_id FROM cached_rfcs WHERE slug = 'human'"
).fetchone()["project_id"]
"SELECT collection_id FROM cached_rfcs WHERE slug = 'human'"
).fetchone()["collection_id"]
assert got == "default"
def test_project_members_table_shape(app_with_fake_gitea):
def test_memberships_table_shape(app_with_fake_gitea):
from app import db
app, _ = app_with_fake_gitea
with TestClient(app):
cols = {r["name"] for r in db.conn().execute(
"PRAGMA table_info(project_members)"
)}
assert cols == {"project_id", "user_id", "role", "granted_by", "granted_at"}
# The role CHECK rejects an unknown role.
# §22 three-tier: project_members generalised into memberships.
assert db.conn().execute(
"SELECT name FROM sqlite_master WHERE type='table' AND name='project_members'"
).fetchone() is None
cols = {r["name"] for r in db.conn().execute("PRAGMA table_info(memberships)")}
assert {"scope_type", "scope_id", "user_id", "role", "granted_by", "granted_at"} <= cols
db.conn().execute(
"INSERT INTO users (id, display_name, role) VALUES (1, 'Ben', 'owner')"
)
db.conn().execute(
"INSERT INTO project_members (project_id, user_id, role) "
"VALUES ('default', 1, 'project_admin')"
"INSERT INTO memberships (scope_type, scope_id, user_id, role) "
"VALUES ('collection', 'default', 1, 'owner')"
)
import sqlite3
try:
db.conn().execute(
"INSERT INTO project_members (project_id, user_id, role) "
"VALUES ('default', 1, 'nonsense')"
"INSERT INTO memberships (scope_type, scope_id, user_id, role) "
"VALUES ('collection', 'default', 1, 'nonsense')"
)
assert False, "CHECK should reject an unknown project role"
assert False, "CHECK should reject an unknown role"
except sqlite3.IntegrityError:
pass
@@ -0,0 +1,79 @@
"""§22.4 (Plan B write): proposing a new entry into a *specific* project lands
it in that project's content repo and surfaces under that project's proposals,
isolated from the default project."""
from __future__ import annotations
from fastapi.testclient import TestClient
from test_propose_vertical import ( # noqa: F401
app_with_fake_gitea, tmp_env, provision_user_row, sign_in_as,
)
def _register_ecomm(fake):
from app import db
db.conn().execute(
"INSERT OR IGNORE INTO projects (id, name, content_repo, visibility) "
"VALUES ('ecomm', 'Ecomm', 'ecomm-content', 'public')"
)
db.conn().execute(
"INSERT OR IGNORE INTO collections (id, project_id, type, subfolder, initial_state, visibility, name) "
"VALUES ('ecomm', 'ecomm', 'document', '', 'super-draft', 'public', 'Ecomm')"
)
fake._seed_repo("wiggleverse", "ecomm-content")
def test_propose_into_second_project_lands_scoped(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_register_ecomm(fake)
provision_user_row(user_id=3, login="alice", role="contributor")
# §22 S3: the grandfathered implicit-public baseline covers only the N=1
# `default` collection; a second project requires an explicit scope grant
# to write. Grant alice contributor at the ecomm project.
from app import db
db.conn().execute(
"INSERT OR REPLACE INTO memberships (scope_type, scope_id, user_id, role) "
"VALUES ('project', 'ecomm', 3, 'contributor')")
sign_in_as(client, user_id=3, gitea_login="alice", display_name="Alice",
role="contributor", email="alice@test")
r = client.post("/api/projects/ecomm/rfcs/propose", json={
"title": "Cart", "slug": "cart", "pitch": "why a cart", "tags": [],
})
assert r.status_code == 200, r.text
# The idea PR shows under ecomm's proposals, not the default's.
e = {i["slug"] for i in client.get("/api/projects/ecomm/proposals").json()["items"]}
d = {i["slug"] for i in client.get("/api/projects/default/proposals").json()["items"]}
assert "cart" in e
assert "cart" not in d
# It landed in ecomm's content repo, not the default 'meta' repo.
assert ("wiggleverse", "ecomm-content") in {
(o, rp) for (o, rp) in fake.branches if rp == "ecomm-content"
}
assert any(
br.startswith("propose/cart")
for br in fake.branches.get(("wiggleverse", "ecomm-content"), {})
)
def test_propose_into_gated_project_404s_for_non_member(app_with_fake_gitea):
app, _ = app_with_fake_gitea
from app import db
with TestClient(app) as client:
db.conn().execute(
"INSERT OR IGNORE INTO projects (id, name, content_repo, visibility) "
"VALUES ('secret', 'Secret', 'secret-content', 'gated')"
)
db.conn().execute(
"INSERT OR IGNORE INTO collections (id, project_id, type, subfolder, initial_state, visibility, name) "
"VALUES ('secret', 'secret', 'document', '', 'super-draft', 'gated', 'Secret')"
)
provision_user_row(user_id=4, login="bob", role="contributor")
sign_in_as(client, user_id=4, gitea_login="bob", display_name="Bob",
role="contributor", email="bob@test")
r = client.post("/api/projects/secret/rfcs/propose", json={
"title": "X", "slug": "x", "pitch": "p", "tags": [],
})
assert r.status_code == 404
+10 -3
View File
@@ -11,18 +11,25 @@ from test_propose_vertical import ( # noqa: F401
def _add_project(pid, name, vis="public"):
# §22 three-tier: a project + its default collection (keyed by the project
# id in tests, so default_collection_id(pid) == pid).
from app import db
db.conn().execute(
"INSERT OR IGNORE INTO projects (id, name, type, content_repo, visibility, initial_state) "
"VALUES (?, ?, 'document', ?, ?, 'super-draft')",
"INSERT OR IGNORE INTO projects (id, name, content_repo, visibility) VALUES (?, ?, ?, ?)",
(pid, name, pid + "-content", vis),
)
db.conn().execute(
"INSERT OR IGNORE INTO collections (id, project_id, type, subfolder, initial_state, visibility, name) "
"VALUES (?, ?, 'document', '', 'super-draft', ?, ?)",
(pid, pid, vis, name),
)
def _add_rfc(slug, title, pid, state="active"):
from app import db
# entries key by the project's default collection (id == pid in these tests)
db.conn().execute(
"INSERT INTO cached_rfcs (slug, title, state, project_id) VALUES (?, ?, ?, ?)",
"INSERT INTO cached_rfcs (slug, title, state, collection_id) VALUES (?, ?, ?, ?)",
(slug, title, state, pid),
)
+60 -11
View File
@@ -54,6 +54,9 @@ class FakeGitea:
self.repos: set[tuple[str, str]] = set()
self._pr_counter = 0
self._commit_counter = 0
# count of batch ChangeFiles commits (one per /contents POST with a
# files[] array) — lets tests assert "N files, one commit" (§22.4a D7).
self.change_files_calls = 0
self._seed_repo("wiggleverse", "meta")
# §22 M3: the deployment's project registry. Startup refresh_registry
# reads projects.yaml here; the single 'default' project's content_repo
@@ -90,6 +93,28 @@ class FakeGitea:
self._commit_counter += 1
return f"sha{self._commit_counter:04d}"
def _dir_listing(self, owner, repo, ref, dirpath):
"""Children directly under `dirpath` on (owner, repo, ref): files as
`type: file` and immediate subdirectories as `type: dir` (the shape real
Gitea returns for a contents listing)."""
prefix = (dirpath.rstrip("/") + "/") if dirpath else ""
files: dict[str, dict] = {}
dirs: set[str] = set()
for (o, r, br, p), data in self.files.items():
if (o, r, br) != (owner, repo, ref) or not p.startswith(prefix):
continue
rest = p[len(prefix):]
if "/" in rest:
dirs.add(rest.split("/", 1)[0])
elif rest:
files[p] = data
children = [{"name": n, "path": prefix + n, "type": "dir"} for n in sorted(dirs)]
children += [
{"name": p.rsplit("/", 1)[-1], "path": p, "type": "file", "sha": d["sha"]}
for p, d in sorted(files.items())
]
return children
def _enrich_pr(self, owner: str, repo: str, pr: dict) -> dict:
"""Return the PR with mergeability fields filled in.
@@ -227,6 +252,16 @@ class FakeGitea:
}
return httpx.Response(201, json={"name": new})
# GET /repos/{owner}/{repo}/contents (root listing, empty path). §22 S2:
# the registry mirror walks the content-repo root for collection
# subfolders, so the simulator models a root directory listing that
# surfaces both file and `dir` children.
m_root = re.fullmatch(r"/repos/([^/]+)/([^/]+)/contents/?", path)
if method == "GET" and m_root:
owner, repo = m_root.groups()
ref = request.url.params.get("ref", "main")
return httpx.Response(200, json=self._dir_listing(owner, repo, ref, ""))
# GET /repos/{owner}/{repo}/contents/{path}?ref=...
m = re.fullmatch(r"/repos/([^/]+)/([^/]+)/contents/(.+)", path)
if method == "GET" and m:
@@ -242,21 +277,35 @@ class FakeGitea:
"sha": f["sha"],
"content": base64.b64encode(f["content"].encode()).decode(),
})
# Directory listing
prefix = fpath.rstrip("/") + "/"
children = []
for (o, r, br, p), data in self.files.items():
if (o, r, br) == (owner, repo, ref) and p.startswith(prefix) and "/" not in p[len(prefix):]:
children.append({
"name": p.rsplit("/", 1)[-1],
"path": p,
"type": "file",
"sha": data["sha"],
})
# Directory listing — both file and subdir children.
children = self._dir_listing(owner, repo, ref, fpath)
if children:
return httpx.Response(200, json=children)
return httpx.Response(404, json={"message": "not found"})
# POST /repos/{owner}/{repo}/contents — ChangeFiles (batch, one commit).
# §22.4a SLICE-1: the frontmatter→sidecar migration writes N files in a
# single commit. Matches the no-path /contents route (the per-path POST
# below needs a /contents/<path> suffix).
m_batch = re.fullmatch(r"/repos/([^/]+)/([^/]+)/contents/?", path)
if method == "POST" and m_batch:
owner, repo = m_batch.groups()
branch = payload["branch"]
self.change_files_calls += 1
sha = self._next_sha()
for f in payload["files"]:
op = f["operation"]
fpath = f["path"]
if op == "delete":
self.files.pop((owner, repo, branch, fpath), None)
else:
content = base64.b64decode(f["content"]).decode()
self.files[(owner, repo, branch, fpath)] = {"content": content, "sha": sha}
br = self.branches[(owner, repo)].setdefault(branch, {})
br["sha"] = sha
br["ts"] = "2026-05-23T00:00:00Z"
return httpx.Response(201, json={"commit": {"sha": sha}})
# POST /repos/{owner}/{repo}/contents/{path}
m = re.fullmatch(r"/repos/([^/]+)/([^/]+)/contents/(.+)", path)
if method == "POST" and m:
@@ -0,0 +1,107 @@
"""§6 hardening — the Gitea-OAuth `auth.provision_user` reconciles by email.
Regression cover for a §9-surfaced fragility: a human who signed in first via
OTC owns an email-only `users` row (`gitea_id` NULL). When they later sign in
via Gitea OAuth with the SAME email, `provision_user` used to match only by
`gitea_id`, miss the OTC row, and INSERT a new row colliding on the
`idx_users_email` unique index and 500-ing the callback. The fix links the OAuth
identity onto the existing email row (the mirror of the OTC linker).
"""
from __future__ import annotations
from test_propose_vertical import ( # noqa: F401
FakeGitea,
app_with_fake_gitea,
provision_user_row,
tmp_env,
)
def _cfg():
from app.config import load_config
return load_config()
def test_oauth_links_onto_existing_otc_email_row(app_with_fake_gitea):
from fastapi.testclient import TestClient
from app import auth, db
app, _fake = app_with_fake_gitea
with TestClient(app):
# An OTC-first user: email-only row, gitea_id NULL, still 'pending'.
db.conn().execute(
"""
INSERT INTO users (gitea_id, gitea_login, email, display_name, avatar_url, role, permission_state)
VALUES (NULL, NULL, ?, ?, '', 'contributor', 'pending')
""",
("dual@example.com", "dual"),
)
otc_id = db.conn().execute(
"SELECT id FROM users WHERE email = ? COLLATE NOCASE", ("dual@example.com",)
).fetchone()["id"]
# Now the same human signs in via Gitea OAuth (new gitea_id, same email).
# Case-different email proves the NOCASE match.
user = auth.provision_user(
_cfg(),
{"id": 9001, "login": "dualgitea", "email": "Dual@Example.com",
"full_name": "Dual User", "avatar_url": "http://x/a.png"},
)
# Linked onto the SAME row — no second row, no 500.
assert user.user_id == otc_id
rows = db.conn().execute(
"SELECT id, gitea_id, gitea_login, permission_state, role FROM users WHERE email = ? COLLATE NOCASE",
("dual@example.com",),
).fetchall()
assert len(rows) == 1
assert rows[0]["id"] == otc_id
assert rows[0]["gitea_id"] == 9001 # OAuth identity attached
assert rows[0]["gitea_login"] == "dualgitea"
assert rows[0]["permission_state"] == "pending" # admission state preserved
assert rows[0]["role"] == "contributor"
def test_oauth_provisions_fresh_user_when_email_matches_no_one(app_with_fake_gitea):
from fastapi.testclient import TestClient
from app import auth, db
app, _fake = app_with_fake_gitea
with TestClient(app):
user = auth.provision_user(
_cfg(),
{"id": 9100, "login": "freshoauth", "email": "fresh@example.com",
"full_name": "Fresh", "avatar_url": ""},
)
row = db.conn().execute(
"SELECT gitea_id, gitea_login, role, permission_state FROM users WHERE id = ?",
(user.user_id,),
).fetchone()
assert row["gitea_id"] == 9100
assert row["gitea_login"] == "freshoauth"
# Reaching provision_user means admission passed → granted (unchanged).
assert row["permission_state"] == "granted"
def test_returning_oauth_user_matched_by_gitea_id_not_duplicated(app_with_fake_gitea):
from fastapi.testclient import TestClient
from app import auth, db
app, _fake = app_with_fake_gitea
with TestClient(app):
provision_user_row(user_id=55, login="returning", role="contributor")
before = db.conn().execute("SELECT COUNT(*) AS n FROM users").fetchone()["n"]
user = auth.provision_user(
_cfg(),
{"id": 55, "login": "returning-renamed", "email": "returning@test",
"full_name": "Returning", "avatar_url": ""},
)
after = db.conn().execute("SELECT COUNT(*) AS n FROM users").fetchone()["n"]
assert user.user_id == 55
assert after == before # matched by gitea_id; no new row
row = db.conn().execute(
"SELECT gitea_login FROM users WHERE id = ?", (55,)
).fetchone()
assert row["gitea_login"] == "returning-renamed" # profile refreshed
+44
View File
@@ -0,0 +1,44 @@
"""The per-IP limiter budgets are env-overridable (test/PPE stacks drive the
auth endpoints repeatedly from one IP); production leaves them unset and keeps
the secure defaults. A non-positive / unparseable value falls back."""
from __future__ import annotations
import importlib
import pytest
def _reload_with(monkeypatch, **env):
for k, v in env.items():
if v is None:
monkeypatch.delenv(k, raising=False)
else:
monkeypatch.setenv(k, v)
import app.ratelimit as ratelimit
return importlib.reload(ratelimit)
@pytest.fixture(autouse=True)
def _restore():
yield
# Leave the module in its default state for other tests.
import app.ratelimit as ratelimit
importlib.reload(ratelimit)
def test_defaults_when_unset(monkeypatch):
rl = _reload_with(monkeypatch, RATELIMIT_OTC_REQUEST_MAX=None, RATELIMIT_VERIFY_MAX=None)
assert rl.otc_request_limiter.max_events == 5
assert rl.verify_limiter.max_events == 10
def test_env_override(monkeypatch):
rl = _reload_with(monkeypatch, RATELIMIT_OTC_REQUEST_MAX="1000", RATELIMIT_VERIFY_MAX="250")
assert rl.otc_request_limiter.max_events == 1000
assert rl.verify_limiter.max_events == 250
def test_bad_value_falls_back_to_default(monkeypatch):
rl = _reload_with(monkeypatch, RATELIMIT_OTC_REQUEST_MAX="nope", RATELIMIT_VERIFY_MAX="0")
assert rl.otc_request_limiter.max_events == 5 # unparseable → default
assert rl.verify_limiter.max_events == 10 # non-positive → default
@@ -0,0 +1,156 @@
"""§22 framework bug — migration-029 vs registry-mirror collection-id
divergence for a multi-project deployment's default project.
Migration 029 seeds the default project's collection id as the literal
'default' only when one project exists at migration time; with 2 projects it
falls back to the *project id*. The registry mirror expects the default
project to own the collection id 'default'. On an upgrade whose DB already
held 2 projects when 029 ran, the default project's collection is therefore
named after the project (e.g. 'ohm'), and the next mirror would INSERT a
second, empty 'default' collection. `reconcile_default_collection_id` heals
the divergence at startup, before the mirror, so the mirror merges instead of
duplicating. This is the collection-grain twin of `restamp_default_project`.
"""
from __future__ import annotations
import tempfile
from pathlib import Path
import app.db as db
from app import projects, registry
_TWO_PROJECT_REGISTRY = """
deployment:
name: Open Human Model
tagline: t
projects:
- id: ohm
name: Open Human Model
type: document
content_repo: ohm-content
visibility: public
- id: ecomm
name: Ecomm
type: bdd
content_repo: ecomm-content
visibility: public
"""
def _collection_ids_for(conn, project_id):
return {r["id"] for r in conn.execute(
"SELECT id FROM collections WHERE project_id = ?", (project_id,))}
class _Cfg:
def __init__(self, path, default_id):
self.database_path = path
self.default_project_id = default_id
def _divergent_multiproject(monkeypatch, default_id="ohm"):
"""A DB in the post-029 divergent state: the default project ('ohm') owns a
collection whose id is the project id (the 029 2-projects seed), a second
project ('ecomm') owns its own collection, and NO 'default' collection
exists. Entry rows point at the divergent collection_id='ohm'."""
path = str(Path(tempfile.mkdtemp()) / "t.db")
cfg = _Cfg(path, default_id)
db.run_migrations(cfg) # seeds bootstrap 'default' project + 'default' collection
monkeypatch.setattr(db, "_CONN", db.connect(path))
conn = db.conn()
# Drop the single-project bootstrap seed and rebuild the divergent
# multi-project state 029 would have produced on an upgrade.
conn.execute("DELETE FROM collections")
conn.execute("DELETE FROM projects")
conn.execute("INSERT INTO projects (id,name,content_repo,visibility) "
"VALUES ('ohm','Open Human Model','ohm-content','public')")
conn.execute("INSERT INTO projects (id,name,content_repo,visibility) "
"VALUES ('ecomm','Ecomm','ecomm-content','public')")
# 029 ≥2-projects seed: collection id == project id.
conn.execute("INSERT INTO collections (id,project_id,type,subfolder,initial_state,visibility,name) "
"VALUES ('ohm','ohm','document','','super-draft','public','Open Human Model')")
conn.execute("INSERT INTO collections (id,project_id,type,subfolder,initial_state,visibility,name) "
"VALUES ('ecomm','ecomm','bdd','ecomm','super-draft','public','Ecomm')")
conn.execute("INSERT INTO users (id,gitea_login,display_name,role) VALUES (1,'a','A','contributor')")
# Entry data for the default project lives in the divergent 'ohm' collection.
conn.execute("INSERT INTO cached_rfcs (slug,title,state,collection_id) VALUES ('human','Human','active','ohm')")
conn.execute("INSERT INTO rfc_collaborators (rfc_slug,user_id,role_in_rfc,collection_id) "
"VALUES ('human',1,'contributor','ohm')")
conn.execute("INSERT INTO stars (user_id,rfc_slug,collection_id) VALUES (1,'human','ohm')")
# The 'ecomm' collection has its own entry.
conn.execute("INSERT INTO cached_rfcs (slug,title,state,collection_id) VALUES ('cart','Cart','active','ecomm')")
return cfg, conn
def test_reconcile_renames_divergent_default_collection_to_default(monkeypatch):
cfg, conn = _divergent_multiproject(monkeypatch, default_id="ohm")
projects.reconcile_default_collection_id(cfg)
# The default project's collection id is now the canonical 'default'.
assert conn.execute("SELECT 1 FROM collections WHERE id='ohm'").fetchone() is None
row = conn.execute("SELECT project_id FROM collections WHERE id='default'").fetchone()
assert row is not None and row["project_id"] == "ohm"
# Entry rows cascaded onto 'default'.
assert conn.execute("SELECT collection_id FROM cached_rfcs WHERE slug='human'").fetchone()["collection_id"] == "default"
assert conn.execute("SELECT collection_id FROM rfc_collaborators WHERE rfc_slug='human'").fetchone()["collection_id"] == "default"
assert conn.execute("SELECT collection_id FROM stars WHERE rfc_slug='human'").fetchone()["collection_id"] == "default"
# The non-default 'ecomm' collection is untouched (mirror + 029 agree on it).
assert conn.execute("SELECT 1 FROM collections WHERE id='ecomm'").fetchone() is not None
assert conn.execute("SELECT collection_id FROM cached_rfcs WHERE slug='cart'").fetchone()["collection_id"] == "ecomm"
# FK integrity intact after the rename.
assert conn.execute("PRAGMA foreign_key_check").fetchall() == []
def test_reconcile_is_idempotent(monkeypatch):
cfg, conn = _divergent_multiproject(monkeypatch, default_id="ohm")
projects.reconcile_default_collection_id(cfg)
projects.reconcile_default_collection_id(cfg) # second call: already aligned → no-op
assert conn.execute("SELECT project_id FROM collections WHERE id='default'").fetchone()["project_id"] == "ohm"
assert conn.execute("SELECT COUNT(*) c FROM cached_rfcs WHERE collection_id='default'").fetchone()["c"] == 1
def test_reconcile_noop_when_default_id_is_literal_default(monkeypatch):
# Single-project deployment, no DEFAULT_PROJECT_ID: 029 already seeded
# 'default' and the mirror agrees — nothing to reconcile.
path = str(Path(tempfile.mkdtemp()) / "t.db")
cfg = _Cfg(path, "") # resolves to 'default'
db.run_migrations(cfg)
monkeypatch.setattr(db, "_CONN", db.connect(path))
conn = db.conn()
projects.reconcile_default_collection_id(cfg)
assert conn.execute("SELECT 1 FROM collections WHERE id='default'").fetchone() is not None
def test_mirror_duplicates_without_reconcile(monkeypatch):
# Demonstrates the bug: the mirror on the divergent state inserts a SECOND
# 'default' collection for the default project (alongside the 029 'ohm').
cfg, conn = _divergent_multiproject(monkeypatch, default_id="ohm")
doc = registry.parse_registry(_TWO_PROJECT_REGISTRY)
registry.apply_registry(doc, registry_sha="s1", default_id="ohm")
assert _collection_ids_for(conn, "ohm") == {"ohm", "default"} # duplicate!
def test_reconcile_then_mirror_merges_no_duplicate(monkeypatch):
# With the fix: reconcile before the mirror → the mirror merges onto the
# canonical 'default' collection; the default project owns exactly one.
cfg, conn = _divergent_multiproject(monkeypatch, default_id="ohm")
projects.reconcile_default_collection_id(cfg)
doc = registry.parse_registry(_TWO_PROJECT_REGISTRY)
registry.apply_registry(doc, registry_sha="s1", default_id="ohm")
assert _collection_ids_for(conn, "ohm") == {"default"}
assert _collection_ids_for(conn, "ecomm") == {"ecomm"}
# The default project's corpus entry is intact under 'default'.
assert conn.execute("SELECT collection_id FROM cached_rfcs WHERE slug='human'").fetchone()["collection_id"] == "default"
# The mirror refreshed the merged collection's metadata (type from registry).
assert conn.execute("SELECT type FROM collections WHERE id='default'").fetchone()["type"] == "document"
def test_reconcile_skips_when_default_collection_already_exists(monkeypatch):
# A prior buggy mirror already created a 'default' collection alongside the
# divergent 'ohm' one: don't auto-merge data — leave both for operator cleanup.
cfg, conn = _divergent_multiproject(monkeypatch, default_id="ohm")
conn.execute("INSERT INTO collections (id,project_id,type,subfolder,initial_state,visibility,name) "
"VALUES ('default','ohm','document','','super-draft','public','dup')")
projects.reconcile_default_collection_id(cfg)
# Both still present (no destructive auto-merge).
assert conn.execute("SELECT 1 FROM collections WHERE id='ohm'").fetchone() is not None
assert conn.execute("SELECT 1 FROM collections WHERE id='default'").fetchone() is not None
+14 -6
View File
@@ -74,13 +74,20 @@ def test_parse_rejects_invalid(bad, msg):
def test_apply_upserts_projects_and_deployment():
_db()
doc = registry.parse_registry(VALID)
registry.apply_registry(doc, registry_sha="regsha1")
registry.apply_registry(doc, registry_sha="regsha1", default_id="default")
# §22 three-tier: the project carries the grouping-tier fields; the
# per-corpus type/initial_state live on its default collection.
prow = db.conn().execute(
"SELECT name, type, content_repo, visibility, initial_state, registry_sha FROM projects WHERE id='default'"
"SELECT name, content_repo, visibility, registry_sha FROM projects WHERE id='default'"
).fetchone()
assert prow["name"] == "Open Human Model"
assert prow["content_repo"] == "meta"
assert prow["registry_sha"] == "regsha1"
crow = db.conn().execute(
"SELECT type, initial_state FROM collections WHERE id='default'"
).fetchone()
assert crow["type"] == "document"
assert crow["initial_state"] == "super-draft"
drow = db.conn().execute("SELECT name, tagline FROM deployment WHERE id=1").fetchone()
assert drow["name"] == "Open Human Model"
assert drow["tagline"] == "A model of human flourishing"
@@ -88,11 +95,12 @@ def test_apply_upserts_projects_and_deployment():
def test_apply_rejects_type_change_on_existing_project():
_db()
registry.apply_registry(registry.parse_registry(VALID), "s1")
registry.apply_registry(registry.parse_registry(VALID), "s1", default_id="default")
changed = VALID.replace("type: document", "type: specification")
registry.apply_registry(registry.parse_registry(changed), "s2") # logged + skipped, no raise
t = db.conn().execute("SELECT type FROM projects WHERE id='default'").fetchone()["type"]
registry.apply_registry(registry.parse_registry(changed), "s2", default_id="default") # skipped
# §22.4a immutable type — now enforced on the collection.
t = db.conn().execute("SELECT type FROM collections WHERE id='default'").fetchone()["type"]
assert t == "document" # immutable — unchanged
# The deployment row IS still advanced even though the project upsert was skipped.
# The deployment row IS still advanced even though the type change was skipped.
drow = db.conn().execute("SELECT registry_sha FROM deployment WHERE id=1").fetchone()
assert drow["registry_sha"] == "s2"
+6 -2
View File
@@ -12,10 +12,14 @@ def test_startup_mirrors_registry_into_projects_and_deployment(app_with_fake_git
app, _ = app_with_fake_gitea
with TestClient(app):
prow = db.conn().execute(
"SELECT content_repo, type, initial_state FROM projects WHERE id='default'"
"SELECT content_repo FROM projects WHERE id='default'"
).fetchone()
assert prow["content_repo"] == "meta" # from the registry, not META_REPO
assert prow["type"] == "document"
# §22 three-tier: type now lives on the default collection.
crow = db.conn().execute(
"SELECT type FROM collections WHERE id='default'"
).fetchone()
assert crow["type"] == "document"
drow = db.conn().execute("SELECT name FROM deployment WHERE id=1").fetchone()
assert drow["name"] # deployment name mirrored from the registry
@@ -0,0 +1,69 @@
"""§22.13 step 1 — the bootstrap-id re-stamp: 'default' → the configured
DEFAULT_PROJECT_ID. §22 three-tier (S1): the entry-corpus tables key on
collection_id now, so the re-stamp renames the *project grain* the
`collections.project_id` link and the denormalised project_id tags while the
entries stay in their collection. The stale 'default' projects row is dropped,
the composite FKs stay intact, and it is idempotent."""
from __future__ import annotations
import tempfile
from pathlib import Path
import app.db as db
from app import projects
class _Cfg:
def __init__(self, path, default_id):
self.database_path = path
self.default_project_id = default_id
def _setup(monkeypatch, default_id="ohm"):
path = str(Path(tempfile.mkdtemp()) / "t.db")
cfg = _Cfg(path, default_id)
db.run_migrations(cfg) # seeds the bootstrap 'default' project + its default collection
monkeypatch.setattr(db, "_CONN", db.connect(path))
conn = db.conn()
# A registry-mirrored 'ohm' project coexists with the bootstrap pre-restamp.
conn.execute("INSERT OR IGNORE INTO projects (id,name,content_repo,visibility) "
"VALUES ('ohm','Open Human Model','ohm-content','public')")
conn.execute("INSERT INTO users (id,gitea_login,display_name,role) VALUES (1,'a','A','contributor')")
# Entry data lives in the default collection (id='default'); the entry grain
# is the collection and does not move on a re-stamp.
conn.execute("INSERT INTO cached_rfcs (slug,title,state,collection_id) VALUES ('human','Human','active','default')")
conn.execute("INSERT INTO rfc_collaborators (rfc_slug,user_id,role_in_rfc,collection_id) "
"VALUES ('human',1,'contributor','default')")
conn.execute("INSERT INTO stars (user_id,rfc_slug,collection_id) VALUES (1,'human','default')")
return cfg, conn
def test_restamp_moves_project_grain_and_drops_bootstrap_row(monkeypatch):
cfg, conn = _setup(monkeypatch, default_id="ohm")
projects.restamp_default_project(cfg)
# The project grain (the collection's parent link) re-stamps to 'ohm'.
assert conn.execute("SELECT COUNT(*) c FROM collections WHERE project_id='default'").fetchone()["c"] == 0
assert conn.execute("SELECT project_id FROM collections WHERE id='default'").fetchone()["project_id"] == "ohm"
# Entries stay in their collection — the collection_id is unchanged.
assert conn.execute("SELECT collection_id FROM cached_rfcs WHERE slug='human'").fetchone()["collection_id"] == "default"
assert conn.execute("SELECT collection_id FROM rfc_collaborators WHERE rfc_slug='human'").fetchone()["collection_id"] == "default"
# stale bootstrap projects row removed; 'ohm' remains
assert conn.execute("SELECT 1 FROM projects WHERE id='default'").fetchone() is None
assert conn.execute("SELECT 1 FROM projects WHERE id='ohm'").fetchone() is not None
# FK integrity intact after the rename
assert conn.execute("PRAGMA foreign_key_check").fetchall() == []
def test_restamp_is_idempotent(monkeypatch):
cfg, conn = _setup(monkeypatch, default_id="ohm")
projects.restamp_default_project(cfg)
projects.restamp_default_project(cfg) # second call: no bootstrap rows left → no-op
assert conn.execute("SELECT project_id FROM collections WHERE id='default'").fetchone()["project_id"] == "ohm"
assert conn.execute("SELECT COUNT(*) c FROM cached_rfcs WHERE collection_id='default'").fetchone()["c"] == 1
def test_restamp_noop_when_default_id_unchanged(monkeypatch):
cfg, conn = _setup(monkeypatch, default_id="") # resolves to 'default'
projects.restamp_default_project(cfg)
# nothing renamed; the default collection still belongs to the bootstrap project
assert conn.execute("SELECT project_id FROM collections WHERE id='default'").fetchone()["project_id"] == "default"
+13 -9
View File
@@ -57,12 +57,15 @@ def test_rfc_owner_can_retire_and_entry_leaves_every_surface(app_with_fake_gitea
assert r.status_code == 200, r.text
assert r.json()["state"] == "retired"
# Meta entry on main: state retired, body + fields kept.
meta = entry_mod.parse(
fake.files[("wiggleverse", "meta", "main", "rfcs/ohm.md")]["content"]
# §22.4a SLICE-4: the state flip lands in the metadata sidecar and the
# `.md` is lazy-migrated to a clean body-only file (INV-2). Fields kept.
import yaml as _yaml
sc = _yaml.safe_load(
fake.files[("wiggleverse", "meta", "main", "rfcs/ohm.meta.yaml")]["content"]
)
assert meta.state == "retired"
assert "carol" in meta.owners
assert sc["state"] == "retired"
assert "carol" in sc["owners"]
assert "---" not in fake.files[("wiggleverse", "meta", "main", "rfcs/ohm.md")]["content"]
# Cache flipped; gone from the catalog.
cached = db.conn().execute(
@@ -177,11 +180,12 @@ def test_site_owner_can_retire_active_and_unretire_restores_active_with_id(app_w
assert r.status_code == 200, r.text
assert r.json()["state"] == "active"
meta = entry_mod.parse(
fake.files[("wiggleverse", "meta", "main", "rfcs/ohm.md")]["content"]
import yaml as _yaml
sc = _yaml.safe_load(
fake.files[("wiggleverse", "meta", "main", "rfcs/ohm.meta.yaml")]["content"]
)
assert meta.state == "active"
assert meta.id == "RFC-0042"
assert sc["state"] == "active"
assert sc["id"] == "RFC-0042"
cached = db.conn().execute(
"SELECT state, rfc_id FROM cached_rfcs WHERE slug = 'ohm'"
@@ -0,0 +1,78 @@
"""@S1 acceptance — the collection grain exists (invisible default) and N=1 is
unchanged.
Part C scenarios C3.7 (single-collection project skips the directory) and C3.8
(single-project deployment skips the directory) are the client-side redirect
contract asserted in the frontend; this module asserts the backend N=1
invariants behind the slice: every entry keys on a real collection_id, the
shipped project-scoped serving still resolves through the default collection,
and the legacy /rfc/<slug> URL 308-redirects through /c/<default>/.
Binding: docs/design/2026-06-05-three-tier-projects-collections.md §A.6 / Part E.
"""
from __future__ import annotations
from fastapi.testclient import TestClient
from test_propose_vertical import ( # noqa: F401
app_with_fake_gitea, tmp_env, provision_user_row, sign_in_as,
)
def _seed_entry(slug, title, collection_id="default", state="active"):
from app import db
db.conn().execute(
"INSERT INTO cached_rfcs (slug, title, state, collection_id) VALUES (?, ?, ?, ?)",
(slug, title, state, collection_id),
)
def test_s1_migration_seeds_one_default_collection_for_the_default_project(app_with_fake_gitea):
from app import db
app, _ = app_with_fake_gitea
with TestClient(app):
row = db.conn().execute(
"SELECT id FROM collections WHERE project_id = 'default'"
).fetchall()
assert len(row) == 1
assert row[0]["id"] == "default"
def test_s1_entry_served_under_default_collection(app_with_fake_gitea):
"""N=1 unchanged: an entry is keyed by collection_id under the hood and the
shipped project-scoped serving endpoint still resolves it."""
from app import db
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_entry("human", "Human")
# the row carries a real collection grain (the default collection)
cid = db.conn().execute(
"SELECT collection_id FROM cached_rfcs WHERE slug='human'"
).fetchone()["collection_id"]
assert cid == "default"
# project-scoped serving (collection = default) still returns it
r = client.get("/api/projects/default/rfcs/human")
assert r.status_code == 200, r.text
assert r.json()["slug"] == "human"
# and it appears in the project catalog
slugs = [i["slug"] for i in client.get("/api/projects/default/rfcs").json()["items"]]
assert "human" in slugs
def test_s1_legacy_rfc_url_redirects_through_collection(app_with_fake_gitea):
"""The shipped /rfc/<slug> now 308s through the default collection segment."""
app, _ = app_with_fake_gitea
with TestClient(app) as client:
r = client.get("/rfc/human", follow_redirects=False)
assert r.status_code == 308
assert r.headers["location"] == "/p/default/c/default/e/human"
def test_s1_deployment_reports_single_project(app_with_fake_gitea):
"""C3.8 precondition: the N=1 deployment reports exactly one visible project
and its default id (the frontend uses this to skip the directory)."""
app, _ = app_with_fake_gitea
with TestClient(app) as client:
body = client.get("/api/deployment").json()
assert body["default_project_id"] == "default"
assert [p["id"] for p in body["projects"]] == ["default"]
@@ -0,0 +1,287 @@
"""Slice S3 — scope-role enforcement + collection-grain visibility (@S3).
The acceptance gate for S3 is "every Part C.1 scenario passes" (the design doc
docs/design/2026-06-05-three-tier-projects-collections.md, §C.1, tagged @S3) plus
the operator's S3 visibility requirements (a collection settable public/hidden;
hidden = visible to project/global scope contributors but not the public; a
collection's visibility may be set only as strict or stricter than its project).
The §B.2 resolver folds four layers global project collection per-entry
most-permissively, with no negative override. The scenarios below are exercised
directly against the resolver/gate helpers, and the visibility ones additionally
through the HTTP surface.
Background (C.1): a deployment with a project "ohm" owning collections "model"
(document) and "features" (bdd); a second project "acme" with collection
"specs". Plus a hidden ("gated") collection "secret" under ohm for the
hidden-from-public scenarios.
"""
from __future__ import annotations
from fastapi.testclient import TestClient
from test_propose_vertical import ( # noqa: F401 — fixtures land via import
app_with_fake_gitea,
provision_user_row,
sign_in_as,
tmp_env,
)
# ---------------------------------------------------------------------------
# Helpers
# ---------------------------------------------------------------------------
def _su(user_id: int, login: str, role: str = "contributor", *, state: str = "granted"):
from app import auth
return auth.SessionUser(
user_id=user_id, gitea_id=user_id, gitea_login=login,
display_name=login.capitalize(), email=f"{login}@test", avatar_url="",
role=role, permission_state=state,
)
def _project(pid: str, visibility: str = "public", content_repo: str = "meta") -> None:
from app import db
db.conn().execute(
"INSERT OR REPLACE INTO projects (id, name, content_repo, visibility, updated_at) "
"VALUES (?, ?, ?, ?, datetime('now'))",
(pid, pid.capitalize(), content_repo, visibility),
)
def _collection(cid: str, project_id: str, *, ctype: str = "document",
visibility: str = "public", subfolder: str | None = None) -> None:
from app import db
db.conn().execute(
"INSERT OR REPLACE INTO collections "
"(id, project_id, type, subfolder, initial_state, visibility, name, created_at, updated_at) "
"VALUES (?, ?, ?, ?, 'super-draft', ?, ?, datetime('now'), datetime('now'))",
(cid, project_id, ctype, subfolder if subfolder is not None else cid,
visibility, cid.capitalize()),
)
def _grant(scope_type: str, scope_id: str, user_id: int, role: str) -> None:
from app import db
db.conn().execute(
"INSERT OR REPLACE INTO memberships (scope_type, scope_id, user_id, role) "
"VALUES (?, ?, ?, ?)",
(scope_type, scope_id, user_id, role),
)
def _seed_world() -> None:
"""The C.1 background plus a hidden collection and a second project."""
_project("ohm", "public")
_collection("ohm", "ohm", subfolder="") # ohm's structural default
_collection("model", "ohm", ctype="document")
_collection("features", "ohm", ctype="bdd")
_collection("secret", "ohm", visibility="gated") # hidden from public
_project("acme", "public")
_collection("specs", "acme", ctype="specification")
# the cast
for uid, login in [(1, "ada"), (2, "ben"), (3, "cleo"), (4, "dan"),
(5, "eve"), (6, "fay"), (7, "gil"), (8, "hana")]:
provision_user_row(user_id=uid, login=login, role="contributor")
_grant("collection", "model", 1, "contributor") # ada
_grant("project", "ohm", 2, "contributor") # ben
_grant("global", "*", 3, "contributor") # cleo
_grant("collection", "features", 4, "owner") # dan
_grant("project", "ohm", 5, "owner") # eve
_grant("collection", "model", 6, "contributor") # fay (+ project owner below)
_grant("project", "ohm", 6, "owner") # fay
_grant("project", "ohm", 7, "contributor") # gil
# hana (8): no grant.
# ---------------------------------------------------------------------------
# C.1 — role usage: inheritance and the most-permissive union
# ---------------------------------------------------------------------------
def test_c1_1_collection_contributor_proposes_only_in_that_collection(app_with_fake_gitea):
from app import auth
app, _ = app_with_fake_gitea
with TestClient(app):
_seed_world()
ada = _su(1, "ada")
# may submit a new entry in ohm/model
assert auth.can_contribute_in_collection(ada, "model") is True
# ohm/features is read-only and propose is not offered
assert auth.can_read_collection(ada, "features") is True
assert auth.can_contribute_in_collection(ada, "features") is False
def test_c1_2_project_contributor_proposes_in_every_collection(app_with_fake_gitea):
from app import auth
app, _ = app_with_fake_gitea
with TestClient(app):
_seed_world()
ben = _su(2, "ben")
assert auth.can_contribute_in_collection(ben, "model") is True
assert auth.can_contribute_in_collection(ben, "features") is True
# a collection added later is writable with no new grant
_collection("roadmap", "ohm", ctype="document")
assert auth.can_contribute_in_collection(ben, "roadmap") is True
def test_c1_3_global_contributor_proposes_everywhere(app_with_fake_gitea):
from app import auth
app, _ = app_with_fake_gitea
with TestClient(app):
_seed_world()
cleo = _su(3, "cleo")
assert auth.can_contribute_in_collection(cleo, "model") is True
assert auth.can_contribute_in_collection(cleo, "specs") is True # acme
def test_c1_4_collection_owner_administers_one_collection_only(app_with_fake_gitea):
from app import auth
app, _ = app_with_fake_gitea
with TestClient(app):
_seed_world()
dan = _su(4, "dan")
# graduate / mark-reviewed / manage membership in ohm/features
assert auth.is_collection_superuser(dan, "features") is True
# but not change ohm project settings
assert auth.is_project_superuser(dan, "ohm") is False
assert auth.can_create_collection(dan, "ohm") is False
# and not act on entries in ohm/model
assert auth.is_collection_superuser(dan, "model") is False
assert auth.can_contribute_in_collection(dan, "model") is False
def test_c1_5_project_owner_administers_all_collections_and_creates_more(app_with_fake_gitea):
from app import auth
app, _ = app_with_fake_gitea
with TestClient(app):
_seed_world()
eve = _su(5, "eve")
assert auth.is_collection_superuser(eve, "model") is True
assert auth.is_collection_superuser(eve, "features") is True
assert auth.is_project_superuser(eve, "ohm") is True # edit project settings
assert auth.can_create_collection(eve, "ohm") is True # create a new collection
def test_c1_6_most_permissive_union_higher_grant_wins(app_with_fake_gitea):
from app import auth
app, _ = app_with_fake_gitea
with TestClient(app):
_seed_world()
fay = _su(6, "fay")
# collection RFC Contributor at model + project Owner at ohm → acts as Owner in model
assert auth.effective_scope_role(fay, "model") == "owner"
assert auth.is_collection_superuser(fay, "model") is True
def test_c1_7_no_negative_override(app_with_fake_gitea):
from app import auth, db
app, _ = app_with_fake_gitea
with TestClient(app):
_seed_world()
gil = _su(7, "gil")
# gil can propose in ohm/model via the project grant…
assert auth.can_contribute_in_collection(gil, "model") is True
# …and there is no collection-scope row to remove at model while keeping
# the project grant (a child cannot subtract a parent grant).
row = db.conn().execute(
"SELECT 1 FROM memberships WHERE user_id = 7 AND scope_type = 'collection' AND scope_id = 'model'"
).fetchone()
assert row is None
def test_c1_8_granted_account_no_role_sees_only_public(app_with_fake_gitea):
from app import auth
app, _ = app_with_fake_gitea
with TestClient(app):
_seed_world()
hana = _su(8, "hana")
# may read public collections
assert auth.can_read_collection(hana, "model") is True
# but is not offered the propose action anywhere (no scope role; the
# grandfathered baseline covers only the N=1 `default` collection)
assert auth.can_contribute_in_collection(hana, "model") is False
assert auth.can_contribute_in_collection(hana, "features") is False
assert auth.can_contribute_in_collection(hana, "specs") is False
# gated (hidden) collections do not appear for her
assert auth.can_read_collection(hana, "secret") is False
# ---------------------------------------------------------------------------
# Collection-grain visibility — the operator's S3 requirements
# ---------------------------------------------------------------------------
def test_hidden_collection_invisible_to_public_visible_to_scope_holder(app_with_fake_gitea):
"""A gated collection is omitted from the directory and 404s on read for the
public, yet is listed + readable for a scope-role contributor."""
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
# anonymous: the gated 'secret' collection is not listed, and 404s.
listed = {c["id"] for c in client.get("/api/projects/ohm/collections").json()["items"]}
assert "secret" not in listed
assert "model" in listed # public ones still listed
assert client.get("/api/projects/ohm/collections/secret").status_code == 404
assert client.get("/api/projects/ohm/collections/secret/rfcs").status_code == 404
# ben (project contributor) sees and reads it.
sign_in_as(client, user_id=2, gitea_login="ben", display_name="Ben", role="contributor")
listed2 = {c["id"] for c in client.get("/api/projects/ohm/collections").json()["items"]}
assert "secret" in listed2
assert client.get("/api/projects/ohm/collections/secret").status_code == 200
assert client.get("/api/projects/ohm/collections/secret/rfcs").status_code == 200
# hana (granted, no role) is back to the public view.
sign_in_as(client, user_id=8, gitea_login="hana", display_name="Hana", role="contributor")
listed3 = {c["id"] for c in client.get("/api/projects/ohm/collections").json()["items"]}
assert "secret" not in listed3
assert client.get("/api/projects/ohm/collections/secret").status_code == 404
def test_collection_visibility_strictness_validated_at_create(app_with_fake_gitea):
"""A collection may be created only as strict or stricter than its project;
a looser request is refused (422). On a gated project, a 'public' collection
is rejected."""
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_project("locked", "gated")
_collection("locked", "locked", subfolder="", visibility="gated")
# eve is a deployment owner here to clear the create-authority gate;
# the strictness check fires regardless.
provision_user_row(user_id=9, login="root", role="owner")
sign_in_as(client, user_id=9, gitea_login="root", display_name="Root", role="owner")
r = client.post("/api/projects/locked/collections", json={
"collection_id": "wideopen", "type": "document", "visibility": "public",
})
assert r.status_code == 422, r.text
assert "looser" in r.json()["detail"]
def test_create_collection_allowed_for_project_owner_not_plain_contributor(app_with_fake_gitea):
"""§B.1: a project-scope Owner may create a collection; a plain granted
contributor with no project/global grant may not (403)."""
app, fake = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
# gil is only a project *contributor* on ohm — per §B.1 a project-scope
# contributor CAN create collections (the project-level create
# affordance). A collection-scope grant cannot.
sign_in_as(client, user_id=7, gitea_login="gil", display_name="Gil", role="contributor")
r_ok = client.post("/api/projects/ohm/collections", json={
"collection_id": "fromgil", "type": "document", "visibility": "public",
})
assert r_ok.status_code in (200, 502), r_ok.text # past the authz gate
# ada holds only a *collection*-scope grant (at model) — no create right.
sign_in_as(client, user_id=1, gitea_login="ada", display_name="Ada", role="contributor")
r_no = client.post("/api/projects/ohm/collections", json={
"collection_id": "fromada", "type": "document", "visibility": "public",
})
assert r_no.status_code == 403, r_no.text
@@ -0,0 +1,331 @@
"""Slice S4 — invitation surfaces + role-aware empty states (@S4).
The acceptance gate for S4 is "every Part C.2 invitation scenario passes" (the
design doc docs/design/2026-06-05-three-tier-projects-collections.md, §C.2,
tagged @S4), plus the capability flags that drive the C.3 (@S4) role-aware
empty states.
An Owner grants {owner, contributor} at a scope their reach covers the
project, or a single collection within it to an existing account looked up by
email. The grant writes a `memberships` row immediately and §15-notifies the
grantee (no accept round-trip). Reach is bounded by the inviter's Owner reach;
re-granting at a broader scope supersedes the narrower row; a `pending`
deployment account's grant is recorded but confers no write.
Background (C.2): project "ohm" owns collections "model" (document) and
"features" (bdd). eve is project Owner of ohm; dan is collection Owner of
features only; ben is a project Contributor of ohm. ivy / jo / jet are
grantees; kim is a pending deployment account.
"""
from __future__ import annotations
from fastapi.testclient import TestClient
from test_propose_vertical import ( # noqa: F401 — fixtures land via import
app_with_fake_gitea,
provision_user_row,
sign_in_as,
tmp_env,
)
# ---------------------------------------------------------------------------
# Helpers (mirror the S3 vertical's world-builders)
# ---------------------------------------------------------------------------
def _su(user_id: int, login: str, role: str = "contributor", *, state: str = "granted"):
from app import auth
return auth.SessionUser(
user_id=user_id, gitea_id=user_id, gitea_login=login,
display_name=login.capitalize(), email=f"{login}@test", avatar_url="",
role=role, permission_state=state,
)
def _project(pid: str, visibility: str = "public", content_repo: str = "meta") -> None:
from app import db
db.conn().execute(
"INSERT OR REPLACE INTO projects (id, name, content_repo, visibility, updated_at) "
"VALUES (?, ?, ?, ?, datetime('now'))",
(pid, pid.capitalize(), content_repo, visibility),
)
def _collection(cid: str, project_id: str, *, ctype: str = "document",
visibility: str = "public", subfolder: str | None = None) -> None:
from app import db
db.conn().execute(
"INSERT OR REPLACE INTO collections "
"(id, project_id, type, subfolder, initial_state, visibility, name, created_at, updated_at) "
"VALUES (?, ?, ?, ?, 'super-draft', ?, ?, datetime('now'), datetime('now'))",
(cid, project_id, ctype, subfolder if subfolder is not None else cid,
visibility, cid.capitalize()),
)
def _grant(scope_type: str, scope_id: str, user_id: int, role: str) -> None:
from app import db
db.conn().execute(
"INSERT OR REPLACE INTO memberships (scope_type, scope_id, user_id, role) "
"VALUES (?, ?, ?, ?)",
(scope_type, scope_id, user_id, role),
)
def _membership(user_id: int):
"""The set of (scope_type, scope_id, role) rows a user holds."""
from app import db
rows = db.conn().execute(
"SELECT scope_type, scope_id, role FROM memberships WHERE user_id = ?",
(user_id,),
).fetchall()
return {(r["scope_type"], r["scope_id"], r["role"]) for r in rows}
def _seed_world() -> None:
_project("ohm", "public")
_collection("model", "ohm", ctype="document")
_collection("features", "ohm", ctype="bdd")
# the cast
provision_user_row(user_id=2, login="ben", role="contributor")
provision_user_row(user_id=4, login="dan", role="contributor")
provision_user_row(user_id=5, login="eve", role="contributor")
provision_user_row(user_id=10, login="ivy", role="contributor")
provision_user_row(user_id=11, login="jo", role="contributor")
provision_user_row(user_id=12, login="jet", role="contributor")
provision_user_row(user_id=13, login="kim", role="contributor")
_grant("project", "ohm", 5, "owner") # eve — project Owner
_grant("collection", "features", 4, "owner") # dan — collection Owner only
_grant("project", "ohm", 2, "contributor") # ben — project Contributor
# kim is a pending deployment account.
from app import db
db.conn().execute(
"UPDATE users SET permission_state = 'pending' WHERE id = 13"
)
def _login_eve(client) -> None:
sign_in_as(client, user_id=5, gitea_login="eve", display_name="Eve", role="contributor")
# ---------------------------------------------------------------------------
# C.2 — invitation: who may invite whom, at which scope
# ---------------------------------------------------------------------------
def test_c2_1_project_owner_invites_at_project_scope(app_with_fake_gitea):
"""A project Owner grants at project scope; the grant covers every
collection, and the grantee is §15-notified naming the project and role."""
from app import auth
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
_login_eve(client)
r = client.post("/api/projects/ohm/members",
json={"email": "ivy@test", "role": "contributor"})
assert r.status_code == 200, r.text
# a membership row is written at scope project "ohm"
assert ("project", "ohm", "contributor") in _membership(10)
# ivy may propose in every collection of "ohm"
ivy = _su(10, "ivy")
assert auth.can_contribute_in_collection(ivy, "model") is True
assert auth.can_contribute_in_collection(ivy, "features") is True
# ivy receives a §15 notification naming the project and role
sign_in_as(client, user_id=10, gitea_login="ivy", display_name="Ivy", role="contributor")
inbox = client.get("/api/notifications").json()["items"]
granted = [n for n in inbox if n["event_kind"] == "scope_role_granted"]
assert granted, inbox
assert "Ohm" in granted[0]["summary"]
assert "RFC Contributor" in granted[0]["summary"]
def test_c2_2_owner_invites_at_specific_collection(app_with_fake_gitea):
"""A grant at a single collection scope reaches that collection only."""
from app import auth
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
_login_eve(client)
r = client.post("/api/projects/ohm/members",
json={"email": "jo@test", "role": "contributor",
"collection_id": "features"})
assert r.status_code == 200, r.text
assert ("collection", "features", "contributor") in _membership(11)
jo = _su(11, "jo")
assert auth.can_contribute_in_collection(jo, "features") is True
assert auth.can_contribute_in_collection(jo, "model") is False
def test_c2_3_invitation_reach_bounded_by_inviter_scope(app_with_fake_gitea):
"""A collection Owner may invite within that collection, but is not offered
(is refused) the control to invite at the project or globally."""
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
# dan is Owner at collection "features" only.
sign_in_as(client, user_id=4, gitea_login="dan", display_name="Dan", role="contributor")
# may grant at his collection
ok = client.post("/api/projects/ohm/members",
json={"email": "ivy@test", "role": "contributor",
"collection_id": "features"})
assert ok.status_code == 200, ok.text
# but not at the project scope
no = client.post("/api/projects/ohm/members",
json={"email": "ivy@test", "role": "contributor"})
assert no.status_code == 403, no.text
# the capability flags the UI reads agree: no project invite, yes collection
proj = client.get("/api/projects/ohm/collections").json()["viewer"]
assert proj["can_invite"] is False
col = client.get("/api/projects/ohm/collections/features").json()["viewer"]
assert col["can_invite"] is True
def test_c2_4_contributors_do_not_manage_membership(app_with_fake_gitea):
"""An RFC Contributor (project- or collection-scoped) holds no invite
capability and the grant endpoints refuse them."""
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
# ben is a project *Contributor* on ohm.
sign_in_as(client, user_id=2, gitea_login="ben", display_name="Ben", role="contributor")
no_proj = client.post("/api/projects/ohm/members",
json={"email": "ivy@test", "role": "contributor"})
assert no_proj.status_code == 403, no_proj.text
no_col = client.post("/api/projects/ohm/members",
json={"email": "ivy@test", "role": "contributor",
"collection_id": "features"})
assert no_col.status_code == 403, no_col.text
# no invite control surfaced anywhere
assert client.get("/api/projects/ohm/collections").json()["viewer"]["can_invite"] is False
assert client.get("/api/projects/ohm/collections/features").json()["viewer"]["can_invite"] is False
def test_c2_5_no_grant_at_parent_revoke_at_child_option(app_with_fake_gitea):
"""The grant surface offers only role + scope (project or one collection);
there is no way to grant at the project yet carve out a child collection
a project grant reaches every collection, full stop."""
from app import auth
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
_login_eve(client)
# the only knobs are role and an optional single collection_id; an
# "exclude" field has no effect (it is not part of the contract).
r = client.post("/api/projects/ohm/members",
json={"email": "ivy@test", "role": "contributor",
"exclude_collection_id": "features"})
assert r.status_code == 200, r.text
# the project grant still reaches the supposedly-excluded collection
ivy = _su(10, "ivy")
assert auth.can_contribute_in_collection(ivy, "features") is True
def test_c2_6_broader_scope_supersedes_narrower(app_with_fake_gitea):
"""Re-granting at a broader scope removes the subsumed narrower row; the
grantee holds the role across the whole project."""
from app import auth
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
_grant("collection", "features", 12, "contributor") # jet starts narrow
_login_eve(client)
r = client.post("/api/projects/ohm/members",
json={"email": "jet@test", "role": "contributor"})
assert r.status_code == 200, r.text
rows = _membership(12)
# the project grant is present…
assert ("project", "ohm", "contributor") in rows
# …and the redundant collection-scope row is gone (subsumed)
assert ("collection", "features", "contributor") not in rows
jet = _su(12, "jet")
assert auth.can_contribute_in_collection(jet, "model") is True
assert auth.can_contribute_in_collection(jet, "features") is True
def test_c2_6b_narrower_stronger_role_is_not_subtracted(app_with_fake_gitea):
"""No negative override: a child Owner grant survives a parent Contributor
grant (the stronger collection role is kept)."""
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
_grant("collection", "features", 12, "owner") # jet is collection Owner
_login_eve(client)
r = client.post("/api/projects/ohm/members",
json={"email": "jet@test", "role": "contributor"})
assert r.status_code == 200, r.text
rows = _membership(12)
assert ("project", "ohm", "contributor") in rows
# the stronger collection-Owner row is NOT pruned by a weaker project grant
assert ("collection", "features", "owner") in rows
def test_c2_7_pending_account_grant_confers_no_write(app_with_fake_gitea):
"""A grant to a pending deployment account is recorded but confers no write
until the account is granted at the deployment (§6)."""
from app import auth
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
_login_eve(client)
r = client.post("/api/projects/ohm/members",
json={"email": "kim@test", "role": "contributor"})
assert r.status_code == 200, r.text
assert r.json()["pending"] is True
# the grant row is recorded…
assert ("project", "ohm", "contributor") in _membership(13)
# …but confers no write while pending (the §6 admission floor)
kim = _su(13, "kim", state="pending")
assert auth.effective_scope_role(kim, "model") is None
assert auth.can_contribute_in_collection(kim, "model") is False
# ---------------------------------------------------------------------------
# C.3 (@S4) — the capability flags behind the role-aware empty states
# ---------------------------------------------------------------------------
def test_c3_3_project_owner_sees_create_first_collection_capability(app_with_fake_gitea):
"""C3.3: a project Owner landing on an empty project may create a
collection the flag the 'Create your first collection' CTA reads."""
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
_login_eve(client)
caps = client.get("/api/projects/ohm/collections").json()["viewer"]
assert caps["can_create_collection"] is True
assert caps["role"] == "owner"
def test_c3_4_contributor_without_create_rights_has_no_create_capability(app_with_fake_gitea):
"""C3.4: a contributor whose only grant is at a collection elsewhere has no
create-collection capability the empty directory shows no create action."""
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
# dan holds only a collection-scope grant (features); no project create right.
sign_in_as(client, user_id=4, gitea_login="dan", display_name="Dan", role="contributor")
caps = client.get("/api/projects/ohm/collections").json()["viewer"]
assert caps["can_create_collection"] is False
def test_c3_5_collection_contributor_sees_propose_first_capability(app_with_fake_gitea):
"""C3.5: a collection contributor landing on an empty collection may propose
the flag the 'Propose the first entry' CTA reads; an anonymous reader may
not (the sign-in prompt path, already shipped in S2)."""
app, _ = app_with_fake_gitea
with TestClient(app) as client:
_seed_world()
_grant("collection", "model", 10, "contributor") # ivy contributes in model
sign_in_as(client, user_id=10, gitea_login="ivy", display_name="Ivy", role="contributor")
caps = client.get("/api/projects/ohm/collections/model").json()["viewer"]
assert caps["can_contribute"] is True
# anonymous reader: no propose capability
client.cookies.clear()
caps_anon = client.get("/api/projects/ohm/collections/model").json()["viewer"]
assert caps_anon["can_contribute"] is False
@@ -0,0 +1,141 @@
"""§22.12 S6 — per-collection model universe.
A collection's `.collection.yaml` may carry an `enabled_models` list that
NARROWS its project's universe, which in turn narrows the deployment
ENABLED_MODELS. The resolution chain (extending §6.6/§6.7) is:
funder per-entry models collection universe project universe
(operator providers = the ceiling)
These tests prove (1) the manifest parser reads `enabled_models`, (2) the
mirror stores it on the collection row, (3) the resolver narrows the base
universe by project then collection, with absent = inherit and [] = opt-out,
and (4) the collection API surfaces the collection's own narrowing.
"""
from __future__ import annotations
import json
from fastapi.testclient import TestClient
from app import db, models_resolver, registry
from test_propose_vertical import ( # noqa: F401
app_with_fake_gitea, provision_user_row, sign_in_as, tmp_env,
)
from test_rfc_view_vertical import FakeProvider, seed_active_rfc # noqa: F401
def _install_two_providers(app) -> None:
app.state.providers.clear()
app.state.providers["claude"] = FakeProvider("TITLE: A\nDESCRIPTION: B")
app.state.providers["gemini"] = FakeProvider("TITLE: G\nDESCRIPTION: H")
def _set_project_models(project_id: str, models) -> None:
cfg = {} if models is None else {"enabled_models": models}
db.conn().execute(
"UPDATE projects SET config_json = ? WHERE id = ?",
(json.dumps(cfg), project_id),
)
def _set_collection_models(collection_id: str, models) -> None:
cfg = None if models is None else json.dumps({"enabled_models": models})
db.conn().execute(
"UPDATE collections SET config_json = ? WHERE id = ?",
(cfg, collection_id),
)
# ── manifest parsing ───────────────────────────────────────────────────────
def test_manifest_parses_enabled_models_into_config():
ce = registry.parse_collection_manifest(
"type: bdd\nname: Features\nenabled_models: [claude, gemini]\n"
)
assert ce.config.get("enabled_models") == ["claude", "gemini"]
def test_manifest_without_enabled_models_has_no_key():
ce = registry.parse_collection_manifest("type: document\nname: Model\n")
assert "enabled_models" not in ce.config
def test_manifest_rejects_non_list_enabled_models():
import pytest
with pytest.raises(registry.RegistryError):
registry.parse_collection_manifest("type: bdd\nenabled_models: claude\n")
# ── resolver narrowing (the §22.12 chain) ──────────────────────────────────
def test_resolver_inherits_operator_universe_when_no_scope_narrowing(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app):
seed_active_rfc(fake, slug="ohm", title="OHM", body="x")
_install_two_providers(app)
# No project/collection narrowing → full operator universe.
resolved = models_resolver.resolve_models_for_rfc("ohm", app.state.providers)
assert resolved == ["claude", "gemini"]
def test_resolver_narrows_by_project_universe(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app):
seed_active_rfc(fake, slug="ohm", title="OHM", body="x")
_install_two_providers(app)
_set_project_models("default", ["gemini"])
resolved = models_resolver.resolve_models_for_rfc("ohm", app.state.providers)
assert resolved == ["gemini"]
def test_resolver_collection_narrows_within_project(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app):
seed_active_rfc(fake, slug="ohm", title="OHM", body="x")
_install_two_providers(app)
# Project allows both; the collection narrows to claude only.
_set_project_models("default", ["claude", "gemini"])
_set_collection_models("default", ["claude"])
resolved = models_resolver.resolve_models_for_rfc("ohm", app.state.providers)
assert resolved == ["claude"]
def test_resolver_collection_empty_list_opts_out(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app):
seed_active_rfc(fake, slug="ohm", title="OHM", body="x")
_install_two_providers(app)
_set_collection_models("default", []) # opt this collection out of AI
resolved = models_resolver.resolve_models_for_rfc("ohm", app.state.providers)
assert resolved == []
def test_resolver_collection_cannot_widen_project(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app):
seed_active_rfc(fake, slug="ohm", title="OHM", body="x")
_install_two_providers(app)
# Project restricts to gemini; the collection naming claude+gemini
# cannot re-add claude (narrowing only).
_set_project_models("default", ["gemini"])
_set_collection_models("default", ["claude", "gemini"])
resolved = models_resolver.resolve_models_for_rfc("ohm", app.state.providers)
assert resolved == ["gemini"]
# ── API surfacing ──────────────────────────────────────────────────────────
def test_collection_api_surfaces_enabled_models(app_with_fake_gitea):
app, fake = app_with_fake_gitea
with TestClient(app) as client:
provision_user_row(user_id=1, login="ben", role="owner")
sign_in_as(client, user_id=1, gitea_login="ben", display_name="Ben",
role="owner", email="ben@test")
_set_collection_models("default", ["claude"])
r = client.get("/api/projects/default/collections/default")
assert r.status_code == 200, r.text
assert r.json().get("enabled_models") == ["claude"]
@@ -0,0 +1,48 @@
"""§22.4a S6 — the type-driven entry noun.
The displayed noun for an entry is a framework concept keyed on the
collection's immutable type: document→"RFC", specification→"Spec", bdd→"Feature".
The chrome reads it from the API rather than hardcoding "RFC", so the propose
CTA + entry chrome name entries correctly per collection type with no
per-deployment config.
"""
from __future__ import annotations
from fastapi.testclient import TestClient
from app import collections, db
from test_propose_vertical import ( # noqa: F401
app_with_fake_gitea, provision_user_row, sign_in_as, tmp_env,
)
def test_entry_noun_map():
assert collections.entry_noun("document") == "RFC"
assert collections.entry_noun("specification") == "Spec"
assert collections.entry_noun("bdd") == "Feature"
# An unknown/future type falls back to the generic noun, never label-less.
assert collections.entry_noun("mystery") == "RFC"
def test_collection_api_surfaces_entry_noun(app_with_fake_gitea):
app, _ = app_with_fake_gitea
with TestClient(app) as client:
provision_user_row(user_id=1, login="ben", role="owner")
sign_in_as(client, user_id=1, gitea_login="ben", display_name="Ben",
role="owner", email="ben@test")
# Flip the default collection to a bdd type and confirm the noun follows.
db.conn().execute("UPDATE collections SET type='bdd' WHERE id='default'")
r = client.get("/api/projects/default/collections/default")
assert r.status_code == 200, r.text
assert r.json()["entry_noun"] == "Feature"
def test_directory_items_carry_entry_noun(app_with_fake_gitea):
app, _ = app_with_fake_gitea
with TestClient(app) as client:
db.conn().execute("UPDATE projects SET visibility='public' WHERE id='default'")
db.conn().execute("UPDATE collections SET type='specification' WHERE id='default'")
r = client.get("/api/deployment")
assert r.status_code == 200, r.text
item = next(p for p in r.json()["projects"] if p["id"] == "default")
assert item["entry_noun"] == "Spec"
@@ -0,0 +1,90 @@
"""§22 S6 — two-project / multi-collection integrative pass.
Proves the S6 surface holds across a deployment with two projects, each owning
distinct collections, with the new per-collection knobs (§22.12 enabled_models,
§22.4a type noun) resolving independently per collection no cross-project or
cross-collection bleed.
Slug note: the resolver is slug-keyed, so this test uses distinct slugs across
collections (the documented limitation same-slug-across-collections model
resolution is part of the broader collection-id-threading follow-up).
"""
from __future__ import annotations
import json
from fastapi.testclient import TestClient
from app import collections, db, models_resolver
from test_propose_vertical import app_with_fake_gitea, tmp_env # noqa: F401
from test_rfc_view_vertical import FakeProvider # noqa: F401
def _project(pid: str, visibility: str = "public") -> None:
db.conn().execute(
"INSERT OR REPLACE INTO projects (id, name, content_repo, visibility, config_json, updated_at) "
"VALUES (?, ?, ?, ?, ?, datetime('now'))",
(pid, pid.capitalize(), f"{pid}-content", visibility, None),
)
def _collection(cid: str, project_id: str, *, ctype: str, enabled_models=None) -> None:
cfg = None if enabled_models is None else json.dumps({"enabled_models": enabled_models})
db.conn().execute(
"INSERT OR REPLACE INTO collections "
"(id, project_id, type, subfolder, initial_state, visibility, name, config_json, "
" created_at, updated_at) "
"VALUES (?, ?, ?, ?, 'super-draft', 'public', ?, ?, datetime('now'), datetime('now'))",
(cid, project_id, ctype, cid, cid.capitalize(), cfg),
)
def _entry(slug: str, collection_id: str) -> None:
db.conn().execute(
"INSERT OR REPLACE INTO cached_rfcs (slug, title, state, collection_id) "
"VALUES (?, ?, 'active', ?)",
(slug, slug.upper(), collection_id),
)
def _three_providers(app) -> None:
app.state.providers.clear()
for k in ("claude", "gemini", "gpt"):
app.state.providers[k] = FakeProvider("TITLE: A\nDESCRIPTION: B")
def test_two_projects_multicollection_model_universe_isolated(app_with_fake_gitea):
app, _ = app_with_fake_gitea
with TestClient(app):
_three_providers(app)
# alpha: a document collection narrowed to claude.
_project("alpha")
_collection("alpha-docs", "alpha", ctype="document", enabled_models=["claude"])
_entry("alpha-intro", "alpha-docs")
# beta: two collections — a spec narrowed to gemini, a bdd left open.
_project("beta")
_collection("beta-specs", "beta", ctype="specification", enabled_models=["gemini"])
_collection("beta-features", "beta", ctype="bdd") # no narrowing → all three
_entry("beta-runtime", "beta-specs")
_entry("beta-login", "beta-features")
# Each entry resolves against ITS collection's universe — no bleed.
assert models_resolver.resolve_models_for_rfc("alpha-intro", app.state.providers) == ["claude"]
assert models_resolver.resolve_models_for_rfc("beta-runtime", app.state.providers) == ["gemini"]
assert models_resolver.resolve_models_for_rfc("beta-login", app.state.providers) == [
"claude", "gemini", "gpt"
]
def test_two_projects_multicollection_type_noun_isolated(app_with_fake_gitea):
app, _ = app_with_fake_gitea
with TestClient(app):
_project("alpha")
_collection("alpha-docs", "alpha", ctype="document")
_project("beta")
_collection("beta-specs", "beta", ctype="specification")
_collection("beta-features", "beta", ctype="bdd")
assert collections.get_collection("alpha-docs")["entry_noun"] == "RFC"
assert collections.get_collection("beta-specs")["entry_noun"] == "Spec"
assert collections.get_collection("beta-features")["entry_noun"] == "Feature"
+1 -1
View File
@@ -5,7 +5,7 @@
#
# while IFS='=' read -r k v; do
# [ -n "$k" ] && case "$k" in \#*) ;; *) \
# ohm-rfc-app-flotilla overlay set ohm-rfc-app "$k=$v" --preview ;; esac
# flotilla-core overlay set <deployment> "$k=$v" --preview ;; esac
# done < deploy/preview/preview.env.example
#
# CRITICAL (§15 / §3 invariant 1): a preview resolves ZERO real secret bytes.
+93
View File
@@ -90,6 +90,99 @@ the framework at that pin against your Gitea; the framework knows
nothing about your specific deployment beyond what your `.env`
files told it.
## The registry: projects and collections (§22)
From **v0.45.0** a deployment is **three tiers** — deployment →
**project** → **RFC collection** — instead of one corpus. Most of
the single-corpus setup above still applies (a deployment that wants
one corpus runs the **N=1 case** unchanged), but two things move into
**git-truth the framework mirrors**, exactly the way RFC bodies do.
The binding model is [`SPEC.md` §22](../SPEC.md); this is the operator
view of the file formats.
### The registry repo (`REGISTRY_REPO`) declares projects
`VITE_APP_NAME` and `META_REPO` are superseded. The framework now reads
a **registry repo** — a dedicated repo under your Gitea org, named
whatever you like, whose location you pass in `backend/.env` as
`REGISTRY_REPO` (the framework fails loudly at startup if it is unset).
Its root holds a `projects.yaml`:
```yaml
# projects.yaml (registry repo root)
deployment:
name: Wiggleverse # deployment display name (was VITE_APP_NAME)
tagline: A substrate for collaborative standardization
projects:
- id: ohm # url-stable slug, unique in the deployment → /p/ohm/
name: Open Human Model
content_repo: ohm-content # ONE repo under your org; collections live inside it
visibility: public # gated | public | unlisted (§22.5)
theme: { accent: "#5b5bd6" } # optional per-project token overrides
enabled_models: [claude, gemini] # optional; falls back to ENABLED_MODELS
```
Each project owns **exactly one content repo**. Adding, reconfiguring,
or archiving a project is a PR against `projects.yaml`; the framework's
webhook + reconciler mirror it into the `projects` cache table. Project
**membership** is app state (it churns at user speed), not registry
data. The deployment's display name and tagline come from here, served
at runtime via `GET /api/deployment` — no rebuild needed to rename.
### The content repo declares collections via `.collection.yaml`
A project's collections are **typed subfolders** of its content repo,
each carrying a `.collection.yaml` manifest the registry mirror reads:
```
ohm-content/
model/
.collection.yaml # type: document
rfcs/intro.md
specs/
.collection.yaml # type: specification
rfcs/runtime.md
features/
.collection.yaml # type: bdd
rfcs/login.md
```
```yaml
# ohm-content/model/.collection.yaml
type: document # document | specification | bdd — IMMUTABLE once set
visibility: gated # defaults to the project's; may only NARROW it
initial_state: super-draft # super-draft | active — defaults from type
name: The Model
# enabled_models: [claude] # optional; may only narrow the project's universe
```
`type` is fixed at creation (the framework refuses to change it on a
later mirror). `visibility` may be as strict as or stricter than the
project's, never looser (`public` < `unlisted` < `gated`); reading or
writing a collection requires passing **both** the project and the
collection gate. `enabled_models`, if present, narrows the project's
model universe for that collection.
### Default project + default collection (the N=1 upgrade)
A deployment upgrading from a pre-§22 version is migrated automatically:
its single corpus becomes one **default project** carrying one **default
collection** (`id` `default`, subfolder = repo root), inheriting the old
`type` / `initial_state` / visibility. Old URLs 308-redirect into the
`/p/<project>/c/default/…` form, so existing links survive. Until you
add a second project or collection the deployment is functionally
identical to before, with one extra path segment. The ordered upgrade
actions are the [`CHANGELOG.md`](../CHANGELOG.md) entry's upgrade-steps
block (§20.4).
### Creating projects and collections in-app
Both tiers can also be created from the UI by an Owner — the action
wraps a bot commit (a new content repo + `projects.yaml` entry for a
project; a new subfolder + `.collection.yaml` for a collection), so
everything still flows from git. Nothing becomes app state the mirror
cannot rebuild.
## The two-repo working pattern
Once a deployment is live, day-to-day changes split across the
@@ -0,0 +1,672 @@
# Draft spec — §22 refactor: three tiers (deployment → project → RFC collection)
> Status: **draft for review.** Binding voice, but not yet merged into
> `SPEC.md`. This doc **revises the §22 model** in
> [`multi-project-spec.md`](./multi-project-spec.md) from two tiers
> (deployment → project, where a "project" *is* a corpus) to **three tiers**
> (deployment → project → RFC collection, where the *collection* is the
> corpus). It supersedes the conflicting parts of that draft; the parts it does
> not touch (the registry-is-git-truth stance, the cache mirror, visibility
> semantics, the `type`/`initial_state`/`unreviewed` machinery) carry over
> unchanged, re-homed onto the collection. Rationale and the decisions behind
> this live in [`multi-project.md`](./multi-project.md) and session 0072.
>
> ⚠️ **CORRECTION (session 0072, after code re-check).** Parts of §0/§A.3/§E
> were drafted on a stale-memory premise that "Plan B (migration 028) and
> M3-frontend have not shipped." **That is false.** As of v0.39.0 the entire
> **two-tier** model is shipped to `main`: migration 028 already rebuilt the
> slug PK to `(project_id, slug)`; v0.35.0 shipped `/p/<project>/` routing and
> the live `/p/<project>/e/<slug>` URLs; v0.37.0/0.38.0 shipped per-project
> read + propose. Inserting the third tier is therefore an **evolution of a
> shipped system**, not a revision of unshipped designs. The migration strategy
> (Part E) was **re-decided on these corrected facts** (session 0072): a new
> **migration 029** adds a *collection* grain *beneath* today's project, plus a
> breaking `/p/<project>/e/<slug>``/p/<project>/c/<collection>/e/<slug>` URL
> change with 308s. The structural model (Parts AD) is unaffected. Target
> release: a further pre-1.0 minor with breaking changes + upgrade steps (§20.2).
---
## 0. Why this revision
The original §22 (`multi-project-spec.md`) gave a deployment **N projects**,
where each project *was* a single typed corpus: one content repo, one `type`,
one slug namespace, one member roster. That conflates two responsibilities —
**organizational grouping** and **a typed body of entries** — into one noun.
This revision splits them. A **project** becomes a pure grouping tier (settings
+ one content repo) that holds **any number of RFC collections**; an **RFC
collection** is the typed corpus the original §22 called a "project." Everything
the original §22 said about a corpus (type, slug namespace, catalog, philosophy,
landing state, review flag, membership) moves down one level to the collection;
the deployment level is unchanged.
⚠️ The two-tier model is **already shipped** (v0.39.0): migration 028 rebuilt
the slug PK to `(project_id, slug)`, and `/p/<project>/e/<slug>` URLs are live
(v0.35.0). So inserting the third tier evolves a shipped system — see Part E
for the decided strategy (a new migration 029 adding a collection grain beneath
today's project, + a breaking URL change with 308s).
---
# Part A — The three-tier model
## A.1 The tiers
```
deployment (= "global" in the UI) one Gitea org, one bot, one account
│ system, one inbox, one running process;
│ the surface a visitor first lands on.
└─ project ◀ NEW a named grouping + project settings;
│ owns exactly ONE content repo. No type.
└─ RFC collection a typed corpus: type, slug namespace,
│ catalog, philosophy, initial_state,
│ unreviewed flag, members. (= what the
│ original §22 called a "project".)
└─ entry an RFC / spec / feature, identified by
its slug within the collection.
```
- **Deployment / "global."** Unchanged top tier. Owns accounts, the §6
admission gate, the §15 inbox, the §1 bot, and the deployment landing
directory. Its management surface is **projects + global settings**.
- **Project** *(new)*. Belongs to exactly one deployment; never moves. Owns one
content repo (§A.2) and carries project settings (name, tagline, theme,
visibility, model universe). Has **no `type`** of its own. Its management
surface is **RFC collections + project settings**.
- **RFC collection.** A typed subfolder of its project's content repo (§A.2).
Carries everything the original §22 pinned on a "project": the immutable
`type` (§22.4a `document` | `specification` | `bdd` | …), the per-collection
slug namespace (§A.3), `initial_state` (§22.4b), the `unreviewed` flag
(§22.4c), catalog, philosophy. This is "closest to what OHM originally
managed as a single corpus."
- **Entry.** Unchanged (§2). Identified by its slug **within its collection**.
A collection belongs to exactly one project; a project to exactly one
deployment. Isolation (§22.1) now holds at the **collection** grain: an RFC,
branch, thread, star, or watch belongs to exactly one collection.
## A.2 Storage and git-truth
Two git sources, both read by the bot, both mirrored into cache tables the §4
way:
1. **The registry repo** (`projects.yaml`, located by `REGISTRY_REPO`, §22.2)
declares **projects**`id`, `name`, `content_repo`, settings, `visibility`,
`theme`, `enabled_models`. `content_repo` moves **up** from the collection
(original §22) to the project: a project owns exactly one content repo.
2. **Each project's content repo** declares its **collections** as typed
subfolders, each carrying a **`.collection.yaml` manifest** (the collection's
`type`, `visibility`, `initial_state`). The registry mirror walks the content
repo and reads these manifests, so collection configuration is git-truth and
survives a cache rebuild — exactly as entry frontmatter does.
```yaml
# projects.yaml (registry repo root)
deployment:
name: Wiggleverse
tagline: ...
projects:
- id: ohm
name: Open Human Model
content_repo: ohm-content # ONE repo; collections live inside it
visibility: public # gated | public | unlisted (§22.5)
theme: { accent: "#5b5bd6" }
enabled_models: [claude, gemini]
```
```yaml
# ohm-content/features/.collection.yaml (one per collection subfolder)
type: bdd # document | specification | bdd — immutable
visibility: gated # defaults to the project's, may narrow
initial_state: active # defaults from type (§22.4b)
name: Feature scenarios
```
```
ohm-content/
model/
.collection.yaml # type: document
intro.md
specs/
.collection.yaml # type: specification
runtime.md
features/
.collection.yaml # type: bdd
login.md
```
**Creation is in-app, wrapping a bot commit, at both tiers:**
- **+ New project** (a global Owner action): the bot **creates a Gitea content
repo** under the deployment org, **commits a project entry** to
`projects.yaml`, and the mirror picks it up.
- **+ New collection** (a project Owner / RFC Contributor-with-create action):
the bot **commits a new subfolder + `.collection.yaml`** to the project's
content repo; the mirror picks it up.
The in-app button is a thin convenience over a git write; nothing becomes app
state that git cannot rebuild. `projects` and `collections` cache rows are never
written from user actions directly — they flow from the mirror only (§22.2).
**Membership** (§B-roles) remains app state, as `rfc_collaborators` always has
been — it churns at user speed and is not document state.
## A.3 Identity and routing
The slug is unique **within a collection**; the fully-qualified identity is
`(project, collection, slug)`. `model/intro` and `specs/intro` coexist. No type
prefix, no numbers (the §22.4 retirement of `RFC-NNNN` allocation stands;
legacy `id` frontmatter remains a frozen, non-identity display label).
Canonical route:
```
/p/<project>/c/<collection>/e/<slug>
```
The `c/` segment keeps collection ids from colliding with reserved
project-level segments (project settings, the collection directory). Reserved
**collection-level** siblings (`proposals`, `philosophy`) sit under
`/p/<project>/c/<collection>/…`. The displayed entry noun ("RFC", "Spec",
"Feature") is the collection type's label (§22.4a), not part of the path.
The root `/` is the deployment landing: a **directory of projects** the visitor
can see (§22.5). `/p/<project>/` is the project landing: a **directory of
collections** in that project the visitor can see. Conveniences:
- `/p/<project>/` redirects to its sole collection when the project has exactly
one visible collection.
- `/` redirects to the sole visible project when there is exactly one (the N=1
case, §A.6).
⚠️ **Backcompat is heavier than first drafted.** `/p/<project>/e/<slug>` URLs
**are live** (v0.35.0), so adding the `/c/<collection>/` segment is a breaking
URL change: the shipped `/p/<project>/e/<slug>` must **308-redirect** to
`/p/<project>/c/<default-collection>/e/<slug>`, alongside the pre-multi-project
`/rfc/<slug>``/p/<default-project>/c/<default-collection>/…` redirect. Both
are handled in the migration (§A.6 / Part E).
---
# Part B — Roles and authorization
## B.1 One role vocabulary, attached at a scope
There is **one role enum — `{owner, contributor}`** — displayed as **Owner**
and **RFC Contributor**. A grant *attaches that role at a scope*: **global**,
**project**, or **collection**. "Owner at all levels, RFC Contributor at all
levels" is therefore literal — the same two words at every tier, not a fresh
pair invented per tier.
| Role | Capabilities within its scope's subtree |
|---|---|
| **Owner** | Superuser: manage settings and membership; create child projects/collections; act on any entry (merge on behalf, graduate, mark-reviewed, withdraw/reopen, set branch visibility). |
| **RFC Contributor** | Propose entries, create branches, open PRs, claim unclaimed super-drafts, participate in discussion. At **project** (or global) scope this additionally includes **creating collections** in that project — the "anyone at the project level with permission to create a collection" affordance. (A *collection*-scope grant cannot create sibling collections; creating one is a project-level action.) |
This **reconciles** the role names the prior drafts accumulated — they were
different words for the same idea:
| Prior spec term | Tier it lived at | Unified role |
|---|---|---|
| deployment `owner` / `admin` (§6.1) | global | **Owner** (global) |
| deployment `contributor` (§6.1) | global | **RFC Contributor** (global) |
| `project_admin` (M2 §22.6) | the corpus → now the **collection** | **Owner** (collection) |
| `project_contributor` (M2 §22.6) | the corpus → now the **collection** | **RFC Contributor** (collection) |
| `project_viewer` (M2 §22.6) | the corpus | *deferred* (read-only grant; not one of this pass's two) |
> **Scope-narrowing, not renaming.** Collapsing `owner`/`admin` into one
> **Owner** and dropping `viewer` for this pass are deliberate deferrals (the
> launch ask: "we don't need to get all permissions right yet"). When they
> return they **re-split out of** Owner / add a tier; they are not aliases of
> the unified roles. The richer set is future work.
## B.2 Inheritance and resolution
Grants inherit **downward**, are **additive**, and admit **no negative
override**:
- A grant at **global** covers every project and collection in the deployment.
- A grant at **project** covers every collection in that project.
- A grant at **collection** covers just that collection.
- You **cannot** grant a role at a parent scope and revoke it at a child (the
launch ask: "too complex"). Resolution never subtracts a parent grant.
Effective authority on an entry generalizes the §22.7 most-permissive union
from three layers to four (global → project → collection → per-entry):
```
effective authority on an entry =
global role (users.role)
project role (membership at the entry's project)
collection role (membership at the entry's collection)
per-entry authority (owners / arbiters / rfc_collaborators — §6.3, §12)
then minus §6.2 write-mute and §22.5 visibility (subtractive, as today)
```
**Per-entry authority is a distinct, finer layer — not a synonym.** `owners` /
`arbiters` / `rfc_collaborators` apply to *one specific entry* (§6.3, §12); the
three named scopes apply to a *subtree*. Per-entry authority is unchanged and
sits beneath collection in the union. `arbiter` is narrower than Owner (one
entry, not a subtree) and stays distinct.
## B.3 Schema impact
- `users.role` continues to carry the **global** role (deployment owner /
contributor).
- M2's `project_members(project_id, role)` rows were attached at what we now
call the **collection**. They generalize into a single polymorphic
**`memberships(scope_type ∈ {project, collection}, scope_id, user_id, role,
granted_by, granted_at)`** table; the M2 rows migrate to
`scope_type='collection'`. The **project** tier gets the same two roles,
freshly grantable.
- The M2 three-role enum (`viewer`/`contributor`/`admin`) collapses to
`{owner, contributor}`: `project_admin → owner`, `project_contributor →
contributor`, `project_viewer →` a read grant (no write) folded into
visibility, not a membership role this pass.
---
# Part C — Behavioral scenarios (BDD)
> These Gherkin scenarios are the behavioral spec for **role usage**,
> **invitation**, and **empty-state** experiences. They attach to the rewritten
> §22 as **§22.6a (role & invitation scenarios)**. They are written so they can
> *also* seed a `bdd`-type collection later (the framework dogfooding its own
> model). "Owner"/"RFC Contributor" are the unified roles (§B.1); a *scope* in
> the `Given` is global / project / collection.
>
> **Each scenario carries a `@S<n>` tag** naming the **slice** (Part E) that
> makes it pass — the "which scenarios are done after this slice" marker. After
> shipping slice S<n>, its acceptance gate is "every `@S<n>` scenario passes"
> (e.g. `--tags @S3`). The Part E table is the inverse index (slice →
> scenarios).
## C.1 Role usage — inheritance and the most-permissive union
```gherkin
Feature: Scope roles grant authority over a subtree
As a member of the deployment
I want a role granted at one tier to apply to everything beneath it
So that I can be invited once and work across the right set of collections
Background:
Given a deployment with a project "ohm"
And "ohm" owns collections "model" (document) and "features" (bdd)
@S3
Scenario: Collection RFC Contributor may propose only in that collection
Given "ada" is RFC Contributor at collection "ohm/model"
When "ada" opens the propose form in "ohm/model"
Then she may submit a new entry
When "ada" opens "ohm/features"
Then she sees it read-only and the propose action is not offered
@S3
Scenario: Project RFC Contributor may propose in every collection of the project
Given "ben" is RFC Contributor at project "ohm"
Then "ben" may propose in "ohm/model"
And "ben" may propose in "ohm/features"
And a collection added to "ohm" later is writable by "ben" with no new grant
@S3
Scenario: Global RFC Contributor may propose in every collection of every project
Given a second project "acme" with collection "acme/specs"
And "cleo" is RFC Contributor at global scope
Then "cleo" may propose in "ohm/model" and "acme/specs"
@S3
Scenario: Collection Owner administers one collection only
Given "dan" is Owner at collection "ohm/features"
Then "dan" may graduate, mark-reviewed, and manage membership in "ohm/features"
But "dan" may not change "ohm" project settings
And "dan" may not act on entries in "ohm/model"
@S3
Scenario: Project Owner administers all collections and may create more
Given "eve" is Owner at project "ohm"
Then "eve" may manage membership in "ohm/model" and "ohm/features"
And "eve" may edit "ohm" project settings
And "eve" may create a new collection in "ohm"
@S3
Scenario: Most-permissive union — the higher grant wins
Given "fay" is RFC Contributor at collection "ohm/model"
And "fay" is Owner at project "ohm"
Then "fay" acts as Owner in "ohm/model"
@S3
Scenario: No negative override — a child cannot subtract a parent grant
Given "gil" is RFC Contributor at project "ohm"
Then there is no control to remove "gil" from "ohm/model" while keeping the project grant
And "gil" can propose in "ohm/model"
@S3
Scenario: A granted account with no scope role sees only public content
Given "hana" has a granted deployment account but no global, project, or collection role
Then "hana" may read public collections under the §6.1 anonymous-read contract
But "hana" is not offered the propose action anywhere
And gated projects and collections do not appear for her
```
## C.2 Invitation — who may invite whom, at which scope
```gherkin
Feature: Inviting users to a scope role
As an Owner of a scope
I want to grant Owner or RFC Contributor at my scope or any scope beneath it
So that collaborators get exactly the reach they need
@S4
Scenario: Project Owner invites at project scope (covers all collections)
Given "eve" is Owner at project "ohm"
When "eve" invites "ivy" as RFC Contributor at project "ohm"
Then a membership row is written at scope project "ohm"
And "ivy" receives a §15 notification naming the project and role
And "ivy" may propose in every collection of "ohm"
@S4
Scenario: Owner invites at a specific collection
When "eve" invites "jo" as RFC Contributor at collection "ohm/features"
Then a membership row is written at scope collection "ohm/features"
And "jo" may propose in "ohm/features" but not "ohm/model"
@S4
Scenario: Invitation reach is bounded by the inviter's scope
Given "dan" is Owner at collection "ohm/features"
Then "dan" may invite users to roles in "ohm/features"
But "dan" is not offered the control to invite at project "ohm" or global scope
@S4
Scenario: RFC Contributors do not manage membership
Given "ben" is RFC Contributor at project "ohm"
Then "ben" may propose and create collections in "ohm"
But "ben" is not offered any invite control (membership is an Owner capability)
@S4
Scenario: The invite UI offers no grant-at-parent-revoke-at-child option
Given "eve" is Owner at project "ohm"
When "eve" opens the invite control for "ivy" at project "ohm"
Then she may choose role Owner or RFC Contributor and scope project or a single collection
But there is no option to grant at "ohm" and exclude a child collection
@S4
Scenario: Re-inviting at a broader scope supersedes the narrower grant
Given "jo" is RFC Contributor at collection "ohm/features"
When "eve" invites "jo" as RFC Contributor at project "ohm"
Then "jo" has the role across all of "ohm"
And the redundant collection-scope row is removed or shown as subsumed
@S4
Scenario: A pending deployment account cannot be granted write
Given "kim" has permission_state "pending" at the deployment
When "eve" invites "kim" as RFC Contributor at project "ohm"
Then the grant is recorded but confers no write capability until "kim" is granted at the deployment (§6)
```
## C.3 Empty-state experiences
```gherkin
Feature: Empty states at each tier
As a viewer of a tier with nothing in it yet
I want a clear, role-appropriate empty state
So that I know whether there is an action to take or simply nothing to see
@S5
Scenario: Global directory with no projects — Owner
Given a deployment with no projects
And "root" is Owner at global scope
When "root" lands on "/"
Then she sees an empty directory with a "Create your first project" call to action
@S5
Scenario: Global directory with no visible projects — non-owner
Given a deployment whose only projects are gated
And "vee" is a granted account with no roles
When "vee" lands on "/"
Then she sees an empty directory with no create action
And a note that there is nothing shared with her yet
@S4
Scenario: Project with no collections — project Owner
Given project "ohm" with no collections
And "eve" is Owner at project "ohm"
When "eve" lands on "/p/ohm/"
Then she sees an empty collection directory with a "Create your first collection" call to action
And the action lets her choose a type and subfolder
@S4
Scenario: Project with no collections — RFC Contributor without create rights
Given project "ohm" with no collections
And "ben" is RFC Contributor at collection scope elsewhere only
When "ben" lands on "/p/ohm/"
Then he sees an empty collection directory with no create action
@S4
Scenario: Collection with no entries — a contributor
Given collection "ohm/model" with no entries
And "ada" is RFC Contributor at collection "ohm/model"
When "ada" lands on "/p/ohm/c/model/"
Then she sees an empty catalog with a "Propose the first entry" call to action
@S2
Scenario: Collection with no entries — an anonymous reader
Given a public collection "ohm/model" with no entries
When an anonymous visitor lands on "/p/ohm/c/model/"
Then they see an empty catalog with no propose action and a sign-in prompt
@S1
Scenario: Single-collection project skips the directory
Given project "ohm" with exactly one visible collection "model"
When a visitor lands on "/p/ohm/"
Then they are redirected to "/p/ohm/c/model/"
@S1
Scenario: Single-project deployment skips the directory
Given a deployment with exactly one visible project "ohm"
When a visitor lands on "/"
Then they are redirected to "/p/ohm/"
```
---
# Part D — Amendments to the original §22 draft
Applied in place when §22 is rewritten; listed here as the change surface.
- **§22 preamble / §22.1.** "A deployment hosts N projects, each a corpus" →
"a deployment hosts N **projects**, each owning one content repo and holding
N **RFC collections**, each collection a typed corpus." Isolation moves to the
collection grain.
- **§22.2 Registry.** `projects.yaml` declares projects with one `content_repo`
each (no per-collection `content_repo`). New: collections are declared by
`.collection.yaml` manifests inside the content repo; the mirror reads them.
In-app create-project / create-collection actions wrap bot commits.
- **§22.3 Content repos.** "One per project" (not per collection); collections
are subfolders within it.
- **§22.4 / §22.4a-c.** Slug is unique **per collection**. `type`,
`initial_state`, and `unreviewed` are **collection** properties (re-homed from
"project"). Unchanged otherwise.
- **§22.5 Visibility.** Applies at **both** project and collection. A collection
defaults to its project's visibility and may narrow it; reading/writing a
collection requires passing both gates.
- **§22.6 Membership and roles → the unified model (Part B).** Replace the three
`project_*` roles with `{owner, contributor}` at `{global, project,
collection}` via a polymorphic `memberships` table. Add **§22.6a** = the
Part C scenarios.
- **§22.7 Composition.** Four-layer most-permissive union (global → project →
collection → per-entry); no negative override.
- **§22.9 / §22.10 Branding & routing.** Routes gain the collection segment:
`/p/<project>/c/<collection>/…`. `GET /api/deployment` lists visible projects;
add `GET /api/projects/:id` (lists visible collections + project settings) and
`GET /api/projects/:id/collections/:cid` (collection settings incl. `type`).
- **§22.11 Notifications / §22.13 migration / §5 amendments.** `project_id`
becomes `collection_id` on every entry-scoped row (the corpus grain is now the
collection); a separate `project_id` exists only on the `collections` table
and project-scoped rows. The §22.13 default project gains a default collection
(§A.6 below).
---
# Part E — Revised slicing plan (the roadmap re-slot)
**Strategy (session 0072, decided on corrected facts).** The two-tier model is
shipped end-to-end (v0.39.0): migration 028 keyed entries `(project_id, slug)`;
v0.35.0 shipped `/p/<project>/` routing + live `/p/<project>/e/<slug>` URLs;
v0.37.0/0.38.0 shipped per-project read + propose. Inserting the third tier is
therefore an **evolution of a shipped system**. The chosen mapping **adds a
collection grain *beneath* today's project** — the shipped `projects` table
stays the grouping tier (it already owns `content_repo`, where §A.2 wants it),
a new `collections` table holds the per-corpus fields, and entries re-key to the
finer `(collection_id, slug)`.
**Slicing principle (session 0072): every slice ends in a *usable* deployment,
and declares the Part C scenarios it makes pass** (its `@S<n>` tag). "Usable"
means the deployment runs and either gains a capability or provably loses none
(N=1 unchanged). A slice is done when its `@S<n>` scenarios are green.
- **Landed, unchanged (v0.39.0):** M1M2, M3-backend Plan A **and** Plan B
(read+propose, mig 028), M3-frontend (`/p/<project>/` routing), §22.13
re-stamp. None of this is rebuilt; it is *evolved* by the slices below.
- **S1 — The collection grain exists (invisible default).** Migration 029 +
backend threading + the default-routing redirect, shipped **together** (they
are coupled — renaming `project_id``collection_id` breaks every reader until
the code is threaded, so a green tree needs both). Migration 029
(`029_collections.sql`): (1) add a `collections` table
`(id, project_id, type, subfolder, initial_state, visibility, name,
registry_sha)`; (2) move the per-corpus fields (`type`, `initial_state`,
visibility) **down** from `projects` (leaving it `(id, content_repo,
visibility, name, tagline, theme, enabled_models, …)`); (3) create one default
collection per project (id `default`, `subfolder` = repo root); (4) re-key
every entry-scoped table `(project_id, slug)``(collection_id, slug)` via the
`028_project_scoped_keys.sql` rebuild pattern (`__new`, copy, drop, rename,
FK-off + `foreign_key_check`); (5) generalize `project_members`
`memberships(scope_type ∈ {project, collection}, …)`, collapsing the role enum
(§B.3). Then thread `collection_id` through `app/auth.py` / `app/projects.py`
/ `app/cache.py` / the `api_*` writers, and **308** `/p/<project>/e/<slug>`
`/p/<project>/c/<default>/e/<slug>`. **Usable end-state:** the deployment runs
exactly as before, now with a real collection layer and one extra path segment.
**Completes:** `@S1` (the single-collection / single-project redirect skips).
- **S2 — Create & navigate a second collection.** *(Shipped v0.41.0.)* Teach the
registry mirror to
read `.collection.yaml`; add the bot-commit-wrapped **create-collection**
endpoint (authorized by existing deployment owner/admin for now — the scoped
role surface lands in S3); the project collection-directory at `/p/<project>/`;
collection-scoped propose/serve under `/p/<project>/c/<collection>/`.
**Usable end-state:** an admin creates a `bdd` collection beside the document
one and it is navigable + proposable. **Completes:** `@S2` (anonymous reader of
an empty collection catalog).
- **S3 — Scope-role enforcement.** *(Shipped v0.42.0.)* The four-layer
most-permissive resolver (§B.2) over `{owner, contributor}` grants at
`{global, project, collection}` (migration 030 adds the `global` scope_type),
with grants applied administratively (DB / the Owner-authorized create
surface); every write gate re-checked under the collection axis. **Plus the
operator's S3 visibility requirements:** collection-grain visibility is
enforced — a `gated` collection is hidden from the public (404, omitted from
the directory) yet visible to scope-role contributors; a collection's
visibility may be set only as strict or stricter than its project's
(`public` < `unlisted` < `gated`). **Keystone reconciliation (session 0076):**
§B.1/§B.3's literal "deployment contributor = global RFC Contributor"
contradicted the C.1 "hana" scenario and the M2 implicit-public baseline;
resolved as — a plain granted account is a granted *account*, not a
write-everywhere global role; "global RFC Contributor" is an explicit
`scope_type='global'` grant; the implicit-public write baseline is
grandfathered onto the migration-seeded `default` collection only (N=1
preserved). *Flag for the SPEC merge (S6): reinterprets §B.1/§B.3.* **Usable
end-state:** a user granted RFC Contributor at a scope can contribute across
exactly that subtree, Owners administer their subtree, and a collection can be
hidden from the public. **Completes:** `@S3` (all of C.1 — role usage,
inheritance, union, no-negative-override).
- **S4 — Invitation surfaces + role-aware empty states.** *(Shipped v0.43.0.)*
The invite UI
(Owner-only) granting Owner/RFC Contributor at a scope or any scope beneath it,
with §15 notifications and the broader-scope-supersedes rule; the
create-first-collection / propose-first empty states keyed to the actor's role.
Modelled as a **direct grant** to an existing account looked up by email (the
C.2 scenarios write the membership row immediately and §15-notify an existing
user — no accept round-trip; inviting a not-yet-account email is out of S4
scope, handled by the admin-create-invite path). **Usable end-state:** an Owner
invites collaborators at the right scope from the UI. **Completes:** `@S4` (all
of C.2 — invitation; plus the project/collection empty states C3.3C3.5).
- **S5 — In-app create-project + the global directory.** *(Shipped v0.44.0.)*
The global-Owner
**create-project** action (bot provisions a Gitea content repo + commits to
`projects.yaml`); the deployment directory empty states. Modelled as a
global-Owner gate (`auth.can_create_project`: a deployment owner/admin or an
explicit `scope_type='global'` Owner grant) over `POST /api/projects`; the
content repo defaults to `<id>-content` and is seeded with a `README.md` so
`main` exists. The deployment payload gains `viewer.can_create_project` +
`default_project_readable` so the directory renders the role-aware empty state
rather than bouncing into an unreadable/absent default. **Usable end-state:**
a global Owner stands up a new project end-to-end from the UI. **Completes:**
`@S5` (the global-directory empty states C3.1C3.2).
- **S6 — Type modules, membership lifecycle, hardening, SPEC merge.** Per-type
frontmatter + surfaces selected on the **collection's** `type`; request-to-join
+ cross-collection inbox; per-collection `enabled_models`; the registry +
manifest format in `docs/DEPLOYMENTS.md`; two-project / multi-collection e2e;
the §20.4 changelog + upgrade-steps; the SPEC merge (Part A applied, Part D in
place). **Usable end-state:** the model is fully realized and merged into
`SPEC.md`. **Completes:** type-specific scenarios (added in S6, beyond Part C's
role focus).
- *Shipped in S6 core (v0.45.0):* the SPEC merge, per-collection
`enabled_models`, the type-driven entry noun (§22.4a item 2).
- *Shipped as the S6 remainder (v0.46.0):* request-to-join + the
cross-collection inbox (§22.8).
- *Spec'd, not yet built — the last S6 item:* the per-type **frontmatter
schemas** (§22.4a item 1) and **surfaces** (§22.4a item 3, the
`specification` release-planning + `bdd` scenario/coverage views). The
discovery/spec pass + BDD scenarios + slicing (S7aS7c) are in
[`2026-06-06-per-type-surfaces.md`](./2026-06-06-per-type-surfaces.md).
### Slice → scenario index (the inverse of the `@S<n>` tags)
| Slice | Usable thing it ships | Completes (`@S<n>`) |
|---|---|---|
| **S1** | collection grain + default + redirects; N=1 unchanged | C3.7, C3.8 (`@S1`) |
| **S2** | create + navigate + propose a 2nd collection | C3.6 (`@S2`) |
| **S3** | scope-role enforcement across global/project/collection | C1.1C1.8 (`@S3`) |
| **S4** | invitation UI + role-aware empty states | C2.1C2.7, C3.3C3.5 (`@S4`) |
| **S5** | in-app create-project + global directory | C3.1, C3.2 (`@S5`) |
| **S6** | type surfaces, lifecycle, hardening, SPEC merge | type-specific (new) |
Each slice is a candidate single session: it lands a usable deployment and a
runnable acceptance gate (`--tags @S<n>`). **S1 is the natural first session**
the coupled migration 029 + threading + redirect, sized as one usable increment
(answering the in-session question: bundled, it is right-sized, not too much).
## E.1 (= §A.6) Migration — the default collection (the N=1 case)
A deployment on the shipped two-tier schema (v0.39.0) is migrated by 029 so it
keeps running unchanged:
1. The existing `projects` row **stays as the project** (it already owns
`content_repo` and its config-derived `id` from §22.13 step 1).
2. A **default collection** (`id='default'`, `subfolder` = repo root) is created
per project, inheriting that project's `type` / `initial_state` / visibility;
those fields are then dropped from `projects`.
3. Every entry-scoped `project_id` row is re-keyed with the default
`collection_id` (PK `(project_id, slug)``(collection_id, slug)`).
4. `project_members` rows migrate to `memberships(scope_type='collection')` on
the default collection, role-collapsed (§B.3).
5. **308 redirects:** the shipped `/p/<project>/e/<slug>`
`/p/<project>/c/<default>/e/<slug>`, and the pre-multi-project `/rfc/<slug>`
/ `/proposals/<n>` → their `/p/<project>/c/<default>/…` equivalents.
Until a second collection is added, the deployment is functionally identical to
before, with one extra path segment. This is the §20.4 upgrade-steps content for
the release.
## E.2 Scope of the first implementation pass
Per the launch ask — "we don't need to get all permissions right yet, just have
Owner at all levels, and RFC Contributor at the global, project, and RFC
collection level" — the **role surface** this pass implements is exactly
`{owner, contributor}` × `{global, project, collection}` (Part B), plus the
unchanged per-entry layer. `viewer`, the owner/admin split, request-to-join
nuances, and per-type role labels are deferred (§B.1 note).
@@ -0,0 +1,456 @@
# Solution Design: Configurable Collection Metadata (clean-doc tagging)
| | |
| --- | --- |
| **Author(s)** | Ben Stull |
| **Reviewers / approvers** | Ben Stull |
| **Status** | `draft` |
| **Version** | v0.1.6 |
| **Source artifacts** | Reference modeled: retired **BDD Release Planner** (`wiggleverse/wiggleverse-ecomm-bdd-release-planner-app`, RETIRED 2026-06-04) · Related: [`2026-06-05-three-tier-projects-collections.md`](./2026-06-05-three-tier-projects-collections.md) (§22) · Corpus: ecomm Shopify-modeled BDD (`wiggleverse-ecomm-meta/research/shopify`, ~1,238 scenarios) · **Supersedes:** [`2026-06-06-per-type-surfaces.md`](./2026-06-06-per-type-surfaces.md) |
**Change log**
| Date | Version | Change | By |
| --- | --- | --- | --- |
| 2026-06-06 | v0.1.0 | Initial draft from discovery session OHM-0079.0 | Ben Stull |
| 2026-06-06 | v0.1.1 | Value-only Executive Summary; add Pain Points | Ben Stull |
| 2026-06-06 | v0.1.2 | Business Outcomes restated as business (adoption/diversity); Business Use Cases → solution-agnostic | Ben Stull |
| 2026-06-06 | v0.1.3 | Supersede per-type-surfaces draft (harvest patterns; bdd coverage future; §22.4a amendment); split Business Actors / Product Personas | Ben Stull |
| 2026-06-06 | v0.1.4 | Two-part restructure: §1 Business Context (solution-agnostic, 1.11.9) + §2 Solution Proposal; renumber | Ben Stull |
| 2026-06-06 | v0.1.5 | Move Business Actors to §1.3 (define roles before Problem/Pain reference them) | Ben Stull |
| 2026-06-06 | v0.1.6 | §7.1 execution convention — each slice is its own writing-plans→executing-plans coding session, plans just-in-time | Ben Stull |
---
## 1. Business Context
*The business lens — solution-agnostic throughout. No mechanism is proposed until §2.*
### 1.1 Executive Summary
A deployment's corpus is only as valuable as the ability of the people running it to prioritise it, navigate it, and act on it — and as valuable as the downstream tools that can read structured signal out of it. Today that value is stranded: operators and contributors can't rank what matters or find content by what matters, and the tools meant to plan and build from the corpus have nothing structured to consume. The value at stake is **lower-friction corpus planning** for operators and contributors, **broader adoption** by teams whose document types the platform couldn't previously serve, and **a corpus external tooling can consume without bespoke glue**. *(Value summary; the solution is proposed in §2.)*
### 1.2 Background
The framework hosts RFC standardization for multiple deployments. One deployment hosts the ecomm BDD corpus — ~1,238 Shopify-modeled scenarios, one markdown file per scenario, slugged by feature ID (`DD-FF-NNNN-slug`). A standalone **BDD Release Planner** previously let operators search that corpus, attach metadata (priority P0P3, owner, status), cluster scenarios into named releases, and emit each release as a roadmap phase. §22 (three-tier projects/collections) absorbed the planner's *corpus hosting* into rfc-app (the corpus now runs as a `bdd` project on the RFC deployment) and the planner was retired — but its *annotation* half (priority/tags on scenarios, filtering, bulk assignment) was never rebuilt. Teams evaluating rfc-app for *other* document types often need structured attributes (a priority, a status, domain tags) the platform can't yet express — so they go elsewhere.
### 1.3 Business Actors / Roles
Real-world roles, **solution-agnostic** — they exist whether or not rfc-app does. They are defined here, before the Problem (§1.4) and Pain Points (§1.5) reference them; the Business Use Cases (§1.9) are about these roles, and the Product Personas (§3) map onto them.
| Role | Responsible for (in the business) |
| --- | --- |
| Standards owner | Owns an organization's RFC / standards / requirements process; decides what's tracked and how |
| Release planner | Decides what work belongs in upcoming releases |
| Requirements author | Proposes and curates the requirements (e.g. BDD scenarios) |
| Requirements consumer | A person or downstream tool that plans or builds from the requirements |
| Reader | Anyone navigating the corpus to find what's relevant to them |
### 1.4 Problem Statement
rfc-app cannot express or surface structured signal about its content. Tags are free-form strings with no filtering; there is no notion of priority or any other collection-defined attribute; the catalog is a flat list; and what little metadata exists is mixed into the top of every document. As a result, a corpus cannot be prioritised, navigated by attribute, planned in bulk, or cleanly consumed by downstream tools — and teams whose workflows depend on such attributes cannot adopt the platform at all.
### 1.5 Pain Points
| # | Pain | Who feels it | Cost / frequency today |
| --- | --- | --- | --- |
| PP-1 | Scenarios carry no priority, so triage and planning happen off-platform, in spreadsheets and memory | Release planner, contributor | Every planning cycle; signal lives off-platform and goes stale |
| PP-2 | The catalog is a flat, unfilterable list — at ~1,200 scenarios, "show me the P0 checkout scenarios" is impractical | Reader, release planner | Every browse/triage; finding the right work is slow and error-prone |
| PP-3 | Tags are free-form with no filtering payoff, so they're decorative and go unmaintained | Contributor | Ongoing; the one existing affordance rots |
| PP-4 | Annotating many scenarios means opening many PRs, so bulk planning has no home in the tool | Release planner | Every batch; the core planning gesture is effectively impossible |
| PP-5 | rfc-app metadata clutters the top of every document, hurting readability and making the corpus awkward to consume cleanly | Reader, downstream consumer | Every read; every downstream integration |
| PP-6 | Downstream tools have no structured signal to read — the retired planner's capability left a gap | Downstream consumer | Continuous since the planner's retirement |
| PP-7 | Teams whose document types need structured attributes can't model them, so they don't adopt rfc-app | Prospective adopter (org/team) | Every evaluation that ends in "not yet" |
### 1.6 Targeted Business Outcomes
Business outcomes for rfc-app as a platform — adoption, reach, and diversity of use — **not** solution outputs. (Whether documents carry a priority is a solution output, tracked as a slice's Definition of Done in §7, not here.)
| Outcome | Success metric | Baseline → Target | Guardrail (must not regress) | How / when measured |
| --- | --- | --- | --- | --- |
| Teams blocked by missing structured attributes now adopt rfc-app | # organizations on rfc-app; # active users | internal deployments only → external orgs onboard | existing deployments don't churn | deployment registry + usage analytics; quarterly |
| The platform hosts a wider variety of workflows and document types | # distinct document/collection types & field schemas in use | today's handful → broader mix | existing types' experience unchanged | type/schema census; quarterly |
| Corpus planning happens on-platform rather than in side tools | share of prioritisation/planning done in rfc-app vs spreadsheets | largely off-platform → on-platform | — | operator interviews + usage signals; quarterly |
### 1.7 Scope (business)
- **In scope:** the corpus can carry per-item importance and categorisation; people can find items by those attributes; the signal is captured durably and is consumable by other people and tools; teams with new document types can express the attributes their workflow needs.
- **Out of scope (business):** deciding *what* a given deployment's priorities or categories should be (that's the deployment's editorial choice); release sequencing and ship tracking as a business process (stays a downstream/operator concern).
- **Non-goals:** modelling "releases" as a first-class business object inside the platform.
*(Solution-specific scope/non-goals are in §2.)*
### 1.8 Assumptions · Constraints · Dependencies
- **Assumptions:** git remains the content source of truth and downstream consumers can read the corpus from git; the BDD grain is one markdown file per scenario (already true for the ecomm corpus).
- **Constraints:** rfc-app is a framework hosting multiple deployments — any change must be **mechanical and non-breaking**, with §20 changelog/upgrade-steps; the hard secrets rule (§6.3) holds; edits must respect scope-role authorization (§22 Part B / S3); the §22.4a "engine unchanged" rule holds (INV-8).
- **Dependencies:** the S3 scope-role resolver (`auth.effective_scope_role`); the existing git write-through used by `edit-meta` (§9.5); the §22 collection model; the binding `SPEC.md` §22.4a contract, which §2's solution amends (§7 SLICE-0).
### 1.9 Business Use Cases
Solution-agnostic: what an actor (§1.3) wants to accomplish, *why* (value), and what *success* looks like — **no reference to any product**. Each could be satisfied by a person by hand before any software. Form: "As a … I can … so that …".
**BUC-1 — As a release planner, I can prioritise the requirements in a body of work, so that I can decide what belongs in upcoming releases.**
```gherkin
Scenario: BUC-1 — Prioritise to plan releases
Given a body of requirements of varying importance
When the planner weighs which matter most
Then they hold a ranking of those requirements by importance
And can decide a release's contents from it
```
- **Acceptance:** the planner can select and justify the next release's contents from the relative importance of the work.
**BUC-2 — As a planner facing a large body of requirements, I can organise and triage it within a normal working session, so that planning actually gets done rather than deferred or improvised.**
```gherkin
Scenario: BUC-2 — Triage at scale
Given more requirements than can be weighed one at a time
When the planner ranks and groups them in bulk
Then the body of work reflects those decisions without per-item drudgery
```
- **Acceptance:** a planner moves from an unsorted corpus to a prioritised plan in one sitting.
**BUC-3 — As a team, I want the importance and categorisation of our requirements captured durably and shareably, so that other people and tools can plan from it without re-deriving it.**
```gherkin
Scenario: BUC-3 — Durable, shareable signal
Given requirements that have been weighed and categorised
When someone or something else needs to plan from them
Then they can read what matters and why without asking the original author
```
- **Acceptance:** a second party — person or tool — can pick up the work and plan from it unaided.
**BUC-4 — As a team with a specialised body of documents, I can capture the attributes that make them actionable (importance, status, category), so that I can manage that work the way my domain requires.**
```gherkin
Scenario: BUC-4 — Manage a domain's work on its own terms
Given documents whose usefulness depends on domain-specific attributes
When the team records and works with those attributes
Then they can run their workflow with the distinctions it depends on
```
- **Acceptance:** the team can capture and act on the distinctions their domain requires — success is them choosing to manage the work this way.
**BUC-5 — As someone consuming a large corpus, I can find the items that matter to my current purpose, so that I act on the right things instead of wading through everything.**
```gherkin
Scenario: BUC-5 — Find what matters
Given a large body of items
When the consumer looks for the important ones for their task
Then they can locate them quickly
```
- **Acceptance:** a person narrows a large corpus to the relevant, important subset for their task.
---
## 2. Solution Proposal
**The solution is to build it into rfc-app.** Give every collection a small, declared **field schema** (in its `.collection.yaml`) so it can carry structured metadata — priority, tags, and any custom fields the deployment defines. Store each entry's values in a **clean sidecar** file so the document body stays pure prose. rfc-app then **renders those fields as forms, filters the catalog by them (faceted, with counts), and lets authorized users tag in single and bulk gestures** committed straight to git; downstream tools read the values from the sidecars directly. It is one generic mechanism — tags and priority are just *fields* — not per-type special-casing and not a bespoke "release" entity.
**Why a software solution (and not a manual one).** A non-build alternative — operators maintaining priorities/tags in a shared spreadsheet — was considered and rejected: it leaves the corpus unfilterable in-tool (PP-2), keeps documents and the side-sheet out of sync, produces no durable git-readable signal for downstream tools (PP-5/PP-6), and does nothing for the adoption outcome (§1.6, PP-7). The value only lands if the structure lives with the content.
**Solution-specific scope.** *Out:* release ordering, ship status, roadmap emission, the `specification` release-planning surface — all downstream, reading sidecars from git. In-app management of field definitions (edit `.collection.yaml` in git for v1); corpus-wide tag rename/merge/delete; sub-document grain; a whole-corpus export endpoint. *Future (recorded, not v1):* a **bdd coverage surface** — a `verifies`-style **`ref` field type** plus a read-derived view mapping features to the spec sections they exercise (harvested from the superseded per-type-surfaces draft); deferred pending §9 Q4. This solution **amends the binding `SPEC.md` §22.4a contract** (§7 SLICE-0).
*(The Product and Engineering sections below — §§37 — elaborate this build. They would be replaced by an operational plan if the chosen solution were non-software.)*
---
## 3. Product Personas
rfc-app's user types — each an embodiment of one or more Business Roles (§1.3). The Product Use Cases (§4) are about these personas.
| Product persona | In rfc-app | Maps to business role(s) |
| --- | --- | --- |
| Collection Owner | scope-role Owner; declares the collection's `fields:` schema (edits `.collection.yaml`) | Standards owner |
| Contributor | scope-role contributor; sets metadata (single + bulk), proposes/curates entries | Requirements author; Release planner |
| Reader | viewer; browses and filters the catalog | Reader |
| Downstream consumer | an external system reading sidecars + `.collection.yaml` from git | Requirements consumer |
## 4. Product Use Cases
```gherkin
Scenario: PUC-1 — Set priority/tags on a scenario (realizes BUC-1, BUC-4)
Given I am a Contributor viewing a scenario whose collection defines priority and tags
When I choose P0 in the priority control and add the tag "checkout"
Then the metadata panel reflects P0 and the checkout tag
And the change is committed directly to the scenario's sidecar
Scenario: PUC-2 — Bulk tag/untag from the catalog (realizes BUC-2)
Given I have multi-selected several scenarios in the catalog
When I choose "Set priority → P1" from the bulk action bar
Then every selected scenario shows P1
And the bulk change is one commit
Scenario: PUC-3 — Filter the catalog by facet (realizes BUC-5, BUC-1)
Given the left pane shows faceted filters generated from the collection schema
When I check Priority P0 and tag "checkout"
Then the catalog shows only scenarios matching both
And each facet value shows its result count
Scenario: PUC-4 — A Collection Owner declares fields (realizes BUC-4)
Given a Collection Owner edits .collection.yaml to add a priority enum field
When the collection is re-ingested
Then the priority filter and the priority form control appear automatically
Scenario: PUC-5 — Migrate a collection to clean docs (product-only; enables BUC-3)
Given a collection whose docs still carry top-of-doc frontmatter
When the operator runs the frontmatter→sidecar migration
Then each doc body becomes pure prose and a sidecar holds its metadata
And rfc-app reads the collection identically before and after
Scenario: PUC-6 — A malformed entry is visibly fixable (realizes BUC-3)
Given a stored entry whose metadata fails its collection's schema
When the catalog renders
Then the entry still loads (read never hard-fails)
And it is flagged "malformed metadata" so a Contributor can fix it
```
## 5. UX Layout
### 5.1 Screen: Catalog (left pane) (serves PUC-3, PUC-6)
- **Purpose:** browse and filter a collection's entries.
- **Layout (top → bottom):** full-text search (existing); **faceted filter groups** (one per schema field + state): each a collapsible group with per-value **result counts** and multi-select checkboxes; `tags`-type fields include a "filter values…" search box to stay usable at 30+ values.
- **States:** happy: facets with counts · empty: "no entries match" + clear-filters · loading: skeleton facets · error: retry · **malformed:** entries failing their schema carry a fixable marker (parallel to §22.4c `unreviewed`) and are filterable.
### 5.2 Screen: Scenario detail — metadata panel (serves PUC-1)
- **Purpose:** view/edit one entry's metadata.
- **Layout:** one control per schema field — `enum` → single-select; `tags` → removable chips + add-tag input (with existing AI suggest); `text` → text input. The body renders below as pure prose; metadata never appears inline.
- **States:** read (no edit role) shows values · edit (authorized) shows controls · saving: spinner · error: field-level validation message.
### 5.3 Screen: Catalog — bulk action bar (serves PUC-2)
- **Purpose:** apply a field value to many entries at once.
- **Layout:** selecting ≥1 row reveals a sticky bar: "*N* selected · Set priority ▾ · Add tag ▾ · Remove tag ▾ · Clear". Applying commits once.
- **States:** none selected: hidden · applying: progress · partial failure: toast naming entries that failed validation, others applied.
## 6. Technical Design
### 6.1 Invariants
- **INV-1:** The sidecar (`<slug>.meta.yaml`) is the source of truth for entry metadata; `cached_rfcs` is a derived index, fully rebuildable from git.
- **INV-2:** A document body (`.md`) never contains rfc-app metadata once migrated; metadata lives only in the sidecar.
- **INV-3:** Reading a collection never hard-fails on bad metadata — an invalid value surfaces as a warning, the entry still loads, and the catalog flags it (§5.1).
- **INV-4:** Metadata writes are authorized by scope-role (contributor+ on the collection) and validated at the write boundary; content-body edits keep their existing PR-review path.
- **INV-5:** A collection with no `fields:` block behaves exactly as today (free-form `tags` only). The §22.13 generated **default collection is `document`** with no fields → **N=1 deployments see zero change**.
- **INV-6:** Dual-read: parser reads the sidecar if present, else legacy top-of-doc frontmatter, with identical resulting in-memory records.
- **INV-7:** Unknown / forward-compat keys in a sidecar **ride along untouched** — never dropped on read or rewrite, never reported as malformed.
- **INV-8:** **Engine unchanged** (§22.4a) — additive and read-mostly; never forks the content write path, the propose→branch→PR→graduate lifecycle, threads/flags/chat, or the storage model. Metadata edits reuse the existing `edit-meta` git write-through.
### 6.2 High-level architecture
```mermaid
flowchart LR
subgraph Git[content repo]
CY[.collection.yaml<br/>fields: schema]
MD[slug.md<br/>prose body]
SC[slug.meta.yaml<br/>values]
end
CY --> ING[ingest / parser<br/>lenient, type-agnostic]
MD --> ING
SC --> ING
ING --> VAL[metadata_schema.validate<br/>advisory at read]
VAL --> DB[(cached_rfcs<br/>values + facet counts + malformed)]
DB --> API[API: schema · list+filter · facets · edit]
API --> FILT[left-pane faceted filters]
API --> PANEL[detail metadata panel]
API --> BULK[bulk select bar]
PANEL -->|validate + direct commit| SC
BULK -->|validate + 1 commit| SC
SC -.read from git.-> CONS[downstream consumers]
```
- **ingest/parser** — reads `.collection.yaml` schema + sidecars (or legacy frontmatter), stays lenient/type-agnostic (INV-7); rebuilds `cached_rfcs`; never authoritative.
- **`metadata_schema.validate(values, fields) → [problems]`** — the one place that knows a collection's required/forbidden fields and each field's shape (modeled on `registry.py`). Advisory at ingest (warn + malformed flag, INV-3); enforced at the write boundary (INV-4).
- **API** — serves the schema, filtered lists with facet counts + malformed flag, and metadata edits; never writes metadata anywhere but the sidecar.
### 6.3 Data model & ownership
| Entity | Owned by | Key fields | System of record |
| --- | --- | --- | --- |
| Collection field schema | Collection Owner | `fields: {name → {type, values?, label}}` in `.collection.yaml` | git |
| Entry metadata values | Contributor | sidecar `<slug>.meta.yaml`: lifecycle + schema fields + forward-compat keys (INV-7) | git (sidecar) |
| Derived index | ingest | per-entry values + facet aggregations + `malformed` flag | `cached_rfcs` (SQLite, derived) |
**Field types (v1):** `enum` (single-select; controlled by required `values:`), `tags` (multi-value; free-form unless `values:` given), `text` (free string). **Future:** `ref` (a typed cross-entry link — basis for the deferred bdd `verifies`/coverage surface; §2, §9 Q4). Unknown types ignored with a warning.
**Sidecar example:**
```yaml
slug: 01-01-0001-view-today-s-key-performance-metrics-at-a-glance
title: View today's key performance metrics at a glance
state: active
owners: [ben.stull]
priority: P1
tags: [dashboard, analytics]
owner: hasan
```
### 6.4 Interfaces & contracts
- **`GET …/collections/<c>`** — out: collection incl. `fields` schema.
- **`GET …/collections/<c>/rfcs`** — in: filter params (`?priority=P0&tags=checkout&state=active`; OR within a field, AND across fields; `?malformed=true`) · out: entries with values + per-entry `malformed` + `facets: {field → {value → count}}`. Errors: 400 unknown field.
- **`POST …/rfcs/<slug>/meta`** — in: `{field: value}` · effect: validate → write sidecar → direct commit → re-ingest. Errors: 403, 422.
- **`POST …/collections/<c>/meta/bulk`** — in: `{slugs, op: set|add|remove, field, value}` · out: `{applied, rejected}` · effect: validate → write N sidecars → one commit → re-ingest. Errors: 403, 422.
### 6.5 PerProduct-Use-Case design
#### PUC-2 — Bulk tag/untag
```mermaid
sequenceDiagram
actor U as Contributor
participant C as Catalog UI
participant A as API
participant V as metadata_schema
participant G as Git
participant D as cached_rfcs
U->>C: select rows, "Set priority P1"
C->>A: POST /meta/bulk {slugs, set, priority, P1}
A->>A: authz (scope-role)
A->>V: validate values vs schema
A->>G: write N sidecars, 1 commit
A->>D: re-ingest affected entries
A-->>C: {applied, rejected}
C-->>U: rows show P1; toast on any rejected
```
- **Implementation:** reuse the `edit-meta` git write-through, extended to target the sidecar and batch N files into one commit. Honors INV-1/INV-4/INV-8.
#### PUC-5 — Migration
- **Implementation:** a tool walks a collection; for each entry with legacy frontmatter it writes `<slug>.meta.yaml` and rewrites `<slug>.md` to the body only — one commit per collection, idempotent, preserving unknown keys (INV-7). Dual-read (INV-6) lets it run anytime; lazy migration converts stragglers on first metadata edit.
### 6.6 Non-functional requirements & cross-cutting concerns
- **Security & privacy:** edits gated by `auth.effective_scope_role`; no secrets in sidecars; git history records authorship.
- **Performance & scale:** facet counts from the derived DB; responsive at ~1.2k entries with dozens of tag values.
- **Availability & resilience:** bad metadata never blocks read (INV-3); failed re-ingest leaves git authoritative, recoverable by rebuild.
- **Observability:** log each metadata commit; warn-log + count schema-validation failures on ingest.
- **Accessibility:** facet groups and form controls keyboard-navigable; checkboxes labelled value + count.
### 6.7 Key decisions & alternatives considered
| Decision | Chosen | Alternatives | Why |
| --- | --- | --- | --- |
| Solution type | Build into rfc-app | Manual (shared spreadsheet) | Manual leaves corpus unfilterable, out of sync, no git-readable signal (§2) |
| Release modeling | Metadata only; releases downstream | First-class release entity | Operator pulled ordering/ship-status out of rfc-app |
| Tag system shape | One generic typed-field system | Releases first-class + simple tags; namespaced facets | Tags/priority/custom are all just fields |
| Schema model (D9) | Pure collection-config | Type-driven hard-coded schemas (per-type-surfaces draft) | Flexible, data-driven |
| Metadata storage | Sidecar per entry | Frontmatter; end-of-doc; index file; DB-only | Clean docs + git-visible + locality |
| Left-pane filtering | Faceted groups with counts | Flat facet chips | Scales to ~1.2k-scenario, many-tag corpus |
| Edit governance | Direct commit for authorized roles | PR per change | Bulk planning impractical via PR-per-toggle |
| bdd coverage (D10) | Future per-type surface over a `ref` field | Build now; drop | Valuable but not v1; needs Q4 |
### 6.8 Testing strategy
Unit: schema parsing (all types, missing block); sidecar round-trip incl. unknown-key preservation (INV-7); dual-read equivalence (INV-6); validation; malformed-flag; facet aggregation; bulk op (single commit, partial-rejection). Two-tier local-Docker→PPE for API + git write-through. "Tested" = PUC acceptance scenarios pass + migration proven idempotent and reversible-on-read.
### 6.9 Failure modes, rollback & flags
- **Invalid value committed out-of-band** → ingest warns + loads with the value flagged malformed (INV-3).
- **Re-ingest fails after commit** → git authoritative; full rebuild recovers.
- **Migration rollback:** dual-read keeps an un-/partly-migrated corpus working; the migration commit is revertible.
- **Feature flag:** inherently opt-in per collection (INV-5) — no global flag.
## 7. Delivery Plan
### 7.1 Approach / strategy
Amend the binding contract first, then build storage/compat, then schema, then read, then write. Each build slice is shippable and non-breaking.
**Execution convention.** Each slice is taken as **its own coding session**`writing-plans → executing-plans → verify → ship/deploy → merge + version bump` — in dependency order, with the slice's implementation plan written **just-in-time** at the start of that session, not up front (later slices' plans depend on the code earlier slices land). `brainstorming` ran once to produce this spec and recurs only if a slice proves the spec wrong. A slice's **Definition of Done** (§7.2) is the signal to advance the `Next /goal:` cursor to the next slice. SLICE-0 is doc-only (no implementation plan).
### 7.2 Slicing plan
#### SLICE-0 — Amend `SPEC.md` §22.4a (contract) → unblocks the rest
- **Depends on:**
- **Definition of done:** §22.4a reframed — item 1 (entry schema) is **collection-configured sidecar fields**, not type-driven frontmatter; item 3 (type surfaces) deferred to a future design (bdd coverage recorded); per-type-surfaces draft marked superseded; §20 changelog. *Doc-only; no code.*
#### SLICE-1 — Sidecar storage + dual-read + migration → completes PUC-5, PUC-6
- **Depends on:** SLICE-0
- **DoD:** parser reads sidecar-else-legacy (INV-6), preserves unknown keys (INV-7); migration tool idempotent; existing collections load byte-identically; malformed flag derived; tests green.
#### SLICE-2 — Collection field schema + central validation → completes PUC-4
- **Depends on:** SLICE-1
- **DoD:** `.collection.yaml fields:` parsed; `metadata_schema.validate` advisory at read / enforced at write; schema served via the collection API; no-`fields:` collections unchanged (INV-5).
#### SLICE-3 — Faceted left-pane filtering (read) → completes PUC-3
- **Depends on:** SLICE-2
- **DoD:** list endpoint returns facet counts + honors filter params (incl. `malformed`); left pane renders faceted groups with counts + tag-value search; filters compose.
#### SLICE-4 — Single-entry metadata edit → completes PUC-1
- **Depends on:** SLICE-2
- **DoD:** detail panel renders schema controls; `POST …/meta` validates, direct-commits, re-ingests; scope-role gated (INV-4); lazy-migrates a legacy entry on first edit.
- **Carried from SLICE-1 (deferred there):** make the **write paths**
sidecar-aware — every site that today does `entry.parse(<slug>.md)` and
serializes back into the `.md` must read/write metadata via the sidecar so a
migrated (body-only) entry doesn't crash or re-grow frontmatter. The known
sites: graduation + claim + `_read_meta_entry` (`api_graduation.py`),
`mark_entry_reviewed` (`bot.py`), body-edit / accept-change wrappers
(`api_branches.py` `_wrap_body`/`_extract_body`), and the PR-replay wrappers
(`api_prs.py`). Only once these are sidecar-aware should the **operator
trigger** for `metadata.migrate_collection` (the Owner-gated migrate endpoint)
ship.
#### SLICE-5 — Bulk tag/untag → completes PUC-2
- **Depends on:** SLICE-3, SLICE-4
- **DoD:** multi-select + bulk bar; `POST …/meta/bulk` applies set/add/remove as one commit; partial-rejection reported.
### 7.3 Rollout / launch plan
Pre-v1, single production: ship slices in order; each minor bump carries §20 changelog + upgrade steps. Opt-in per collection (INV-5): a deployment adopts it only by declaring a `fields:` block and (optionally) running the migration.
### 7.4 Risks & mitigations
| Risk | L/I | Mitigation |
| --- | --- | --- |
| Amending binding §22.4a destabilises a shipped contract | M/M | SLICE-0 doc-only, reviewed; dual-read keeps runtime non-breaking; supersede note preserves rationale |
| Frontmatter→sidecar migration corrupts content | L/H | Dual-read; idempotent, revertible migration; body-byte-identity + unknown-key tests |
| Doubling file count (sidecars) clutters corpus | M/L | Docs stay clean; sidecars small/co-located |
| Direct-commit metadata edits bypass review | M/M | Scope-role gate (INV-4); content-body edits still PR'd; git audit trail |
| Facet aggregation slow at scale | L/M | Compute from indexed derived DB; measure at ~1.2k entries |
## 8. Traceability matrix
| Pain | Business UC | Product UC | Slice | Tests |
| --- | --- | --- | --- | --- |
| — (contract) | — | — | SLICE-0 | (doc review) |
| PP-5 | BUC-3 | PUC-5, PUC-6 | SLICE-1 | `test_dual_read_equiv`, `test_migration_idempotent`, `test_unknown_keys_preserved` |
| PP-7 | BUC-4 | PUC-4 | SLICE-2 | `test_schema_parse`, `test_validate` |
| PP-2 | BUC-5, BUC-1 | PUC-3 | SLICE-3 | `test_facet_counts`, `test_filter_compose` |
| PP-1, PP-3 | BUC-1, BUC-4 | PUC-1 | SLICE-4 | `test_single_meta_commit`, `test_authz` |
| PP-4 | BUC-2 | PUC-2 | SLICE-5 | `test_bulk_one_commit`, `test_partial_reject` |
| PP-6 | BUC-3 | (consumer reads git) | — | `test_sidecar_schema_stable` |
## 9. Open Questions & Decisions log
**Open**
| # | Question | Owner | Blocks |
| --- | --- | --- | --- |
| Q1 | Do downstream consumers read sidecars from git, via API, or both? (leaning git) | Ben | nothing v1 |
| Q2 | Ship `multi-enum` (multi-select controlled) in v1 or later? | Ben | SLICE-2 scope |
| Q3 | Exact §22.4a amendment wording + the future-surfaces home | Ben | SLICE-0 |
| Q4 | bdd coverage: `ref` field grammar + a coverage view honoring §22's no-cross-collection-join rule as hyperlinks | Ben | future surface |
**Resolved**
| # | Decision | Resolution | Date |
| --- | --- | --- | --- |
| D1 | Release behaviors | Out of rfc-app; downstream | 2026-06-06 |
| D2 | Tag system shape | Approach A — one generic typed-field system | 2026-06-06 |
| D3 | Metadata grain | Per entry (corpus already one file per scenario) | 2026-06-06 |
| D4 | Schema location | `.collection.yaml` `fields:` block | 2026-06-06 |
| D5 | Value storage | Sidecar `<slug>.meta.yaml`; doc body pure prose | 2026-06-06 |
| D6 | Left-pane filtering | Faceted groups with counts | 2026-06-06 |
| D7 | Edit governance | Direct commit for authorized roles; bulk = 1 commit | 2026-06-06 |
| D8 | Management scope | Deferred; edit `.collection.yaml` in git for v1 | 2026-06-06 |
| D9 | Schema model | Pure collection-config; not type-driven | 2026-06-06 |
| D10 | bdd coverage | Future per-type surface over a `ref` field; not v1 | 2026-06-06 |
| D11 | per-type-surfaces draft | Superseded; §22.4a to be amended (SLICE-0) | 2026-06-06 |
## 10. Glossary & References
- **Sidecar**`<slug>.meta.yaml`, the per-entry metadata file that is the source of truth; keeps the `.md` body pure prose.
- **Field schema** — the `fields:` block in `.collection.yaml` declaring a collection's typed metadata fields.
- **Facet** — a schema field surfaced as a left-pane filter group with per-value counts.
- **Malformed metadata** — stored values that fail their collection's schema; flagged in the catalog, never a hard read failure (INV-3).
- **Downstream consumer** — an external tool that reads corpus metadata from git; rfc-app does not model releases.
- **References:** retired BDD Release Planner; superseded per-type-surfaces draft (`2026-06-06-per-type-surfaces.md`); §22 three-tier design; `SPEC.md` §7.1 (left-pane filter), §9.5 (edit-meta), §20 (versioning), §22.4a (per-type contract — to be amended), §22 Part B / S3 (scope-role).
```
+366
View File
@@ -0,0 +1,366 @@
# Draft spec — §22.4a per-type surfaces (the last S6 item)
> # ⛔ SUPERSEDED (2026-06-06)
>
> This draft is **superseded by**
> [`2026-06-06-configurable-collection-metadata.md`](./2026-06-06-configurable-collection-metadata.md),
> which reframes §22.4a item 1 as **collection-configured** metadata in
> **sidecars** (not type-driven frontmatter) and defers item 3's surfaces.
> Harvested into the successor: the validation seam (A.1), the malformed-metadata
> catalog flag (A.5), unknown-fields-ride-along (C.1), the engine-unchanged rule
> (§0), and the N=1 `document` backcompat anchor (A.2). The **bdd coverage**
> capability (`feature`/`verifies` → coverage view, Part B.2) is preserved there
> as a *future* per-type surface over a generic `ref` field. The binding
> `SPEC.md` §22.4a contract is to be amended by the successor's SLICE-0. Kept for
> historical rationale; do not build from this document.
> **Status:** discovery/spec pass — *not yet sliced into a shipped release.*
> Author session: 0083 (2026-06-06). This document is the spec pass the §22 S6
> remainder called for: it specifies **§22.4a item 1** (the per-type entry
> **frontmatter schema**) and **§22.4a item 3** (the per-type **surfaces**) for
> the `specification` and `bdd` collection types, with BDD-style acceptance
> scenarios and a delivery slicing, so a later coding session can build them
> against a written contract rather than improvising.
>
> It is the sibling of
> [`2026-06-05-three-tier-projects-collections.md`](./2026-06-05-three-tier-projects-collections.md)
> (the three-tier model, S1S6 core, shipped through v0.46.0) and refines, in
> implementable detail, what `SPEC.md` §22.4a states at the contract level.
> §22.4a **item 2** (the type-driven entry noun / terminology) shipped in the S6
> core (v0.45.0) and is out of scope here.
## 0. Why this is its own pass
`SPEC.md` §22.4a says a collection's immutable `type` selects exactly three
things: (1) the entry **frontmatter schema**, (2) the **terminology**, and (3)
the **type-specific surfaces**. The S6 core shipped (2) and merged the contract
into `SPEC.md`; it also shipped the type *plumbing*`collections.type` is
immutable, validated (`registry.VALID_TYPES = {document, specification, bdd}`),
and drives the entry noun. What it did **not** ship is any *behavior* keyed on
type beyond the noun: every type today parses the same §2 baseline frontmatter
(`backend/app/entry.py` is type-agnostic and lenient about unknown keys) and
renders the same §7 catalog with no type-specific surface.
Items 1 and 3 were deliberately deferred at v0.45.0 (CHANGELOG: "they want a
discovery/spec pass first, lacking BDD scenarios in Part C"). The three-tier
design doc's Part C scenarios are all about **roles** (C.1C.3); there are no
scenarios describing what a `specification` release-planning view *does* or what
a `bdd` coverage view *shows*. This document supplies them.
The governing constraint from §22.4a, which every proposal below honors:
> Type does not change the **engine** — every type uses the same content repo
> (§22.3), the same propose→branch→PR→discuss→graduate lifecycle (§§913), the
> same threads, flags, and chat. … the engine itself treats every entry as
> markdown + frontmatter regardless of type.
So a per-type surface is **additive and read-mostly**: it reads the (now
type-aware) frontmatter and presents a derived view. It never forks the write
path, the PR lifecycle, or the storage model.
---
# Part A — Item 1: the per-type frontmatter schema
## A.1 Where validation lives today, and where it should land
`entry.py:parse()` reads a fixed set of §2 baseline fields and is **lenient**:
unknown keys are ignored, future fields ride along untouched (its own docstring
says so). That leniency is the seam. The per-type schema is layered as a
**validator**, not a parser rewrite:
- `entry.py` keeps parsing the union of all known fields into the `Entry`
dataclass (add the new optional fields below; absent → `None`/default, exactly
as `models`/`funder`/`unreviewed` already do). The parser stays type-agnostic.
- A new **`entry_schema.py`** module exposes `validate(entry, collection_type)
-> list[str]` returning human-readable problems (empty = valid). It is the one
place that knows which fields a type **requires**, which it **forbids**, and
the **enum/shape** of each.
- Validation is **advisory at parse, enforced at the write boundary.** The
propose/edit/PR-merge paths (§9.1, §22.4b) call `entry_schema.validate` and
surface problems the way the propose modal already surfaces field errors. A
malformed historical file still *parses* (we never hard-fail a read — a
deployment's existing corpus must keep loading), but the catalog flags it
(§A.4) and the next write must fix it.
This mirrors how visibility/initial_state are validated centrally in
`registry.py` rather than at each call site.
## A.2 `document` — unchanged (the §2 baseline)
`document` is the baseline: the §2 fields exactly as today
(`slug, title, state, id, repo, proposed_by, proposed_at, owners, arbiters,
tags`, plus the §6.6/§6.7 `models`/`funder` and §22.4c `unreviewed`/`reviewed_*`).
No new fields, no type-specific surface. The §22.13 generated default collection
is `document`, so **N=1 deployments see zero change** — the load-bearing
backcompat guarantee.
## A.3 `specification` — versioned-spec metadata
A `specification` entry is a versioned technical spec (the archetype is this
framework's own `SPEC.md`). Frontmatter **adds** (all optional at parse,
required/validated per A.1 at write):
| Field | Shape | Meaning | Required when |
|---|---|---|---|
| `spec_version` | semver string (`MAJOR.MINOR.PATCH`) | the entry's own version | `state = active` |
| `lifecycle` | enum `draft \| active \| superseded` | spec lifecycle, **orthogonal to** the §2.4 entry `state` | always (defaults `draft`) |
| `supersedes` | list of slugs (in this collection) | specs this one replaces | optional |
Notes / decisions:
- **`lifecycle``state`.** The §2.4 `state` (super-draft/active/withdrawn) is
the *engine's* workflow position; `lifecycle` is the *spec's* editorial status.
An `active` (graduated) entry can be `lifecycle: draft` (published but not yet
ratified) or `superseded`. Keeping them orthogonal avoids overloading the
shared state machine (the §22.4a "engine unchanged" rule).
- **`supersedes` is validated as in-collection slugs** (§22.14 §2: slugs are
unique *per collection*). A `superseded` lifecycle with no inbound
`supersedes` from a newer entry is a soft warning in the surface, not a write
error (the replacement may land later).
- `spec_version` uses the same semver vocabulary as the framework `VERSION`/§20
so the release surface (A.5 / Part B) can sort and group.
## A.4 `bdd` — feature/scenario metadata
A `bdd` entry states a feature as Given/When/Then scenarios. Frontmatter
**adds**:
| Field | Shape | Meaning | Required when |
|---|---|---|---|
| `feature` | string | the feature's one-line statement (the "In order to / As a / I want" intent) | `state = active` |
| `verifies` | list of refs | the `specification` entries/sections this feature exercises | optional |
| `scenarios` | derived, **not** frontmatter | count/list parsed from the body's `Scenario:` blocks | n/a |
Notes / decisions:
- **`verifies` is a cross-collection ref.** A ref is `"<collection>/<slug>"` or
`"<collection>/<slug>#<anchor>"`. The default `<collection>` is a sibling
`specification` collection in the same project; an unqualified `<slug>` means
"a spec slug in this project's specification collection" (resolved at render).
This is the one place a `bdd` surface reaches across collections — and §22's
"no app surface joins across collections" rule (`SPEC.md` line 5016) is
**honored**: `verifies` is a *declared link rendered as a hyperlink*, not a
query that fuses two corpora. The coverage view (B.2) aggregates these links
but each entry still lives in exactly one collection.
- **Scenarios are parsed from the body, not frontmatter.** Gherkin-style
`Scenario:` / `Given`/`When`/`Then` lines in the markdown body are the source
of truth; the surface counts and lists them. This keeps the authoring
experience plain-markdown (the engine's invariant) — no structured
scenario-editor write path.
- `bdd` collections default `initial_state: active` (§22.4b) so a feature lands
active-but-`unreviewed`; the schema validator therefore requires `feature` for
active entries, which is every freshly-landed `bdd` entry.
## A.5 Schema surfacing in the existing chrome
Item-1 work is mostly invisible plumbing, but two small surfaces make it real
without waiting for Part B:
1. **Propose/edit validation** — the propose modal and edit-branch flow run
`entry_schema.validate` for the collection's type and block submit on errors
(e.g. proposing into a `specification` collection without a `lifecycle`).
2. **A "malformed frontmatter" catalog flag** — the §7 catalog marks entries
whose stored frontmatter fails its type's schema (parallel to the §22.4c
`unreviewed` filter), so a corpus migrated from `document`→… or hand-edited
is visibly fixable.
---
# Part B — Item 3: the type-specific surfaces
A surface is an **additional view** layered on the shared §7 catalog +
§8 entry view, selected on `collection.type`. It is read-derived from
frontmatter + body; it adds no write path the engine doesn't already have.
## B.1 `specification` → the release-planning surface
§22.4a: "group entries/changes into versioned releases with a changelog +
§20-style upgrade-steps per release." Concretely, a per-collection
**Releases** view at `/p/<project>/c/<collection>/releases`:
- **A release** is a named, ordered version (e.g. `0.46.0`) with: the set of
spec entries at a given `spec_version`/`lifecycle`, a changelog body, and an
optional upgrade-steps block (the §20.4 RFC-2119 convention reused verbatim).
- **Source of truth = the content repo**, per §22.2/§22.3. A release is a file
in the collection's subfolder (proposal: `releases/<version>.md`,
frontmatter `version` + `released_at` + `entries: [slug@spec_version, …]`,
body = changelog + upgrade-steps). The registry mirror caches a `releases`
table the way it caches `collections` — git is truth, the table is a cache
(§22.2 "never written except by the mirror").
- **The view** lists releases newest-first; each expands to its changelog +
upgrade-steps and the entries it cut. An Owner (scope-role, §22.6) can cut a
new release (a bot-committed file, exactly like create-collection commits a
manifest — §22 S5 pattern); contributors read.
- **Reuse, don't reinvent:** the changelog + upgrade-steps renderer is the same
markdown the framework's own `CHANGELOG.md`/§20.4 uses; the "cut a release"
write is the §22 S5 bot-commit-then-mirror pattern.
Deliberately **out of this surface** (deferred): cross-release diffing, automated
version bumping, dependency graphs between specs. The MVP is "see the releases,
their changelog, their upgrade-steps, and what each contained."
## B.2 `bdd` → the scenario/acceptance + coverage surfaces
§22.4a: "a scenario/acceptance view and a coverage view mapping features to the
spec sections they exercise." Two read-derived views:
1. **Scenario/acceptance view** (per entry, on the §8 entry page): renders the
body's parsed `Scenario:` blocks as a structured checklist — each scenario's
Given/When/Then, plus the entry's `feature` line as the header. No new
storage; pure body parse. This is the `bdd` analogue of the `document`
entry's prose view.
2. **Coverage view** (per collection, at
`/p/<project>/c/<collection>/coverage`): a matrix of **features → the spec
entries/sections they `verifies`**. Rows are this collection's `bdd` entries;
columns (or grouped rows) are the referenced `specification` entries. Cells
show "covered / declared-but-spec-missing / spec-section-with-no-feature".
The view aggregates the `verifies` links (A.4) across the collection but
renders each as a hyperlink into the spec collection — it does not fuse the
corpora (the §22 cross-collection rule, B/A.4).
Deliberately **out of this surface** (deferred): executing scenarios, CI/test
result ingestion, auto-detecting coverage from code. The MVP maps *declared*
coverage (`verifies`), surfacing gaps for humans to close.
## B.3 How a surface is selected and routed
- The collection payload already carries `type` and `entry_noun`
(`GET /api/projects/:id/collections/:cid`). The frontend's `ProjectLayout` /
collection chrome reads `type` and mounts the type's surface routes
(`releases` for `specification`; `coverage` for `bdd`) alongside the shared
catalog. `document` mounts none.
- Backend: a per-type router group (`api_releases.py`, `api_coverage.py`)
guarded by the same §22.5 read gates as the rest of the collection; the
release-cut write reuses `auth.is_collection_superuser` / the S5 bot pattern.
- The "per-type module the framework selects on `collection.type`" (§22.4a) is
realized as: backend `entry_schema.py` (item 1) + the two router groups
(item 3), and frontend a `typeModules[type]` map of `{ schema, surfaces }`.
Adding a future type = a new map entry + enum value, no rebuild (§22.4a "open
set").
---
# Part C — Behavioral scenarios (BDD)
> Tagged for the proposed slices in Part D (`@S7a` = item 1 schemas; `@S7b` =
> specification releases; `@S7c` = bdd surfaces). These are the acceptance gate
> the implementing session writes tests against, in the Part C style of the
> three-tier doc.
## C.1 Per-type frontmatter schema (`@S7a`)
```gherkin
Scenario: document collection is unchanged
Given a "document" collection
When a contributor proposes an entry with the §2 baseline frontmatter only
Then the proposal is accepted with no schema error
Scenario: specification entry requires a lifecycle
Given a "specification" collection
When a contributor proposes an entry with no `lifecycle`
Then it defaults to lifecycle "draft" and is accepted
And when an Owner graduates it to active without a `spec_version`
Then the write is blocked with "spec_version is required for an active specification"
Scenario: bdd entry requires a feature statement once active
Given a "bdd" collection whose initial_state is "active"
When a contributor proposes an entry with no `feature`
Then the write is blocked with "feature is required for a bdd entry"
Scenario: unknown future field still rides along
Given any collection
When an entry carries a frontmatter key no schema names
Then it parses unchanged and is not reported as malformed
Scenario: malformed existing entry loads but is flagged
Given a stored specification entry missing a required field
When the catalog renders
Then the entry still loads (read never hard-fails)
And the catalog marks it "malformed frontmatter"
```
## C.2 specification release-planning surface (`@S7b`)
```gherkin
Scenario: an Owner cuts a release
Given a "specification" collection with two active entries
When the collection Owner cuts release "1.0.0" with a changelog and upgrade-steps
Then a releases/1.0.0.md file is committed to the content repo
And the registry mirror caches the release
And the Releases view lists "1.0.0" newest-first with its changelog + upgrade-steps
Scenario: a contributor reads releases but cannot cut one
Given a contributor (not Owner) in the collection
Then the Releases view is read-only (no "Cut release" control)
Scenario: upgrade-steps render with the §20.4 convention
Given a release whose body uses MUST/SHOULD/MAY upgrade-steps
Then they render with the same normative-language styling as CHANGELOG.md
```
## C.3 bdd scenario + coverage surfaces (`@S7c`)
```gherkin
Scenario: an entry's scenarios render as an acceptance checklist
Given a "bdd" entry whose body has two Scenario: blocks
When the entry page renders
Then it shows the `feature` header and both scenarios' Given/When/Then
Scenario: coverage maps features to the specs they verify
Given a "bdd" entry that `verifies: ["spec/auth#sessions"]`
And a sibling "specification" collection "spec" containing entry "auth"
When the coverage view renders
Then a row links the feature to spec/auth#sessions as covered
Scenario: a declared ref to a missing spec is surfaced as a gap
Given a "bdd" entry that `verifies: ["spec/ghost"]` where no such spec exists
Then the coverage view marks that ref "declared but spec missing"
Scenario: coverage does not fuse corpora
Then each cell is a hyperlink into the spec collection
And no entry from the spec collection is listed as if it belonged to the bdd collection
```
---
# Part D — Delivery slicing
Each slice lands a usable increment + a runnable acceptance gate (`--tags @S7x`),
in the three-tier doc's slicing style. Suggested order (item 1 first — the
surfaces read its fields):
- **S7a — per-type frontmatter schema (item 1).** `entry_schema.py` +
the new optional `Entry` fields + write-boundary validation + the catalog
"malformed" flag. **Usable:** proposing into a typed collection is validated;
N=1 `document` unchanged. **Completes:** `@S7a` (C.1). *Non-breaking, additive.*
- **S7b — specification release planning (item 3a).** `releases/<v>.md` storage
+ registry mirror + `api_releases.py` + the Releases view + cut-release write.
**Usable:** a spec collection has versioned releases with changelog +
upgrade-steps. **Completes:** `@S7b` (C.2).
- **S7c — bdd scenario + coverage surfaces (item 3b).** Body scenario parser +
the per-entry acceptance view + the per-collection coverage view +
`api_coverage.py`. **Usable:** a bdd collection shows scenarios and declared
coverage. **Completes:** `@S7c` (C.3).
All three are **additive** (new optional frontmatter, new tables that are pure
caches, new read views): each is a minor, non-breaking release, and a `document`
N=1 deployment is unaffected by any of them. None touches the engine, the PR
lifecycle, or the role model — they consume the §22 three-tier + §22.6 role
work already shipped.
## D.1 Open questions for the implementing session
1. **Release identity vs. entry `spec_version`.** Should a release's `entries`
pin exact `slug@spec_version` (immutable snapshot) or just slugs (live)? This
doc proposes the pinned snapshot; confirm against a real spec-collection
workflow before building S7b.
2. **`verifies` ref grammar.** `"<collection>/<slug>#<anchor>"` is proposed;
anchor resolution into a spec entry's section needs the spec body to carry
stable anchors. May want a lightweight `## §n` anchor convention on
`specification` entries first.
3. **Whether `lifecycle` belongs in the shared state machine after all.** Kept
orthogonal here; revisit if product wants `superseded` to gate the catalog.
These are genuine product decisions a discovery/spec session (or the operator)
should settle before S7b/S7c code; S7a (schemas) is unblocked and buildable now.
@@ -0,0 +1,121 @@
# Deployed-environment E2E harness (PPE)
**Date:** 2026-06-07 · **Version:** v0.52.0 · **Status:** implemented
## Why
The §9 deployment pipeline is `localhost + E2E → PPE + E2E → prod`. The
middle stage — running the Playwright E2E suite against a *deployed*
pre-prod host (`https://rfc-ppe.wiggleverse.org`) — was unreachable
because the suite (`e2e/metadata.spec.js`, SLICE-3/4/5 of the
configurable-collection-metadata work) was bound to three scaffolds that
exist only in the local Tier-1 docker stack:
1. **A faceted `bdd` collection**, seeded into a throwaway Gitea by
`testing/seed-gitea.sh`.
2. **A granted-owner identity** (`e2e-owner@example.test`), injected
directly into SQLite by the docker-compose `backend-seed` step.
3. **Mailpit**, the SMTP sink the OTC sign-in reads the one-time code
from.
PPE has none of these: it runs against the real `git.wiggleverse.org`
(shared with prod), has no direct DB access, and has no mail sink. This
note records how each coupling is replaced so the *same* spec runs green
against both localhost and PPE.
## The three seams
### 1. Auth — a gated test-login endpoint (the framework change)
A new backend route, `POST /auth/test/login`, replaces both the Mailpit
OTC dance *and* the SQLite owner injection with one gesture: it mints an
authenticated **owner** session for a single pre-configured identity.
It is the framework's only auth bypass, so it is **fail-closed** and must
never function in production:
- **Off by default.** It returns `404` unless **both**
`E2E_TEST_AUTH_SECRET` and `E2E_TEST_AUTH_EMAIL` are set. A production
deployment sets neither, so the route is invisible and inert.
- **Secret-gated.** The caller must present `E2E_TEST_AUTH_SECRET` in the
`X-Test-Auth-Secret` header, compared in constant time. A wrong/absent
secret returns `404` (it does not advertise the route's existence).
- **Single identity.** It will only mint the one configured
`E2E_TEST_AUTH_EMAIL` (case-insensitive); any other address is `403`.
So an enabled PPE exposes exactly one throwaway owner, with the secret
as the trust boundary.
- **Loud at startup.** When enabled, the app logs a `WARNING` at boot, so
an accidental prod enablement is visible rather than silent.
On success it provision-or-links the row (reusing `otc.provision_or_link_user`),
forces it to `role='owner', permission_state='granted'` (the deployed
equivalent of the Tier-1 owner-seed), and stores the session exactly like
the OTC verify path.
**Why an endpoint rather than alternatives.** Reading the OTC code from
the VM's journald (the email adapter logs the envelope to stdout when
SMTP is unconfigured) would couple the test harness to `gcloud` SSH at
runtime — slow, brittle, and operator-cred-bound. Running Mailpit on the
VM and exposing its API publicly is more infra and its own exposure
surface. A default-off, secret-gated endpoint is the portable engineering
seam: it works for *any* deployed environment, needs no SSH, and the
secrets rule (§6.3) is honored — the secret is a Secret Manager ref
injected as VM env, never a literal.
The hard-secrets caveat: the E2E runner presents the secret by resolving
it from Secret Manager at runtime (command substitution), never echoing
it.
### 2. Content — a dedicated PPE registry + content repo
PPE shares the prod Gitea org (`wiggleverse`) and, until now, prod's
registry (`rfc-registry`) and default project (`ohm`). Seeding a faceted
test collection into that shared registry would surface it on **prod**.
So PPE gets its **own**, prod-untouching fixtures:
- `wiggleverse/rfc-registry-ppe` — PPE's project registry. Prod keeps
`rfc-registry`, so prod is never affected.
- `wiggleverse/rfc-app-ppe-content` — one project `ohm` (document) with a
default collection entry plus a faceted `bdd` named collection
(`priority` enum + `tags`) and three entries, mirroring the Tier-1
seed. The E2E path `/p/ohm/c/bdd` therefore resolves identically on
both environments.
PPE is pointed at it with `overlay set rfc-app-ppe
REGISTRY_REPO=rfc-registry-ppe`. The startup reconciler sweep loads the
content into `cached_rfcs` (incl. `meta_json` for facets) on the next
deploy — no webhook needed for the initial load. The seed is scripted in
`testing/seed-ppe.sh` (idempotent; `RESEED=1` restores entry values for a
re-run). Repo *creation* is a one-time operator gesture (the
`write:repository` Keychain token cannot create org repos; create the two
empty repos in the Gitea UI or re-scope the PAT).
### 3. Parameterization — one spec, two environments
- `e2e/playwright.config.js` already honors `BASE_URL`
(default `http://localhost:8080`); PPE sets
`BASE_URL=https://rfc-ppe.wiggleverse.org`.
- `e2e/lib/auth.js` branches on `E2E_TEST_AUTH_SECRET`: set → use
`/auth/test/login`; unset → the original Mailpit OTC path. `OWNER_EMAIL`
reads `E2E_OWNER_EMAIL` (PPE points it at `E2E_TEST_AUTH_EMAIL`) or the
Tier-1 default. The spec itself is unchanged, so the localhost Tier-1
path keeps working.
## PPE version
The harness *requires* the test-login endpoint to exist in the deployed
build, so PPE must run a framework version that contains it — **v0.52.0**,
not v0.51.1. PPE is pinned ahead of prod via its own
`ben/ohm-rfc/.rfc-app-version.ppe` (prod stays on `.rfc-app-version`),
realizing the "PPE stages newer versions first" note the
`deployment.ppe.toml` always anticipated.
## Known limitations
- **Re-runnability.** SLICE-4/5 mutate the seeded entries (commit
sidecars). A clean run needs seed-state preconditions; re-run after
`RESEED=1` + a cache refresh (next reconciler sweep or a redeploy).
Unlike Tier-1's `make e2e-fresh`, PPE has no per-run teardown.
- **smoke.spec.js** stays Tier-1-only (anonymous OTC smoke through
Mailpit); only `metadata.spec.js` runs against PPE.
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,108 @@
# SLICE-1 plan — sidecar storage + dual-read + migration + malformed flag
Just-in-time implementation plan for **SLICE-1** of
[Configurable Collection Metadata](../2026-06-06-configurable-collection-metadata.md)
(§7.2). Authored at the start of the SLICE-1 coding session (session 0084,
2026-06-07), against the code SLICE-0 (v0.46.2) landed.
## Scope (and non-scope)
**In:** the storage/compat layer only — sidecar files become the source of
truth for entry metadata, with a dual-read parser, an idempotent
frontmatter→sidecar migration tool, and a derived `metadata_malformed` flag.
**Out (later slices):** the `.collection.yaml` `fields:` schema + validation
(SLICE-2), faceted filtering (SLICE-3), the edit/bulk UIs (SLICE-4/5). SLICE-1
maps sidecar values onto the **existing** typed `cached_rfcs` columns; it does
not add per-field schema columns or facet aggregation.
## Invariants honored
- **INV-6 dual-read:** parser reads the sidecar if present, else legacy
top-of-doc frontmatter, with identical resulting in-memory records.
- **INV-7 unknown keys ride along:** preserved through parse→serialize and
through the migration (never dropped).
- **INV-2:** a migrated `.md` body contains no metadata.
- **INV-1:** `cached_rfcs` stays a derived, rebuildable index; the sidecar in
git is the source of truth.
- **INV-3:** bad metadata never hard-fails a read — the entry still loads and
the catalog flags it (`metadata_malformed`).
- **INV-5 / byte-identity:** a collection with no sidecars behaves exactly as
today (legacy frontmatter path); existing entries load identically.
## Components
1. **`app/entry.py` — unknown-key preservation (INV-7).** Add
`extra: dict[str, Any]` to `Entry`. `parse()` collects frontmatter keys
outside the known set into `extra`; `serialize()` re-emits them after the
known keys. Makes frontmatter round-trips lossless.
2. **`app/metadata.py` — new module (sidecar concerns).**
- `SIDECAR_SUFFIX = ".meta.yaml"`; `sidecar_name(slug)`,
`is_sidecar(name)`, `slug_of_sidecar(name)`.
- `metadata_dict(entry) -> dict` — the full metadata mapping (known
emit-rules + `extra`), shared by the sidecar writer and the frontmatter
serializer.
- `sidecar_yaml(entry) -> str` — canonical YAML for a sidecar from
`metadata_dict`.
- `strip_frontmatter(md_text) -> str` — body-only (drops a leading
`---…---` block if present; whole text otherwise).
- `parse_sidecar(text) -> tuple[dict, bool]` — lenient: `(values, malformed)`;
non-mapping / YAML error → `({}, True)`.
- `read_entry(md_text, sidecar_text|None) -> tuple[Entry, bool]` — dual-read:
sidecar present → metadata from sidecar values, body from
`strip_frontmatter(md_text)`, `malformed` from `parse_sidecar`; absent →
`entry.parse(md_text)`, `malformed=False`.
3. **`app/gitea.py``change_files(...)` batch commit.** `POST
/repos/{owner}/{repo}/contents` (Gitea ChangeFiles) with a `files[]` array
of `{operation, path, content(b64), sha?}` — one commit for N files. Backs
the migration's "one commit per collection".
4. **`metadata.migrate_collection(gitea, org, repo, subfolder, actor)`.**
Lists `<subfolder>/rfcs`; for each `<slug>.md` **without** a `<slug>.meta.yaml`
sibling and **with** legacy frontmatter, batch: create the sidecar
(`metadata_dict` → YAML) + update the `.md` to body-only. One ChangeFiles
commit per collection. Idempotent (skip entries already migrated; no-op when
none remain). Returns a summary (`migrated`, `skipped`, `committed`).
> **Deferred (decided mid-slice, after code review):** the **operator
> trigger** for this tool (an Owner-gated endpoint) is held back to SLICE-4.
> The propose/graduate/mark-reviewed/edit write paths still `entry.parse` the
> `.md` directly, so migrating a corpus to body-only `.md`s before those
> paths are sidecar-aware would break them (crash / re-introduce
> frontmatter). SLICE-1 ships the tool as tested groundwork; SLICE-4 makes
> the write paths sidecar-aware (and adds lazy migration) and is where the
> trigger belongs (INV-8: engine write paths unchanged this slice).
5. **`app/cache.py` — dual-read in `_refresh_collection_corpus`.** Build a
sidecar-by-stem map from the dir listing; for each `.md`, read its sidecar
sibling (if any), `metadata.read_entry(...)`, thread `malformed` into
`_upsert_cached_rfc(metadata_malformed=…)`.
6. **`backend/migrations/033_metadata_malformed.sql`** — additive
`ALTER TABLE cached_rfcs ADD COLUMN metadata_malformed INTEGER NOT NULL DEFAULT 0`.
7. **`app/api.py` — surface the flag.** Add `metadata_malformed` (bool) to the
two catalog list dicts and `get_rfc`/`_get_rfc_for_collection`. (Frontend
badge + `?malformed=` filter are SLICE-3.)
## Tests (TDD — write first)
- `test_metadata.py` (unit, pure): dual-read equivalence (sidecar vs legacy →
identical Entry); unknown-key preservation through parse→serialize and
through `metadata_dict`; `strip_frontmatter` (with/without frontmatter);
`parse_sidecar` malformed cases.
- `test_metadata_migration.py` (integration, FakeGitea): migrate a collection
→ sidecars written + `.md` bodies stripped + one commit; **idempotent**
(second run is a no-op); unknown keys preserved in the sidecar.
- extend the cache/propose vertical: a collection with a sidecar mirrors from
the sidecar; a malformed sidecar sets `metadata_malformed` and still loads
the entry (INV-3); a no-sidecar collection is byte-identical to today.
## Release
Minor bump **0.46.2 → 0.47.0** (new functionality: sidecar storage + migration
tool; non-breaking — additive migration 033, dual-read keeps legacy corpora
working, opt-in). §20 CHANGELOG + upgrade-steps: migration 033 auto-applies;
running the migration tool per collection is optional (**MAY**).
@@ -0,0 +1,109 @@
# Implementation plan — SLICE-2: collection field schema + central validation
**Slice:** SLICE-2 of
[`docs/design/2026-06-06-configurable-collection-metadata.md`](../2026-06-06-configurable-collection-metadata.md)
§7.2. **Session:** OHM-0085. **Branch:** `worktree-metadata-slice2-schema`.
## Goal / Definition of Done (from the design)
- `.collection.yaml` `fields:` block is parsed and stored.
- `metadata_schema.validate` exists — **advisory at read**, the enforcement
point **at write** (write endpoints land in SLICE-4/5; this slice supplies and
read-wires the function).
- The schema is served via the collection API (`GET …/collections/{id}`).
- A collection with **no `fields:`** behaves exactly as today (INV-5) — the
§22.13 default `document` collection sees zero change.
## Design decisions (this slice)
- **Field-def shape** (design §6.3): `fields:` is an ordered mapping
`{name → {type, values?, label?}}`. `type ∈ {enum, tags, text}` (v1).
Order is preserved for facet display (SLICE-3).
- `enum` — single scalar; requires a non-empty `values:` list.
- `tags` — list; `values:` optional (controlled when present, free-form
otherwise).
- `text` — free string scalar.
- **`ref` and `multi-enum` are out of v1** (design §2 future; Q2 leans "later").
Unknown field types are **ignored with a warning** (design §6.3), not fatal.
- **Lenient schema parsing.** A malformed `fields:` block or an individual bad
field def is skipped with a warning, never raised — a typo in one field must
not nuke the collection mirror (INV-3 spirit). Structural manifest errors
(`type`, `visibility`) keep raising `RegistryError` as before.
- **No DB migration.** The normalized `fields` schema rides in the existing
`collections.config_json` column, exactly like `enabled_models`.
- **`values` source for validation.** An entry's full metadata mapping
(`metadata.metadata_dict(entry)` = `to_frontmatter_dict`) — `tags` come from
the known `Entry.tags`; custom fields (e.g. `priority`) come from
`Entry.extra`. Undeclared keys are forward-compat (INV-7) and never flagged.
## Tasks
### 1. `app/metadata_schema.py` (new) — the one place that knows field shapes
- `VALID_FIELD_TYPES = {"enum", "tags", "text"}`.
- `parse_fields(raw) -> dict[str, dict]` — lenient/normalizing. Returns an
ordered mapping of `{name: {"type", "values"?, "label"?}}`. Skips: non-mapping
block; non-mapping field def; unknown/missing type; `enum` without a non-empty
`values:` list. Logs a warning per skip. Pure (no I/O).
- `@dataclass Problem(field, code, message)` + `as_dict()`.
- `validate(values: dict, fields: dict[str, dict]) -> list[Problem]` — for each
**declared** field, validate the entry's value (absent is OK):
- `enum`: scalar ∈ `values:` else `not-in-values`; a list/non-scalar →
`wrong-type`.
- `tags`: must be a list (`wrong-type` otherwise); if controlled, each member
`values:` else `not-in-values`.
- `text`: must be a scalar string (`wrong-type` otherwise).
Undeclared keys ignored (INV-7). Never raises.
- **Tests** `tests/test_metadata_schema.py`: parse (each type, missing block,
enum-without-values skipped, unknown type skipped, order preserved); validate
(happy, enum bad value, tags uncontrolled-ok, tags controlled-bad, text wrong
type, undeclared key ignored, empty schema → no problems).
### 2. `registry.py` — parse `fields:` into the collection config
- In `parse_collection_manifest`, after the existing keys: if `raw.get("fields")`
present, `cfg["fields"] = metadata_schema.parse_fields(raw["fields"])` (only
set when non-empty). Flows into `CollectionEntry.config`
`config_json` via the existing `json.dumps(ce.config)` in
`_upsert_named_collection`.
- **Default collection** (`apply_registry`) currently writes the `collections`
row **without** `config_json`. A default collection's `fields:` would live in
`projects.yaml`? No — the design says `fields:` is a `.collection.yaml` block.
The default collection has no `.collection.yaml`. **Decision:** default-
collection field schemas are out of this slice's happy path (the N=1 default is
`document` with no fields, INV-5). Leave `apply_registry` untouched; only
named collections (with a `.collection.yaml`) carry `fields:`. Documented as a
known limitation (a deployment wanting fields on its primary corpus declares a
named collection — consistent with the design's opt-in story).
- **Tests** extend `tests/test_collection_registry.py`: a manifest with a
`fields:` block round-trips into `get_collection(...)["fields"]`; a manifest
with a bad field def still upserts (lenient).
### 3. `collections.py` — unpack `fields` on read + serve via API
- Add `_fields_from_config(config_json) -> dict | None` (mirror
`_enabled_models_from_config`).
- `get_collection` sets `out["fields"] = _fields_from_config(config_json)`
(alongside `enabled_models`). `api_collections.get_col` then serves it with no
change. `list_collections` left as-is (facets are SLICE-3).
- **Tests** extend `tests/test_collection_helpers.py`: `get_collection` exposes
`fields`; a no-fields collection → `fields is None` (INV-5).
### 4. `cache.py` — advisory validation at ingest (INV-3)
- In `_refresh_collection_corpus`, fetch the collection's `fields` schema once
(`collections.get_collection(collection_id)`); for each entry, if a schema is
present, `problems = metadata_schema.validate(metadata.metadata_dict(entry),
fields)` and OR any problems into `metadata_malformed` (warn-log a summary).
No schema → behavior identical to today (INV-5).
- **Tests** `tests/test_metadata_cache.py` (extend): an entry violating an enum
field ingests with `metadata_malformed = 1`; a conforming entry → `0`; a
collection with no schema → `0` regardless of extra keys.
### 5. Verify · version · ship
- `pytest` full backend suite green (575 baseline + new).
- Bump `VERSION` + `frontend/package.json`**0.48.0**; CHANGELOG minor entry
(§20) — non-breaking, opt-in, N=1 unchanged; note write-enforcement lands with
SLICE-4/5.
- Commit (cite design §7.2 SLICE-2), PR on Gitea `origin`, merge to `main`.
## Invariants honored
INV-3 (read never hard-fails — advisory malformed), INV-5 (no-`fields:`
unchanged), INV-7 (undeclared keys ride along, never flagged), INV-8 (additive,
read-mostly; no write-path fork — write enforcement is a later slice).

Some files were not shown because too many files have changed in this diff Show More