Solution Design for a generic per-collection metadata system — typed fields (enum/tags/text) declared in .collection.yaml, per-entry values in a clean <slug>.meta.yaml sidecar, schema-derived faceted left-pane filtering, and single/bulk direct-commit tag/untag. Restores the retired BDD Release Planner's corpus-annotation half inside the framework; release planning stays downstream. Output of discovery session OHM-0079.0. Proposed as §23; flag for reconciliation with §22 S6 type-modules at the SPEC merge. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
24 KiB
Solution Design: Configurable Collection Metadata (clean-doc tagging)
| Author(s) | Ben Stull |
| Reviewers / approvers | Ben Stull |
| Status | draft |
| Version | v0.1.0 |
| Source artifacts | Reference modeled: retired BDD Release Planner (wiggleverse/wiggleverse-ecomm-bdd-release-planner-app, RETIRED 2026-06-04) · Related spec: docs/design/2026-06-05-three-tier-projects-collections.md (§22) · Corpus: ecomm Shopify-modeled BDD (wiggleverse-ecomm-meta/research/shopify, ~1,238 scenarios) · Supersedes: — |
Change log
| Date | Version | Change | By |
|---|---|---|---|
| 2026-06-06 | v0.1.0 | Initial draft from discovery session OHM-0079.0 | Ben Stull |
1. Executive Summary
rfc-app gains a generic, per-collection-configurable metadata system. A collection declares its fields — priority, tags, and arbitrary custom fields — in its .collection.yaml; each entry's values live in a sidecar (<slug>.meta.yaml) so the document body stays pure prose. rfc-app renders schema-driven forms and faceted left-pane filters, and supports fast single + bulk tag/untag via direct commits for authorized roles. This restores the corpus-annotation capability of the retired BDD Release Planner inside the framework, while keeping release planning (ordering, ship status, roadmap emission) a downstream concern that reads the sidecars straight from git.
2. Business Context
The framework hosts RFC standardization for multiple deployments. One deployment hosts the ecomm BDD corpus — ~1,238 Shopify-modeled scenarios, one markdown file per scenario, slugged by feature ID (DD-FF-NNNN-slug). A standalone BDD Release Planner previously let operators search that corpus, attach metadata (priority P0–P3, owner, status), cluster scenarios into named releases, and emit each release as a roadmap phase. §22 (three-tier projects/collections) absorbed the planner's corpus hosting into rfc-app (the corpus now runs as a bdd project on the RFC deployment) and the planner was retired — but its annotation half (priority/tags on scenarios, filtering, bulk assignment) was never rebuilt. Operators planning work, and downstream tools consuming the corpus, currently have no structured, framework-native way to set or read that metadata.
3. Problem Statement
Today rfc-app metadata is weak and intrusive:
- Tags are free-form strings with no filtering — §7.1 specifies a
Tag:filter chip that was never implemented. - No priority, and no way for a collection to declare any other structured field.
- No per-collection schema — every collection gets the same fixed frontmatter shape; a
bddcollection cannot say "scenarios have a P0–P3 priority." - Metadata pollutes the document body — a large YAML frontmatter block sits at the top of every doc, mixing rfc-app lifecycle bookkeeping with content metadata.
- No bulk workflow — annotating hundreds of scenarios one PR at a time is impractical, so the planner's core gesture has no analog.
Result: corpus authors can't meaningfully prioritise/tag, downstream consumers have nothing structured to read, and documents aren't clean.
4. Stakeholders / Personas / Actors
| Actor | Type | Goal in this design |
|---|---|---|
| Corpus contributor | persona | Set priority and tags on a scenario when proposing/curating it |
| Release planner (operator) | persona | Bulk-prioritise/tag many scenarios quickly to shape a plan |
| Collection owner | operator | Declare the metadata fields a collection supports |
| Downstream consumer | system | Read structured per-scenario metadata from the git corpus (e.g. an external release planner) |
| Reader | persona | Read clean scenario docs; filter the catalog by priority/tag |
5. Scope
- In scope: per-collection field schema in
.collection.yaml(enum,tags,text); per-entry sidecar (<slug>.meta.yaml) as source of truth with pure-prose doc bodies; schema-derived faceted left-pane filtering with counts; schema-derived metadata form (detail panel); single + bulk tag/untag with direct commit for authorized roles; dual-read compatibility + a one-shot frontmatter→sidecar migration tool + lazy migration on write. - Out of scope: release ordering, ship status, roadmap emission,
RELEASE-PLAN.mdgeneration — these stay downstream, consuming sidecars from git. In-app management of field definitions / controlled vocabularies; corpus-wide tag rename/merge/delete. Sub-document (per-scenario-within-a-file) grain. A whole-corpus metadata export endpoint. - Non-goals: rebuilding a bespoke "release" entity in rfc-app — releases are not modeled here at all; metadata is the only primitive.
6. Assumptions · Constraints · Dependencies
- Assumptions: git remains the content source of truth and downstream consumers can read the corpus from git; the BDD grain is one markdown file per scenario (already true for the ecomm corpus), so no sub-document parsing is needed.
- Constraints: rfc-app is a framework hosting multiple deployments — the upgrade must be mechanical and non-breaking, with §20 changelog/upgrade-steps; the hard secrets rule (§6.3) holds; metadata edits must respect scope-role authorization (§22 Part B / S3).
- Dependencies: the S3 scope-role resolver (
auth.effective_scope_role) for edit authorization; the existing git write-through path used byedit-meta(§9.5); the §22 collection model (.collection.yaml,cached_rfcskeyed by(collection_id, slug)).
7. Targeted Business Outcomes
| Outcome | Success metric | Baseline → Target | Guardrail (must not regress) | How / when measured |
|---|---|---|---|---|
| Authors can prioritise/tag scenarios | scenarios carrying a priority | 0 → corpus-wide | doc bodies stay prose-clean | corpus inspection after rollout |
| Corpus is filterable | left-pane facet filters available | none → priority+tags+state | filter latency acceptable at ~1.2k entries | manual + perf check |
| Downstream tools can consume metadata | sidecars readable from git | none → all migrated entries | sidecar schema stable | consumer integration |
| Clean docs | docs with no rfc-app frontmatter | 0% → 100% (post-migration) | dual-read keeps legacy working | post-migration inspection |
8. Business Use Cases
Scenario: BUC-1 — A contributor prioritises a scenario
Given a bdd collection whose schema defines a priority field
When a contributor marks a scenario as P0
Then the scenario's metadata records priority P0
And the document body is unchanged prose
Scenario: BUC-2 — A planner batch-prioritises a set of scenarios
Given a contributor has selected 40 scenarios
When they set priority P1 on the selection
Then all 40 carry priority P1
And the change lands as a single auditable commit
Scenario: BUC-3 — A downstream tool consumes priorities
Given scenarios carry priority and tags in their sidecars
When an external release planner reads the corpus from git
Then it can group and order scenarios using that metadata
And rfc-app did not need to model releases at all
Scenario: BUC-4 — Documents stay clean
Given a migrated collection
When a reader opens a scenario document
Then they see only prose, with no rfc-app metadata block
- BUC-1 acceptance: the scenario's
priorityvalue isP0and its.mdbody is byte-identical to before. - BUC-2 acceptance: 40 sidecars updated; exactly one commit; authorship recorded.
- BUC-3 acceptance: sidecars +
.collection.yamlare sufficient for a consumer to read priority/tags without rfc-app's API. - BUC-4 acceptance: no
---frontmatter remains in migrated docs.
9. Product Use Cases
Scenario: PUC-1 — Set priority/tags on a scenario (realizes BUC-1)
Given I am a contributor viewing a scenario whose collection defines priority and tags
When I choose P0 in the priority control and add the tag "checkout"
Then the metadata panel reflects P0 and the checkout tag
And the change is committed directly to the scenario's sidecar
Scenario: PUC-2 — Bulk tag/untag from the catalog (realizes BUC-2)
Given I have multi-selected several scenarios in the catalog
When I choose "Set priority → P1" from the bulk action bar
Then every selected scenario shows P1
And the bulk change is one commit
Scenario: PUC-3 — Filter the catalog by facet (realizes BUC-1/BUC-3)
Given the left pane shows faceted filters generated from the collection schema
When I check Priority P0 and tag "checkout"
Then the catalog shows only scenarios matching both
And each facet value shows its result count
Scenario: PUC-4 — A collection declares its fields (product-only)
Given a collection owner edits .collection.yaml to add a priority enum field
When the collection is re-ingested
Then the priority filter and the priority form control appear automatically
Scenario: PUC-5 — Migrate a collection to clean docs (product-only)
Given a collection whose docs still carry top-of-doc frontmatter
When the operator runs the frontmatter→sidecar migration
Then each doc body becomes pure prose and a sidecar holds its metadata
And rfc-app reads the collection identically before and after
10. UX Layout
10.1 Screen: Catalog (left pane) (serves PUC-3)
- Purpose: browse and filter a collection's entries.
- Layout (top → bottom):
- Search: full-text box (existing).
- Faceted filter groups (one per schema field + state): each is a collapsible group showing per-value result counts and multi-select checkboxes;
tags-type fields include a "filter values…" search box to stay usable at 30+ values. (Chosen layout: faceted groups with counts — validated in brainstorming over flat chips.)
- States: happy: facets with counts · empty: "no entries match these filters" with a clear-filters action · loading: skeleton facets · error: "couldn't load facets" with retry.
10.2 Screen: Scenario detail — metadata panel (serves PUC-1)
- Purpose: view/edit one entry's metadata.
- Layout: a panel rendering one control per schema field —
enum→ single-select;tags→ removable chips + add-tag input (with existing AI suggest where applicable);text→ text input. The document body renders below as pure prose; metadata never appears inline in the body. - States: read (role without edit) shows values, no controls · edit (authorized) shows controls · saving: inline spinner · error: field-level validation message (e.g. "P5 is not an allowed priority").
10.3 Screen: Catalog — bulk action bar (serves PUC-2)
- Purpose: apply a field value to many entries at once.
- Layout: selecting ≥1 row reveals a sticky action bar: "N selected · Set priority ▾ · Add tag ▾ · Remove tag ▾ · Clear". Each action targets one field; applying commits once.
- States: none selected: bar hidden · applying: bar shows progress · partial failure: toast naming entries that failed validation, others applied.
11. Technical Design
11.1 Invariants
- INV-1: The sidecar (
<slug>.meta.yaml) is the source of truth for entry metadata;cached_rfcsis a derived index, fully rebuildable from git. - INV-2: A document body (
.md) never contains rfc-app metadata once migrated; metadata lives only in the sidecar. - INV-3: Reading a collection never hard-fails on bad metadata — an invalid value against the schema surfaces as a warning and the entry still loads.
- INV-4: Metadata writes are authorized by scope-role (contributor+ on the collection); content-body edits keep their existing PR-review path.
- INV-5: A collection with no
fields:block behaves exactly as today (free-formtagsonly) — the feature is additive and opt-in per collection. - INV-6: Dual-read holds throughout: parser reads the sidecar if present, else legacy top-of-doc frontmatter, with identical resulting in-memory records.
11.2 High-level architecture
flowchart LR
subgraph Git[content repo]
CY[.collection.yaml<br/>fields: schema]
MD[slug.md<br/>prose body]
SC[slug.meta.yaml<br/>values]
end
CY --> ING[ingest / parser<br/>validate vs schema]
MD --> ING
SC --> ING
ING --> DB[(cached_rfcs<br/>values + facet counts)]
DB --> API[API: schema · list+filter · facets · edit]
API --> FILT[left-pane faceted filters]
API --> PANEL[detail metadata panel]
API --> BULK[bulk select bar]
PANEL -->|direct commit| SC
BULK -->|1 commit| SC
SC -.read from git.-> CONS[downstream consumers]
- ingest/parser — owns reading
.collection.yamlschema + sidecars (or legacy frontmatter), validating values, and rebuildingcached_rfcs; must never treat the DB as authoritative. - API — owns serving the schema, filtered lists with facet counts, and metadata edits; must never write metadata anywhere but the sidecar in git.
11.3 Data model & ownership
| Entity | Owned by | Key fields | System of record |
|---|---|---|---|
| Collection field schema | collection owner | fields: {name → {type, values?, label}} in .collection.yaml |
git |
| Entry metadata values | contributor | sidecar <slug>.meta.yaml: lifecycle (slug,title,state,owners,…) + schema fields (priority,tags,…) |
git (sidecar) |
| Derived index | ingest | per-entry values + facet aggregations | cached_rfcs (SQLite, derived) |
Field types (v1): enum (single-select; controlled by required values:), tags (multi-value; free-form unless values: given), text (free string). Unknown types are ignored with a warning (forward-compat).
Sidecar example:
slug: 01-01-0001-view-today-s-key-performance-metrics-at-a-glance
title: View today's key performance metrics at a glance
state: active
owners: [ben.stull]
priority: P1
tags: [dashboard, analytics]
owner: hasan
11.4 Interfaces & contracts
GET /api/projects/<p>/collections/<c>— out: collection incl.fieldsschema. Errors: 404.GET /api/projects/<p>/collections/<c>/rfcs— in: filter params (?priority=P0&tags=checkout&state=active, repeatable for multi-value/OR-within-field, AND across fields) · out: entries with metadata values +facets: {field → {value → count}}. Errors: 400 on unknown field.POST /api/projects/<p>/collections/<c>/rfcs/<slug>/meta— in:{field: value, …}· out: updated values · effect: validate vs schema → write sidecar → commit directly → re-ingest entry. Errors: 403 (role), 422 (invalid value).POST /api/projects/<p>/collections/<c>/meta/bulk— in:{slugs: […], op: set|add|remove, field, value}· out:{applied: […], rejected: [{slug, reason}]}· effect: validate → write N sidecars → one commit → re-ingest. Errors: 403, 422.
11.5 Per–Product-Use-Case design
PUC-2 — Bulk tag/untag
sequenceDiagram
actor U as Contributor
participant C as Catalog UI
participant A as API
participant G as Git
participant D as cached_rfcs
U->>C: select rows, "Set priority P1"
C->>A: POST /meta/bulk {slugs, set, priority, P1}
A->>A: authz (scope-role) + validate vs schema
A->>G: write N sidecars, 1 commit
A->>D: re-ingest affected entries
A-->>C: {applied, rejected}
C-->>U: rows show P1; toast on any rejected
- Implementation: reuse the
edit-metagit write-through, extended to (a) target the sidecar rather than frontmatter and (b) batch N files into one commit. Honors INV-1/INV-4.
PUC-5 — Migration
- Implementation: a tool (framework CLI verb /
tools/script) walks a collection, and for each entry with legacy frontmatter, writes<slug>.meta.yamlfrom the frontmatter and rewrites<slug>.mdto the body only — one commit per collection, idempotent. Dual-read (INV-6) means this can run anytime; lazy migration converts stragglers on their first metadata edit.
11.6 Non-functional requirements & cross-cutting concerns
- Security & privacy: metadata edits gated by
auth.effective_scope_role(contributor+ on the collection); no secrets in sidecars; git history records authorship. - Performance & scale: facet counts computed from the derived DB; must stay responsive at ~1.2k entries with dozens of tag values (indexed value columns / aggregation query).
- Availability & resilience: bad metadata never blocks read (INV-3); a failed re-ingest leaves git authoritative and is recoverable by full rebuild.
- Observability: log each metadata commit (actor, field, entry count); warn-log schema validation failures encountered on ingest.
- Accessibility: facet groups and form controls keyboard-navigable; checkboxes labelled with value + count.
11.7 Key decisions & alternatives considered
| Decision | Chosen | Alternatives | Why |
|---|---|---|---|
| Release modeling | Metadata only; releases downstream | First-class release entity in rfc-app; release-typed tags with behavior | Operator pulled ordering/ship-status out of rfc-app; metadata is the only needed primitive |
| Tag system shape | One generic typed-field system (Approach A) | Releases first-class + simple tags; namespaced facets | Tags/priority/custom are all just fields; one mechanism |
| Metadata storage | Sidecar per entry | Top-of-doc frontmatter (today); end-of-doc block; collection index file; DB-only | Clean docs + git-visible to consumers + locality per scenario |
| Left-pane filtering | Faceted groups with counts | Flat facet chips | Scales to the ~1.2k-scenario, many-tag ecomm corpus |
| Edit governance | Direct commit for authorized roles (bulk = 1 commit) | PR per change | Bulk planning is impractical via PR-per-toggle |
| Mgmt UI | Deferred; edit .collection.yaml in git |
In-app field/vocab management in v1 | Smallest coherent v1 |
11.8 Testing strategy
Unit tests for: schema parsing (fields: block, all types, missing block); sidecar read/write round-trip; dual-read equivalence (frontmatter vs sidecar produce identical records); schema validation (reject bad enum, accept free-form tag); facet aggregation; bulk op (set/add/remove, single commit, partial-rejection). Two-tier local-Docker→PPE for the API + git write-through. "Tested" = the PUC acceptance scenarios pass plus the migration is proven idempotent and reversible-on-read.
11.9 Failure modes, rollback & flags
- Failure mode: invalid value committed out-of-band → on ingest, warn + load entry with the raw value flagged (INV-3); not surfaced as a filter facet count error.
- Failure mode: re-ingest fails after commit → git is authoritative; full rebuild recovers.
- Migration rollback: dual-read means an un-migrated or partially-migrated corpus still works; the migration commit is revertible.
- Feature flag: the feature is inherently opt-in per collection (INV-5) — no global flag needed; absent a
fields:block, behavior is unchanged.
12. Delivery Plan
12.1 Approach / strategy
Build the storage/compat foundation first (sidecars + dual-read + migration) so nothing breaks, then the schema, then read (filtering), then write (single, bulk). Each slice is shippable and non-breaking.
12.2 Slicing plan
SLICE-1 — Sidecar storage + dual-read + migration → completes PUC-5
- Depends on: —
- DoD: parser reads sidecar-if-present else legacy frontmatter (INV-6); migration tool splits frontmatter→sidecar idempotently; existing collections load byte-identically; tests green.
SLICE-2 — Collection field schema + validation → completes PUC-4
- Depends on: SLICE-1
- DoD:
.collection.yaml fields:parsed (enum/tags/text); values validated on read (warn) and on write (reject); schema served via the collection API; no-fields:collections unchanged (INV-5).
SLICE-3 — Faceted left-pane filtering (read) → completes PUC-3
- Depends on: SLICE-2
- DoD: list endpoint returns facet counts + honors filter params; left pane renders faceted groups with counts and tag-value search; filters compose (AND across fields).
SLICE-4 — Single-entry metadata edit → completes PUC-1
- Depends on: SLICE-2
- DoD: detail metadata panel renders schema controls;
POST …/metavalidates, direct-commits the sidecar, re-ingests; authorized by scope-role (INV-4); lazy-migrates a legacy entry on first edit.
SLICE-5 — Bulk tag/untag → completes PUC-2
- Depends on: SLICE-3, SLICE-4
- DoD: multi-select + bulk action bar;
POST …/meta/bulkapplies set/add/remove as one commit; partial-rejection reported.
12.3 Rollout / launch plan
Pre-v1, single production: ship slices in order to the RFC deployment; each minor-version bump carries §20 changelog + upgrade steps. The opt-in-per-collection nature (INV-5) means a deployment adopts it only when it declares a fields: block and (optionally) runs the migration.
12.4 Risks & mitigations
| Risk | L/I | Mitigation |
|---|---|---|
| Frontmatter→sidecar migration corrupts content | L/H | Dual-read; idempotent, revertible migration; body-byte-identity test |
| Doubling file count (sidecars) clutters corpus | M/L | Docs stay clean; sidecars are small/co-located; acceptable for one-file-per-scenario corpora |
| Direct-commit metadata edits bypass review | M/M | Scope-role gate (INV-4); content-body edits still PR'd; full git audit trail |
| Overlap/conflict with §22 S6 "type modules" | M/M | Position as §23, generalizing S6's per-type frontmatter; reconcile at the S6 SPEC merge |
| Facet aggregation slow at scale | L/M | Compute from indexed derived DB; measure at ~1.2k entries |
13. Traceability matrix
| Business UC | Product UC | Slice | Tests |
|---|---|---|---|
| BUC-4 | PUC-5 | SLICE-1 | test_dual_read_equiv, test_migration_idempotent |
| — (product-only) | PUC-4 | SLICE-2 | test_schema_parse, test_validation |
| BUC-1/BUC-3 | PUC-3 | SLICE-3 | test_facet_counts, test_filter_compose |
| BUC-1 | PUC-1 | SLICE-4 | test_single_meta_commit, test_authz |
| BUC-2 | PUC-2 | SLICE-5 | test_bulk_one_commit, test_partial_reject |
| BUC-3 | (consumer reads git) | — | test_sidecar_schema_stable |
14. Open Questions & Decisions log
Open
| # | Question | Owner | Blocks |
|---|---|---|---|
| Q1 | Do downstream consumers read sidecars from git, via rfc-app API, or both? (leaning git) | Ben | nothing v1 |
| Q2 | Should a multi-enum type (multi-select controlled) ship in v1 or later? |
Ben | SLICE-2 scope |
| Q3 | Exact §23 placement / reconciliation with §22 S6 type-modules | Ben | SPEC merge |
Resolved
| # | Decision | Resolution | Date |
|---|---|---|---|
| D1 | Release behaviors (ordering, ship status, roadmap emit) | Out of rfc-app; downstream | 2026-06-06 |
| D2 | Tag system shape | Approach A — one generic typed-field system | 2026-06-06 |
| D3 | Metadata grain | Per entry (corpus already one file per scenario) | 2026-06-06 |
| D4 | Schema location | .collection.yaml fields: block |
2026-06-06 |
| D5 | Value storage | Sidecar <slug>.meta.yaml; doc body pure prose |
2026-06-06 |
| D6 | Left-pane filtering | Faceted groups with counts | 2026-06-06 |
| D7 | Edit governance | Direct commit for authorized roles; bulk = 1 commit | 2026-06-06 |
| D8 | Management scope | Deferred; edit .collection.yaml in git for v1 |
2026-06-06 |
15. Glossary & References
- Sidecar —
<slug>.meta.yaml, the per-entry metadata file that is the source of truth; keeps the.mdbody pure prose. - Field schema — the
fields:block in.collection.yamldeclaring a collection's typed metadata fields. - Facet — a schema field surfaced as a left-pane filter group with per-value counts.
- Downstream consumer — an external tool (e.g. a release planner) that reads corpus metadata from git; rfc-app does not model releases.
- References: retired BDD Release Planner (
wiggleverse-ecomm-bdd-release-planner-app); §22 three-tier design (docs/design/2026-06-05-three-tier-projects-collections.md); SPEC §7.1 (left-pane filter), §9.5 (edit-meta), §20 (versioning), §22 Part B / S3 (scope-role).