Files
rfc-app/docs/design/2026-06-06-configurable-collection-metadata.md
T
Ben Stull e86fc65643 Discovery spec: configurable collection metadata (clean-doc tagging)
Solution Design for a generic per-collection metadata system — typed fields
(enum/tags/text) declared in .collection.yaml, per-entry values in a clean
<slug>.meta.yaml sidecar, schema-derived faceted left-pane filtering, and
single/bulk direct-commit tag/untag. Restores the retired BDD Release Planner's
corpus-annotation half inside the framework; release planning stays downstream.

Output of discovery session OHM-0079.0. Proposed as §23; flag for reconciliation
with §22 S6 type-modules at the SPEC merge.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 01:52:26 -07:00

24 KiB
Raw Blame History

Solution Design: Configurable Collection Metadata (clean-doc tagging)

Author(s) Ben Stull
Reviewers / approvers Ben Stull
Status draft
Version v0.1.0
Source artifacts Reference modeled: retired BDD Release Planner (wiggleverse/wiggleverse-ecomm-bdd-release-planner-app, RETIRED 2026-06-04) · Related spec: docs/design/2026-06-05-three-tier-projects-collections.md (§22) · Corpus: ecomm Shopify-modeled BDD (wiggleverse-ecomm-meta/research/shopify, ~1,238 scenarios) · Supersedes: —

Change log

Date Version Change By
2026-06-06 v0.1.0 Initial draft from discovery session OHM-0079.0 Ben Stull

1. Executive Summary

rfc-app gains a generic, per-collection-configurable metadata system. A collection declares its fields — priority, tags, and arbitrary custom fields — in its .collection.yaml; each entry's values live in a sidecar (<slug>.meta.yaml) so the document body stays pure prose. rfc-app renders schema-driven forms and faceted left-pane filters, and supports fast single + bulk tag/untag via direct commits for authorized roles. This restores the corpus-annotation capability of the retired BDD Release Planner inside the framework, while keeping release planning (ordering, ship status, roadmap emission) a downstream concern that reads the sidecars straight from git.

2. Business Context

The framework hosts RFC standardization for multiple deployments. One deployment hosts the ecomm BDD corpus — ~1,238 Shopify-modeled scenarios, one markdown file per scenario, slugged by feature ID (DD-FF-NNNN-slug). A standalone BDD Release Planner previously let operators search that corpus, attach metadata (priority P0P3, owner, status), cluster scenarios into named releases, and emit each release as a roadmap phase. §22 (three-tier projects/collections) absorbed the planner's corpus hosting into rfc-app (the corpus now runs as a bdd project on the RFC deployment) and the planner was retired — but its annotation half (priority/tags on scenarios, filtering, bulk assignment) was never rebuilt. Operators planning work, and downstream tools consuming the corpus, currently have no structured, framework-native way to set or read that metadata.

3. Problem Statement

Today rfc-app metadata is weak and intrusive:

  • Tags are free-form strings with no filtering — §7.1 specifies a Tag: filter chip that was never implemented.
  • No priority, and no way for a collection to declare any other structured field.
  • No per-collection schema — every collection gets the same fixed frontmatter shape; a bdd collection cannot say "scenarios have a P0P3 priority."
  • Metadata pollutes the document body — a large YAML frontmatter block sits at the top of every doc, mixing rfc-app lifecycle bookkeeping with content metadata.
  • No bulk workflow — annotating hundreds of scenarios one PR at a time is impractical, so the planner's core gesture has no analog.

Result: corpus authors can't meaningfully prioritise/tag, downstream consumers have nothing structured to read, and documents aren't clean.

4. Stakeholders / Personas / Actors

Actor Type Goal in this design
Corpus contributor persona Set priority and tags on a scenario when proposing/curating it
Release planner (operator) persona Bulk-prioritise/tag many scenarios quickly to shape a plan
Collection owner operator Declare the metadata fields a collection supports
Downstream consumer system Read structured per-scenario metadata from the git corpus (e.g. an external release planner)
Reader persona Read clean scenario docs; filter the catalog by priority/tag

5. Scope

  • In scope: per-collection field schema in .collection.yaml (enum, tags, text); per-entry sidecar (<slug>.meta.yaml) as source of truth with pure-prose doc bodies; schema-derived faceted left-pane filtering with counts; schema-derived metadata form (detail panel); single + bulk tag/untag with direct commit for authorized roles; dual-read compatibility + a one-shot frontmatter→sidecar migration tool + lazy migration on write.
  • Out of scope: release ordering, ship status, roadmap emission, RELEASE-PLAN.md generation — these stay downstream, consuming sidecars from git. In-app management of field definitions / controlled vocabularies; corpus-wide tag rename/merge/delete. Sub-document (per-scenario-within-a-file) grain. A whole-corpus metadata export endpoint.
  • Non-goals: rebuilding a bespoke "release" entity in rfc-app — releases are not modeled here at all; metadata is the only primitive.

6. Assumptions · Constraints · Dependencies

  • Assumptions: git remains the content source of truth and downstream consumers can read the corpus from git; the BDD grain is one markdown file per scenario (already true for the ecomm corpus), so no sub-document parsing is needed.
  • Constraints: rfc-app is a framework hosting multiple deployments — the upgrade must be mechanical and non-breaking, with §20 changelog/upgrade-steps; the hard secrets rule (§6.3) holds; metadata edits must respect scope-role authorization (§22 Part B / S3).
  • Dependencies: the S3 scope-role resolver (auth.effective_scope_role) for edit authorization; the existing git write-through path used by edit-meta (§9.5); the §22 collection model (.collection.yaml, cached_rfcs keyed by (collection_id, slug)).

7. Targeted Business Outcomes

Outcome Success metric Baseline → Target Guardrail (must not regress) How / when measured
Authors can prioritise/tag scenarios scenarios carrying a priority 0 → corpus-wide doc bodies stay prose-clean corpus inspection after rollout
Corpus is filterable left-pane facet filters available none → priority+tags+state filter latency acceptable at ~1.2k entries manual + perf check
Downstream tools can consume metadata sidecars readable from git none → all migrated entries sidecar schema stable consumer integration
Clean docs docs with no rfc-app frontmatter 0% → 100% (post-migration) dual-read keeps legacy working post-migration inspection

8. Business Use Cases

Scenario: BUC-1 — A contributor prioritises a scenario
  Given a bdd collection whose schema defines a priority field
  When a contributor marks a scenario as P0
  Then the scenario's metadata records priority P0
  And the document body is unchanged prose

Scenario: BUC-2 — A planner batch-prioritises a set of scenarios
  Given a contributor has selected 40 scenarios
  When they set priority P1 on the selection
  Then all 40 carry priority P1
  And the change lands as a single auditable commit

Scenario: BUC-3 — A downstream tool consumes priorities
  Given scenarios carry priority and tags in their sidecars
  When an external release planner reads the corpus from git
  Then it can group and order scenarios using that metadata
  And rfc-app did not need to model releases at all

Scenario: BUC-4 — Documents stay clean
  Given a migrated collection
  When a reader opens a scenario document
  Then they see only prose, with no rfc-app metadata block
  • BUC-1 acceptance: the scenario's priority value is P0 and its .md body is byte-identical to before.
  • BUC-2 acceptance: 40 sidecars updated; exactly one commit; authorship recorded.
  • BUC-3 acceptance: sidecars + .collection.yaml are sufficient for a consumer to read priority/tags without rfc-app's API.
  • BUC-4 acceptance: no --- frontmatter remains in migrated docs.

9. Product Use Cases

Scenario: PUC-1 — Set priority/tags on a scenario (realizes BUC-1)
  Given I am a contributor viewing a scenario whose collection defines priority and tags
  When I choose P0 in the priority control and add the tag "checkout"
  Then the metadata panel reflects P0 and the checkout tag
  And the change is committed directly to the scenario's sidecar

Scenario: PUC-2 — Bulk tag/untag from the catalog (realizes BUC-2)
  Given I have multi-selected several scenarios in the catalog
  When I choose "Set priority → P1" from the bulk action bar
  Then every selected scenario shows P1
  And the bulk change is one commit

Scenario: PUC-3 — Filter the catalog by facet (realizes BUC-1/BUC-3)
  Given the left pane shows faceted filters generated from the collection schema
  When I check Priority P0 and tag "checkout"
  Then the catalog shows only scenarios matching both
  And each facet value shows its result count

Scenario: PUC-4 — A collection declares its fields (product-only)
  Given a collection owner edits .collection.yaml to add a priority enum field
  When the collection is re-ingested
  Then the priority filter and the priority form control appear automatically

Scenario: PUC-5 — Migrate a collection to clean docs (product-only)
  Given a collection whose docs still carry top-of-doc frontmatter
  When the operator runs the frontmatter→sidecar migration
  Then each doc body becomes pure prose and a sidecar holds its metadata
  And rfc-app reads the collection identically before and after

10. UX Layout

10.1 Screen: Catalog (left pane) (serves PUC-3)

  • Purpose: browse and filter a collection's entries.
  • Layout (top → bottom):
    • Search: full-text box (existing).
    • Faceted filter groups (one per schema field + state): each is a collapsible group showing per-value result counts and multi-select checkboxes; tags-type fields include a "filter values…" search box to stay usable at 30+ values. (Chosen layout: faceted groups with counts — validated in brainstorming over flat chips.)
  • States: happy: facets with counts · empty: "no entries match these filters" with a clear-filters action · loading: skeleton facets · error: "couldn't load facets" with retry.

10.2 Screen: Scenario detail — metadata panel (serves PUC-1)

  • Purpose: view/edit one entry's metadata.
  • Layout: a panel rendering one control per schema field — enum → single-select; tags → removable chips + add-tag input (with existing AI suggest where applicable); text → text input. The document body renders below as pure prose; metadata never appears inline in the body.
  • States: read (role without edit) shows values, no controls · edit (authorized) shows controls · saving: inline spinner · error: field-level validation message (e.g. "P5 is not an allowed priority").

10.3 Screen: Catalog — bulk action bar (serves PUC-2)

  • Purpose: apply a field value to many entries at once.
  • Layout: selecting ≥1 row reveals a sticky action bar: "N selected · Set priority ▾ · Add tag ▾ · Remove tag ▾ · Clear". Each action targets one field; applying commits once.
  • States: none selected: bar hidden · applying: bar shows progress · partial failure: toast naming entries that failed validation, others applied.

11. Technical Design

11.1 Invariants

  • INV-1: The sidecar (<slug>.meta.yaml) is the source of truth for entry metadata; cached_rfcs is a derived index, fully rebuildable from git.
  • INV-2: A document body (.md) never contains rfc-app metadata once migrated; metadata lives only in the sidecar.
  • INV-3: Reading a collection never hard-fails on bad metadata — an invalid value against the schema surfaces as a warning and the entry still loads.
  • INV-4: Metadata writes are authorized by scope-role (contributor+ on the collection); content-body edits keep their existing PR-review path.
  • INV-5: A collection with no fields: block behaves exactly as today (free-form tags only) — the feature is additive and opt-in per collection.
  • INV-6: Dual-read holds throughout: parser reads the sidecar if present, else legacy top-of-doc frontmatter, with identical resulting in-memory records.

11.2 High-level architecture

flowchart LR
  subgraph Git[content repo]
    CY[.collection.yaml<br/>fields: schema]
    MD[slug.md<br/>prose body]
    SC[slug.meta.yaml<br/>values]
  end
  CY --> ING[ingest / parser<br/>validate vs schema]
  MD --> ING
  SC --> ING
  ING --> DB[(cached_rfcs<br/>values + facet counts)]
  DB --> API[API: schema · list+filter · facets · edit]
  API --> FILT[left-pane faceted filters]
  API --> PANEL[detail metadata panel]
  API --> BULK[bulk select bar]
  PANEL -->|direct commit| SC
  BULK -->|1 commit| SC
  SC -.read from git.-> CONS[downstream consumers]
  • ingest/parser — owns reading .collection.yaml schema + sidecars (or legacy frontmatter), validating values, and rebuilding cached_rfcs; must never treat the DB as authoritative.
  • API — owns serving the schema, filtered lists with facet counts, and metadata edits; must never write metadata anywhere but the sidecar in git.

11.3 Data model & ownership

Entity Owned by Key fields System of record
Collection field schema collection owner fields: {name → {type, values?, label}} in .collection.yaml git
Entry metadata values contributor sidecar <slug>.meta.yaml: lifecycle (slug,title,state,owners,…) + schema fields (priority,tags,…) git (sidecar)
Derived index ingest per-entry values + facet aggregations cached_rfcs (SQLite, derived)

Field types (v1): enum (single-select; controlled by required values:), tags (multi-value; free-form unless values: given), text (free string). Unknown types are ignored with a warning (forward-compat).

Sidecar example:

slug: 01-01-0001-view-today-s-key-performance-metrics-at-a-glance
title: View today's key performance metrics at a glance
state: active
owners: [ben.stull]
priority: P1
tags: [dashboard, analytics]
owner: hasan

11.4 Interfaces & contracts

  • GET /api/projects/<p>/collections/<c> — out: collection incl. fields schema. Errors: 404.
  • GET /api/projects/<p>/collections/<c>/rfcs — in: filter params (?priority=P0&tags=checkout&state=active, repeatable for multi-value/OR-within-field, AND across fields) · out: entries with metadata values + facets: {field → {value → count}}. Errors: 400 on unknown field.
  • POST /api/projects/<p>/collections/<c>/rfcs/<slug>/meta — in: {field: value, …} · out: updated values · effect: validate vs schema → write sidecar → commit directly → re-ingest entry. Errors: 403 (role), 422 (invalid value).
  • POST /api/projects/<p>/collections/<c>/meta/bulk — in: {slugs: […], op: set|add|remove, field, value} · out: {applied: […], rejected: [{slug, reason}]} · effect: validate → write N sidecars → one commit → re-ingest. Errors: 403, 422.

11.5 PerProduct-Use-Case design

PUC-2 — Bulk tag/untag

sequenceDiagram
  actor U as Contributor
  participant C as Catalog UI
  participant A as API
  participant G as Git
  participant D as cached_rfcs
  U->>C: select rows, "Set priority P1"
  C->>A: POST /meta/bulk {slugs, set, priority, P1}
  A->>A: authz (scope-role) + validate vs schema
  A->>G: write N sidecars, 1 commit
  A->>D: re-ingest affected entries
  A-->>C: {applied, rejected}
  C-->>U: rows show P1; toast on any rejected
  • Implementation: reuse the edit-meta git write-through, extended to (a) target the sidecar rather than frontmatter and (b) batch N files into one commit. Honors INV-1/INV-4.

PUC-5 — Migration

  • Implementation: a tool (framework CLI verb / tools/ script) walks a collection, and for each entry with legacy frontmatter, writes <slug>.meta.yaml from the frontmatter and rewrites <slug>.md to the body only — one commit per collection, idempotent. Dual-read (INV-6) means this can run anytime; lazy migration converts stragglers on their first metadata edit.

11.6 Non-functional requirements & cross-cutting concerns

  • Security & privacy: metadata edits gated by auth.effective_scope_role (contributor+ on the collection); no secrets in sidecars; git history records authorship.
  • Performance & scale: facet counts computed from the derived DB; must stay responsive at ~1.2k entries with dozens of tag values (indexed value columns / aggregation query).
  • Availability & resilience: bad metadata never blocks read (INV-3); a failed re-ingest leaves git authoritative and is recoverable by full rebuild.
  • Observability: log each metadata commit (actor, field, entry count); warn-log schema validation failures encountered on ingest.
  • Accessibility: facet groups and form controls keyboard-navigable; checkboxes labelled with value + count.

11.7 Key decisions & alternatives considered

Decision Chosen Alternatives Why
Release modeling Metadata only; releases downstream First-class release entity in rfc-app; release-typed tags with behavior Operator pulled ordering/ship-status out of rfc-app; metadata is the only needed primitive
Tag system shape One generic typed-field system (Approach A) Releases first-class + simple tags; namespaced facets Tags/priority/custom are all just fields; one mechanism
Metadata storage Sidecar per entry Top-of-doc frontmatter (today); end-of-doc block; collection index file; DB-only Clean docs + git-visible to consumers + locality per scenario
Left-pane filtering Faceted groups with counts Flat facet chips Scales to the ~1.2k-scenario, many-tag ecomm corpus
Edit governance Direct commit for authorized roles (bulk = 1 commit) PR per change Bulk planning is impractical via PR-per-toggle
Mgmt UI Deferred; edit .collection.yaml in git In-app field/vocab management in v1 Smallest coherent v1

11.8 Testing strategy

Unit tests for: schema parsing (fields: block, all types, missing block); sidecar read/write round-trip; dual-read equivalence (frontmatter vs sidecar produce identical records); schema validation (reject bad enum, accept free-form tag); facet aggregation; bulk op (set/add/remove, single commit, partial-rejection). Two-tier local-Docker→PPE for the API + git write-through. "Tested" = the PUC acceptance scenarios pass plus the migration is proven idempotent and reversible-on-read.

11.9 Failure modes, rollback & flags

  • Failure mode: invalid value committed out-of-band → on ingest, warn + load entry with the raw value flagged (INV-3); not surfaced as a filter facet count error.
  • Failure mode: re-ingest fails after commit → git is authoritative; full rebuild recovers.
  • Migration rollback: dual-read means an un-migrated or partially-migrated corpus still works; the migration commit is revertible.
  • Feature flag: the feature is inherently opt-in per collection (INV-5) — no global flag needed; absent a fields: block, behavior is unchanged.

12. Delivery Plan

12.1 Approach / strategy

Build the storage/compat foundation first (sidecars + dual-read + migration) so nothing breaks, then the schema, then read (filtering), then write (single, bulk). Each slice is shippable and non-breaking.

12.2 Slicing plan

SLICE-1 — Sidecar storage + dual-read + migration → completes PUC-5

  • Depends on:
  • DoD: parser reads sidecar-if-present else legacy frontmatter (INV-6); migration tool splits frontmatter→sidecar idempotently; existing collections load byte-identically; tests green.

SLICE-2 — Collection field schema + validation → completes PUC-4

  • Depends on: SLICE-1
  • DoD: .collection.yaml fields: parsed (enum/tags/text); values validated on read (warn) and on write (reject); schema served via the collection API; no-fields: collections unchanged (INV-5).

SLICE-3 — Faceted left-pane filtering (read) → completes PUC-3

  • Depends on: SLICE-2
  • DoD: list endpoint returns facet counts + honors filter params; left pane renders faceted groups with counts and tag-value search; filters compose (AND across fields).

SLICE-4 — Single-entry metadata edit → completes PUC-1

  • Depends on: SLICE-2
  • DoD: detail metadata panel renders schema controls; POST …/meta validates, direct-commits the sidecar, re-ingests; authorized by scope-role (INV-4); lazy-migrates a legacy entry on first edit.

SLICE-5 — Bulk tag/untag → completes PUC-2

  • Depends on: SLICE-3, SLICE-4
  • DoD: multi-select + bulk action bar; POST …/meta/bulk applies set/add/remove as one commit; partial-rejection reported.

12.3 Rollout / launch plan

Pre-v1, single production: ship slices in order to the RFC deployment; each minor-version bump carries §20 changelog + upgrade steps. The opt-in-per-collection nature (INV-5) means a deployment adopts it only when it declares a fields: block and (optionally) runs the migration.

12.4 Risks & mitigations

Risk L/I Mitigation
Frontmatter→sidecar migration corrupts content L/H Dual-read; idempotent, revertible migration; body-byte-identity test
Doubling file count (sidecars) clutters corpus M/L Docs stay clean; sidecars are small/co-located; acceptable for one-file-per-scenario corpora
Direct-commit metadata edits bypass review M/M Scope-role gate (INV-4); content-body edits still PR'd; full git audit trail
Overlap/conflict with §22 S6 "type modules" M/M Position as §23, generalizing S6's per-type frontmatter; reconcile at the S6 SPEC merge
Facet aggregation slow at scale L/M Compute from indexed derived DB; measure at ~1.2k entries

13. Traceability matrix

Business UC Product UC Slice Tests
BUC-4 PUC-5 SLICE-1 test_dual_read_equiv, test_migration_idempotent
— (product-only) PUC-4 SLICE-2 test_schema_parse, test_validation
BUC-1/BUC-3 PUC-3 SLICE-3 test_facet_counts, test_filter_compose
BUC-1 PUC-1 SLICE-4 test_single_meta_commit, test_authz
BUC-2 PUC-2 SLICE-5 test_bulk_one_commit, test_partial_reject
BUC-3 (consumer reads git) test_sidecar_schema_stable

14. Open Questions & Decisions log

Open

# Question Owner Blocks
Q1 Do downstream consumers read sidecars from git, via rfc-app API, or both? (leaning git) Ben nothing v1
Q2 Should a multi-enum type (multi-select controlled) ship in v1 or later? Ben SLICE-2 scope
Q3 Exact §23 placement / reconciliation with §22 S6 type-modules Ben SPEC merge

Resolved

# Decision Resolution Date
D1 Release behaviors (ordering, ship status, roadmap emit) Out of rfc-app; downstream 2026-06-06
D2 Tag system shape Approach A — one generic typed-field system 2026-06-06
D3 Metadata grain Per entry (corpus already one file per scenario) 2026-06-06
D4 Schema location .collection.yaml fields: block 2026-06-06
D5 Value storage Sidecar <slug>.meta.yaml; doc body pure prose 2026-06-06
D6 Left-pane filtering Faceted groups with counts 2026-06-06
D7 Edit governance Direct commit for authorized roles; bulk = 1 commit 2026-06-06
D8 Management scope Deferred; edit .collection.yaml in git for v1 2026-06-06

15. Glossary & References

  • Sidecar<slug>.meta.yaml, the per-entry metadata file that is the source of truth; keeps the .md body pure prose.
  • Field schema — the fields: block in .collection.yaml declaring a collection's typed metadata fields.
  • Facet — a schema field surfaced as a left-pane filter group with per-value counts.
  • Downstream consumer — an external tool (e.g. a release planner) that reads corpus metadata from git; rfc-app does not model releases.
  • References: retired BDD Release Planner (wiggleverse-ecomm-bdd-release-planner-app); §22 three-tier design (docs/design/2026-06-05-three-tier-projects-collections.md); SPEC §7.1 (left-pane filter), §9.5 (edit-meta), §20 (versioning), §22 Part B / S3 (scope-role).