Compare commits

...

4 Commits

Author SHA1 Message Date
BenStullsBets 5fae4f9005 feat(sim): replace ring buttons with a turnable "Altitude" knob
The scale ring's ⊖ out / in ⊕ buttons are replaced by a circular **Altitude**
knob (the endless rotary encoder, drawn as a dial). cosmos (highest) sits at the
top; turning clockwise descends cosmos → orbit → coast → reef → abyss and past
the deepest **wraps back to the top** — the same endless ring the server already
models (advance + wrap unchanged).

- Built dynamically from `ring.scales` (labels + ticks per scale, a needle that
  points at the current scale, an ALTITUDE caption).
- Drag to turn: the rotation accumulates and commits whole detents on release, so
  a big spin becomes one fast blended pass (reuses the proven wheel/detent path);
  scroll the knob to step; tap a label to jump the shortest way around.
- Pure frontend — no engine/API change. JS syntax-checked; live server probed
  (assets serve, dial present, abyss→cosmos wrap intact). 23 sim-API tests green.
  By-eye review deferred to the operator (no Chrome on this box).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 21:44:51 -07:00
benstull c4452396a1 Merge pull request 'tools(pipeline): hybrid track.py + simulator label author mode' (#17) from feature/content-pipeline-track-author into main 2026-06-25 01:14:59 +00:00
BenStullsBets 70fd367c70 tools(pipeline): hybrid track.py + simulator label author mode
Content-pipeline Increment 2, part 2 (the tooling). Implements design §11.5
(docs/superpowers/specs/2026-06-24-content-pipeline-design.md):

- tools/pipeline/track.py — the stage-5 geometry pass. Classical path: a
  hand-seeded normalized box propagated by OpenCV Lucas-Kanade optical flow
  (CSRT used instead when a contrib build provides it), sampled to a SPARSE
  loop-normalized keyframed track + an appear/disappear window. Optional ML
  detect+track path lazy-imports ultralytics (clear ImportError if absent;
  base install needs only cv2). Pure helpers (normalize/denormalize, loop_t,
  infer_window, sample_track, track_to_annotation) are unit-tested; the cv2
  propagation is an opt-in integration test on real abyss_wow footage.
  Semantics are never produced here — geometry only.

- Author mode — /author.html + author.js/.css reuse the preview stage: pick a
  pool clip, scrub, drag a seed box, Run tracker (or Add as static box), author
  the LEFT detail tiers (general -> scientific+fact) + salience, shift-click to
  place affect anchors with RIGHT emotion tiers, and Save to manifest. Backend:
  POST /api/author/track (runs track_seed on the clip's base) + POST
  /api/author/clip (idempotent upsert via tools.pipeline.manifest — keeps media
  + provenance, replaces only authored content, reloads in place). The tracker
  propagates box geometry; all strings/scientific names/facts are hand-typed.

- Repointed the pipeline integration test off the retired forest base to the
  cosmos pool primary. USER_GUIDE simulator section brought current (pools,
  coast, 3-knob Mood, real-time dream, progressive tiers, author mode).

267 passed / 2 skipped (+ track pure + opt-in real-footage + author endpoint
tests). Author UI by-eye review deferred to the operator (no Chrome on this box).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 18:14:32 -07:00
benstull 12408e505e Merge pull request 'content(sim): rotating clip pools + progressive tracked labels & emotions' (#16) from feature/content-pipeline-increment-2 into main 2026-06-25 01:05:58 +00:00
13 changed files with 1164 additions and 62 deletions
+48 -37
View File
@@ -296,24 +296,17 @@ exists. (The earlier selection-era "curator's X-ray" view was retired when the
piece moved from *selecting* clips to *altering* them.)
**One-time setup — populate the sample footage** (look-tuning only; not shipped
content):
content). The ring uses real strict-PD footage in a **rotating pool** per scale
(`docs/content-candidate-pool.md`); regenerate the manifest + new transition
placeholders with:
python simulator/setup_sample_media.py # stage the forest neutral base
python -m simulator.bake_right_variants --scale forest # real Right variants (SD on MPS, ~11 min)
python simulator/setup_scales_media.py # the cosmos + abyss scales + ring transitions
python simulator/build_pool_manifest.py --media # write the pool manifest + coast-edge transitions
`setup_sample_media.py` stages the forest neutral base from the session-0008 POC
artifacts (`~/hef-poc/out/neutral.mp4`). `bake_right_variants.py` then produces
the **real** flow-stabilized painterly Right variants (strengths 14) for the
forest scale — SD img2img keyframes + Farneback optical-flow tweens on Apple MPS,
the temporally-calm POC algorithm, productionized (scales design §1/§4). The
keyframe-strength ramp per Right level is by-eye tunable in
`simulator/bake_right_variants.py`. `setup_scales_media.py` makes the scale
**ring** demonstrable: cheap synthetic placeholder bases for the two true-PD
scales (`cosmos` = NASA/Hubble, `abyss` = NOAA Ocean Exploration) plus the
per-edge zoom/warp transition clips — these scales carry a raw base only (no real
Right variants yet; the baker can extend to them once their real footage is
sourced). The `.mp4` binaries are gitignored.
`build_pool_manifest.py` holds the hand-authored label/affect content and emits
`simulator/sample_media/manifest.json` (the 19-clip pool baseline). The pool clips
themselves are sourced via the content pipeline (`tools/pipeline/run.py
process_clip`); the `.mp4` binaries are gitignored. (`setup_scales_media.py` is the
older one-base-per-scale generator, now superseded.)
**Run it (Docker):**
@@ -331,27 +324,45 @@ then open http://localhost:8000.
- **Content dial** — picks audio/video channel; "off" and audio-only positions go
to black walls.
- **Scale ring (endless encoder)** — `⊖ out` / `in ⊕` (or scroll the stage) walk a
*closed ring* of neutral "scales of nature" clips — cosmos → forest → abyss and
back around (diving past the smallest **wraps** to the largest). It is *relative*
(an endless encoder), distinct from the absolute 04 knobs: each step plays a
short placeholder zoom/warp **transition**, then settles on the next scale, with
the current knob alteration still applied. The current scale is named beside the
buttons (`name (i/N)`). A **fast spin** (scroll several detents at once, ≥3)
collapses to a single quick **blended pass** straight to the destination scale
instead of grinding through every transition (scales design §3); slow single
steps still chain one full transition per scale crossed.
- **Four experience knobs (04):**
- **Dark / Light** — a live runtime color grade (cool/dark ↔ warm/bright; equal
or zero = the raw footage).
- **Right (dreamlike)** — selects a discrete pre-baked, flow-stabilized restyle
variant and crossfades to it (strength 0 = raw base).
- **Left (analytical)** — a live overlay: labelled boxes from the clip's authored
annotation track, with more annotations appearing at higher levels. Text is
shaped live (the simulator analogue of the Pi's Pango/HarfBuzz path).
- **Calibration sliders** — adjust the grade/overlay gain curves live; once a look
is liked, bake the values into `DEFAULT_CALIBRATION` in `player/alteration.py`.
*closed ring* of neutral "scales of nature" — **cosmos → orbit → coast → reef →
abyss** and back around (diving past the smallest **wraps** to the largest). Each
scale is a **rotating pool** of vetted clips: landing on a scale plays a random
pool member, so the ring feels fresh each pass. It is *relative* (an endless
encoder), distinct from the absolute 04 knobs: each step plays a short
placeholder zoom/warp **transition**, then settles, with the current alteration
still applied. The current scale + chosen member is named beside the buttons
(`scale · member (i/N · pool N)`). A **fast spin** (≥3 detents at once) collapses
to one quick **blended pass** instead of grinding through every transition.
- **Three experience knobs:**
- **Mood (4 dark .. 0 .. +4 light)** — a live runtime color grade (cool/dark ↔
warm/bright; 0 = the raw footage).
- **Right (dreamlike, 04)** — a **deterministic real-time painterly dream** (a
WebGL Kuwahara restyle of the live frames; holds still across the loop, goes
trippy at max). It also drives **progressive emotion tiers**: when both knobs
are up, the affect words escalate from basic to compound as Right rises.
- **Left (analytical, 04)** — a live HUD of labelled boxes from the clip's
authored annotation track. Labels use **progressive detail tiers**: low Left
shows fewer, more general labels (high-salience objects only); raising Left
brings in more objects *and* escalates each label general → specific →
scientific → +fact. Time-windowed tracked labels appear only while their
subject is on screen and follow it.
- **RenderPlan readout** — always shows the exact numbers the engine produced (the
project's honesty "X-ray," now over the alteration model).
project's honesty "X-ray," over the alteration model).
The base clips, Right variants, Left annotation track, and string tables come from
The base clips, the Left annotation tracks (with tiers + appear/disappear
windows), affect anchors, and per-tier string tables all come from
`simulator/sample_media/manifest.json`.
### Authoring labels — the author mode
Open **http://localhost:8000/author.html** to author a clip's labels by eye
(content-pipeline §11.5). Pick a pool clip, **scrub** to where a subject is on
screen, **drag a box** over it, set the label key + salience + the four detail
tiers, then **Run tracker** to propagate the box into a keyframed track (classical
OpenCV optical flow) with an appear/disappear window — or **Add as static box**
for a fixed label. Shift-click the stage to place an **affect anchor** and type its
emotion tiers. **Save to manifest** writes the entry back via the pipeline's
idempotent upsert (keeping the clip's media + provenance, replacing only the
authored content). The tracker only propagates **box geometry** — every label
string, scientific name, and fact is **hand-authored**, because a generic detector
can't produce "*Gymnothorax*" or a lifespan.
@@ -1,7 +1,7 @@
# HEF — Content Pipeline (source → process → label → manifest)
**Date:** 2026-06-24
**Status:** Graduated — operator-approved (session 0014); **Increment 1 built & merged** (PR #13); real PD sourcing + rotating pool curated (PR #14/#15). **Increment 2 schema pinned + building** (session 0016 — see §11).
**Status:** Graduated — operator-approved (session 0014); **Increment 1 built & merged** (PR #13); real PD sourcing + rotating pool curated (PR #14/#15). **Increment 2 BUILT** (session 0016 — pool model + tiers + windows in PR #16; `track.py` + author mode in PR B; see §11).
**Repo:** `human-experience-filter-art`
**Anchor:** `docs/ROADMAP.md` sub-project 3 (Player Runtime) content needs; the
spiritual successor to sub-project 2 (Ingest & Tagging tools), retargeted from the
+75
View File
@@ -8,6 +8,7 @@ endpoints (/api/select, /api/catalog/meta) and the X-ray are retired.
from __future__ import annotations
import hashlib
import json
import os
import random
import time
@@ -84,6 +85,29 @@ class RingAdvanceRequest(BaseModel):
delta: int
class AuthorTrackRequest(BaseModel):
"""Author mode: propagate a hand-seeded box on a clip into a keyframed track
(content-pipeline §11.5). `seed_box` + `seed_t` are normalized."""
clip_id: str
key: str
seed_box: list[float] = Field(min_length=4, max_length=4)
seed_t: float = Field(default=0.0, ge=0.0, le=1.0)
salience: int = Field(default=4, ge=1, le=4)
n_keyframes: int = Field(default=5, ge=2, le=20)
max_frames: Optional[int] = None
class AuthorClipRequest(BaseModel):
"""Author mode: persist the authored labels/affect/strings for an existing
pool clip into the manifest (geometry from the tracker, semantics by hand)."""
clip_id: str
annotations: list = []
affect: list = []
strings: dict = {}
def _load_clips(manifest_path: Optional[Path]):
path = Path(manifest_path) if manifest_path else DEFAULT_MANIFEST
if path.exists():
@@ -100,6 +124,7 @@ def _load_ring(manifest_path: Optional[Path]):
def create_app(manifest_path: Optional[Path] = None) -> FastAPI:
app = FastAPI(title="HEF Alteration Simulator")
app.state.manifest_path = Path(manifest_path) if manifest_path else DEFAULT_MANIFEST
app.state.clips = _load_clips(manifest_path)
app.state.ring = _load_ring(manifest_path)
@@ -158,6 +183,56 @@ def create_app(manifest_path: Optional[Path] = None) -> FastAPI:
chosen = pick_clip_id(landed, random.random())
return ring_move_to_dict(move, app.state.ring, chosen)
@app.post("/api/author/track")
def api_author_track(req: AuthorTrackRequest):
# Author mode (§11.5): run the classical optical-flow tracker on a clip's
# base footage, returning the keyframed track geometry for the author to
# accept/correct. Semantics (key/strings/tiers) stay the author's.
clip = next((c for c in app.state.clips if c.id == req.clip_id), None)
if clip is None:
raise HTTPException(status_code=404, detail=f"unknown clip {req.clip_id!r}")
base = MEDIA_DIR / clip.base_file
if not base.exists():
raise HTTPException(status_code=404, detail=f"media absent: {clip.base_file}")
try:
import cv2
except ImportError:
raise HTTPException(status_code=503, detail="opencv-python not installed")
from tools.pipeline.track import track_seed
cap = cv2.VideoCapture(str(base))
total = int(cap.get(cv2.CAP_PROP_FRAME_COUNT) or 0)
cap.release()
seed_frame = max(0, min(total - 1, round(req.seed_t * total))) if total else 0
return track_seed(
base, req.key, tuple(req.seed_box), seed_frame=seed_frame,
salience=req.salience, n_keyframes=req.n_keyframes, max_frames=req.max_frames,
)
@app.post("/api/author/clip")
def api_author_clip(req: AuthorClipRequest):
# Author mode (§11.5 / stage 6): persist authored labels/affect/strings for
# an EXISTING pool clip — idempotent upsert that keeps the clip's media +
# provenance and replaces only the authored content. Reloads in place so
# the preview reflects the edit without a restart.
from tools.pipeline.manifest import upsert_clip
path = app.state.manifest_path
data = json.loads(path.read_text())
existing = next((c for c in data.get("clips", []) if c["id"] == req.clip_id), None)
if existing is None:
raise HTTPException(status_code=404, detail=f"unknown clip {req.clip_id!r}")
merged = dict(existing)
merged["annotations"] = req.annotations
merged["affect"] = req.affect
if req.strings:
merged["strings"] = req.strings
data = upsert_clip(data, merged)
path.write_text(json.dumps(data, indent=2, ensure_ascii=False) + "\n")
app.state.clips = load_manifest(path)
app.state.ring = load_ring(path)
return {"ok": True, "clip": merged}
@app.middleware("http")
async def _no_cache(request, call_next):
# Dev preview server: never let a browser serve a stale app.js/style.css.
+119 -2
View File
@@ -389,6 +389,7 @@ function renderScaleReadout() {
// scale id · the chosen pool member · position on the ring (pool size if >1)
const poolTag = poolN > 1 ? ` · pool ${poolN}` : "";
$("scale-name").textContent = `${s.id} · ${member} (${ringIndex + 1}/${ring.scales.length}${poolTag})`;
renderDial();
}
function controls() {
@@ -500,6 +501,118 @@ function onWheel(e) {
}, 90);
}
// --- Altitude knob: the endless rotary encoder, drawn as a turnable dial ---
// cosmos (highest) sits at the top; turning CLOCKWISE descends through the scales
// (cosmos → orbit → coast → reef → abyss) and past the deepest WRAPS back to the
// top — the same endless ring the server already models. The drag accumulates a
// rotation and commits whole detents on release (so a big spin becomes one fast
// blended pass, exactly like the wheel); a tap on a label jumps to that scale.
const dial = $("dial");
const DIAL_C = 50, DIAL_LABEL_R = 41, DIAL_BODY_R = 27;
let needleDeg = 0;
let dialDrag = null; // {lastAng, accum, moved} while turning
function dialStep() { return ring && ring.scales.length ? 360 / ring.scales.length : 360; }
function _xy(r, deg) {
const a = (deg - 90) * Math.PI / 180; // -90° so 0° points UP
return [DIAL_C + r * Math.cos(a), DIAL_C + r * Math.sin(a)];
}
function buildDial() {
if (!dial || !ring) return;
dial.innerHTML = "";
const n = ring.scales.length, step = 360 / n;
svg("circle", { cx: DIAL_C, cy: DIAL_C, r: DIAL_LABEL_R + 6, class: "dial-rim" }, dial);
svg("circle", { cx: DIAL_C, cy: DIAL_C, r: DIAL_BODY_R, class: "dial-body" }, dial);
for (let i = 0; i < n; i++) {
const deg = i * step;
const [tx0, ty0] = _xy(DIAL_BODY_R - 1.5, deg);
const [tx1, ty1] = _xy(DIAL_BODY_R + 2.5, deg);
svg("line", { x1: tx0, y1: ty0, x2: tx1, y2: ty1, class: "dial-tick" }, dial);
const [lx, ly] = _xy(DIAL_LABEL_R, deg);
const t = svg("text", {
x: lx, y: ly, "text-anchor": "middle", "dominant-baseline": "central",
class: "dial-label", "data-index": String(i),
}, dial);
t.textContent = ring.scales[i].id;
}
svg("text", { x: DIAL_C, y: DIAL_C - 7, "text-anchor": "middle",
"dominant-baseline": "central", class: "dial-caption" }, dial)
.textContent = "ALTITUDE";
const needle = svg("g", { id: "needle", class: "dial-needle" }, dial);
const tip = DIAL_C - (DIAL_BODY_R - 4);
svg("polygon", { points: `${DIAL_C - 2.4},${DIAL_C} ${DIAL_C + 2.4},${DIAL_C} ${DIAL_C},${tip}` }, needle);
svg("circle", { cx: DIAL_C, cy: DIAL_C, r: 3, class: "dial-hub" }, dial);
renderDial();
}
function setNeedle(deg) {
const needle = $("needle");
if (needle) needle.setAttribute("transform", `rotate(${deg} ${DIAL_C} ${DIAL_C})`);
}
function renderDial() {
if (!dial || !ring) return;
needleDeg = ringIndex * dialStep();
setNeedle(needleDeg);
for (const el of dial.querySelectorAll(".dial-label")) {
el.classList.toggle("active", +el.getAttribute("data-index") === ringIndex);
}
}
function dialAngle(e) {
const r = dial.getBoundingClientRect();
const dx = e.clientX - (r.left + r.width / 2);
const dy = e.clientY - (r.top + r.height / 2);
return (Math.atan2(dx, -dy) * 180 / Math.PI + 360) % 360; // 0 at top, clockwise +
}
function angDelta(a, b) {
let d = a - b;
while (d > 180) d -= 360;
while (d < -180) d += 360;
return d;
}
function onDialDown(e) {
if (busy || !ring || ring.scales.length < 2) return;
e.preventDefault();
dialDrag = { lastAng: dialAngle(e), accum: 0, moved: 0, target: e.target };
}
function onDialMove(e) {
if (!dialDrag) return;
const a = dialAngle(e);
const d = angDelta(a, dialDrag.lastAng);
dialDrag.lastAng = a;
dialDrag.accum += d;
dialDrag.moved += Math.abs(d);
setNeedle(ringIndex * dialStep() + dialDrag.accum); // live feedback while turning
}
function onDialUp(e) {
if (!dialDrag) return;
const { accum, moved, target } = dialDrag;
dialDrag = null;
if (moved < 6) { // a tap, not a turn
if (target && target.classList && target.classList.contains("dial-label")) {
jumpToScale(+target.getAttribute("data-index"));
}
renderDial();
return;
}
const detents = Math.round(accum / dialStep());
if (detents) advance(detents); else renderDial(); // snap back if it didn't cross a detent
}
// Click a label → travel the SHORTEST signed way around the ring to that scale.
function jumpToScale(idx) {
if (!ring) return;
const n = ring.scales.length;
let d = (idx - ringIndex) % n;
if (d > n / 2) d -= n;
if (d < -n / 2) d += n;
if (d) advance(d);
}
// Dev live-reload: poll the asset version and reload when it changes, so an open
// tab never keeps running a stale renderer while we iterate (the readout updates
// from the live API and masks it otherwise). Reloads on my edits AND on a server
@@ -520,12 +633,16 @@ async function main() {
try { initPaint(); } catch (e) { paintOK = false; paint.style.display = "none"; showError("WebGL init: " + e.message); }
await loadData();
await landScale(); // pick the initial scale's pool member before first render
buildDial(); // draw the altitude knob from the ring's scales
renderScaleReadout();
for (const id of ["content", "left", "right", "mood"]) {
$(id).addEventListener("input", debounced);
}
$("zoom-in").addEventListener("click", () => advance(1));
$("zoom-out").addEventListener("click", () => advance(-1));
// Altitude knob: drag to turn (commit detents on release), scroll to step, tap a label to jump.
dial.addEventListener("pointerdown", onDialDown);
window.addEventListener("pointermove", onDialMove);
window.addEventListener("pointerup", onDialUp);
dial.addEventListener("wheel", onWheel, { passive: false });
$("stage").addEventListener("wheel", onWheel, { passive: false });
update();
}
+64
View File
@@ -0,0 +1,64 @@
/* Author-mode-only styling (content-pipeline §11.5). Reuses style.css for the
shared stage/panel chrome; this adds the box-draw layer + editor widgets. */
.author #draw {
position: absolute;
inset: 0;
width: 100%;
height: 100%;
cursor: crosshair;
}
.author #draw .seed-box {
fill: rgba(80, 200, 255, 0.12);
stroke: #50c8ff;
stroke-width: 0.4;
vector-effect: non-scaling-stroke;
}
.author #draw .track-box {
fill: rgba(120, 255, 160, 0.10);
stroke: #78ffa0;
stroke-width: 0.4;
stroke-dasharray: 1.5 1;
vector-effect: non-scaling-stroke;
}
.author #draw .affect-dot {
fill: rgba(190, 150, 255, 0.85);
stroke: #fff;
stroke-width: 0.25;
vector-effect: non-scaling-stroke;
}
.author .scrub {
display: flex;
align-items: center;
gap: 8px;
margin-top: 8px;
}
.author .scrub input[type="range"] { flex: 1; }
.author .mono { font-family: ui-monospace, monospace; font-size: 12px; opacity: 0.85; }
.author .row { display: flex; gap: 8px; margin-top: 6px; flex-wrap: wrap; }
.author fieldset label { display: block; margin: 4px 0; font-size: 13px; }
.author fieldset input[type="text"],
.author fieldset input[type="number"] { width: 100%; box-sizing: border-box; }
.author fieldset input[type="number"] { width: 5em; }
.author ul.authored { list-style: none; padding: 0; margin: 6px 0; }
.author ul.authored li {
font-size: 12px;
padding: 3px 6px;
margin: 2px 0;
background: rgba(255, 255, 255, 0.05);
border-radius: 3px;
display: flex;
justify-content: space-between;
gap: 6px;
}
.author ul.authored .rm {
background: none;
border: none;
color: #f88;
cursor: pointer;
font-size: 12px;
}
.author header .hint { font-size: 12px; opacity: 0.8; max-width: 70ch; }
+79
View File
@@ -0,0 +1,79 @@
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8" />
<meta name="viewport" content="width=device-width, initial-scale=1" />
<title>HEF — Label Author Mode</title>
<link rel="stylesheet" href="/style.css" />
<link rel="stylesheet" href="/author.css" />
</head>
<body class="author">
<header><h1>HEF — Label Author Mode</h1>
<p class="hint">Scrub a pool clip · drag a box on a subject · run the tracker · author the tiers · save to the manifest. Geometry is propagated; <strong>semantics are hand-authored</strong> (content-pipeline §11.5). <a href="/">← back to preview</a></p>
</header>
<main>
<section class="stage" id="stage">
<div class="screen">
<video id="vid" muted playsinline></video>
<svg id="draw" viewBox="0 0 100 100" preserveAspectRatio="none"></svg>
</div>
<div class="scrub">
<button type="button" id="play">▶/⏸</button>
<input type="range" id="seek" min="0" max="1000" value="0" />
<span id="time" class="mono">0.00</span>
</div>
</section>
<section class="panel">
<fieldset>
<legend>Clip</legend>
<select id="clip"></select>
<p class="hint" id="clip-meta"></p>
</fieldset>
<fieldset>
<legend>Seed box (drag on the video)</legend>
<p class="mono" id="seed-readout">no box — drag on the stage</p>
<label>label key <input type="text" id="ann-key" placeholder="detected.jelly" /></label>
<label>salience (14; 4 = shows first/at low Left)
<input type="number" id="ann-salience" min="1" max="4" value="4" /></label>
<label>tier 1 — general <input type="text" id="t1" placeholder="jelly" /></label>
<label>tier 2 — specific <input type="text" id="t2" placeholder="comb jelly" /></label>
<label>tier 3 — scientific <input type="text" id="t3" placeholder="Ctenophora" /></label>
<label>tier 4 — +fact <input type="text" id="t4" placeholder="Ctenophora · beats rows of cilia" /></label>
<label>keyframes <input type="number" id="ann-keyframes" min="2" max="20" value="5" /></label>
<div class="row">
<button type="button" id="run-track">Run tracker → track</button>
<button type="button" id="add-static">Add as static box</button>
</div>
<p class="mono" id="track-readout"></p>
<button type="button" id="add-ann" disabled>+ Add label</button>
</fieldset>
<fieldset>
<legend>Affect anchor (click the stage to place)</legend>
<p class="mono" id="affect-readout">no point — shift-click the stage</p>
<label>feel key <input type="text" id="aff-key" placeholder="feel.unease" /></label>
<label>min strength (14) <input type="number" id="aff-min" min="1" max="4" value="1" /></label>
<label>tier 1 — basic <input type="text" id="a1" placeholder="uh" /></label>
<label>tier 2 <input type="text" id="a2" placeholder="unease" /></label>
<label>tier 3 <input type="text" id="a3" placeholder="disquiet" /></label>
<label>tier 4 — compound <input type="text" id="a4" placeholder="a creeping disquiet" /></label>
<button type="button" id="add-aff" disabled>+ Add emotion</button>
</fieldset>
<fieldset>
<legend>Authored for this clip</legend>
<ul id="ann-list" class="authored"></ul>
<ul id="aff-list" class="authored"></ul>
<div class="row">
<button type="button" id="load-existing">Load existing</button>
<button type="button" id="save">Save to manifest</button>
</div>
<p class="mono" id="save-readout"></p>
</fieldset>
</section>
</main>
<script src="/author.js"></script>
</body>
</html>
+254
View File
@@ -0,0 +1,254 @@
// Label author mode (content-pipeline §11.5): scrub a pool clip, drag a seed box
// on a subject, run the classical tracker to propagate a keyframed track, author
// the LEFT detail tiers + RIGHT emotion tiers, and save the entry to the manifest.
// Geometry comes from the tracker; SEMANTICS (keys, strings, tiers) are hand-typed.
const $ = (id) => document.getElementById(id);
const vid = $("vid"), draw = $("draw");
const SVGNS = "http://www.w3.org/2000/svg";
let clips = [];
let currentClip = null;
let seedBox = null; // [x,y,w,h] normalized, from a drag
let pendingTrack = null; // {track, appear, disappear} from the tracker (or static)
let affectPoint = null; // [x,y] normalized, from a shift-click
let annotations = []; // authored annotation dicts (carry _tiers)
let affect = []; // authored affect dicts (carry _tiers)
function svg(tag, attrs, parent) {
const el = document.createElementNS(SVGNS, tag);
for (const k in attrs) el.setAttribute(k, attrs[k]);
if (parent) parent.appendChild(el);
return el;
}
const mono = (n) => Number(n).toFixed(3);
async function loadClips() {
clips = (await (await fetch("/api/clips")).json()).clips || [];
const sel = $("clip");
sel.innerHTML = "";
for (const c of clips) {
const o = document.createElement("option");
o.value = c.id; o.textContent = c.id;
sel.appendChild(o);
}
if (clips.length) selectClip(clips[0].id);
}
function selectClip(id) {
currentClip = clips.find((c) => c.id === id) || null;
if (!currentClip) return;
$("clip-meta").textContent = `${currentClip.title}${currentClip.license}`;
vid.src = "/media/" + currentClip.base_file;
vid.currentTime = 0;
annotations = []; affect = []; seedBox = null; pendingTrack = null; affectPoint = null;
renderLists(); renderDrawLayer();
$("seed-readout").textContent = "no box — drag on the stage";
$("track-readout").textContent = "";
$("affect-readout").textContent = "no point — shift-click the stage";
}
// --- scrub ---
function seekT() { return vid.duration ? vid.currentTime / vid.duration : 0; }
$("play").addEventListener("click", () => { vid.paused ? vid.play() : vid.pause(); });
$("seek").addEventListener("input", () => {
if (vid.duration) vid.currentTime = (+$("seek").value / 1000) * vid.duration;
});
vid.addEventListener("timeupdate", () => {
if (!vid.duration) return;
$("seek").value = String(Math.round(seekT() * 1000));
$("time").textContent = seekT().toFixed(3);
});
// --- draw a seed box / place an affect point on the stage ---
function evtNorm(e) {
const r = draw.getBoundingClientRect();
return [(e.clientX - r.left) / r.width, (e.clientY - r.top) / r.height];
}
let dragStart = null;
draw.addEventListener("mousedown", (e) => {
if (e.shiftKey) { // shift-click = affect anchor
affectPoint = evtNorm(e).map((v) => Math.min(Math.max(v, 0), 1));
$("affect-readout").textContent = `affect @ ${affectPoint.map(mono).join(", ")}`;
$("add-aff").disabled = false;
renderDrawLayer();
return;
}
dragStart = evtNorm(e);
});
draw.addEventListener("mousemove", (e) => {
if (!dragStart) return;
seedBox = boxFrom(dragStart, evtNorm(e));
renderDrawLayer();
});
window.addEventListener("mouseup", (e) => {
if (!dragStart) return;
seedBox = boxFrom(dragStart, evtNorm(e));
dragStart = null;
pendingTrack = null;
$("seed-readout").textContent = `seed [${seedBox.map(mono).join(", ")}] @ t=${seekT().toFixed(3)}`;
$("add-ann").disabled = false;
renderDrawLayer();
});
function boxFrom(a, b) {
const x = Math.min(a[0], b[0]), y = Math.min(a[1], b[1]);
const w = Math.abs(b[0] - a[0]), h = Math.abs(b[1] - a[1]);
const cl = (v) => Math.min(Math.max(v, 0), 1);
return [cl(x), cl(y), Math.min(w, 1 - cl(x)), Math.min(h, 1 - cl(y))];
}
function renderDrawLayer() {
draw.innerHTML = "";
if (seedBox) {
const [x, y, w, h] = seedBox.map((n) => n * 100);
svg("rect", { x, y, width: w, height: h, class: "seed-box" }, draw);
}
if (pendingTrack && pendingTrack.track && vid.duration) {
const b = boxAt(pendingTrack.track, seekT());
if (b) {
const [x, y, w, h] = b.map((n) => n * 100);
svg("rect", { x, y, width: w, height: h, class: "track-box" }, draw);
}
}
if (affectPoint) {
const [x, y] = affectPoint.map((n) => n * 100);
svg("circle", { cx: x, cy: y, r: 1.2, class: "affect-dot" }, draw);
}
}
// Interpolate a track box at loop-normalized t (mirrors the preview renderer).
function boxAt(track, t) {
if (!track || !track.length) return null;
if (t <= track[0].t) return track[0].box;
if (t >= track[track.length - 1].t) return track[track.length - 1].box;
for (let i = 1; i < track.length; i++) {
if (t <= track[i].t) {
const a = track[i - 1], b = track[i], f = (t - a.t) / (b.t - a.t);
return a.box.map((v, j) => v + (b.box[j] - v) * f);
}
}
return track[track.length - 1].box;
}
// Animate the pending track preview while playing.
function previewLoop() { renderDrawLayer(); requestAnimationFrame(previewLoop); }
// --- run the tracker on the seed box ---
$("run-track").addEventListener("click", async () => {
if (!seedBox || !currentClip) { $("track-readout").textContent = "draw a seed box first"; return; }
const key = $("ann-key").value.trim() || "detected.object";
$("track-readout").textContent = "tracking…";
const resp = await fetch("/api/author/track", {
method: "POST", headers: { "content-type": "application/json" },
body: JSON.stringify({
clip_id: currentClip.id, key, seed_box: seedBox, seed_t: seekT(),
salience: +$("ann-salience").value, n_keyframes: +$("ann-keyframes").value,
}),
});
if (!resp.ok) { $("track-readout").textContent = "tracker error " + resp.status; return; }
const ann = await resp.json();
pendingTrack = { track: ann.track, appear: ann.appear, disappear: ann.disappear };
$("track-readout").textContent =
`track: ${ann.track.length} keyframes · window ${mono(ann.appear)}${mono(ann.disappear)}`;
$("add-ann").disabled = false;
});
// Add the current seed as a STATIC (untracked) box at the scrub position.
$("add-static").addEventListener("click", () => {
if (!seedBox) return;
pendingTrack = { static: true, box: seedBox.slice() };
$("track-readout").textContent = "static box (no track) at the drawn position";
$("add-ann").disabled = false;
});
function tiers(ids) {
const vals = ids.map((id) => $(id).value.trim());
while (vals.length && !vals[vals.length - 1]) vals.pop(); // drop empty trailing tiers
return vals.length ? vals : null;
}
$("add-ann").addEventListener("click", () => {
const key = $("ann-key").value.trim();
if (!key) { $("track-readout").textContent = "a label key is required"; return; }
const t = tiers(["t1", "t2", "t3", "t4"]);
const ann = { key, salience: +$("ann-salience").value, _tiers: t || key };
if (pendingTrack && pendingTrack.track) {
ann.track = pendingTrack.track;
ann.appear = pendingTrack.appear; ann.disappear = pendingTrack.disappear;
} else if (pendingTrack && pendingTrack.static) {
ann.box = pendingTrack.box;
} else if (seedBox) {
ann.box = seedBox.slice();
} else { $("track-readout").textContent = "draw + (optionally) track a box first"; return; }
annotations.push(ann);
seedBox = null; pendingTrack = null; $("add-ann").disabled = true;
renderLists(); renderDrawLayer();
});
$("add-aff").addEventListener("click", () => {
const key = $("aff-key").value.trim();
if (!key || !affectPoint) return;
const t = tiers(["a1", "a2", "a3", "a4"]);
affect.push({ key, at: affectPoint.slice(), min_level: +$("aff-min").value, _tiers: t || key });
affectPoint = null; $("add-aff").disabled = true;
renderLists(); renderDrawLayer();
});
function renderLists() {
const al = $("ann-list"); al.innerHTML = "";
annotations.forEach((a, i) => {
const li = document.createElement("li");
const kind = a.track ? `track ${a.track.length}kf ${mono(a.appear)}${mono(a.disappear)}` : "static";
const tip = Array.isArray(a._tiers) ? a._tiers.join(" → ") : a._tiers;
li.textContent = `${a.key} (s${a.salience}, ${kind}) — ${tip}`;
li.appendChild(rm(() => { annotations.splice(i, 1); renderLists(); }));
al.appendChild(li);
});
const fl = $("aff-list"); fl.innerHTML = "";
affect.forEach((f, i) => {
const li = document.createElement("li");
const tip = Array.isArray(f._tiers) ? f._tiers.join(" → ") : f._tiers;
li.textContent = `${f.key} (min ${f.min_level}) — ${tip}`;
li.appendChild(rm(() => { affect.splice(i, 1); renderLists(); }));
fl.appendChild(li);
});
}
function rm(fn) {
const b = document.createElement("button");
b.textContent = "✕"; b.className = "rm"; b.onclick = fn;
return b;
}
// Load whatever is already authored for this clip into the editor.
$("load-existing").addEventListener("click", () => {
if (!currentClip) return;
const strings = (currentClip.strings && currentClip.strings.en) || {};
annotations = (currentClip.annotations || []).map((a) => ({ ...a, _tiers: strings[a.key] || a.key }));
affect = (currentClip.affect || []).map((f) => ({ ...f, _tiers: strings[f.key] || f.key }));
renderLists();
});
// Assemble strings.en from each entry's _tiers and POST to the manifest.
$("save").addEventListener("click", async () => {
if (!currentClip) return;
const en = {};
const strip = (x) => { const { _tiers, ...rest } = x; return rest; };
const anns = annotations.map((a) => { en[a.key] = a._tiers; return strip(a); });
const affs = affect.map((f) => { en[f.key] = f._tiers; return strip(f); });
const resp = await fetch("/api/author/clip", {
method: "POST", headers: { "content-type": "application/json" },
body: JSON.stringify({ clip_id: currentClip.id, annotations: anns, affect: affs, strings: { en } }),
});
$("save-readout").textContent = resp.ok
? `saved ${anns.length} labels + ${affs.length} emotions to the manifest ✓`
: "save error " + resp.status;
if (resp.ok) { // refresh the in-memory clip copy
clips = (await (await fetch("/api/clips")).json()).clips || [];
currentClip = clips.find((c) => c.id === currentClip.id) || currentClip;
}
});
$("clip").addEventListener("change", (e) => selectClip(e.target.value));
async function main() {
await loadClips();
requestAnimationFrame(previewLoop);
}
main();
+5 -6
View File
@@ -35,13 +35,12 @@
</fieldset>
<fieldset>
<legend>Scale ring (endless encoder)</legend>
<div class="ring-control">
<button type="button" id="zoom-out" title="zoom out (toward cosmos)">⊖ out</button>
<span id="scale-name" class="scale-name"></span>
<button type="button" id="zoom-in" title="zoom in (toward microscopic)">in ⊕</button>
<legend>Altitude</legend>
<div class="dial-wrap">
<svg id="dial" viewBox="0 0 100 100" aria-label="Altitude knob (turn to change scale)"></svg>
</div>
<p class="hint">Relative, endless — wraps past the smallest back to the largest. Scroll the stage too.</p>
<span id="scale-name" class="scale-name"></span>
<p class="hint">Turn the knob (drag it, or scroll) to change altitude — endless: past the deepest it wraps back up to the highest. Click a label to jump there.</p>
</fieldset>
<fieldset>
+17 -6
View File
@@ -48,10 +48,21 @@ label { display: block; margin: 0.4rem 0; }
input[type=range], select { width: 100%; }
#readout { background: #000; padding: 0.5rem; border-radius: 4px; font-size: 12px;
white-space: pre-wrap; max-height: 240px; overflow: auto; }
.ring-control { display: flex; align-items: center; gap: 0.4rem; }
.ring-control button { flex: 0 0 auto; background: #1a2436; color: #9af;
border: 1px solid #345; border-radius: 4px; padding: 0.3rem 0.5rem;
cursor: pointer; font-size: 13px; }
.ring-control button:hover { background: #243352; }
.scale-name { flex: 1 1 auto; text-align: center; font-size: 12px; color: #cde; }
/* Altitude knob */
.dial-wrap { display: flex; justify-content: center; padding: 0.3rem 0 0.1rem; }
#dial { width: 190px; height: 190px; touch-action: none; cursor: grab; user-select: none; }
#dial:active { cursor: grabbing; }
.dial-rim { fill: #0d1320; stroke: #243352; stroke-width: 1.2; }
.dial-body { fill: #16203200; stroke: #2c3c5c; stroke-width: 1; }
.dial-tick { stroke: #3a4d70; stroke-width: 0.8; }
.dial-label { fill: #789ac0; font-size: 6px; font-family: ui-monospace, monospace;
letter-spacing: 0.2px; cursor: pointer; }
.dial-label:hover { fill: #cde; }
.dial-label.active { fill: #9cf; font-weight: 700; }
.dial-caption { fill: #4d6184; font-size: 4.4px; letter-spacing: 1.2px;
font-family: ui-monospace, monospace; }
.dial-needle { fill: #9cf; }
.dial-needle polygon { filter: drop-shadow(0 0 1px #9cf); }
.dial-hub { fill: #2c3c5c; stroke: #9cf; stroke-width: 0.6; }
.scale-name { display: block; text-align: center; font-size: 12px; color: #cde; }
.hint { margin: 0.3rem 0 0; font-size: 11px; color: #789; }
+10 -10
View File
@@ -4,7 +4,9 @@ import pytest
from tools.pipeline.run import probe_duration, process_clip, resolve_ffmpeg
_FOREST = Path("simulator/sample_media/forest/base.mp4")
# A real pool base (forest was retired with the rotating-pool rename, session 0016);
# cosmos/base.mp4 is the cosmos pool primary and the most stable PD clip on disk.
_BASE = Path("simulator/sample_media/cosmos/base.mp4")
def _ffmpeg_available() -> bool:
@@ -18,20 +20,18 @@ def _ffmpeg_available() -> bool:
_HAS_FF = _ffmpeg_available()
@pytest.mark.skipif(not _FOREST.exists(),
reason="forest POC base absent (run simulator/setup_sample_media.py)")
def test_probe_duration_reads_forest_base():
if not _FOREST.exists():
pytest.skip("forest POC base absent")
assert probe_duration(_FOREST) > 0
@pytest.mark.skipif(not _BASE.exists(),
reason="cosmos base absent (run simulator/build_pool_manifest.py)")
def test_probe_duration_reads_real_base():
assert probe_duration(_BASE) > 0
@pytest.mark.skipif(not _HAS_FF,
reason="neither system ffmpeg nor imageio-ffmpeg available")
@pytest.mark.skipif(not _FOREST.exists(),
reason="forest POC base absent (run simulator/setup_sample_media.py)")
@pytest.mark.skipif(not _BASE.exists(),
reason="cosmos base absent (run simulator/build_pool_manifest.py)")
def test_process_clip_emits_proxy_and_master(tmp_path):
out = process_clip(_FOREST, tmp_path, overlap=0.5, start=0.0, duration=4.0)
out = process_clip(_BASE, tmp_path, overlap=0.5, start=0.0, duration=4.0)
assert out["proxy"].exists() and out["master"].exists()
# Proxy is exactly 1920×1080 — verified with cv2, no ffprobe dependency.
import cv2
+120
View File
@@ -0,0 +1,120 @@
"""tools/pipeline/track.py — the hybrid motion-track pass (content-pipeline §11.5).
Pure geometry helpers are unit-tested here; the cv2 optical-flow propagation is an
opt-in integration test that runs on a real pool clip when present."""
from pathlib import Path
import pytest
from tools.pipeline.track import (
clamp01_box,
denormalize_box,
infer_window,
loop_t,
normalize_box,
sample_track,
track_to_annotation,
)
# --- pure helpers -----------------------------------------------------------
def test_clamp01_box_keeps_box_on_screen():
assert clamp01_box((-0.2, 0.1, 0.5, 0.5)) == [0.0, 0.1, 0.5, 0.5]
# size trimmed so the box never runs past the right/bottom edge
assert clamp01_box((0.8, 0.8, 0.5, 0.5)) == [0.8, 0.8, pytest.approx(0.2), pytest.approx(0.2)]
def test_normalize_denormalize_round_trip():
px = (192, 108, 384, 216)
norm = normalize_box(px, 1920, 1080)
assert norm == [pytest.approx(0.1), pytest.approx(0.1), pytest.approx(0.2), pytest.approx(0.2)]
assert denormalize_box(norm, 1920, 1080) == (192, 108, 384, 216)
def test_loop_t_is_frame_over_total():
assert loop_t(0, 100) == 0.0
assert loop_t(50, 100) == 0.5
assert loop_t(5, 0) == 0.0 # guard divide-by-zero
def test_infer_window_spans_tracked_frames():
appear, disappear = infer_window([10, 11, 30, 31], 100)
assert appear == pytest.approx(0.1)
assert disappear == pytest.approx(0.31)
# no frames -> full clip
assert infer_window([], 100) == (0.0, 1.0)
def test_sample_track_downsamples_to_sparse_keyframes():
# 21 dense frames -> at most n_keyframes, including the first and last
per_frame = {i: [i / 100, 0.2, 0.1, 0.1] for i in range(0, 21)}
track = sample_track(per_frame, total_frames=100, n_keyframes=5)
assert len(track) <= 5
ts = [k["t"] for k in track]
assert ts == sorted(ts) # sorted by t
assert ts[0] == pytest.approx(0.0) # first tracked frame
assert ts[-1] == pytest.approx(0.2) # last tracked frame (frame 20 / 100)
assert len(set(ts)) == len(ts) # de-duplicated on t
def test_sample_track_keeps_all_when_already_sparse():
per_frame = {0: [0.1, 0.1, 0.1, 0.1], 60: [0.5, 0.3, 0.1, 0.1]}
track = sample_track(per_frame, total_frames=120, n_keyframes=5)
assert [k["t"] for k in track] == [pytest.approx(0.0), pytest.approx(0.5)]
def test_sample_track_empty_is_empty():
assert sample_track({}, total_frames=100) == []
def test_track_to_annotation_assembles_geometry_only():
track = [{"t": 0.0, "box": [0.1, 0.2, 0.1, 0.1]}, {"t": 0.5, "box": [0.4, 0.2, 0.1, 0.1]}]
ann = track_to_annotation("detected.jelly", track, salience=3, appear=0.0, disappear=0.55)
assert ann["key"] == "detected.jelly"
assert ann["salience"] == 3
assert ann["track"] == track
assert ann["appear"] == 0.0 and ann["disappear"] == 0.55
# no semantics leaked in — strings/tiers are the author's, added elsewhere
assert "strings" not in ann and "_tiers" not in ann
def test_track_to_annotation_omits_window_when_absent():
ann = track_to_annotation("detected.x", [], salience=4)
assert "appear" not in ann and "disappear" not in ann
# --- opt-in: classical optical-flow propagation on a real clip --------------
_CLIP = Path("simulator/sample_media/abyss_wow/base.mp4")
def _cv2_available() -> bool:
try:
import cv2 # noqa: F401
return True
except Exception:
return False
@pytest.mark.skipif(not _CLIP.exists(),
reason="abyss_wow base absent (run simulator/build_pool_manifest.py)")
@pytest.mark.skipif(not _cv2_available(), reason="opencv-python not available")
def test_track_seed_produces_a_sparse_annotation_on_real_footage():
from tools.pipeline.track import track_box, track_seed
seed = (0.2, 0.3, 0.15, 0.18)
per_frame, total = track_box(_CLIP, seed, seed_frame=0, max_frames=40)
assert total > 0
assert per_frame[0] == pytest.approx(list(seed), abs=0.02)
assert len(per_frame) > 1 # propagated beyond the seed frame
ann = track_seed(_CLIP, "detected.jelly", seed, seed_frame=0, n_keyframes=5, max_frames=40)
assert ann["key"] == "detected.jelly"
assert 2 <= len(ann["track"]) <= 5 # sparse keyframes
for kf in ann["track"]:
assert 0.0 <= kf["t"] <= 1.0
assert all(0.0 <= v <= 1.0 for v in kf["box"])
assert 0.0 <= ann["appear"] <= ann["disappear"] <= 1.0
+75
View File
@@ -250,3 +250,78 @@ def test_delta_zero_is_an_initial_pool_pick(pool_client):
assert data["to_index"] == 1
assert data["steps"] == []
assert data["target_clip_id"] in members
# --- author mode (content-pipeline §11.5) ---
@pytest.fixture
def author_client(tmp_path):
p = tmp_path / "manifest.json"
p.write_text(json.dumps({
"clips": [{
"id": "abyss_wow", "title": "World of Water", "base_file": "abyss_wow/base.mp4",
"license": "PD", "source": "NOAA", "right_variants": {},
"annotations": [], "affect": [], "strings": {"en": {}},
}],
}))
return TestClient(create_app(manifest_path=p)), p
def test_author_clip_persists_labels_and_keeps_provenance(author_client):
client, path = author_client
body = {
"clip_id": "abyss_wow",
"annotations": [{"key": "detected.jelly", "salience": 4,
"appear": 0.0, "disappear": 0.5,
"track": [{"t": 0.0, "box": [0.1, 0.2, 0.1, 0.1]}]}],
"affect": [{"key": "feel.unease", "at": [0.5, 0.5], "min_level": 1}],
"strings": {"en": {"detected.jelly": ["jelly", "comb jelly", "Ctenophora", "Ctenophora · cilia rows"],
"feel.unease": ["uh", "unease", "disquiet", "a creeping disquiet"]}},
}
resp = client.post("/api/author/clip", json=body)
assert resp.status_code == 200 and resp.json()["ok"] is True
# persisted to disk
saved = json.loads(path.read_text())["clips"][0]
assert saved["annotations"][0]["key"] == "detected.jelly"
assert saved["affect"][0]["key"] == "feel.unease"
assert saved["strings"]["en"]["detected.jelly"][2] == "Ctenophora"
# provenance preserved (only authored content replaced)
assert saved["base_file"] == "abyss_wow/base.mp4"
assert saved["license"] == "PD"
# reflected by /api/clips without a restart
served = client.get("/api/clips").json()["clips"][0]
assert served["annotations"][0]["key"] == "detected.jelly"
def test_author_clip_unknown_clip_404(author_client):
client, _ = author_client
resp = client.post("/api/author/clip", json={"clip_id": "nope", "annotations": []})
assert resp.status_code == 404
def test_author_track_unknown_clip_404(author_client):
client, _ = author_client
resp = client.post("/api/author/track", json={
"clip_id": "nope", "key": "detected.x", "seed_box": [0.1, 0.1, 0.1, 0.1]})
assert resp.status_code == 404
_REAL_CLIP = __import__("pathlib").Path("simulator/sample_media/abyss_wow/base.mp4")
@pytest.mark.skipif(not _REAL_CLIP.exists(),
reason="abyss_wow base absent (run simulator/build_pool_manifest.py)")
def test_author_track_on_real_clip_returns_keyframes():
# Against the real default manifest + footage: the tracker yields a sparse track.
client = TestClient(create_app())
resp = client.post("/api/author/track", json={
"clip_id": "abyss_wow", "key": "detected.jelly",
"seed_box": [0.2, 0.3, 0.15, 0.18], "seed_t": 0.0,
"salience": 4, "n_keyframes": 5, "max_frames": 30,
})
assert resp.status_code == 200
ann = resp.json()
assert ann["key"] == "detected.jelly"
assert 2 <= len(ann["track"]) <= 5
assert 0.0 <= ann["appear"] <= ann["disappear"] <= 1.0
+297
View File
@@ -0,0 +1,297 @@
"""Stage 5 geometry — the hybrid motion-track pass (content-pipeline §3.5 / §11.5).
Turns a hand-seeded box (or an ML detection) on one keyframe into the sparse,
loop-normalized keyframed `track` the manifest stores (design §4) plus an
appear/disappear window (§11.2). SEMANTICS are never produced here the author
owns the `key`, strings, `salience` and tiers; this module only propagates BOX
GEOMETRY.
Two geometry paths, same artifact:
- **classical** (always available): a hand-seeded box propagated by OpenCV
Lucas-Kanade optical flow (a CSRT tracker is used instead when a contrib
build provides one). Translation model right for calm drifting subjects
(the abyss/reef creatures). Re-seeds features when too few survive.
- **ML detect+track** (optional, lazy-imported): a YOLO+tracker pass proposing
boxes the author maps to a key. Imported only on demand so the base install
needs no torch/ultralytics; a clear ImportError tells you what to `pip
install` if you ask for it without it present.
The PURE helpers (normalization, keyframe sampling, window inference, annotation
assembly) are unit-tested; the cv2/ML propagation is covered by the opt-in
integration test on real footage.
"""
from __future__ import annotations
from typing import Optional
Box = tuple[float, float, float, float] # (x, y, w, h)
# --- pure geometry helpers --------------------------------------------------
def clamp01_box(box: Box) -> list[float]:
"""Clamp a normalized box into the visible [0,1] frame, keeping it on-screen
(origin clamped, then size trimmed so x+w<=1, y+h<=1)."""
x, y, w, h = box
x = min(max(x, 0.0), 1.0)
y = min(max(y, 0.0), 1.0)
w = min(max(w, 0.0), 1.0 - x)
h = min(max(h, 0.0), 1.0 - y)
return [x, y, w, h]
def normalize_box(box_px: Box, width: int, height: int) -> list[float]:
"""Pixel box -> normalized [0,1] box (the manifest's resolution-independent
form). Clamped on-screen."""
x, y, w, h = box_px
return clamp01_box((x / width, y / height, w / width, h / height))
def denormalize_box(box_norm: Box, width: int, height: int) -> tuple[int, int, int, int]:
"""Normalized [0,1] box -> integer pixel box for the cv2 tracker."""
x, y, w, h = box_norm
return (round(x * width), round(y * height), round(w * width), round(h * height))
def loop_t(frame_idx: int, total_frames: int) -> float:
"""Loop-normalized time of a frame in [0,1), matching the client's
`currentTime / duration` clock (design §4)."""
if total_frames <= 0:
return 0.0
return frame_idx / total_frames
def infer_window(tracked_frames: list[int], total_frames: int) -> tuple[float, float]:
"""The appear/disappear window (§11.2) covering the frames where a box was
tracked, as loop-normalized t. The first/last tracked frame bound it."""
if not tracked_frames:
return (0.0, 1.0)
lo, hi = min(tracked_frames), max(tracked_frames)
return (loop_t(lo, total_frames), loop_t(hi, total_frames))
def sample_track(
per_frame: dict[int, Box], total_frames: int, n_keyframes: int = 5
) -> list[dict]:
"""Down-sample a dense per-frame box map to a SPARSE keyframed track
(design §4: sparse keyframes interpolated at runtime, never dense per-frame).
Picks `n_keyframes` frames evenly across the tracked span (always including
the first and last tracked frame) and emits `[{t, box}, ...]` sorted by t,
de-duplicated on t. Pure takes already-normalized boxes."""
if not per_frame:
return []
frames = sorted(per_frame)
n = max(2, n_keyframes)
if len(frames) <= n:
picks = frames
else:
lo, hi = frames[0], frames[-1]
step = (hi - lo) / (n - 1)
targets = [lo + round(step * i) for i in range(n)]
# snap each target to the nearest actually-tracked frame
picks = sorted({min(frames, key=lambda f: abs(f - t)) for t in targets})
out, seen = [], set()
for f in picks:
t = round(loop_t(f, total_frames), 4)
if t in seen:
continue
seen.add(t)
out.append({"t": t, "box": clamp01_box(per_frame[f])})
return out
def track_to_annotation(
key: str,
track: list[dict],
*,
salience: int = 4,
appear: Optional[float] = None,
disappear: Optional[float] = None,
) -> dict:
"""Assemble a manifest annotation from a propagated track + the author's
semantics. Geometry (track + window) comes from this module; `key`,
`salience` and the tiered strings are the author's (added separately)."""
ann: dict = {"key": key, "salience": salience, "track": track}
if appear is not None:
ann["appear"] = round(appear, 4)
if disappear is not None:
ann["disappear"] = round(disappear, 4)
return ann
# --- classical propagation (cv2; covered by the opt-in integration test) -----
def _new_csrt():
"""A CSRT tracker if this OpenCV build ships one (contrib), else None — then
the optical-flow path is used. Kept tiny so the import stays lazy."""
import cv2
for factory in ("TrackerCSRT_create",):
if hasattr(cv2, factory):
return getattr(cv2, factory)()
legacy = getattr(cv2, "legacy", None)
if legacy is not None and hasattr(legacy, "TrackerCSRT_create"):
return legacy.TrackerCSRT_create()
return None
def _features_in_box(gray, box_px, max_corners: int = 60):
"""Good-feature points inside a pixel box, in full-frame coords (for LK)."""
import cv2
import numpy as np
x, y, w, h = (int(v) for v in box_px)
h_img, w_img = gray.shape[:2]
x0, y0 = max(x, 0), max(y, 0)
x1, y1 = min(x + w, w_img), min(y + h, h_img)
if x1 - x0 < 4 or y1 - y0 < 4:
return None
roi = gray[y0:y1, x0:x1]
pts = cv2.goodFeaturesToTrack(roi, max_corners, 0.01, 5)
if pts is None:
return None
pts = pts.reshape(-1, 2) + np.array([x0, y0], dtype="float32")
return pts.reshape(-1, 1, 2)
def track_box(
video_path,
seed_box_norm: Box,
*,
seed_frame: int = 0,
max_frames: Optional[int] = None,
min_features: int = 8,
) -> tuple[dict[int, list[float]], int]:
"""Propagate a hand-seeded normalized box FORWARD from `seed_frame` with
optical flow (or CSRT when available). Returns `(per_frame_norm_boxes,
total_frames)` feed `per_frame` to `sample_track`. Translation model: the
box follows the median displacement of features inside it, re-seeding when too
few survive. Impure (reads the video) opt-in integration test."""
import cv2
import numpy as np
cap = cv2.VideoCapture(str(video_path))
total = int(cap.get(cv2.CAP_PROP_FRAME_COUNT) or 0)
if seed_frame:
cap.set(cv2.CAP_PROP_POS_FRAMES, seed_frame)
ok, frame = cap.read()
if not ok:
cap.release()
raise ValueError(f"could not read frame {seed_frame} of {video_path}")
h_img, w_img = frame.shape[:2]
box = list(denormalize_box(seed_box_norm, w_img, h_img)) # px [x,y,w,h]
csrt = _new_csrt()
if csrt is not None:
csrt.init(frame, tuple(int(v) for v in box))
prev_gray = cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY)
pts = None if csrt is not None else _features_in_box(prev_gray, box)
out: dict[int, list[float]] = {seed_frame: normalize_box(box, w_img, h_img)}
fidx = seed_frame
while True:
ok, frame = cap.read()
if not ok:
break
fidx += 1
if csrt is not None:
ok2, b = csrt.update(frame)
if ok2:
box = [b[0], b[1], b[2], b[3]]
else:
gray = cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY)
if pts is None or len(pts) < min_features:
pts = _features_in_box(prev_gray, box)
if pts is not None and len(pts):
nxt, status, _ = cv2.calcOpticalFlowPyrLK(prev_gray, gray, pts, None)
if nxt is not None:
good_old = pts[status == 1]
good_new = nxt[status == 1]
if len(good_new) >= 3:
dx = float(np.median(good_new[:, 0] - good_old[:, 0]))
dy = float(np.median(good_new[:, 1] - good_old[:, 1]))
box[0] += dx
box[1] += dy
pts = good_new.reshape(-1, 1, 2)
else:
pts = None
prev_gray = gray
out[fidx] = normalize_box(box, w_img, h_img)
if max_frames is not None and fidx - seed_frame >= max_frames:
break
cap.release()
return out, total
def track_seed(
video_path,
key: str,
seed_box_norm: Box,
*,
seed_frame: int = 0,
salience: int = 4,
n_keyframes: int = 5,
max_frames: Optional[int] = None,
) -> dict:
"""End-to-end classical pass: seed -> propagate -> sample -> annotation
(geometry only; the author adds strings/tiers). Impure (reads the video)."""
per_frame, total = track_box(
video_path, seed_box_norm, seed_frame=seed_frame, max_frames=max_frames
)
track = sample_track(per_frame, total, n_keyframes=n_keyframes)
appear, disappear = infer_window(list(per_frame), total)
return track_to_annotation(
key, track, salience=salience, appear=appear, disappear=disappear
)
# --- optional ML detect+track (lazy-imported) --------------------------------
def detect_and_track(video_path, *, model: str = "yolov8n.pt", classes=None) -> list[dict]:
"""Optional ML path (content-pipeline §3.5): run a detect+track model and
return candidate tracks `[{track_id, label, salience, track:[{t,box}],
appear, disappear}]` for the author to map to keys. Heavy deps are
lazy-imported so the base install stays light.
Raises a clear ImportError naming the install if `ultralytics` is absent
this path is opt-in by design; the classical `track_seed` needs only cv2."""
try:
from ultralytics import YOLO # noqa: F401
except ImportError as e: # pragma: no cover - exercised only without the dep
raise ImportError(
"the ML detect+track path needs ultralytics — "
"`pip install ultralytics` (optional; the classical track_seed path "
"uses only opencv-python)."
) from e
import cv2 # noqa: F401 (pragma: no cover below — needs the real model)
yolo = YOLO(model) # pragma: no cover
cap = cv2.VideoCapture(str(video_path)) # pragma: no cover
total = int(cap.get(cv2.CAP_PROP_FRAME_COUNT) or 0) # pragma: no cover
cap.release() # pragma: no cover
per_track: dict[int, dict[int, list[float]]] = {} # pragma: no cover
labels: dict[int, str] = {} # pragma: no cover
for fidx, res in enumerate( # pragma: no cover
yolo.track(source=str(video_path), persist=True, classes=classes, stream=True)
):
if res.boxes is None or res.boxes.id is None: # pragma: no cover
continue
w_img, h_img = res.orig_shape[1], res.orig_shape[0] # pragma: no cover
for box, tid, cls in zip( # pragma: no cover
res.boxes.xywh.tolist(), res.boxes.id.tolist(), res.boxes.cls.tolist()
):
cx, cy, bw, bh = box # pragma: no cover
nb = normalize_box((cx - bw / 2, cy - bh / 2, bw, bh), w_img, h_img)
per_track.setdefault(int(tid), {})[fidx] = nb # pragma: no cover
labels[int(tid)] = yolo.names.get(int(cls), str(cls))
out = [] # pragma: no cover
for tid, frames in per_track.items(): # pragma: no cover
track = sample_track(frames, total) # pragma: no cover
appear, disappear = infer_window(list(frames), total) # pragma: no cover
out.append({ # pragma: no cover
"track_id": tid, "label": labels.get(tid, ""), "salience": 4,
"track": track, "appear": round(appear, 4), "disappear": round(disappear, 4),
})
return out # pragma: no cover