From 70fd367c70e6165e82ed42559b5086793a90b94c Mon Sep 17 00:00:00 2001 From: BenStullsBets Date: Wed, 24 Jun 2026 18:14:32 -0700 Subject: [PATCH] tools(pipeline): hybrid track.py + simulator label author mode MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Content-pipeline Increment 2, part 2 (the tooling). Implements design §11.5 (docs/superpowers/specs/2026-06-24-content-pipeline-design.md): - tools/pipeline/track.py — the stage-5 geometry pass. Classical path: a hand-seeded normalized box propagated by OpenCV Lucas-Kanade optical flow (CSRT used instead when a contrib build provides it), sampled to a SPARSE loop-normalized keyframed track + an appear/disappear window. Optional ML detect+track path lazy-imports ultralytics (clear ImportError if absent; base install needs only cv2). Pure helpers (normalize/denormalize, loop_t, infer_window, sample_track, track_to_annotation) are unit-tested; the cv2 propagation is an opt-in integration test on real abyss_wow footage. Semantics are never produced here — geometry only. - Author mode — /author.html + author.js/.css reuse the preview stage: pick a pool clip, scrub, drag a seed box, Run tracker (or Add as static box), author the LEFT detail tiers (general -> scientific+fact) + salience, shift-click to place affect anchors with RIGHT emotion tiers, and Save to manifest. Backend: POST /api/author/track (runs track_seed on the clip's base) + POST /api/author/clip (idempotent upsert via tools.pipeline.manifest — keeps media + provenance, replaces only authored content, reloads in place). The tracker propagates box geometry; all strings/scientific names/facts are hand-typed. - Repointed the pipeline integration test off the retired forest base to the cosmos pool primary. USER_GUIDE simulator section brought current (pools, coast, 3-knob Mood, real-time dream, progressive tiers, author mode). 267 passed / 2 skipped (+ track pure + opt-in real-footage + author endpoint tests). Author UI by-eye review deferred to the operator (no Chrome on this box). Co-Authored-By: Claude Opus 4.8 (1M context) --- docs/USER_GUIDE.md | 85 ++--- .../2026-06-24-content-pipeline-design.md | 2 +- simulator/app.py | 75 +++++ simulator/static/author.css | 64 ++++ simulator/static/author.html | 79 +++++ simulator/static/author.js | 254 +++++++++++++++ tests/test_pipeline_integration.py | 20 +- tests/test_pipeline_track.py | 120 +++++++ tests/test_simulator_api.py | 75 +++++ tools/pipeline/track.py | 297 ++++++++++++++++++ 10 files changed, 1023 insertions(+), 48 deletions(-) create mode 100644 simulator/static/author.css create mode 100644 simulator/static/author.html create mode 100644 simulator/static/author.js create mode 100644 tests/test_pipeline_track.py create mode 100644 tools/pipeline/track.py diff --git a/docs/USER_GUIDE.md b/docs/USER_GUIDE.md index 59f737d..b77a611 100644 --- a/docs/USER_GUIDE.md +++ b/docs/USER_GUIDE.md @@ -296,24 +296,17 @@ exists. (The earlier selection-era "curator's X-ray" view was retired when the piece moved from *selecting* clips to *altering* them.) **One-time setup — populate the sample footage** (look-tuning only; not shipped -content): +content). The ring uses real strict-PD footage in a **rotating pool** per scale +(`docs/content-candidate-pool.md`); regenerate the manifest + new transition +placeholders with: - python simulator/setup_sample_media.py # stage the forest neutral base - python -m simulator.bake_right_variants --scale forest # real Right variants (SD on MPS, ~11 min) - python simulator/setup_scales_media.py # the cosmos + abyss scales + ring transitions + python simulator/build_pool_manifest.py --media # write the pool manifest + coast-edge transitions -`setup_sample_media.py` stages the forest neutral base from the session-0008 POC -artifacts (`~/hef-poc/out/neutral.mp4`). `bake_right_variants.py` then produces -the **real** flow-stabilized painterly Right variants (strengths 1–4) for the -forest scale — SD img2img keyframes + Farneback optical-flow tweens on Apple MPS, -the temporally-calm POC algorithm, productionized (scales design §1/§4). The -keyframe-strength ramp per Right level is by-eye tunable in -`simulator/bake_right_variants.py`. `setup_scales_media.py` makes the scale -**ring** demonstrable: cheap synthetic placeholder bases for the two true-PD -scales (`cosmos` = NASA/Hubble, `abyss` = NOAA Ocean Exploration) plus the -per-edge zoom/warp transition clips — these scales carry a raw base only (no real -Right variants yet; the baker can extend to them once their real footage is -sourced). The `.mp4` binaries are gitignored. +`build_pool_manifest.py` holds the hand-authored label/affect content and emits +`simulator/sample_media/manifest.json` (the 19-clip pool baseline). The pool clips +themselves are sourced via the content pipeline (`tools/pipeline/run.py +process_clip`); the `.mp4` binaries are gitignored. (`setup_scales_media.py` is the +older one-base-per-scale generator, now superseded.) **Run it (Docker):** @@ -331,27 +324,45 @@ then open http://localhost:8000. - **Content dial** — picks audio/video channel; "off" and audio-only positions go to black walls. - **Scale ring (endless encoder)** — `⊖ out` / `in ⊕` (or scroll the stage) walk a - *closed ring* of neutral "scales of nature" clips — cosmos → forest → abyss and - back around (diving past the smallest **wraps** to the largest). It is *relative* - (an endless encoder), distinct from the absolute 0–4 knobs: each step plays a - short placeholder zoom/warp **transition**, then settles on the next scale, with - the current knob alteration still applied. The current scale is named beside the - buttons (`name (i/N)`). A **fast spin** (scroll several detents at once, ≥3) - collapses to a single quick **blended pass** straight to the destination scale - instead of grinding through every transition (scales design §3); slow single - steps still chain one full transition per scale crossed. -- **Four experience knobs (0–4):** - - **Dark / Light** — a live runtime color grade (cool/dark ↔ warm/bright; equal - or zero = the raw footage). - - **Right (dreamlike)** — selects a discrete pre-baked, flow-stabilized restyle - variant and crossfades to it (strength 0 = raw base). - - **Left (analytical)** — a live overlay: labelled boxes from the clip's authored - annotation track, with more annotations appearing at higher levels. Text is - shaped live (the simulator analogue of the Pi's Pango/HarfBuzz path). -- **Calibration sliders** — adjust the grade/overlay gain curves live; once a look - is liked, bake the values into `DEFAULT_CALIBRATION` in `player/alteration.py`. + *closed ring* of neutral "scales of nature" — **cosmos → orbit → coast → reef → + abyss** and back around (diving past the smallest **wraps** to the largest). Each + scale is a **rotating pool** of vetted clips: landing on a scale plays a random + pool member, so the ring feels fresh each pass. It is *relative* (an endless + encoder), distinct from the absolute 0–4 knobs: each step plays a short + placeholder zoom/warp **transition**, then settles, with the current alteration + still applied. The current scale + chosen member is named beside the buttons + (`scale · member (i/N · pool N)`). A **fast spin** (≥3 detents at once) collapses + to one quick **blended pass** instead of grinding through every transition. +- **Three experience knobs:** + - **Mood (−4 dark .. 0 .. +4 light)** — a live runtime color grade (cool/dark ↔ + warm/bright; 0 = the raw footage). + - **Right (dreamlike, 0–4)** — a **deterministic real-time painterly dream** (a + WebGL Kuwahara restyle of the live frames; holds still across the loop, goes + trippy at max). It also drives **progressive emotion tiers**: when both knobs + are up, the affect words escalate from basic to compound as Right rises. + - **Left (analytical, 0–4)** — a live HUD of labelled boxes from the clip's + authored annotation track. Labels use **progressive detail tiers**: low Left + shows fewer, more general labels (high-salience objects only); raising Left + brings in more objects *and* escalates each label general → specific → + scientific → +fact. Time-windowed tracked labels appear only while their + subject is on screen and follow it. - **RenderPlan readout** — always shows the exact numbers the engine produced (the - project's honesty "X-ray," now over the alteration model). + project's honesty "X-ray," over the alteration model). -The base clips, Right variants, Left annotation track, and string tables come from +The base clips, the Left annotation tracks (with tiers + appear/disappear +windows), affect anchors, and per-tier string tables all come from `simulator/sample_media/manifest.json`. + +### Authoring labels — the author mode + +Open **http://localhost:8000/author.html** to author a clip's labels by eye +(content-pipeline §11.5). Pick a pool clip, **scrub** to where a subject is on +screen, **drag a box** over it, set the label key + salience + the four detail +tiers, then **Run tracker** to propagate the box into a keyframed track (classical +OpenCV optical flow) with an appear/disappear window — or **Add as static box** +for a fixed label. Shift-click the stage to place an **affect anchor** and type its +emotion tiers. **Save to manifest** writes the entry back via the pipeline's +idempotent upsert (keeping the clip's media + provenance, replacing only the +authored content). The tracker only propagates **box geometry** — every label +string, scientific name, and fact is **hand-authored**, because a generic detector +can't produce "*Gymnothorax*" or a lifespan. diff --git a/docs/superpowers/specs/2026-06-24-content-pipeline-design.md b/docs/superpowers/specs/2026-06-24-content-pipeline-design.md index f121d65..08a276f 100644 --- a/docs/superpowers/specs/2026-06-24-content-pipeline-design.md +++ b/docs/superpowers/specs/2026-06-24-content-pipeline-design.md @@ -1,7 +1,7 @@ # HEF — Content Pipeline (source → process → label → manifest) **Date:** 2026-06-24 -**Status:** Graduated — operator-approved (session 0014); **Increment 1 built & merged** (PR #13); real PD sourcing + rotating pool curated (PR #14/#15). **Increment 2 schema pinned + building** (session 0016 — see §11). +**Status:** Graduated — operator-approved (session 0014); **Increment 1 built & merged** (PR #13); real PD sourcing + rotating pool curated (PR #14/#15). **Increment 2 BUILT** (session 0016 — pool model + tiers + windows in PR #16; `track.py` + author mode in PR B; see §11). **Repo:** `human-experience-filter-art` **Anchor:** `docs/ROADMAP.md` sub-project 3 (Player Runtime) content needs; the spiritual successor to sub-project 2 (Ingest & Tagging tools), retargeted from the diff --git a/simulator/app.py b/simulator/app.py index 658bedf..df0f24a 100644 --- a/simulator/app.py +++ b/simulator/app.py @@ -8,6 +8,7 @@ endpoints (/api/select, /api/catalog/meta) and the X-ray are retired. from __future__ import annotations import hashlib +import json import os import random import time @@ -84,6 +85,29 @@ class RingAdvanceRequest(BaseModel): delta: int +class AuthorTrackRequest(BaseModel): + """Author mode: propagate a hand-seeded box on a clip into a keyframed track + (content-pipeline §11.5). `seed_box` + `seed_t` are normalized.""" + + clip_id: str + key: str + seed_box: list[float] = Field(min_length=4, max_length=4) + seed_t: float = Field(default=0.0, ge=0.0, le=1.0) + salience: int = Field(default=4, ge=1, le=4) + n_keyframes: int = Field(default=5, ge=2, le=20) + max_frames: Optional[int] = None + + +class AuthorClipRequest(BaseModel): + """Author mode: persist the authored labels/affect/strings for an existing + pool clip into the manifest (geometry from the tracker, semantics by hand).""" + + clip_id: str + annotations: list = [] + affect: list = [] + strings: dict = {} + + def _load_clips(manifest_path: Optional[Path]): path = Path(manifest_path) if manifest_path else DEFAULT_MANIFEST if path.exists(): @@ -100,6 +124,7 @@ def _load_ring(manifest_path: Optional[Path]): def create_app(manifest_path: Optional[Path] = None) -> FastAPI: app = FastAPI(title="HEF Alteration Simulator") + app.state.manifest_path = Path(manifest_path) if manifest_path else DEFAULT_MANIFEST app.state.clips = _load_clips(manifest_path) app.state.ring = _load_ring(manifest_path) @@ -158,6 +183,56 @@ def create_app(manifest_path: Optional[Path] = None) -> FastAPI: chosen = pick_clip_id(landed, random.random()) return ring_move_to_dict(move, app.state.ring, chosen) + @app.post("/api/author/track") + def api_author_track(req: AuthorTrackRequest): + # Author mode (§11.5): run the classical optical-flow tracker on a clip's + # base footage, returning the keyframed track geometry for the author to + # accept/correct. Semantics (key/strings/tiers) stay the author's. + clip = next((c for c in app.state.clips if c.id == req.clip_id), None) + if clip is None: + raise HTTPException(status_code=404, detail=f"unknown clip {req.clip_id!r}") + base = MEDIA_DIR / clip.base_file + if not base.exists(): + raise HTTPException(status_code=404, detail=f"media absent: {clip.base_file}") + try: + import cv2 + except ImportError: + raise HTTPException(status_code=503, detail="opencv-python not installed") + from tools.pipeline.track import track_seed + + cap = cv2.VideoCapture(str(base)) + total = int(cap.get(cv2.CAP_PROP_FRAME_COUNT) or 0) + cap.release() + seed_frame = max(0, min(total - 1, round(req.seed_t * total))) if total else 0 + return track_seed( + base, req.key, tuple(req.seed_box), seed_frame=seed_frame, + salience=req.salience, n_keyframes=req.n_keyframes, max_frames=req.max_frames, + ) + + @app.post("/api/author/clip") + def api_author_clip(req: AuthorClipRequest): + # Author mode (§11.5 / stage 6): persist authored labels/affect/strings for + # an EXISTING pool clip — idempotent upsert that keeps the clip's media + + # provenance and replaces only the authored content. Reloads in place so + # the preview reflects the edit without a restart. + from tools.pipeline.manifest import upsert_clip + + path = app.state.manifest_path + data = json.loads(path.read_text()) + existing = next((c for c in data.get("clips", []) if c["id"] == req.clip_id), None) + if existing is None: + raise HTTPException(status_code=404, detail=f"unknown clip {req.clip_id!r}") + merged = dict(existing) + merged["annotations"] = req.annotations + merged["affect"] = req.affect + if req.strings: + merged["strings"] = req.strings + data = upsert_clip(data, merged) + path.write_text(json.dumps(data, indent=2, ensure_ascii=False) + "\n") + app.state.clips = load_manifest(path) + app.state.ring = load_ring(path) + return {"ok": True, "clip": merged} + @app.middleware("http") async def _no_cache(request, call_next): # Dev preview server: never let a browser serve a stale app.js/style.css. diff --git a/simulator/static/author.css b/simulator/static/author.css new file mode 100644 index 0000000..8229fca --- /dev/null +++ b/simulator/static/author.css @@ -0,0 +1,64 @@ +/* Author-mode-only styling (content-pipeline §11.5). Reuses style.css for the + shared stage/panel chrome; this adds the box-draw layer + editor widgets. */ + +.author #draw { + position: absolute; + inset: 0; + width: 100%; + height: 100%; + cursor: crosshair; +} +.author #draw .seed-box { + fill: rgba(80, 200, 255, 0.12); + stroke: #50c8ff; + stroke-width: 0.4; + vector-effect: non-scaling-stroke; +} +.author #draw .track-box { + fill: rgba(120, 255, 160, 0.10); + stroke: #78ffa0; + stroke-width: 0.4; + stroke-dasharray: 1.5 1; + vector-effect: non-scaling-stroke; +} +.author #draw .affect-dot { + fill: rgba(190, 150, 255, 0.85); + stroke: #fff; + stroke-width: 0.25; + vector-effect: non-scaling-stroke; +} + +.author .scrub { + display: flex; + align-items: center; + gap: 8px; + margin-top: 8px; +} +.author .scrub input[type="range"] { flex: 1; } +.author .mono { font-family: ui-monospace, monospace; font-size: 12px; opacity: 0.85; } + +.author .row { display: flex; gap: 8px; margin-top: 6px; flex-wrap: wrap; } +.author fieldset label { display: block; margin: 4px 0; font-size: 13px; } +.author fieldset input[type="text"], +.author fieldset input[type="number"] { width: 100%; box-sizing: border-box; } +.author fieldset input[type="number"] { width: 5em; } + +.author ul.authored { list-style: none; padding: 0; margin: 6px 0; } +.author ul.authored li { + font-size: 12px; + padding: 3px 6px; + margin: 2px 0; + background: rgba(255, 255, 255, 0.05); + border-radius: 3px; + display: flex; + justify-content: space-between; + gap: 6px; +} +.author ul.authored .rm { + background: none; + border: none; + color: #f88; + cursor: pointer; + font-size: 12px; +} +.author header .hint { font-size: 12px; opacity: 0.8; max-width: 70ch; } diff --git a/simulator/static/author.html b/simulator/static/author.html new file mode 100644 index 0000000..a4f486f --- /dev/null +++ b/simulator/static/author.html @@ -0,0 +1,79 @@ + + + + + + HEF — Label Author Mode + + + + +

HEF — Label Author Mode

+

Scrub a pool clip · drag a box on a subject · run the tracker · author the tiers · save to the manifest. Geometry is propagated; semantics are hand-authored (content-pipeline §11.5). ← back to preview

+
+
+
+
+ + +
+
+ + + 0.00 +
+
+ +
+
+ Clip + +

+
+ +
+ Seed box (drag on the video) +

no box — drag on the stage

+ + + + + + + +
+ + +
+

+ +
+ +
+ Affect anchor (click the stage to place) +

no point — shift-click the stage

+ + + + + + + +
+ +
+ Authored for this clip +
    +
      +
      + + +
      +

      +
      +
      +
      + + + diff --git a/simulator/static/author.js b/simulator/static/author.js new file mode 100644 index 0000000..7dbb7e5 --- /dev/null +++ b/simulator/static/author.js @@ -0,0 +1,254 @@ +// Label author mode (content-pipeline §11.5): scrub a pool clip, drag a seed box +// on a subject, run the classical tracker to propagate a keyframed track, author +// the LEFT detail tiers + RIGHT emotion tiers, and save the entry to the manifest. +// Geometry comes from the tracker; SEMANTICS (keys, strings, tiers) are hand-typed. +const $ = (id) => document.getElementById(id); +const vid = $("vid"), draw = $("draw"); +const SVGNS = "http://www.w3.org/2000/svg"; + +let clips = []; +let currentClip = null; +let seedBox = null; // [x,y,w,h] normalized, from a drag +let pendingTrack = null; // {track, appear, disappear} from the tracker (or static) +let affectPoint = null; // [x,y] normalized, from a shift-click +let annotations = []; // authored annotation dicts (carry _tiers) +let affect = []; // authored affect dicts (carry _tiers) + +function svg(tag, attrs, parent) { + const el = document.createElementNS(SVGNS, tag); + for (const k in attrs) el.setAttribute(k, attrs[k]); + if (parent) parent.appendChild(el); + return el; +} +const mono = (n) => Number(n).toFixed(3); + +async function loadClips() { + clips = (await (await fetch("/api/clips")).json()).clips || []; + const sel = $("clip"); + sel.innerHTML = ""; + for (const c of clips) { + const o = document.createElement("option"); + o.value = c.id; o.textContent = c.id; + sel.appendChild(o); + } + if (clips.length) selectClip(clips[0].id); +} + +function selectClip(id) { + currentClip = clips.find((c) => c.id === id) || null; + if (!currentClip) return; + $("clip-meta").textContent = `${currentClip.title} — ${currentClip.license}`; + vid.src = "/media/" + currentClip.base_file; + vid.currentTime = 0; + annotations = []; affect = []; seedBox = null; pendingTrack = null; affectPoint = null; + renderLists(); renderDrawLayer(); + $("seed-readout").textContent = "no box — drag on the stage"; + $("track-readout").textContent = ""; + $("affect-readout").textContent = "no point — shift-click the stage"; +} + +// --- scrub --- +function seekT() { return vid.duration ? vid.currentTime / vid.duration : 0; } +$("play").addEventListener("click", () => { vid.paused ? vid.play() : vid.pause(); }); +$("seek").addEventListener("input", () => { + if (vid.duration) vid.currentTime = (+$("seek").value / 1000) * vid.duration; +}); +vid.addEventListener("timeupdate", () => { + if (!vid.duration) return; + $("seek").value = String(Math.round(seekT() * 1000)); + $("time").textContent = seekT().toFixed(3); +}); + +// --- draw a seed box / place an affect point on the stage --- +function evtNorm(e) { + const r = draw.getBoundingClientRect(); + return [(e.clientX - r.left) / r.width, (e.clientY - r.top) / r.height]; +} +let dragStart = null; +draw.addEventListener("mousedown", (e) => { + if (e.shiftKey) { // shift-click = affect anchor + affectPoint = evtNorm(e).map((v) => Math.min(Math.max(v, 0), 1)); + $("affect-readout").textContent = `affect @ ${affectPoint.map(mono).join(", ")}`; + $("add-aff").disabled = false; + renderDrawLayer(); + return; + } + dragStart = evtNorm(e); +}); +draw.addEventListener("mousemove", (e) => { + if (!dragStart) return; + seedBox = boxFrom(dragStart, evtNorm(e)); + renderDrawLayer(); +}); +window.addEventListener("mouseup", (e) => { + if (!dragStart) return; + seedBox = boxFrom(dragStart, evtNorm(e)); + dragStart = null; + pendingTrack = null; + $("seed-readout").textContent = `seed [${seedBox.map(mono).join(", ")}] @ t=${seekT().toFixed(3)}`; + $("add-ann").disabled = false; + renderDrawLayer(); +}); +function boxFrom(a, b) { + const x = Math.min(a[0], b[0]), y = Math.min(a[1], b[1]); + const w = Math.abs(b[0] - a[0]), h = Math.abs(b[1] - a[1]); + const cl = (v) => Math.min(Math.max(v, 0), 1); + return [cl(x), cl(y), Math.min(w, 1 - cl(x)), Math.min(h, 1 - cl(y))]; +} + +function renderDrawLayer() { + draw.innerHTML = ""; + if (seedBox) { + const [x, y, w, h] = seedBox.map((n) => n * 100); + svg("rect", { x, y, width: w, height: h, class: "seed-box" }, draw); + } + if (pendingTrack && pendingTrack.track && vid.duration) { + const b = boxAt(pendingTrack.track, seekT()); + if (b) { + const [x, y, w, h] = b.map((n) => n * 100); + svg("rect", { x, y, width: w, height: h, class: "track-box" }, draw); + } + } + if (affectPoint) { + const [x, y] = affectPoint.map((n) => n * 100); + svg("circle", { cx: x, cy: y, r: 1.2, class: "affect-dot" }, draw); + } +} +// Interpolate a track box at loop-normalized t (mirrors the preview renderer). +function boxAt(track, t) { + if (!track || !track.length) return null; + if (t <= track[0].t) return track[0].box; + if (t >= track[track.length - 1].t) return track[track.length - 1].box; + for (let i = 1; i < track.length; i++) { + if (t <= track[i].t) { + const a = track[i - 1], b = track[i], f = (t - a.t) / (b.t - a.t); + return a.box.map((v, j) => v + (b.box[j] - v) * f); + } + } + return track[track.length - 1].box; +} +// Animate the pending track preview while playing. +function previewLoop() { renderDrawLayer(); requestAnimationFrame(previewLoop); } + +// --- run the tracker on the seed box --- +$("run-track").addEventListener("click", async () => { + if (!seedBox || !currentClip) { $("track-readout").textContent = "draw a seed box first"; return; } + const key = $("ann-key").value.trim() || "detected.object"; + $("track-readout").textContent = "tracking…"; + const resp = await fetch("/api/author/track", { + method: "POST", headers: { "content-type": "application/json" }, + body: JSON.stringify({ + clip_id: currentClip.id, key, seed_box: seedBox, seed_t: seekT(), + salience: +$("ann-salience").value, n_keyframes: +$("ann-keyframes").value, + }), + }); + if (!resp.ok) { $("track-readout").textContent = "tracker error " + resp.status; return; } + const ann = await resp.json(); + pendingTrack = { track: ann.track, appear: ann.appear, disappear: ann.disappear }; + $("track-readout").textContent = + `track: ${ann.track.length} keyframes · window ${mono(ann.appear)}–${mono(ann.disappear)}`; + $("add-ann").disabled = false; +}); + +// Add the current seed as a STATIC (untracked) box at the scrub position. +$("add-static").addEventListener("click", () => { + if (!seedBox) return; + pendingTrack = { static: true, box: seedBox.slice() }; + $("track-readout").textContent = "static box (no track) at the drawn position"; + $("add-ann").disabled = false; +}); + +function tiers(ids) { + const vals = ids.map((id) => $(id).value.trim()); + while (vals.length && !vals[vals.length - 1]) vals.pop(); // drop empty trailing tiers + return vals.length ? vals : null; +} + +$("add-ann").addEventListener("click", () => { + const key = $("ann-key").value.trim(); + if (!key) { $("track-readout").textContent = "a label key is required"; return; } + const t = tiers(["t1", "t2", "t3", "t4"]); + const ann = { key, salience: +$("ann-salience").value, _tiers: t || key }; + if (pendingTrack && pendingTrack.track) { + ann.track = pendingTrack.track; + ann.appear = pendingTrack.appear; ann.disappear = pendingTrack.disappear; + } else if (pendingTrack && pendingTrack.static) { + ann.box = pendingTrack.box; + } else if (seedBox) { + ann.box = seedBox.slice(); + } else { $("track-readout").textContent = "draw + (optionally) track a box first"; return; } + annotations.push(ann); + seedBox = null; pendingTrack = null; $("add-ann").disabled = true; + renderLists(); renderDrawLayer(); +}); + +$("add-aff").addEventListener("click", () => { + const key = $("aff-key").value.trim(); + if (!key || !affectPoint) return; + const t = tiers(["a1", "a2", "a3", "a4"]); + affect.push({ key, at: affectPoint.slice(), min_level: +$("aff-min").value, _tiers: t || key }); + affectPoint = null; $("add-aff").disabled = true; + renderLists(); renderDrawLayer(); +}); + +function renderLists() { + const al = $("ann-list"); al.innerHTML = ""; + annotations.forEach((a, i) => { + const li = document.createElement("li"); + const kind = a.track ? `track ${a.track.length}kf ${mono(a.appear)}–${mono(a.disappear)}` : "static"; + const tip = Array.isArray(a._tiers) ? a._tiers.join(" → ") : a._tiers; + li.textContent = `${a.key} (s${a.salience}, ${kind}) — ${tip}`; + li.appendChild(rm(() => { annotations.splice(i, 1); renderLists(); })); + al.appendChild(li); + }); + const fl = $("aff-list"); fl.innerHTML = ""; + affect.forEach((f, i) => { + const li = document.createElement("li"); + const tip = Array.isArray(f._tiers) ? f._tiers.join(" → ") : f._tiers; + li.textContent = `${f.key} (min ${f.min_level}) — ${tip}`; + li.appendChild(rm(() => { affect.splice(i, 1); renderLists(); })); + fl.appendChild(li); + }); +} +function rm(fn) { + const b = document.createElement("button"); + b.textContent = "✕"; b.className = "rm"; b.onclick = fn; + return b; +} + +// Load whatever is already authored for this clip into the editor. +$("load-existing").addEventListener("click", () => { + if (!currentClip) return; + const strings = (currentClip.strings && currentClip.strings.en) || {}; + annotations = (currentClip.annotations || []).map((a) => ({ ...a, _tiers: strings[a.key] || a.key })); + affect = (currentClip.affect || []).map((f) => ({ ...f, _tiers: strings[f.key] || f.key })); + renderLists(); +}); + +// Assemble strings.en from each entry's _tiers and POST to the manifest. +$("save").addEventListener("click", async () => { + if (!currentClip) return; + const en = {}; + const strip = (x) => { const { _tiers, ...rest } = x; return rest; }; + const anns = annotations.map((a) => { en[a.key] = a._tiers; return strip(a); }); + const affs = affect.map((f) => { en[f.key] = f._tiers; return strip(f); }); + const resp = await fetch("/api/author/clip", { + method: "POST", headers: { "content-type": "application/json" }, + body: JSON.stringify({ clip_id: currentClip.id, annotations: anns, affect: affs, strings: { en } }), + }); + $("save-readout").textContent = resp.ok + ? `saved ${anns.length} labels + ${affs.length} emotions to the manifest ✓` + : "save error " + resp.status; + if (resp.ok) { // refresh the in-memory clip copy + clips = (await (await fetch("/api/clips")).json()).clips || []; + currentClip = clips.find((c) => c.id === currentClip.id) || currentClip; + } +}); + +$("clip").addEventListener("change", (e) => selectClip(e.target.value)); + +async function main() { + await loadClips(); + requestAnimationFrame(previewLoop); +} +main(); diff --git a/tests/test_pipeline_integration.py b/tests/test_pipeline_integration.py index 6cd2b35..d8a8b75 100644 --- a/tests/test_pipeline_integration.py +++ b/tests/test_pipeline_integration.py @@ -4,7 +4,9 @@ import pytest from tools.pipeline.run import probe_duration, process_clip, resolve_ffmpeg -_FOREST = Path("simulator/sample_media/forest/base.mp4") +# A real pool base (forest was retired with the rotating-pool rename, session 0016); +# cosmos/base.mp4 is the cosmos pool primary and the most stable PD clip on disk. +_BASE = Path("simulator/sample_media/cosmos/base.mp4") def _ffmpeg_available() -> bool: @@ -18,20 +20,18 @@ def _ffmpeg_available() -> bool: _HAS_FF = _ffmpeg_available() -@pytest.mark.skipif(not _FOREST.exists(), - reason="forest POC base absent (run simulator/setup_sample_media.py)") -def test_probe_duration_reads_forest_base(): - if not _FOREST.exists(): - pytest.skip("forest POC base absent") - assert probe_duration(_FOREST) > 0 +@pytest.mark.skipif(not _BASE.exists(), + reason="cosmos base absent (run simulator/build_pool_manifest.py)") +def test_probe_duration_reads_real_base(): + assert probe_duration(_BASE) > 0 @pytest.mark.skipif(not _HAS_FF, reason="neither system ffmpeg nor imageio-ffmpeg available") -@pytest.mark.skipif(not _FOREST.exists(), - reason="forest POC base absent (run simulator/setup_sample_media.py)") +@pytest.mark.skipif(not _BASE.exists(), + reason="cosmos base absent (run simulator/build_pool_manifest.py)") def test_process_clip_emits_proxy_and_master(tmp_path): - out = process_clip(_FOREST, tmp_path, overlap=0.5, start=0.0, duration=4.0) + out = process_clip(_BASE, tmp_path, overlap=0.5, start=0.0, duration=4.0) assert out["proxy"].exists() and out["master"].exists() # Proxy is exactly 1920×1080 — verified with cv2, no ffprobe dependency. import cv2 diff --git a/tests/test_pipeline_track.py b/tests/test_pipeline_track.py new file mode 100644 index 0000000..6b09b0d --- /dev/null +++ b/tests/test_pipeline_track.py @@ -0,0 +1,120 @@ +"""tools/pipeline/track.py — the hybrid motion-track pass (content-pipeline §11.5). + +Pure geometry helpers are unit-tested here; the cv2 optical-flow propagation is an +opt-in integration test that runs on a real pool clip when present.""" + +from pathlib import Path + +import pytest + +from tools.pipeline.track import ( + clamp01_box, + denormalize_box, + infer_window, + loop_t, + normalize_box, + sample_track, + track_to_annotation, +) + + +# --- pure helpers ----------------------------------------------------------- + + +def test_clamp01_box_keeps_box_on_screen(): + assert clamp01_box((-0.2, 0.1, 0.5, 0.5)) == [0.0, 0.1, 0.5, 0.5] + # size trimmed so the box never runs past the right/bottom edge + assert clamp01_box((0.8, 0.8, 0.5, 0.5)) == [0.8, 0.8, pytest.approx(0.2), pytest.approx(0.2)] + + +def test_normalize_denormalize_round_trip(): + px = (192, 108, 384, 216) + norm = normalize_box(px, 1920, 1080) + assert norm == [pytest.approx(0.1), pytest.approx(0.1), pytest.approx(0.2), pytest.approx(0.2)] + assert denormalize_box(norm, 1920, 1080) == (192, 108, 384, 216) + + +def test_loop_t_is_frame_over_total(): + assert loop_t(0, 100) == 0.0 + assert loop_t(50, 100) == 0.5 + assert loop_t(5, 0) == 0.0 # guard divide-by-zero + + +def test_infer_window_spans_tracked_frames(): + appear, disappear = infer_window([10, 11, 30, 31], 100) + assert appear == pytest.approx(0.1) + assert disappear == pytest.approx(0.31) + # no frames -> full clip + assert infer_window([], 100) == (0.0, 1.0) + + +def test_sample_track_downsamples_to_sparse_keyframes(): + # 21 dense frames -> at most n_keyframes, including the first and last + per_frame = {i: [i / 100, 0.2, 0.1, 0.1] for i in range(0, 21)} + track = sample_track(per_frame, total_frames=100, n_keyframes=5) + assert len(track) <= 5 + ts = [k["t"] for k in track] + assert ts == sorted(ts) # sorted by t + assert ts[0] == pytest.approx(0.0) # first tracked frame + assert ts[-1] == pytest.approx(0.2) # last tracked frame (frame 20 / 100) + assert len(set(ts)) == len(ts) # de-duplicated on t + + +def test_sample_track_keeps_all_when_already_sparse(): + per_frame = {0: [0.1, 0.1, 0.1, 0.1], 60: [0.5, 0.3, 0.1, 0.1]} + track = sample_track(per_frame, total_frames=120, n_keyframes=5) + assert [k["t"] for k in track] == [pytest.approx(0.0), pytest.approx(0.5)] + + +def test_sample_track_empty_is_empty(): + assert sample_track({}, total_frames=100) == [] + + +def test_track_to_annotation_assembles_geometry_only(): + track = [{"t": 0.0, "box": [0.1, 0.2, 0.1, 0.1]}, {"t": 0.5, "box": [0.4, 0.2, 0.1, 0.1]}] + ann = track_to_annotation("detected.jelly", track, salience=3, appear=0.0, disappear=0.55) + assert ann["key"] == "detected.jelly" + assert ann["salience"] == 3 + assert ann["track"] == track + assert ann["appear"] == 0.0 and ann["disappear"] == 0.55 + # no semantics leaked in — strings/tiers are the author's, added elsewhere + assert "strings" not in ann and "_tiers" not in ann + + +def test_track_to_annotation_omits_window_when_absent(): + ann = track_to_annotation("detected.x", [], salience=4) + assert "appear" not in ann and "disappear" not in ann + + +# --- opt-in: classical optical-flow propagation on a real clip -------------- + +_CLIP = Path("simulator/sample_media/abyss_wow/base.mp4") + + +def _cv2_available() -> bool: + try: + import cv2 # noqa: F401 + return True + except Exception: + return False + + +@pytest.mark.skipif(not _CLIP.exists(), + reason="abyss_wow base absent (run simulator/build_pool_manifest.py)") +@pytest.mark.skipif(not _cv2_available(), reason="opencv-python not available") +def test_track_seed_produces_a_sparse_annotation_on_real_footage(): + from tools.pipeline.track import track_box, track_seed + + seed = (0.2, 0.3, 0.15, 0.18) + per_frame, total = track_box(_CLIP, seed, seed_frame=0, max_frames=40) + assert total > 0 + assert per_frame[0] == pytest.approx(list(seed), abs=0.02) + assert len(per_frame) > 1 # propagated beyond the seed frame + + ann = track_seed(_CLIP, "detected.jelly", seed, seed_frame=0, n_keyframes=5, max_frames=40) + assert ann["key"] == "detected.jelly" + assert 2 <= len(ann["track"]) <= 5 # sparse keyframes + for kf in ann["track"]: + assert 0.0 <= kf["t"] <= 1.0 + assert all(0.0 <= v <= 1.0 for v in kf["box"]) + assert 0.0 <= ann["appear"] <= ann["disappear"] <= 1.0 diff --git a/tests/test_simulator_api.py b/tests/test_simulator_api.py index fcd227a..9acdeb8 100644 --- a/tests/test_simulator_api.py +++ b/tests/test_simulator_api.py @@ -250,3 +250,78 @@ def test_delta_zero_is_an_initial_pool_pick(pool_client): assert data["to_index"] == 1 assert data["steps"] == [] assert data["target_clip_id"] in members + + +# --- author mode (content-pipeline §11.5) --- + + +@pytest.fixture +def author_client(tmp_path): + p = tmp_path / "manifest.json" + p.write_text(json.dumps({ + "clips": [{ + "id": "abyss_wow", "title": "World of Water", "base_file": "abyss_wow/base.mp4", + "license": "PD", "source": "NOAA", "right_variants": {}, + "annotations": [], "affect": [], "strings": {"en": {}}, + }], + })) + return TestClient(create_app(manifest_path=p)), p + + +def test_author_clip_persists_labels_and_keeps_provenance(author_client): + client, path = author_client + body = { + "clip_id": "abyss_wow", + "annotations": [{"key": "detected.jelly", "salience": 4, + "appear": 0.0, "disappear": 0.5, + "track": [{"t": 0.0, "box": [0.1, 0.2, 0.1, 0.1]}]}], + "affect": [{"key": "feel.unease", "at": [0.5, 0.5], "min_level": 1}], + "strings": {"en": {"detected.jelly": ["jelly", "comb jelly", "Ctenophora", "Ctenophora · cilia rows"], + "feel.unease": ["uh", "unease", "disquiet", "a creeping disquiet"]}}, + } + resp = client.post("/api/author/clip", json=body) + assert resp.status_code == 200 and resp.json()["ok"] is True + # persisted to disk + saved = json.loads(path.read_text())["clips"][0] + assert saved["annotations"][0]["key"] == "detected.jelly" + assert saved["affect"][0]["key"] == "feel.unease" + assert saved["strings"]["en"]["detected.jelly"][2] == "Ctenophora" + # provenance preserved (only authored content replaced) + assert saved["base_file"] == "abyss_wow/base.mp4" + assert saved["license"] == "PD" + # reflected by /api/clips without a restart + served = client.get("/api/clips").json()["clips"][0] + assert served["annotations"][0]["key"] == "detected.jelly" + + +def test_author_clip_unknown_clip_404(author_client): + client, _ = author_client + resp = client.post("/api/author/clip", json={"clip_id": "nope", "annotations": []}) + assert resp.status_code == 404 + + +def test_author_track_unknown_clip_404(author_client): + client, _ = author_client + resp = client.post("/api/author/track", json={ + "clip_id": "nope", "key": "detected.x", "seed_box": [0.1, 0.1, 0.1, 0.1]}) + assert resp.status_code == 404 + + +_REAL_CLIP = __import__("pathlib").Path("simulator/sample_media/abyss_wow/base.mp4") + + +@pytest.mark.skipif(not _REAL_CLIP.exists(), + reason="abyss_wow base absent (run simulator/build_pool_manifest.py)") +def test_author_track_on_real_clip_returns_keyframes(): + # Against the real default manifest + footage: the tracker yields a sparse track. + client = TestClient(create_app()) + resp = client.post("/api/author/track", json={ + "clip_id": "abyss_wow", "key": "detected.jelly", + "seed_box": [0.2, 0.3, 0.15, 0.18], "seed_t": 0.0, + "salience": 4, "n_keyframes": 5, "max_frames": 30, + }) + assert resp.status_code == 200 + ann = resp.json() + assert ann["key"] == "detected.jelly" + assert 2 <= len(ann["track"]) <= 5 + assert 0.0 <= ann["appear"] <= ann["disappear"] <= 1.0 diff --git a/tools/pipeline/track.py b/tools/pipeline/track.py new file mode 100644 index 0000000..a93af46 --- /dev/null +++ b/tools/pipeline/track.py @@ -0,0 +1,297 @@ +"""Stage 5 geometry — the hybrid motion-track pass (content-pipeline §3.5 / §11.5). + +Turns a hand-seeded box (or an ML detection) on one keyframe into the sparse, +loop-normalized keyframed `track` the manifest stores (design §4) plus an +appear/disappear window (§11.2). SEMANTICS are never produced here — the author +owns the `key`, strings, `salience` and tiers; this module only propagates BOX +GEOMETRY. + +Two geometry paths, same artifact: + - **classical** (always available): a hand-seeded box propagated by OpenCV + Lucas-Kanade optical flow (a CSRT tracker is used instead when a contrib + build provides one). Translation model — right for calm drifting subjects + (the abyss/reef creatures). Re-seeds features when too few survive. + - **ML detect+track** (optional, lazy-imported): a YOLO+tracker pass proposing + boxes the author maps to a key. Imported only on demand so the base install + needs no torch/ultralytics; a clear ImportError tells you what to `pip + install` if you ask for it without it present. + +The PURE helpers (normalization, keyframe sampling, window inference, annotation +assembly) are unit-tested; the cv2/ML propagation is covered by the opt-in +integration test on real footage. +""" + +from __future__ import annotations + +from typing import Optional + +Box = tuple[float, float, float, float] # (x, y, w, h) + + +# --- pure geometry helpers -------------------------------------------------- + +def clamp01_box(box: Box) -> list[float]: + """Clamp a normalized box into the visible [0,1] frame, keeping it on-screen + (origin clamped, then size trimmed so x+w<=1, y+h<=1).""" + x, y, w, h = box + x = min(max(x, 0.0), 1.0) + y = min(max(y, 0.0), 1.0) + w = min(max(w, 0.0), 1.0 - x) + h = min(max(h, 0.0), 1.0 - y) + return [x, y, w, h] + + +def normalize_box(box_px: Box, width: int, height: int) -> list[float]: + """Pixel box -> normalized [0,1] box (the manifest's resolution-independent + form). Clamped on-screen.""" + x, y, w, h = box_px + return clamp01_box((x / width, y / height, w / width, h / height)) + + +def denormalize_box(box_norm: Box, width: int, height: int) -> tuple[int, int, int, int]: + """Normalized [0,1] box -> integer pixel box for the cv2 tracker.""" + x, y, w, h = box_norm + return (round(x * width), round(y * height), round(w * width), round(h * height)) + + +def loop_t(frame_idx: int, total_frames: int) -> float: + """Loop-normalized time of a frame in [0,1), matching the client's + `currentTime / duration` clock (design §4).""" + if total_frames <= 0: + return 0.0 + return frame_idx / total_frames + + +def infer_window(tracked_frames: list[int], total_frames: int) -> tuple[float, float]: + """The appear/disappear window (§11.2) covering the frames where a box was + tracked, as loop-normalized t. The first/last tracked frame bound it.""" + if not tracked_frames: + return (0.0, 1.0) + lo, hi = min(tracked_frames), max(tracked_frames) + return (loop_t(lo, total_frames), loop_t(hi, total_frames)) + + +def sample_track( + per_frame: dict[int, Box], total_frames: int, n_keyframes: int = 5 +) -> list[dict]: + """Down-sample a dense per-frame box map to a SPARSE keyframed track + (design §4: sparse keyframes interpolated at runtime, never dense per-frame). + + Picks `n_keyframes` frames evenly across the tracked span (always including + the first and last tracked frame) and emits `[{t, box}, ...]` sorted by t, + de-duplicated on t. Pure — takes already-normalized boxes.""" + if not per_frame: + return [] + frames = sorted(per_frame) + n = max(2, n_keyframes) + if len(frames) <= n: + picks = frames + else: + lo, hi = frames[0], frames[-1] + step = (hi - lo) / (n - 1) + targets = [lo + round(step * i) for i in range(n)] + # snap each target to the nearest actually-tracked frame + picks = sorted({min(frames, key=lambda f: abs(f - t)) for t in targets}) + out, seen = [], set() + for f in picks: + t = round(loop_t(f, total_frames), 4) + if t in seen: + continue + seen.add(t) + out.append({"t": t, "box": clamp01_box(per_frame[f])}) + return out + + +def track_to_annotation( + key: str, + track: list[dict], + *, + salience: int = 4, + appear: Optional[float] = None, + disappear: Optional[float] = None, +) -> dict: + """Assemble a manifest annotation from a propagated track + the author's + semantics. Geometry (track + window) comes from this module; `key`, + `salience` and the tiered strings are the author's (added separately).""" + ann: dict = {"key": key, "salience": salience, "track": track} + if appear is not None: + ann["appear"] = round(appear, 4) + if disappear is not None: + ann["disappear"] = round(disappear, 4) + return ann + + +# --- classical propagation (cv2; covered by the opt-in integration test) ----- + +def _new_csrt(): + """A CSRT tracker if this OpenCV build ships one (contrib), else None — then + the optical-flow path is used. Kept tiny so the import stays lazy.""" + import cv2 + for factory in ("TrackerCSRT_create",): + if hasattr(cv2, factory): + return getattr(cv2, factory)() + legacy = getattr(cv2, "legacy", None) + if legacy is not None and hasattr(legacy, "TrackerCSRT_create"): + return legacy.TrackerCSRT_create() + return None + + +def _features_in_box(gray, box_px, max_corners: int = 60): + """Good-feature points inside a pixel box, in full-frame coords (for LK).""" + import cv2 + import numpy as np + + x, y, w, h = (int(v) for v in box_px) + h_img, w_img = gray.shape[:2] + x0, y0 = max(x, 0), max(y, 0) + x1, y1 = min(x + w, w_img), min(y + h, h_img) + if x1 - x0 < 4 or y1 - y0 < 4: + return None + roi = gray[y0:y1, x0:x1] + pts = cv2.goodFeaturesToTrack(roi, max_corners, 0.01, 5) + if pts is None: + return None + pts = pts.reshape(-1, 2) + np.array([x0, y0], dtype="float32") + return pts.reshape(-1, 1, 2) + + +def track_box( + video_path, + seed_box_norm: Box, + *, + seed_frame: int = 0, + max_frames: Optional[int] = None, + min_features: int = 8, +) -> tuple[dict[int, list[float]], int]: + """Propagate a hand-seeded normalized box FORWARD from `seed_frame` with + optical flow (or CSRT when available). Returns `(per_frame_norm_boxes, + total_frames)` — feed `per_frame` to `sample_track`. Translation model: the + box follows the median displacement of features inside it, re-seeding when too + few survive. Impure (reads the video) — opt-in integration test.""" + import cv2 + import numpy as np + + cap = cv2.VideoCapture(str(video_path)) + total = int(cap.get(cv2.CAP_PROP_FRAME_COUNT) or 0) + if seed_frame: + cap.set(cv2.CAP_PROP_POS_FRAMES, seed_frame) + ok, frame = cap.read() + if not ok: + cap.release() + raise ValueError(f"could not read frame {seed_frame} of {video_path}") + h_img, w_img = frame.shape[:2] + box = list(denormalize_box(seed_box_norm, w_img, h_img)) # px [x,y,w,h] + + csrt = _new_csrt() + if csrt is not None: + csrt.init(frame, tuple(int(v) for v in box)) + + prev_gray = cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY) + pts = None if csrt is not None else _features_in_box(prev_gray, box) + out: dict[int, list[float]] = {seed_frame: normalize_box(box, w_img, h_img)} + + fidx = seed_frame + while True: + ok, frame = cap.read() + if not ok: + break + fidx += 1 + if csrt is not None: + ok2, b = csrt.update(frame) + if ok2: + box = [b[0], b[1], b[2], b[3]] + else: + gray = cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY) + if pts is None or len(pts) < min_features: + pts = _features_in_box(prev_gray, box) + if pts is not None and len(pts): + nxt, status, _ = cv2.calcOpticalFlowPyrLK(prev_gray, gray, pts, None) + if nxt is not None: + good_old = pts[status == 1] + good_new = nxt[status == 1] + if len(good_new) >= 3: + dx = float(np.median(good_new[:, 0] - good_old[:, 0])) + dy = float(np.median(good_new[:, 1] - good_old[:, 1])) + box[0] += dx + box[1] += dy + pts = good_new.reshape(-1, 1, 2) + else: + pts = None + prev_gray = gray + out[fidx] = normalize_box(box, w_img, h_img) + if max_frames is not None and fidx - seed_frame >= max_frames: + break + cap.release() + return out, total + + +def track_seed( + video_path, + key: str, + seed_box_norm: Box, + *, + seed_frame: int = 0, + salience: int = 4, + n_keyframes: int = 5, + max_frames: Optional[int] = None, +) -> dict: + """End-to-end classical pass: seed -> propagate -> sample -> annotation + (geometry only; the author adds strings/tiers). Impure (reads the video).""" + per_frame, total = track_box( + video_path, seed_box_norm, seed_frame=seed_frame, max_frames=max_frames + ) + track = sample_track(per_frame, total, n_keyframes=n_keyframes) + appear, disappear = infer_window(list(per_frame), total) + return track_to_annotation( + key, track, salience=salience, appear=appear, disappear=disappear + ) + + +# --- optional ML detect+track (lazy-imported) -------------------------------- + +def detect_and_track(video_path, *, model: str = "yolov8n.pt", classes=None) -> list[dict]: + """Optional ML path (content-pipeline §3.5): run a detect+track model and + return candidate tracks `[{track_id, label, salience, track:[{t,box}], + appear, disappear}]` for the author to map to keys. Heavy deps are + lazy-imported so the base install stays light. + + Raises a clear ImportError naming the install if `ultralytics` is absent — + this path is opt-in by design; the classical `track_seed` needs only cv2.""" + try: + from ultralytics import YOLO # noqa: F401 + except ImportError as e: # pragma: no cover - exercised only without the dep + raise ImportError( + "the ML detect+track path needs ultralytics — " + "`pip install ultralytics` (optional; the classical track_seed path " + "uses only opencv-python)." + ) from e + import cv2 # noqa: F401 (pragma: no cover below — needs the real model) + + yolo = YOLO(model) # pragma: no cover + cap = cv2.VideoCapture(str(video_path)) # pragma: no cover + total = int(cap.get(cv2.CAP_PROP_FRAME_COUNT) or 0) # pragma: no cover + cap.release() # pragma: no cover + per_track: dict[int, dict[int, list[float]]] = {} # pragma: no cover + labels: dict[int, str] = {} # pragma: no cover + for fidx, res in enumerate( # pragma: no cover + yolo.track(source=str(video_path), persist=True, classes=classes, stream=True) + ): + if res.boxes is None or res.boxes.id is None: # pragma: no cover + continue + w_img, h_img = res.orig_shape[1], res.orig_shape[0] # pragma: no cover + for box, tid, cls in zip( # pragma: no cover + res.boxes.xywh.tolist(), res.boxes.id.tolist(), res.boxes.cls.tolist() + ): + cx, cy, bw, bh = box # pragma: no cover + nb = normalize_box((cx - bw / 2, cy - bh / 2, bw, bh), w_img, h_img) + per_track.setdefault(int(tid), {})[fidx] = nb # pragma: no cover + labels[int(tid)] = yolo.names.get(int(cls), str(cls)) + out = [] # pragma: no cover + for tid, frames in per_track.items(): # pragma: no cover + track = sample_track(frames, total) # pragma: no cover + appear, disappear = infer_window(list(frames), total) # pragma: no cover + out.append({ # pragma: no cover + "track_id": tid, "label": labels.get(tid, ""), "salience": 4, + "track": track, "appear": round(appear, 4), "disappear": round(disappear, 4), + }) + return out # pragma: no cover