Compare commits

...

10 Commits

Author SHA1 Message Date
Ben Stull 1e2a4dd966 docs: user guide for ingest + review tools
Per sub-project-2 plan Task 14. Updates the scope banner (catalog core + tools
now built), adds an 'Ingesting & reviewing media' section (ffmpeg prereqs,
--media-root, first-ship archives, ingest_cli/review_cli usage, Freesound token
note, opt-in dominant_color), and reconciles the absolute-vs-relative file_path
note (spec §6.3).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 06:36:15 -07:00
Ben Stull deb26b2575 test: end-to-end ingest+review integration (+ opt-in ffprobe)
Per sub-project-2 plan Task 13 / spec §10 test 8. Hermetic e2e: fake fetcher ->
ingest_candidate -> proposed record; approve flips it; reload + validate_catalog;
select(approved_only=True) finds it. Opt-in real-ffprobe/ffmpeg tests generate a
lavfi clip and assert mode/duration/resolution + dominant color, skipped when the
binaries are absent.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 06:35:17 -07:00
Ben Stull 376edd7819 feat: interactive review CLI
Per sub-project-2 plan Task 12 / spec §8.2. Walks proposed records (fields +
coords + rationale + best-effort ffmpeg preview), prompts accept/edit/skip/quit,
and persists each approval via save_catalog rewrite. I/O seams (input_fn/now_fn/
out) and --no-preview make the walk testable hermetically; decision logic lives
in the tested review core.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 06:34:34 -07:00
Ben Stull 3e07e86187 feat: review transition core (proposed -> approved)
Per sub-project-2 plan Task 11 / spec §8.1. proposed_records filter and approve()
return an approved copy via dataclasses.replace (no in-place mutation), with
optional coordinate/rationale override and an injected reviewed_at timestamp.
Pure, no I/O, no clock.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 06:33:27 -07:00
Ben Stull 1d7f821ab8 feat: ingest CLI entry point
Per sub-project-2 plan Task 10. argparse entry wiring named fetcher +
HeuristicProposer into ingest_search/ingest_candidate; --query/--resolve,
--limit, --catalog, --media-root (env HEF_MEDIA_ROOT), --dominant-color.
Unknown archive -> exit 2, deferred archive -> exit 3. Secrets env-only.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 06:33:01 -07:00
Ben Stull cecc5a0f61 feat: LibriVox/NASA/Internet Archive fetchers (+ deferred stubs)
Per sub-project-2 plan Task 9 / spec §6.4. Three keyless first-ship fetchers
parse documented JSON APIs via an injected HttpClient; license/attribution go
through tools.licensing; NASA third-party + IA no-license cases are flagged in
notes for review. musopen/fma/freesound are explicit deferred stubs raising
NotImplementedError; freesound documents FREESOUND_API_TOKEN (secret, env-only).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 06:32:18 -07:00
Ben Stull 2db2b7ac16 feat: ingest pipeline (Candidate/Fetcher/ingest_candidate)
Per sub-project-2 plan Task 8 / spec §6. ingest_candidate dedupes, downloads,
mechanically tags, drafts coordinates, builds a proposed Record and appends it;
ingest_search loops over fetcher hits. All boundaries (download/prober/color_fn)
injectable, so the pipeline is fully hermetic.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 06:29:09 -07:00
Ben Stull 3d5c34491c feat: heuristic coordinate proposer (drafting)
Per sub-project-2 plan Task 7 / spec §7. HeuristicProposer seeds the brain plane
from archive priors and the mood plane from title/description keyword nudges,
clamps to 0..4, and emits a one-line rationale. Deterministic, no I/O, no ML
(honors design §11) — only a DRAFT a human reviews.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 06:27:37 -07:00
Ben Stull 222421e773 feat: per-archive license normalization
Per sub-project-2 plan Task 6 / spec §5.3. normalize_license maps CC URLs /
identifiers and public-domain markers to the LICENSES vocab, builds non-empty
attribution for cc_by/cc_by_nc, and rejects unmappable licenses at ingest.
librivox_license() helper returns public_domain.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 06:26:21 -07:00
Ben Stull d2a97823f5 feat: ffmpeg frame extraction + optional dominant_color
Per sub-project-2 plan Task 5 / spec §5.2. ffmpeg-only dominant color (no image
library), opt-in by design; extract_frame for review previews. Runners injectable
so the unit suite never shells out.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 06:25:11 -07:00
24 changed files with 1963 additions and 6 deletions
+95 -6
View File
@@ -1,10 +1,11 @@
# Human Experience Filter — User Guide
> **Scope (current build state).** Only the catalog + selection core is built so
> far. This guide covers the one thing you can do today: **configure media by
> hand-authoring and validating catalog records.** The automated ingest tool,
> the tagging/review CLI, and the room player are separate, not-yet-built
> sub-projects and are not covered here.
> **Scope (current build state).** Two pieces are built: the catalog + selection
> core, and the **`tools/` ingest & review pipeline** (sub-project 2). You can
> populate the catalog two ways — **hand-author records** (below) or **assisted
> ingest** that fetches from public-domain archives, mechanically tags, drafts
> coordinates, and lets you review them to `approved` (see *Ingesting & reviewing
> media*). The room **player** is a separate, not-yet-built sub-project.
---
@@ -66,7 +67,7 @@ defaults and may be omitted.
| `review_status` | `"proposed"` | `proposed` or `approved` — whether a human has blessed it. |
| `attribution` | `""` | **Required text when `license` is `cc_by` or `cc_by_nc`.** |
| `resolution` | `""` | e.g. `1920x1080` (video). |
| `dominant_color` | `""` | Hex color, used by the side walls later. |
| `dominant_color` | `""` | Hex color. **Optional / opt-in** — computed only with `--dominant-color`. |
| `rationale` | `""` | One-line note on why you chose the coordinate. |
| `reviewed_at` | `null` | Timestamp when approved. |
| `notes` | `""` | Free text. |
@@ -146,6 +147,94 @@ print("added", r.id)
`append_record` validates the record and raises `CatalogError` (printing what is
wrong) before it writes, so a bad record never lands in the file.
### Option C: assisted ingest (fetch + tag + draft, then review)
See the next section.
---
## Ingesting & reviewing media
Instead of hand-authoring every record, the `tools/` pipeline can fetch a piece
from a public-domain archive, fill in the mechanical fields for you, **draft** a
coordinate, and write the record as `review_status: proposed`. You then walk the
proposed records and bless each one to `approved`. The four coordinates are still
a human act — the tool only *drafts* a starting point; nothing is selectable until
you approve it.
### Prerequisites
- **Python 3.11+**, run from the repo root (as everywhere in this guide).
- **`ffmpeg` and `ffprobe`** on your `PATH` — used to read media properties
(`mode`/`duration_s`/`resolution`), render review previews, and (opt-in)
compute `dominant_color`. Install via your package manager (e.g.
`brew install ffmpeg`). The unit tests don't need them; the live ingest/review
do.
### Where downloaded media goes (`--media-root`)
Ingest downloads each file to `<media-root>/<archive>/<id>.<ext>` and records a
`file_path` **relative to the media root** (e.g. `nasa/nasa-earthrise.mp4`).
Relative paths are portable: the tagging workstation and the Pi's drive store the
same tree under different mounts, and the player joins `file_path` with its own
mount. The media root is `--media-root DIR` (or `HEF_MEDIA_ROOT`), default
`./media/`, which is gitignored — **media never enters the repo**, only metadata.
> **Note on `file_path` style.** Hand-authored records (Option A/B above) often use
> absolute paths like `/media/earthrise.mp4`; ingested records use archive-relative
> paths. Both load and validate fine (`validate()` does not constrain the format).
> Keep a catalog internally consistent where you can; the player spec will define
> how mounts are resolved.
### First-ship archives
Three keyless, clean-license archives are wired today:
| `archive` arg | Pool | License |
|--------------------|---------------------------------------|-------------------|
| `librivox` | Public-domain audiobook recordings | `public_domain` |
| `nasa` | NASA imagery / video | `public_domain` |
| `internet_archive` | Internet Archive / Prelinger | per item (PD / CC)|
`musopen`, `fma`, and `freesound` are defined but **deferred** (auth or
API-stability tax) — invoking them exits with a "deferred" message. **Freesound**
will need an API token supplied via the `FREESOUND_API_TOKEN` environment variable
(a secret — never put it on the command line or into a record).
### Running ingest
```bash
.venv/bin/python -m tools.ingest_cli nasa --query "earthrise" --limit 5
.venv/bin/python -m tools.ingest_cli internet_archive --resolve prelinger_blast
.venv/bin/python -m tools.ingest_cli librivox --query "meditations" --media-root ./media
```
Flags: `--query` (search) or `--resolve <identifier>` (one item); `--limit N`;
`--catalog` (default `catalog/library.jsonl`); `--media-root` (default `./media`,
or `HEF_MEDIA_ROOT`); `--dominant-color` to compute `dominant_color` for
video/`av` records (off by default — its only consumer, the procedural side walls,
was dropped in the single-panoramic-projector design change). Re-running is
idempotent: a candidate whose id already exists is skipped.
### Reviewing proposed records
```bash
.venv/bin/python -m tools.review_cli --catalog catalog/library.jsonl --media-root ./media
```
For each `proposed` record it prints the id/title/source/license, the mechanical
fields, the **drafted coordinates + rationale**, and opens a preview frame
(video/`av`) or waveform image (audio). Then it prompts:
- **`a`** — accept the drafted coordinates and mark `approved`.
- **`e`** — edit `left/right/dark/light` (enter blank to keep a value), then approve.
- **`s`** — skip; leave it `proposed`.
- **`q`** — save and quit.
Each approval stamps `reviewed_at` and is written immediately (full rewrite), so
an interrupted session keeps its progress. **How far from done** is just the count
of records still `proposed`. Add `--no-preview` to skip frame rendering.
---
## Validating the whole catalog
+79
View File
@@ -0,0 +1,79 @@
from hef.selection import Coordinate
from tools.drafting import Draft, HeuristicProposer, Signals
def _sig(**o):
base = dict(
title="",
description="",
source_archive="nasa",
mode="video",
duration_s=600,
)
base.update(o)
return Signals(**base)
def test_returns_draft_with_coordinate_and_rationale():
d = HeuristicProposer().propose(_sig())
assert isinstance(d, Draft)
assert isinstance(d.coordinate, Coordinate)
assert d.rationale and "\n" not in d.rationale
def test_librivox_seeds_left():
d = HeuristicProposer().propose(
_sig(title="Meditations", source_archive="librivox", mode="audio")
)
assert d.coordinate.left >= 3 and d.rationale
def test_music_archives_seed_right():
for arch in ("musopen", "fma", "nasa", "freesound"):
d = HeuristicProposer().propose(_sig(source_archive=arch, mode="audio"))
assert d.coordinate.right >= 3, arch
def test_internet_archive_seeds_left():
d = HeuristicProposer().propose(_sig(source_archive="internet_archive"))
assert d.coordinate.left >= 3
def test_storm_seeds_dark():
d = HeuristicProposer().propose(
_sig(title="Thunderstorm at Night", source_archive="nasa")
)
assert d.coordinate.dark >= 2
def test_sunrise_seeds_light():
d = HeuristicProposer().propose(
_sig(title="Sunrise over the Garden", source_archive="nasa")
)
assert d.coordinate.light >= 2
def test_coordinates_clamped_to_0_4():
d = HeuristicProposer().propose(
_sig(
title="storm night war funeral decay requiem minor",
description="noir death grief",
source_archive="librivox",
)
)
for v in (d.coordinate.left, d.coordinate.right, d.coordinate.dark, d.coordinate.light):
assert 0 <= v <= 4
def test_rationale_cites_a_signal():
d = HeuristicProposer().propose(
_sig(title="Storm", source_archive="librivox", mode="audio")
)
assert "librivox" in d.rationale.lower()
def test_deterministic():
s = _sig(title="Storm at Dawn", source_archive="nasa")
a = HeuristicProposer().propose(s)
b = HeuristicProposer().propose(s)
assert a == b
+181
View File
@@ -0,0 +1,181 @@
import json
import pytest
from tools.http import HttpClient
from tools.ingest.internet_archive import InternetArchiveFetcher
from tools.ingest.librivox import LibriVoxFetcher
from tools.ingest.nasa import NasaFetcher
class _Resp:
def __init__(self, data: bytes):
self._data = data
def read(self):
return self._data
def __enter__(self):
return self
def __exit__(self, *a):
return False
def fake_opener(mapping):
"""Dispatch a urllib Request to a canned payload by URL substring."""
def opener(req, timeout=None):
url = req.full_url
for key, payload in mapping.items():
if key in url:
if isinstance(payload, (bytes, bytearray)):
return _Resp(bytes(payload))
return _Resp(json.dumps(payload).encode("utf-8"))
raise AssertionError(f"unexpected url: {url}")
return opener
def test_librivox_fetcher():
payload = {
"books": [
{
"id": "123",
"title": "Meditations",
"url_librivox": "https://librivox.org/meditations/",
"url_zip_file": "https://archive.org/download/meditations/meditations_mp3.zip",
"authors": [{"first_name": "Marcus", "last_name": "Aurelius"}],
"description": "Stoic philosophy.",
}
]
}
client = HttpClient(opener=fake_opener({"librivox.org/api": payload}))
cands = LibriVoxFetcher(client).search("medit", limit=5)
c = cands[0]
assert c.source_archive == "librivox"
assert c.license == "public_domain" and c.attribution == ""
assert c.suggested_id == "librivox-meditations"
assert c.media_url.endswith(".zip")
assert c.media_ext == "zip"
assert "Aurelius" in c.description
def test_nasa_fetcher():
search_payload = {
"collection": {
"items": [
{
"data": [
{
"nasa_id": "as08-14-2383",
"title": "Earthrise",
"description": "View of Earth from the Moon",
"media_type": "video",
}
],
"href": "https://images-assets.nasa.gov/video/as08-14-2383/collection.json",
}
]
}
}
asset_payload = [
"https://images-assets.nasa.gov/video/as08-14-2383/as08-14-2383~orig.mp4",
"https://images-assets.nasa.gov/video/as08-14-2383/as08-14-2383~thumb.jpg",
]
client = HttpClient(
opener=fake_opener(
{
"images-api.nasa.gov/search": search_payload,
"collection.json": asset_payload,
}
)
)
c = NasaFetcher(client).search("earth", limit=3)[0]
assert c.suggested_id == "nasa-as08-14-2383"
assert c.license == "public_domain" and c.attribution == ""
assert c.media_url.endswith("orig.mp4")
assert c.media_ext == "mp4"
assert "Earth" in c.description
def test_internet_archive_cc_by():
meta_payload = {
"metadata": {
"identifier": "earthrise",
"title": "Earthrise",
"licenseurl": "http://creativecommons.org/licenses/by/4.0/",
"creator": "NASA",
"description": "Apollo 8 footage.",
},
"files": [
{"name": "earthrise.mp4", "format": "h.264", "source": "original"},
{"name": "earthrise.png", "format": "PNG", "source": "derivative"},
],
"server": "ia800100.us.archive.org",
"dir": "/12/items/earthrise",
}
client = HttpClient(opener=fake_opener({"archive.org/metadata/earthrise": meta_payload}))
c = InternetArchiveFetcher(client).resolve("earthrise")
assert c.suggested_id == "ia-earthrise"
assert c.license == "cc_by" and "NASA" in c.attribution
assert c.media_url == "https://ia800100.us.archive.org/12/items/earthrise/earthrise.mp4"
assert c.media_ext == "mp4"
assert c.source_url == "https://archive.org/details/earthrise"
def test_internet_archive_no_license_assumed_public_domain_and_flagged():
meta_payload = {
"metadata": {"identifier": "oldfilm", "title": "Old Film"},
"files": [{"name": "oldfilm.mp4", "source": "original"}],
"server": "ia.example.org",
"dir": "/x/items/oldfilm",
}
client = HttpClient(opener=fake_opener({"archive.org/metadata/oldfilm": meta_payload}))
c = InternetArchiveFetcher(client).resolve("oldfilm")
assert c.license == "public_domain"
assert "verify" in c.description.lower()
def test_internet_archive_search_resolves_each_hit():
search_payload = {"response": {"docs": [{"identifier": "earthrise"}]}}
meta_payload = {
"metadata": {
"identifier": "earthrise",
"title": "Earthrise",
"licenseurl": "https://creativecommons.org/publicdomain/mark/1.0/",
},
"files": [{"name": "earthrise.mp4", "source": "original"}],
"server": "ia.example.org",
"dir": "/x",
}
client = HttpClient(
opener=fake_opener(
{
"advancedsearch.php": search_payload,
"archive.org/metadata/earthrise": meta_payload,
}
)
)
cands = InternetArchiveFetcher(client).search("earthrise", limit=5)
assert [c.suggested_id for c in cands] == ["ia-earthrise"]
assert cands[0].license == "public_domain"
def test_deferred_stubs_raise_not_implemented():
from tools.ingest.fma import FmaFetcher
from tools.ingest.freesound import FreesoundFetcher
from tools.ingest.musopen import MusopenFetcher
with pytest.raises(NotImplementedError):
MusopenFetcher(None).search("x", limit=1)
with pytest.raises(NotImplementedError):
FmaFetcher(None).search("x", limit=1)
with pytest.raises(NotImplementedError):
FreesoundFetcher(None).search("x", limit=1)
def test_fetchers_expose_archive_label():
assert LibriVoxFetcher(None).archive == "librivox"
assert NasaFetcher(None).archive == "nasa"
assert InternetArchiveFetcher(None).archive == "internet_archive"
+25
View File
@@ -0,0 +1,25 @@
import pytest
from tools.ingest_cli import main
def test_help_parses():
with pytest.raises(SystemExit) as e:
main(["--help"])
assert e.value.code == 0
def test_unknown_archive_errors_cleanly(capsys):
rc = main(["bogus", "--query", "x"])
assert rc == 2
assert "unknown archive" in capsys.readouterr().err
def test_missing_query_and_resolve_errors():
assert main(["nasa"]) == 2
def test_deferred_archive_reports_not_implemented(capsys):
rc = main(["musopen", "--query", "bach"])
assert rc == 3
assert "deferred" in capsys.readouterr().err
+193
View File
@@ -0,0 +1,193 @@
import pytest
from hef.catalog import CatalogError, load_catalog, validate_catalog
from tools.drafting import HeuristicProposer
from tools.ingest.base import Candidate, ingest_candidate, ingest_search
from tools.probe import Probe
def make_candidate(**o):
base = dict(
source_archive="nasa",
source_url="https://example.org/landing/x",
media_url="https://example.org/media/x.mp4",
title="Earthrise",
license="public_domain",
attribution="",
suggested_id="nasa-earthrise",
media_ext="mp4",
description="a view of earth",
)
base.update(o)
return Candidate(**base)
def video_probe():
return Probe(
streams=[
{
"codec_type": "video",
"width": 1920,
"height": 1080,
"disposition": {"attached_pic": 0},
},
{"codec_type": "audio"},
],
format={"duration": "10.0"},
)
def make_prober(probe):
def prober(path):
return probe
return prober
def make_downloader(content=b"FAKEMEDIA"):
def downloader(url, dest):
dest.write_bytes(content)
return downloader
def test_ingest_appends_one_proposed_record(tmp_path):
catalog = tmp_path / "library.jsonl"
media_root = tmp_path / "media"
rec = ingest_candidate(
make_candidate(),
catalog_path=catalog,
media_root=media_root,
proposer=HeuristicProposer(),
prober=make_prober(video_probe()),
downloader=make_downloader(),
)
assert rec is not None
records = load_catalog(catalog)
assert len(records) == 1
r = records[0]
assert r.id == "nasa-earthrise"
assert r.mode == "av"
assert r.resolution == "1920x1080"
assert r.duration_s == 10
assert r.review_status == "proposed"
assert r.reviewed_at is None
assert r.rationale != ""
assert r.file_path == "nasa/nasa-earthrise.mp4"
assert r.notes == "a view of earth"
# media written under media_root/<archive>/<id>.<ext>
assert (media_root / "nasa" / "nasa-earthrise.mp4").exists()
def test_ingest_is_idempotent(tmp_path):
catalog = tmp_path / "library.jsonl"
media_root = tmp_path / "media"
kw = dict(
catalog_path=catalog,
media_root=media_root,
proposer=HeuristicProposer(),
prober=make_prober(video_probe()),
downloader=make_downloader(),
)
first = ingest_candidate(make_candidate(), **kw)
second = ingest_candidate(make_candidate(), **kw)
assert first is not None
assert second is None # skipped as duplicate
assert len(load_catalog(catalog)) == 1
def test_dominant_color_off_by_default(tmp_path):
catalog = tmp_path / "library.jsonl"
rec = ingest_candidate(
make_candidate(),
catalog_path=catalog,
media_root=tmp_path / "media",
proposer=HeuristicProposer(),
prober=make_prober(video_probe()),
downloader=make_downloader(),
)
assert rec.dominant_color == ""
def test_dominant_color_computed_when_enabled(tmp_path):
catalog = tmp_path / "library.jsonl"
color_calls = []
def fake_color_fn(path, *, midpoint_s=0.0):
color_calls.append(midpoint_s)
return "#abcdef"
rec = ingest_candidate(
make_candidate(),
catalog_path=catalog,
media_root=tmp_path / "media",
proposer=HeuristicProposer(),
prober=make_prober(video_probe()),
downloader=make_downloader(),
compute_color=True,
color_fn=fake_color_fn,
)
assert rec.dominant_color == "#abcdef"
assert color_calls # was invoked
def test_audio_keeps_empty_color_even_when_enabled(tmp_path):
audio = Probe(streams=[{"codec_type": "audio"}], format={"duration": "30.0"})
def boom(path, *, midpoint_s=0.0):
raise AssertionError("must not compute color for audio")
rec = ingest_candidate(
make_candidate(media_ext="mp3", source_archive="librivox", license="public_domain"),
catalog_path=tmp_path / "library.jsonl",
media_root=tmp_path / "media",
proposer=HeuristicProposer(),
prober=make_prober(audio),
downloader=make_downloader(),
compute_color=True,
color_fn=boom,
)
assert rec.mode == "audio" and rec.dominant_color == ""
def test_unmappable_license_raises_and_writes_nothing(tmp_path):
catalog = tmp_path / "library.jsonl"
with pytest.raises(CatalogError):
ingest_candidate(
make_candidate(license="all_rights_reserved"),
catalog_path=catalog,
media_root=tmp_path / "media",
proposer=HeuristicProposer(),
prober=make_prober(video_probe()),
downloader=make_downloader(),
)
assert not catalog.exists() or load_catalog(catalog) == []
def test_ingest_search_runs_over_hits(tmp_path):
catalog = tmp_path / "library.jsonl"
class FakeFetcher:
archive = "nasa"
def search(self, query, *, limit):
return [
make_candidate(suggested_id="nasa-a", media_url="u/a.mp4"),
make_candidate(suggested_id="nasa-b", media_url="u/b.mp4"),
][:limit]
def resolve(self, identifier): # pragma: no cover - not used here
raise NotImplementedError
recs = ingest_search(
FakeFetcher(),
"earth",
limit=2,
catalog_path=catalog,
media_root=tmp_path / "media",
proposer=HeuristicProposer(),
prober=make_prober(video_probe()),
downloader=make_downloader(),
)
assert [r.id for r in recs] == ["nasa-a", "nasa-b"]
validate_catalog(load_catalog(catalog))
+84
View File
@@ -0,0 +1,84 @@
import pytest
from hef.catalog import LICENSES
from tools.licensing import librivox_license, normalize_license
def test_cc_by_url_maps_and_requires_attribution():
lic, attr = normalize_license(
"https://creativecommons.org/licenses/by/4.0/", creator="Jane Doe"
)
assert lic == "cc_by" and "Jane Doe" in attr
def test_cc_by_identifier_maps():
lic, attr = normalize_license("CC BY 4.0", creator="Sam")
assert lic == "cc_by" and attr
def test_cc_by_nc_maps():
lic, attr = normalize_license(
"https://creativecommons.org/licenses/by-nc/4.0/", creator="Artist"
)
assert lic == "cc_by_nc" and "Artist" in attr
def test_cc_by_without_creator_still_has_nonempty_attribution():
# validate() rejects an attribution license with empty attribution, so the
# normalizer must always produce a non-empty string for cc_by/cc_by_nc.
lic, attr = normalize_license("https://creativecommons.org/licenses/by/4.0/")
assert lic == "cc_by" and attr != ""
def test_cc0_maps_with_empty_attribution():
lic, attr = normalize_license("https://creativecommons.org/publicdomain/zero/1.0/")
assert lic == "cc0" and attr == ""
def test_cc0_identifier():
lic, attr = normalize_license("CC0")
assert lic == "cc0" and attr == ""
def test_public_domain_mark():
lic, attr = normalize_license(
"https://creativecommons.org/publicdomain/mark/1.0/"
)
assert lic == "public_domain" and attr == ""
def test_no_known_copyright_is_public_domain():
lic, attr = normalize_license("No known copyright")
assert lic == "public_domain" and attr == ""
def test_public_domain_literal():
lic, attr = normalize_license("public_domain")
assert lic == "public_domain" and attr == ""
def test_unmappable_rejected():
with pytest.raises(ValueError):
normalize_license("All Rights Reserved")
def test_empty_rejected():
with pytest.raises(ValueError):
normalize_license("")
def test_librivox_license_helper():
lic, attr = librivox_license()
assert lic == "public_domain" and attr == ""
def test_all_outputs_in_vocab():
samples = [
"https://creativecommons.org/licenses/by/4.0/",
"https://creativecommons.org/licenses/by-nc/4.0/",
"CC0",
"public domain",
]
for raw in samples:
lic, _ = normalize_license(raw, creator="x")
assert lic in LICENSES
+38
View File
@@ -0,0 +1,38 @@
from tools.mediatools import (
compute_dominant_color,
dominant_color_from_rgb,
extract_frame,
)
def test_dominant_color_from_rgb_pure_red():
assert dominant_color_from_rgb(b"\xff\x00\x00") == "#ff0000"
def test_dominant_color_from_rgb_arbitrary():
assert dominant_color_from_rgb(b"\x12\xab\x0f") == "#12ab0f"
def test_compute_dominant_color_uses_runner():
captured = {}
def fake_runner(args):
captured["args"] = args
return b"\x00\x80\xff"
color = compute_dominant_color("clip.mp4", midpoint_s=5.0, runner=fake_runner)
assert color == "#0080ff"
assert "ffmpeg" in captured["args"]
assert "5.0" in captured["args"]
def test_extract_frame_invokes_runner_and_returns_dest(tmp_path):
dest = tmp_path / "frame.png"
calls = []
def fake_runner(args):
calls.append(args)
out = extract_frame("clip.mp4", dest, midpoint_s=2.0, runner=fake_runner)
assert out == dest
assert calls and "ffmpeg" in calls[0] and str(dest) in calls[0]
+59
View File
@@ -0,0 +1,59 @@
from hef.catalog import Record, validate
from hef.selection import Coordinate
from tools.review import approve, proposed_records
def make_record(**o):
base = dict(
id="a",
title="t",
source_url="u",
source_archive="nasa",
license="public_domain",
mode="video",
left=0,
right=0,
dark=0,
light=0,
duration_s=1,
file_path="p",
)
base.update(o)
return Record(**base)
def test_proposed_records_filters():
a = make_record(id="a", review_status="proposed")
b = make_record(id="b", review_status="approved")
assert [r.id for r in proposed_records([a, b])] == ["a"]
def test_approve_sets_status_and_timestamp():
r = make_record(review_status="proposed")
out = approve(r, reviewed_at="2026-06-04T13:00:00+00:00")
assert out.review_status == "approved"
assert out.reviewed_at == "2026-06-04T13:00:00+00:00"
def test_approve_can_override_coordinates():
r = make_record(left=0, right=0, dark=0, light=0, review_status="proposed")
out = approve(r, reviewed_at="t", coordinate=Coordinate(4, 1, 2, 3))
assert (out.left, out.right, out.dark, out.light) == (4, 1, 2, 3)
def test_approve_does_not_mutate_input():
r = make_record(review_status="proposed")
approve(r, reviewed_at="t")
assert r.review_status == "proposed" and r.reviewed_at is None
def test_approve_can_set_rationale():
r = make_record(review_status="proposed", rationale="auto")
out = approve(r, reviewed_at="t", rationale="human note")
assert out.rationale == "human note"
def test_approved_record_revalidates():
r = make_record(review_status="proposed")
out = approve(r, reviewed_at="t", coordinate=Coordinate(2, 2, 1, 1))
validate(out) # must not raise
+103
View File
@@ -0,0 +1,103 @@
import io
import pytest
from hef.catalog import Record, load_catalog, save_catalog
from tools.review_cli import main
def make_record(**o):
base = dict(
id="a",
title="t",
source_url="u",
source_archive="nasa",
license="public_domain",
mode="video",
left=0,
right=0,
dark=0,
light=0,
duration_s=1,
file_path="nasa/a.mp4",
)
base.update(o)
return Record(**base)
def test_help_parses():
with pytest.raises(SystemExit) as e:
main(["--help"])
assert e.value.code == 0
def test_no_proposed_records(tmp_path):
catalog = tmp_path / "library.jsonl"
save_catalog([make_record(id="x", review_status="approved", reviewed_at="t")], catalog)
out = io.StringIO()
rc = main(["--catalog", str(catalog), "--no-preview"], out=out)
assert rc == 0
assert "no proposed records" in out.getvalue()
def test_accept_flips_to_approved(tmp_path):
catalog = tmp_path / "library.jsonl"
save_catalog([make_record(id="a", review_status="proposed")], catalog)
out = io.StringIO()
rc = main(
["--catalog", str(catalog), "--no-preview"],
input_fn=lambda prompt: "a",
now_fn=lambda: "2026-06-04T13:00:00+00:00",
out=out,
)
assert rc == 0
r = load_catalog(catalog)[0]
assert r.review_status == "approved"
assert r.reviewed_at == "2026-06-04T13:00:00+00:00"
def test_skip_leaves_proposed(tmp_path):
catalog = tmp_path / "library.jsonl"
save_catalog([make_record(id="a", review_status="proposed")], catalog)
rc = main(
["--catalog", str(catalog), "--no-preview"],
input_fn=lambda prompt: "s",
out=io.StringIO(),
)
assert rc == 0
assert load_catalog(catalog)[0].review_status == "proposed"
def test_edit_overrides_coordinates(tmp_path):
catalog = tmp_path / "library.jsonl"
save_catalog([make_record(id="a", review_status="proposed")], catalog)
answers = iter(["e", "4", "1", "2", "3"])
rc = main(
["--catalog", str(catalog), "--no-preview"],
input_fn=lambda prompt: next(answers),
now_fn=lambda: "t",
out=io.StringIO(),
)
assert rc == 0
r = load_catalog(catalog)[0]
assert (r.left, r.right, r.dark, r.light) == (4, 1, 2, 3)
assert r.review_status == "approved"
def test_quit_stops_walk(tmp_path):
catalog = tmp_path / "library.jsonl"
save_catalog(
[
make_record(id="a", review_status="proposed"),
make_record(id="b", review_status="proposed"),
],
catalog,
)
rc = main(
["--catalog", str(catalog), "--no-preview"],
input_fn=lambda prompt: "q",
out=io.StringIO(),
)
assert rc == 0
statuses = {r.id: r.review_status for r in load_catalog(catalog)}
assert statuses == {"a": "proposed", "b": "proposed"}
+134
View File
@@ -0,0 +1,134 @@
"""End-to-end: ingest (mocked boundaries) -> review -> select.
Plus opt-in tests that invoke real ffprobe/ffmpeg, skipped when absent.
"""
import re
import shutil
import subprocess
import pytest
from hef.catalog import load_catalog, validate_catalog
from hef.selection import Coordinate, select
from tools.drafting import HeuristicProposer
from tools.ingest.base import Candidate, ingest_candidate
from tools.mediatools import compute_dominant_color
from tools.probe import probe_file
from tools.review import approve, proposed_records
from tools.tagging import derive_tags
class _FakeFetcher:
archive = "nasa"
def search(self, query, *, limit):
return [
Candidate(
source_archive="nasa",
source_url="https://example.org/landing",
media_url="https://example.org/x.mp4",
title="Earthrise",
license="public_domain",
attribution="",
suggested_id="nasa-earthrise",
media_ext="mp4",
description="earth from the moon",
)
][:limit]
def resolve(self, identifier):
raise NotImplementedError
def _av_prober(path):
from tools.probe import Probe
return Probe(
streams=[
{"codec_type": "video", "width": 1920, "height": 1080, "disposition": {"attached_pic": 0}},
{"codec_type": "audio"},
],
format={"duration": "10.0"},
)
def _fake_downloader(url, dest):
dest.write_bytes(b"FAKE")
def test_end_to_end_ingest_review_select(tmp_path):
catalog = tmp_path / "library.jsonl"
media_root = tmp_path / "media"
# Ingest one candidate (all boundaries faked).
candidate = _FakeFetcher().search("earth", limit=1)[0]
rec = ingest_candidate(
candidate,
catalog_path=catalog,
media_root=media_root,
proposer=HeuristicProposer(),
prober=_av_prober,
downloader=_fake_downloader,
)
assert rec is not None and rec.review_status == "proposed"
# The proposed record can't be selected with approved_only yet.
loaded = load_catalog(catalog)
validate_catalog(loaded)
coord = Coordinate(loaded[0].left, loaded[0].right, loaded[0].dark, loaded[0].light)
assert select(loaded, coord, "av", approved_only=True) is None
# Review: approve it, persist.
pending = proposed_records(loaded)
assert len(pending) == 1
approved = approve(pending[0], reviewed_at="2026-06-04T13:00:00+00:00")
from hef.catalog import save_catalog
idx = {r.id: i for i, r in enumerate(loaded)}
loaded[idx[approved.id]] = approved
save_catalog(loaded, catalog)
# Reload, validate, and now select() finds it.
final = load_catalog(catalog)
validate_catalog(final)
picked = select(final, coord, "av", approved_only=True)
assert picked is not None and picked.id == "nasa-earthrise"
assert picked.review_status == "approved"
_HAS_FFMPEG = shutil.which("ffmpeg") is not None
_HAS_FFPROBE = shutil.which("ffprobe") is not None
@pytest.mark.skipif(not (_HAS_FFMPEG and _HAS_FFPROBE), reason="ffmpeg/ffprobe not installed")
def test_real_ffprobe_on_generated_clip(tmp_path):
clip = tmp_path / "clip.mp4"
subprocess.run(
[
"ffmpeg", "-v", "quiet", "-y",
"-f", "lavfi", "-i", "testsrc=duration=1:size=320x240:rate=30",
"-f", "lavfi", "-i", "sine=frequency=440:duration=1",
"-shortest", str(clip),
],
check=True,
)
tags = derive_tags(probe_file(clip))
assert tags.mode == "av"
assert tags.resolution == "320x240"
assert tags.duration_s == 1
@pytest.mark.skipif(not _HAS_FFMPEG, reason="ffmpeg not installed")
def test_real_dominant_color_on_generated_clip(tmp_path):
clip = tmp_path / "clip.mp4"
subprocess.run(
[
"ffmpeg", "-v", "quiet", "-y",
"-f", "lavfi", "-i", "color=c=red:duration=1:size=320x240:rate=30",
str(clip),
],
check=True,
)
color = compute_dominant_color(clip, midpoint_s=0.5)
assert re.fullmatch(r"#[0-9a-f]{6}", color)
+90
View File
@@ -0,0 +1,90 @@
"""Heuristic coordinate proposer — a deterministic, rule-based DRAFT seed.
The four curatorial coordinates are a human act (design §11: no automatic ML
coordinate tagging). This proposer only suggests a starting point written as
review_status='proposed'; a human blesses every coordinate before selection.
There is no ML model here just archive priors (brain plane) and keyword
nudges (mood plane), grounded in the §8 sourcingaxis map.
"""
from __future__ import annotations
from dataclasses import dataclass
from typing import Protocol
from hef.selection import Coordinate
COORD_MIN, COORD_MAX = 0, 4
# Brain plane (left = verbal/spoken/educational, right = music/wordless/abstract).
_ARCHIVE_PRIORS = {
"librivox": (4, 0), # spoken readings
"internet_archive": (3, 1), # educational / industrial / Prelinger
"musopen": (0, 4), # classical music
"fma": (0, 4), # music
"nasa": (0, 4), # wordless awe
"freesound": (0, 4), # abstract field recordings
}
_ARCHIVE_DEFAULT = (2, 2)
# Mood plane keywords (substring match so "thunderstorm" trips "storm").
_DARK_WORDS = (
"storm", "night", "noir", "requiem", "minor", "war", "funeral",
"decay", "death", "grief", "dusk", "winter", "mourning", "shadow",
)
_LIGHT_WORDS = (
"sunrise", "dawn", "garden", "spring", "joy", "hope", "major",
"bright", "bloom", "summer", "smile", "light", "morning",
)
def _clamp(v: int) -> int:
return max(COORD_MIN, min(COORD_MAX, v))
@dataclass
class Signals:
title: str
description: str
source_archive: str
mode: str
duration_s: int
@dataclass
class Draft:
coordinate: Coordinate
rationale: str
class Proposer(Protocol):
def propose(self, signals: Signals) -> Draft: ...
class HeuristicProposer:
def propose(self, signals: Signals) -> Draft:
left, right = _ARCHIVE_PRIORS.get(
signals.source_archive, _ARCHIVE_DEFAULT
)
text = f"{signals.title} {signals.description}".lower()
dark_hits = [w for w in _DARK_WORDS if w in text]
light_hits = [w for w in _LIGHT_WORDS if w in text]
dark = _clamp(len(dark_hits))
light = _clamp(len(light_hits))
coordinate = Coordinate(_clamp(left), _clamp(right), dark, light)
# Rationale: name the dominant brain prior + any mood keywords.
brain = (
f"{signals.source_archive}"
+ ("strong left" if left >= right else "strong right")
)
mood_bits = []
if dark_hits:
mood_bits.append(f"{dark_hits[0]!r} → dark")
if light_hits:
mood_bits.append(f"{light_hits[0]!r} → light")
rationale = "; ".join([brain, *mood_bits])
return Draft(coordinate=coordinate, rationale=rationale)
View File
+148
View File
@@ -0,0 +1,148 @@
"""Ingest seam: a Candidate shape, the Fetcher protocol, and the pipeline.
Every external boundary archive HTTP (download), ffprobe (prober), and the
optional ffmpeg dominant-color (color_fn) is injected, so the whole pipeline
is exercised hermetically (no network, no binaries, no disk beyond tmp).
"""
from __future__ import annotations
import logging
from dataclasses import dataclass
from pathlib import Path
from typing import Protocol
from hef.catalog import (
Record,
append_record,
load_catalog,
validate,
validate_catalog,
)
from tools.drafting import Signals
from tools.http import HttpClient
from tools.mediatools import compute_dominant_color
from tools.probe import probe_file
from tools.tagging import derive_tags
log = logging.getLogger(__name__)
_VIDEO_MODES = {"video", "av"}
@dataclass
class Candidate:
source_archive: str # origin label, e.g. internet_archive | librivox | nasa
source_url: str # human/landing URL recorded in the record
media_url: str # direct download URL of the chosen file
title: str
license: str # normalized (tools.licensing) -> hef LICENSES vocab
attribution: str # "" unless the license requires it
suggested_id: str # stable id derived from archive + archive-identifier
media_ext: str # file extension for the download target
description: str = "" # free text used by drafting signals; lands in notes
class Fetcher(Protocol):
archive: str
def search(self, query: str, *, limit: int) -> list[Candidate]: ...
def resolve(self, identifier: str) -> Candidate: ...
def _default_downloader(url, dest) -> None:
client = HttpClient()
dest.write_bytes(client.get_bytes(url))
def ingest_candidate(
candidate: Candidate,
*,
catalog_path,
media_root,
proposer,
prober=probe_file,
downloader=_default_downloader,
compute_color: bool = False,
color_fn=compute_dominant_color,
):
"""Ingest one candidate into the catalog as a `proposed` record.
Returns the appended Record, or None if it was skipped as a duplicate.
Raises CatalogError if the resulting record is invalid (e.g. an unmappable
license) nothing is appended in that case.
"""
catalog_path = Path(catalog_path)
media_root = Path(media_root)
# 1. Dedupe.
existing = load_catalog(catalog_path) if catalog_path.exists() else []
validate_catalog(existing)
if any(r.id == candidate.suggested_id for r in existing):
log.info("already in catalog, skipping: %s", candidate.suggested_id)
return None
# 2. Download (skip if already present).
rel_path = f"{candidate.source_archive}/{candidate.suggested_id}.{candidate.media_ext}"
dest = media_root / rel_path
dest.parent.mkdir(parents=True, exist_ok=True)
if not dest.exists():
downloader(candidate.media_url, dest)
# 3. Mechanical tagging.
tags = derive_tags(prober(dest))
dominant_color = ""
if compute_color and tags.mode in _VIDEO_MODES:
dominant_color = color_fn(dest, midpoint_s=tags.duration_s / 2)
# 4. Draft coordinates.
draft = proposer.propose(
Signals(
title=candidate.title,
description=candidate.description,
source_archive=candidate.source_archive,
mode=tags.mode,
duration_s=tags.duration_s,
)
)
coord = draft.coordinate
# 5. Build the record — proposed/reviewed_at=None falls out of the defaults.
record = Record(
id=candidate.suggested_id,
title=candidate.title,
source_url=candidate.source_url,
source_archive=candidate.source_archive,
license=candidate.license,
mode=tags.mode,
left=coord.left,
right=coord.right,
dark=coord.dark,
light=coord.light,
duration_s=tags.duration_s,
file_path=rel_path,
attribution=candidate.attribution,
resolution=tags.resolution,
dominant_color=dominant_color,
rationale=draft.rationale,
notes=candidate.description,
)
# 6. Validate + re-check uniqueness + append.
validate(record)
if any(r.id == record.id for r in existing): # racey re-check
log.info("already in catalog (race), skipping: %s", record.id)
return None
append_record(record, catalog_path)
return record
def ingest_search(fetcher: Fetcher, query: str, *, limit: int, **kw):
"""Run ingest_candidate over each search hit. Returns the appended records."""
appended = []
for candidate in fetcher.search(query, limit=limit):
rec = ingest_candidate(candidate, **kw)
if rec is not None:
appended.append(rec)
return appended
+26
View File
@@ -0,0 +1,26 @@
"""Free Music Archive fetcher — DEFERRED (see spec §6.4).
CC-licensed music (cc_by / cc_by_nc / cc0). Highest API-stability risk: the
public FMA API has been deprecated/changed, so a real implementation resolves
from a track URL + page metadata. The Fetcher seam is wired; it raises until
implemented.
"""
from __future__ import annotations
from tools.ingest.base import Candidate
_DEFERRED = "deferred — see spec §6.4 (FMA public API deprecated/changed)"
class FmaFetcher:
archive = "fma"
def __init__(self, client):
self.client = client
def search(self, query: str, *, limit: int) -> list[Candidate]:
raise NotImplementedError(_DEFERRED)
def resolve(self, identifier: str) -> Candidate:
raise NotImplementedError(_DEFERRED)
+34
View File
@@ -0,0 +1,34 @@
"""Freesound fetcher — DEFERRED (see spec §6.4).
CC-licensed sound effects / field recordings (cc0 / cc_by / cc_by_nc) via API
v2 (https://freesound.org/apiv2). Requires an API token supplied through the
FREESOUND_API_TOKEN environment variable (or the macOS Keychain). That token is
a SECRET: it is read from the environment only and must NEVER be written into a
record, a log line, or a transcript (wgl hard secrets rule, spec §9).
The Fetcher seam is wired; it raises until implemented.
"""
from __future__ import annotations
import os
from tools.ingest.base import Candidate
_DEFERRED = "deferred — see spec §6.4 (Freesound requires FREESOUND_API_TOKEN)"
TOKEN_ENV = "FREESOUND_API_TOKEN"
class FreesoundFetcher:
archive = "freesound"
def __init__(self, client, token: str | None = None):
self.client = client
# Read the secret from the environment only; never log or store it.
self._token = token or os.environ.get(TOKEN_ENV)
def search(self, query: str, *, limit: int) -> list[Candidate]:
raise NotImplementedError(_DEFERRED)
def resolve(self, identifier: str) -> Candidate:
raise NotImplementedError(_DEFERRED)
+91
View File
@@ -0,0 +1,91 @@
"""Internet Archive fetcher (incl. Prelinger).
Metadata API: https://archive.org/metadata/<id> (keyless); search via
advancedsearch.php. License is read from `licenseurl`/`rights`/`license` and
normalized through tools.licensing. When no explicit license is present the
item is assumed public_domain and flagged in notes for review (§6.4); an
explicitly non-free license (e.g. All Rights Reserved) raises and is not
ingested.
"""
from __future__ import annotations
from tools.ingest.base import Candidate
from tools.licensing import normalize_license
META = "https://archive.org/metadata/"
SEARCH = "https://archive.org/advancedsearch.php"
_NO_LICENSE_CAVEAT = "[IA: no explicit license — assumed public_domain, verify before use]"
_MEDIA_EXTS = (
".mp4", ".mov", ".m4v", ".ogv", ".webm",
".mp3", ".m4a", ".wav", ".flac", ".ogg",
)
def _ext_from_name(name: str) -> str:
return name.rsplit(".", 1)[-1].lower() if "." in name else ""
def _pick_file(files):
originals = [f for f in files if f.get("source") == "original"]
for pool in (originals, files):
for f in pool:
if str(f.get("name", "")).lower().endswith(_MEDIA_EXTS):
return f
return None
class InternetArchiveFetcher:
archive = "internet_archive"
def __init__(self, client):
self.client = client
def resolve(self, identifier: str) -> Candidate:
data = self.client.get_json(f"{META}{identifier}")
meta = data.get("metadata", {})
files = data.get("files", []) or []
server = data.get("server", "")
directory = data.get("dir", "")
chosen = _pick_file(files)
media_url = ""
media_ext = ""
if chosen and server:
name = chosen["name"]
media_url = f"https://{server}{directory}/{name}"
media_ext = _ext_from_name(name)
raw_license = (
meta.get("licenseurl") or meta.get("rights") or meta.get("license") or ""
)
creator = meta.get("creator", "") or ""
if isinstance(creator, list):
creator = ", ".join(str(c) for c in creator)
description = (meta.get("description", "") or "").strip()
ambiguous = (not raw_license) or ("no known copyright" in raw_license.lower())
if raw_license:
lic, attr = normalize_license(raw_license, creator=creator)
else:
lic, attr = "public_domain", ""
if ambiguous:
description = f"{description} {_NO_LICENSE_CAVEAT}".strip()
return Candidate(
source_archive=self.archive,
source_url=f"https://archive.org/details/{identifier}",
media_url=media_url,
title=meta.get("title", identifier),
license=lic,
attribution=attr,
suggested_id=f"ia-{identifier}",
media_ext=media_ext,
description=description,
)
def search(self, query: str, *, limit: int) -> list[Candidate]:
url = f"{SEARCH}?q={query}&fl[]=identifier&rows={limit}&output=json"
data = self.client.get_json(url)
docs = (data.get("response", {}).get("docs") or [])[:limit]
return [self.resolve(doc["identifier"]) for doc in docs]
+71
View File
@@ -0,0 +1,71 @@
"""LibriVox fetcher — public-domain audiobook recordings.
JSON API: https://librivox.org/api/feed/audiobooks (keyless). LibriVox is
public domain by charter; the reader/author is credited in notes but
attribution is not required. Section-level (per-track) file resolution is a
future refinement; this resolves the audiobook's zip as the media target.
"""
from __future__ import annotations
import re
from tools.ingest.base import Candidate
from tools.licensing import librivox_license
BASE = "https://librivox.org/api/feed/audiobooks"
def _slug(text: str) -> str:
s = re.sub(r"[^a-z0-9]+", "-", (text or "").lower()).strip("-")
return s or "untitled"
def _ext_from_url(url: str, default: str) -> str:
tail = url.rsplit("/", 1)[-1]
return tail.rsplit(".", 1)[-1].lower() if "." in tail else default
class LibriVoxFetcher:
archive = "librivox"
def __init__(self, client):
self.client = client
def _candidate(self, book) -> Candidate:
title = book.get("title", "")
authors = book.get("authors") or []
credit = ", ".join(
f"{a.get('first_name', '')} {a.get('last_name', '')}".strip()
for a in authors
).strip(", ")
media_url = book.get("url_zip_file", "")
description = (book.get("description", "") or "").strip()
if credit:
description = (description + f" (author: {credit})").strip()
lic, attr = librivox_license()
return Candidate(
source_archive=self.archive,
source_url=book.get("url_librivox", ""),
media_url=media_url,
title=title,
license=lic,
attribution=attr,
suggested_id=f"librivox-{_slug(title)}",
media_ext=_ext_from_url(media_url, "zip"),
description=description,
)
def search(self, query: str, *, limit: int) -> list[Candidate]:
url = f"{BASE}/title/^{query}?format=json&limit={limit}"
data = self.client.get_json(url)
books = (data.get("books") or [])[:limit]
return [self._candidate(b) for b in books]
def resolve(self, identifier: str) -> Candidate:
url = f"{BASE}/id/{identifier}?format=json"
data = self.client.get_json(url)
books = data.get("books") or []
if not books:
raise ValueError(f"librivox: no audiobook for {identifier!r}")
return self._candidate(books[0])
+25
View File
@@ -0,0 +1,25 @@
"""Musopen fetcher — DEFERRED (see spec §6.4).
Public-domain / CC classical music. Access has historically been gated and may
require an API key (a secret never logged, see spec §9). The Fetcher seam is
wired so enabling it later is purely additive; it raises until implemented.
"""
from __future__ import annotations
from tools.ingest.base import Candidate
_DEFERRED = "deferred — see spec §6.4 (Musopen access may require an API key)"
class MusopenFetcher:
archive = "musopen"
def __init__(self, client):
self.client = client
def search(self, query: str, *, limit: int) -> list[Candidate]:
raise NotImplementedError(_DEFERRED)
def resolve(self, identifier: str) -> Candidate:
raise NotImplementedError(_DEFERRED)
+79
View File
@@ -0,0 +1,79 @@
"""NASA fetcher — public-domain imagery/video.
JSON API: https://images-api.nasa.gov/search (keyless). Each search item links
an asset-collection JSON (the item's `href`) listing the concrete file URLs;
the chosen media file is resolved from there. NASA media is public domain, but
some items embed third-party content flagged in notes for review (§6.4).
"""
from __future__ import annotations
from tools.ingest.base import Candidate
SEARCH = "https://images-api.nasa.gov/search"
_THIRD_PARTY_CAVEAT = "[NASA: may embed third-party content — verify before use]"
_VIDEO_EXTS = (".mp4", ".mov", ".m4v", ".webm")
_AUDIO_EXTS = (".mp3", ".m4a", ".wav", ".flac", ".ogg")
def _ext_from_url(url: str) -> str:
tail = url.rsplit("/", 1)[-1].split("?", 1)[0]
return tail.rsplit(".", 1)[-1].lower() if "." in tail else ""
def _pick_asset(assets, media_type: str) -> str:
if not isinstance(assets, list):
return ""
prefs = _AUDIO_EXTS if media_type == "audio" else _VIDEO_EXTS
for ext in prefs:
for url in assets:
if isinstance(url, str) and url.lower().endswith(ext):
return url
for url in assets:
if isinstance(url, str) and not url.lower().endswith((".jpg", ".png", ".json")):
return url
return ""
class NasaFetcher:
archive = "nasa"
def __init__(self, client):
self.client = client
def _candidate(self, item) -> Candidate:
data0 = (item.get("data") or [{}])[0]
nasa_id = data0.get("nasa_id", "")
title = data0.get("title", "")
media_type = data0.get("media_type", "")
description = (data0.get("description", "") or "").strip()
description = f"{description} {_THIRD_PARTY_CAVEAT}".strip()
href = item.get("href", "")
assets = self.client.get_json(href) if href else []
media_url = _pick_asset(assets, media_type)
return Candidate(
source_archive=self.archive,
source_url=f"https://images.nasa.gov/details-{nasa_id}",
media_url=media_url,
title=title,
license="public_domain",
attribution="",
suggested_id=f"nasa-{nasa_id}",
media_ext=_ext_from_url(media_url),
description=description,
)
def search(self, query: str, *, limit: int) -> list[Candidate]:
url = f"{SEARCH}?q={query}&media_type=video"
data = self.client.get_json(url)
items = (data.get("collection", {}).get("items") or [])[:limit]
return [self._candidate(i) for i in items]
def resolve(self, identifier: str) -> Candidate:
url = f"{SEARCH}?nasa_id={identifier}"
data = self.client.get_json(url)
items = data.get("collection", {}).get("items") or []
if not items:
raise ValueError(f"nasa: no item for {identifier!r}")
return self._candidate(items[0])
+99
View File
@@ -0,0 +1,99 @@
"""CLI entry point for ingest: `python -m tools.ingest_cli <archive> ...`.
Wires a named fetcher + the HeuristicProposer into the ingest pipeline. Secrets
(e.g. FREESOUND_API_TOKEN) are read from the environment only, never as flags.
"""
from __future__ import annotations
import argparse
import os
import sys
from tools.drafting import HeuristicProposer
from tools.http import HttpClient
from tools.ingest.base import ingest_candidate, ingest_search
from tools.ingest.fma import FmaFetcher
from tools.ingest.freesound import FreesoundFetcher
from tools.ingest.internet_archive import InternetArchiveFetcher
from tools.ingest.librivox import LibriVoxFetcher
from tools.ingest.musopen import MusopenFetcher
from tools.ingest.nasa import NasaFetcher
FETCHERS = {
"librivox": LibriVoxFetcher,
"nasa": NasaFetcher,
"internet_archive": InternetArchiveFetcher,
"musopen": MusopenFetcher,
"fma": FmaFetcher,
"freesound": FreesoundFetcher,
}
DEFAULT_CATALOG = "catalog/library.jsonl"
DEFAULT_MEDIA_ROOT = "./media"
def _build_parser() -> argparse.ArgumentParser:
p = argparse.ArgumentParser(
prog="tools.ingest_cli",
description="Ingest candidates from an archive into the catalog as proposed records.",
)
p.add_argument("archive", help=f"source archive ({', '.join(sorted(FETCHERS))})")
p.add_argument("--query", help="search query")
p.add_argument("--resolve", metavar="IDENTIFIER", help="ingest a single item by id/URL")
p.add_argument("--limit", type=int, default=5, help="max search hits (default 5)")
p.add_argument("--catalog", default=DEFAULT_CATALOG, help=f"catalog JSONL (default {DEFAULT_CATALOG})")
p.add_argument(
"--media-root",
default=os.environ.get("HEF_MEDIA_ROOT", DEFAULT_MEDIA_ROOT),
help="download root (env HEF_MEDIA_ROOT; default ./media)",
)
p.add_argument(
"--dominant-color",
action="store_true",
help="compute dominant_color for video/av (opt-in, ffmpeg-only)",
)
return p
def main(argv=None) -> int:
args = _build_parser().parse_args(argv)
cls = FETCHERS.get(args.archive)
if cls is None:
print(
f"error: unknown archive {args.archive!r}; "
f"choose one of {', '.join(sorted(FETCHERS))}",
file=sys.stderr,
)
return 2
if not args.query and not args.resolve:
print("error: provide --query or --resolve", file=sys.stderr)
return 2
fetcher = cls(HttpClient())
kw = dict(
catalog_path=args.catalog,
media_root=args.media_root,
proposer=HeuristicProposer(),
compute_color=args.dominant_color,
)
try:
if args.resolve:
rec = ingest_candidate(fetcher.resolve(args.resolve), **kw)
appended = [rec] if rec is not None else []
else:
appended = ingest_search(fetcher, args.query, limit=args.limit, **kw)
except NotImplementedError as exc:
print(f"error: {exc}", file=sys.stderr)
return 3
print(f"ingested {len(appended)} proposed record(s) into {args.catalog}")
for rec in appended:
print(f" {rec.id} [{rec.mode}] {rec.title}")
return 0
if __name__ == "__main__": # pragma: no cover
raise SystemExit(main())
+63
View File
@@ -0,0 +1,63 @@
"""Normalize per-archive origin license metadata to the hef.catalog LICENSES vocab.
Maps Creative Commons URLs / identifiers and public-domain markers to
{public_domain, cc0, cc_by, cc_by_nc}. For the attribution licenses (cc_by,
cc_by_nc) it builds a non-empty attribution string hef.catalog.validate()
rejects those licenses with an empty attribution. Anything unmappable is
rejected (a piece whose license can't be established is not ingested).
"""
from __future__ import annotations
import re
def _attribution(creator: str, license_name: str, raw: str, default_label: str) -> str:
label = license_name or default_label or raw
parts = []
if creator:
parts.append(creator)
parts.append(label)
return "".join(parts)
def normalize_license(raw, *, creator="", license_name=""):
"""Return (license, attribution) for an origin license string.
Raises ValueError when `raw` does not map to an allowed license.
"""
text = (raw or "").strip().lower()
if not text:
raise ValueError(f"unmappable license: {raw!r}")
# CC0 / public-domain dedication (check before the generic publicdomain
# markers, since the CC0 URL also contains "publicdomain").
if "publicdomain/zero" in text or re.search(r"\bcc[ _-]?0\b", text):
return ("cc0", "")
# Public Domain Mark / explicit public domain / no-known-copyright.
pd_markers = (
"publicdomain/mark",
"public_domain",
"public domain",
"publicdomain",
"no known copyright",
"pdm",
)
if any(m in text for m in pd_markers):
return ("public_domain", "")
# CC BY-NC must be checked before the bare CC BY.
if re.search(r"by[ _-]?nc", text):
return ("cc_by_nc", _attribution(creator, license_name, raw, "CC BY-NC"))
# CC BY.
if re.search(r"\bcc[ _-]?by\b", text) or "licenses/by" in text:
return ("cc_by", _attribution(creator, license_name, raw, "CC BY"))
raise ValueError(f"unmappable license: {raw!r}")
def librivox_license():
"""LibriVox recordings are public domain by charter (attribution not required)."""
return ("public_domain", "")
+58
View File
@@ -0,0 +1,58 @@
"""ffmpeg helpers: representative frame; OPTIONAL dominant color. Runner injectable."""
from __future__ import annotations
import subprocess
def _run_bytes(args) -> bytes:
return subprocess.run(args, capture_output=True, check=True).stdout
def dominant_color_from_rgb(rgb: bytes) -> str:
r, g, b = rgb[0], rgb[1], rgb[2]
return f"#{r:02x}{g:02x}{b:02x}"
def compute_dominant_color(path, *, midpoint_s=0.0, runner=_run_bytes) -> str:
"""ffmpeg-only single-color palette of a mid-segment frame -> #rrggbb."""
args = [
"ffmpeg",
"-v",
"quiet",
"-ss",
str(midpoint_s),
"-i",
str(path),
"-vf",
"thumbnail,palettegen=max_colors=1",
"-frames:v",
"1",
"-f",
"rawvideo",
"-pix_fmt",
"rgb24",
"-",
]
return dominant_color_from_rgb(runner(args))
def extract_frame(path, dest, *, midpoint_s=0.0, runner=None):
"""Write one representative frame to dest (PNG) for the review preview."""
runner = runner or (lambda a: subprocess.run(a, check=True))
runner(
[
"ffmpeg",
"-v",
"quiet",
"-y",
"-ss",
str(midpoint_s),
"-i",
str(path),
"-frames:v",
"1",
str(dest),
]
)
return dest
+34
View File
@@ -0,0 +1,34 @@
"""Review state-transition core (pure, tested): proposed -> approved.
No I/O, no clock the timestamp is injected so the logic stays deterministic.
The CLI persists the result via hef.catalog.save_catalog after re-validating.
"""
from __future__ import annotations
from dataclasses import replace
def proposed_records(records):
"""Records still awaiting review."""
return [r for r in records if r.review_status == "proposed"]
def approve(record, *, reviewed_at, coordinate=None, rationale=None):
"""Return an approved copy of `record` (the input is not mutated).
review_status -> 'approved' and reviewed_at is stamped. Optionally override
the four coordinates with a human correction and/or replace the rationale.
The caller re-validates before saving.
"""
changes = {"review_status": "approved", "reviewed_at": reviewed_at}
if coordinate is not None:
changes.update(
left=coordinate.left,
right=coordinate.right,
dark=coordinate.dark,
light=coordinate.light,
)
if rationale is not None:
changes["rationale"] = rationale
return replace(record, **changes)
+154
View File
@@ -0,0 +1,154 @@
"""Interactive review walk: `python -m tools.review_cli`.
Thin shell over the tested review core (tools.review). Walks each proposed
record showing its mechanical fields, the proposed coordinates + rationale, and
a best-effort preview frame, then prompts accept/edit/skip/quit. Approvals are
persisted (full rewrite) after each one, so an interrupted session keeps its
progress.
"""
from __future__ import annotations
import argparse
import os
import platform
import shutil
import subprocess
import sys
import tempfile
from datetime import datetime, timezone
from pathlib import Path
from hef.catalog import load_catalog, save_catalog, validate_catalog
from tools.mediatools import extract_frame
from tools.review import approve, proposed_records
DEFAULT_CATALOG = "catalog/library.jsonl"
DEFAULT_MEDIA_ROOT = "./media"
def _now() -> str:
return datetime.now(timezone.utc).isoformat()
def _build_parser() -> argparse.ArgumentParser:
p = argparse.ArgumentParser(
prog="tools.review_cli",
description="Walk proposed records and approve/correct their coordinates.",
)
p.add_argument("--catalog", default=DEFAULT_CATALOG)
p.add_argument(
"--media-root",
default=os.environ.get("HEF_MEDIA_ROOT", DEFAULT_MEDIA_ROOT),
)
p.add_argument("--no-preview", action="store_true", help="skip preview rendering")
return p
def _show(rec, out) -> None:
print(f"\n{'=' * 60}", file=out)
print(f" id {rec.id}", file=out)
print(f" title {rec.title}", file=out)
print(f" source {rec.source_archive} {rec.source_url}", file=out)
attr = f" ({rec.attribution})" if rec.attribution else ""
print(f" license {rec.license}{attr}", file=out)
print(f" media mode={rec.mode} {rec.duration_s}s {rec.resolution}", file=out)
if rec.dominant_color:
print(f" color {rec.dominant_color}", file=out)
print(
f" proposed left={rec.left} right={rec.right} dark={rec.dark} light={rec.light}",
file=out,
)
print(f" rationale {rec.rationale}", file=out)
if rec.notes:
print(f" notes {rec.notes}", file=out)
def _open_cmd():
return "open" if platform.system() == "Darwin" else "xdg-open"
def _preview(rec, media_root, out) -> None:
"""Best-effort preview; never fatal. Requires ffmpeg for video frames."""
if not shutil.which("ffmpeg"):
return
src = Path(media_root) / rec.file_path
if not src.exists():
print(f" (preview: media not found at {src})", file=out)
return
try:
midpoint = rec.duration_s / 2 if rec.duration_s else 0.0
if rec.mode in ("video", "av"):
dest = Path(tempfile.gettempdir()) / f"hef-preview-{rec.id}.png"
extract_frame(src, dest, midpoint_s=midpoint)
else: # audio -> waveform thumbnail
dest = Path(tempfile.gettempdir()) / f"hef-preview-{rec.id}.png"
subprocess.run(
[
"ffmpeg", "-v", "quiet", "-y", "-i", str(src),
"-filter_complex", "showwavespic=s=640x240", "-frames:v", "1",
str(dest),
],
check=True,
)
opener = shutil.which(_open_cmd())
if opener:
subprocess.run([opener, str(dest)], check=False)
except Exception as exc: # noqa: BLE001 - preview is best-effort
print(f" (preview failed: {exc})", file=out)
def _prompt_coords(rec, input_fn, out):
from hef.selection import Coordinate
def _ask(axis, current):
raw = input_fn(f" {axis} [{current}]: ").strip()
if not raw:
return current
try:
return max(0, min(4, int(raw)))
except ValueError:
print(f" (not an int, keeping {current})", file=out)
return current
return Coordinate(
_ask("left", rec.left),
_ask("right", rec.right),
_ask("dark", rec.dark),
_ask("light", rec.light),
)
def main(argv=None, *, input_fn=input, now_fn=_now, out=sys.stdout) -> int:
args = _build_parser().parse_args(argv)
records = load_catalog(args.catalog)
validate_catalog(records)
pending = proposed_records(records)
if not pending:
print("no proposed records to review", file=out)
return 0
pos = {r.id: i for i, r in enumerate(records)}
print(f"{len(pending)} proposed record(s) to review", file=out)
for rec in pending:
_show(rec, out)
if not args.no_preview:
_preview(rec, args.media_root, out)
choice = input_fn("[a]ccept / [e]dit / [s]kip / [q]uit > ").strip().lower()
if choice == "q":
break
if choice == "s" or choice not in ("a", "e"):
continue
coordinate = _prompt_coords(rec, input_fn, out) if choice == "e" else None
approved = approve(records[pos[rec.id]], reviewed_at=now_fn(), coordinate=coordinate)
records[pos[rec.id]] = approved
save_catalog(records, args.catalog)
print(f" approved {rec.id}", file=out)
return 0
if __name__ == "__main__": # pragma: no cover
raise SystemExit(main())