Compare commits
7 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| fe044ed3db | |||
| bd3ef269d4 | |||
| 698821f065 | |||
| e794523079 | |||
| 28015ed1a2 | |||
| 376a6daddc | |||
| fb9b4fa422 |
+198
@@ -23,6 +23,204 @@ skip versions are the composition of each intervening adjacent
|
|||||||
release's steps in order — no A-to-B path is pre-computed beyond
|
release's steps in order — no A-to-B path is pre-computed beyond
|
||||||
that.
|
that.
|
||||||
|
|
||||||
|
## 0.27.0 — 2026-05-28
|
||||||
|
|
||||||
|
**Minor — security-hardening release (Session-0026 audit remediation).
|
||||||
|
One auto-applied migration (023); one behavior change that re-prompts
|
||||||
|
device-trust; deployments MUST re-apply the nginx + systemd files.**
|
||||||
|
This is the work cut as the "v0.25.0 security-hardening" branch; it
|
||||||
|
reversioned to 0.27.0 because v0.26.0 (#28) took the next slot while it
|
||||||
|
was in flight. Shipped from driver session 0030.0.
|
||||||
|
|
||||||
|
- **C1 (Critical) — stored-XSS closed.** Every markdown→HTML sink now
|
||||||
|
routes through one chokepoint, `frontend/src/lib/sanitizeHtml.js`
|
||||||
|
(DOMPurify), before any `innerHTML` / `dangerouslySetInnerHTML` write:
|
||||||
|
`MarkdownPreview`, both `ProposalView` sinks (entry body +
|
||||||
|
`proposed_use_case`), and `Editor`. A hook adds
|
||||||
|
`rel="noopener noreferrer"` to `target=_blank` links. `marked` no
|
||||||
|
longer passes raw HTML / `javascript:` URIs to the DOM, so a
|
||||||
|
contributor can no longer plant a payload that runs in an admin/owner
|
||||||
|
session during review.
|
||||||
|
- **H1 — OTC verify is rate-limited.** New `backend/app/ratelimit.py`
|
||||||
|
(per-IP token buckets) gates `/auth/otc/verify`, `/auth/otc/request`,
|
||||||
|
and the passcode check/verify paths; a per-account OTC-verify lockout
|
||||||
|
(migration `023_otc_verify_lockout.sql`) mirrors the passcode lockout.
|
||||||
|
- **M1 — device-trust lookup no longer table-scans.** The device-trust
|
||||||
|
cookie value is now `"<row_id>.<raw_token>"`; `device_trust.lookup`
|
||||||
|
reads the one indexed row and bcrypt-checks only it, instead of
|
||||||
|
bcrypt-checking every row in the table on each unauthenticated
|
||||||
|
`/auth/device-trust/start`.
|
||||||
|
- **M2 — HTTP security headers** (CSP, HSTS, X-Frame-Options,
|
||||||
|
X-Content-Type-Options, Referrer-Policy) added to the nginx server
|
||||||
|
block. **L8/I1**: `server_tokens off` + legacy TLS1.0/1.1 removed.
|
||||||
|
- **M4 — session cookie `Secure` by default** (`SESSION_COOKIE_SECURE`,
|
||||||
|
defaults on; a dev box on plain http sets it `false`).
|
||||||
|
- **M5 — bounce webhook fails closed.** An unset
|
||||||
|
`WEBHOOK_EMAIL_BOUNCE_SECRET` now **disables** `/api/webhooks/email-bounce`
|
||||||
|
(503) instead of leaving it open; a dev opts back in with
|
||||||
|
`RFC_APP_INSECURE_BOUNCE_WEBHOOK=1`.
|
||||||
|
- **L2/L3** per-IP cooldown + check-endpoint throttle. **L4** systemd
|
||||||
|
sandbox knobs (`CapabilityBoundingSet=`, `ProtectKernel*`,
|
||||||
|
`RestrictAddressFamilies`, `SystemCallFilter`, …).
|
||||||
|
|
||||||
|
Upgrade steps:
|
||||||
|
|
||||||
|
1. **Migration** — none manual; `023_otc_verify_lockout.sql` auto-applies
|
||||||
|
at startup via `db.run_migrations`.
|
||||||
|
2. **Device trust (MUST expect re-prompt)** — the cookie format changed,
|
||||||
|
so existing "trusted device" cookies no longer match; affected users
|
||||||
|
are re-prompted for device verification once. No data migration; old
|
||||||
|
rows are simply never matched and age out.
|
||||||
|
3. **nginx + systemd (MUST apply out-of-band)** — the deploy gesture does
|
||||||
|
**not** install `deploy/nginx/ohm.wiggleverse.org.conf` or
|
||||||
|
`deploy/systemd/rfc-app.service`. After deploying the code, copy both
|
||||||
|
to their system locations, then `nginx -t && systemctl reload nginx`
|
||||||
|
and `systemctl daemon-reload && systemctl restart <unit>`. (M2 headers
|
||||||
|
and L4 sandboxing do not take effect until this is done.)
|
||||||
|
4. **Bounce webhook (SHOULD)** — bind `WEBHOOK_EMAIL_BOUNCE_SECRET` (or
|
||||||
|
set `RFC_APP_INSECURE_BOUNCE_WEBHOOK=1` for dev). Unset → the endpoint
|
||||||
|
returns 503 (closed). No legitimate bounce source is wired today, so
|
||||||
|
503 is the safe default.
|
||||||
|
5. **Session cookie (SHOULD, dev only)** — a deployment served over plain
|
||||||
|
http MUST set `SESSION_COOKIE_SECURE=false` or the session cookie
|
||||||
|
won't be sent. Production over HTTPS leaves it unset (Secure on).
|
||||||
|
|
||||||
|
## 0.26.0 — 2026-05-28
|
||||||
|
|
||||||
|
**Minor — no schema migration, no new secret, no config, no upgrade
|
||||||
|
steps. Additive read-time enrichment; deployments inherit it on deploy
|
||||||
|
with nothing to set.** Roadmap item #28 **Part 1**: references to existing
|
||||||
|
**accepted** RFCs inside PR descriptions and comment text now render as
|
||||||
|
inline links to the referenced RFC. Shipped from driver session 0029.0,
|
||||||
|
in parallel with the v0.25.0 security-hardening session — hence the
|
||||||
|
version-slot gap (0.25.0 is that session's; this took the next free slot
|
||||||
|
per the roadmap's "claims the next available version number" rule).
|
||||||
|
|
||||||
|
Parts 2 (offer-to-create-RFC for strong-candidate terms) and 3
|
||||||
|
(offer-to-contribute-to-a-pending-RFC) of item #28 are deliberately **not**
|
||||||
|
in this release — Part 1 ships first as the easy win, exactly as the
|
||||||
|
roadmap row anticipates.
|
||||||
|
|
||||||
|
- **Where it links.** The PR review page's description and review
|
||||||
|
comments (`GET /api/rfcs/{slug}/prs/{n}`), and the PR-less per-RFC
|
||||||
|
discussion comments (`GET /api/rfcs/{slug}/discussion/threads/{id}/messages`).
|
||||||
|
Branch-chat (the AI-collaboration surface) is intentionally out of
|
||||||
|
scope for Part 1 — a noted follow-up.
|
||||||
|
- **Read-time, not submit-time.** The roadmap row phrases the scan as
|
||||||
|
happening "at submit/post time"; this ships it as **read-time**
|
||||||
|
enrichment instead. The intent the row actually names — "not as live
|
||||||
|
compose preview" — holds (drafts are never scanned, only submitted
|
||||||
|
content on read). Read-time was chosen because (a) links track the
|
||||||
|
*live* accepted-RFC set — a newly-accepted RFC starts linking in older
|
||||||
|
comments, a withdrawn one stops linking everywhere — rather than
|
||||||
|
freezing stale at submit; (b) it needs **no migration** (no derived
|
||||||
|
data to persist); (c) the active-RFC corpus is small and cache-resident,
|
||||||
|
so per-read scanning is cheap. (Recorded as a §19.3-rule-2 spec note in
|
||||||
|
the session transcript.)
|
||||||
|
- **XSS-safe by construction.** The backend returns structured *segments*
|
||||||
|
(a list of `{type:"text"}` / `{type:"rfc", slug, label, title}` items),
|
||||||
|
never HTML. The new `LinkedText` frontend component maps segments onto
|
||||||
|
React text nodes and anchors — no `dangerouslySetInnerHTML` — so a
|
||||||
|
comment author cannot inject markup through this path. Every enriched
|
||||||
|
field keeps its raw `text`/`description` alongside the `*_segments`, so
|
||||||
|
any non-segment-aware caller is unaffected.
|
||||||
|
- **Conservative matching (false-positive-averse).** A reference links
|
||||||
|
only when it is unlikely to be coincidental: an `rfc_id` token
|
||||||
|
(`RFC-0001`), a **multi-word** title (`Open Human Model`), or a
|
||||||
|
**hyphenated** slug (`open-human-model`). A single common-word title or
|
||||||
|
slug (a hypothetical RFC titled "Human") is **not** auto-linked — that
|
||||||
|
would turn every prose "human" into a link. Matching is case-insensitive,
|
||||||
|
word-boundary-anchored, longest-match-wins, and suppresses an RFC's
|
||||||
|
self-references inside its own PR/discussion. The roadmap's "curated
|
||||||
|
canonical-terms list" remains an explicit future opt-in rather than a
|
||||||
|
guessed-at default. New module: `backend/app/rfc_links.py`.
|
||||||
|
|
||||||
|
Tests: 12 new (`test_rfc_links_vertical.py`) — 9 scanner units (gating
|
||||||
|
rules, word boundaries, longest-match, case-insensitivity, casing
|
||||||
|
preservation, the rfc_id/multi-word/hyphenated gates) plus 3 end-to-end
|
||||||
|
(PR description + review comment + discussion comment all surface
|
||||||
|
`*_segments`; self-reference suppression). Full suite 363 green; frontend
|
||||||
|
builds clean.
|
||||||
|
|
||||||
|
Upgrade steps:
|
||||||
|
|
||||||
|
1. None. This release is purely additive: no migration, no new
|
||||||
|
environment variable, no secret, no overlay. A deployment picks up RFC
|
||||||
|
auto-linking the moment it deploys this version. (RFCs only link once
|
||||||
|
they are in the `active` state — proposed/super-draft and withdrawn
|
||||||
|
RFCs are never link targets, matching the §11.3 universal-public read
|
||||||
|
rule.)
|
||||||
|
|
||||||
|
## 0.24.0 — 2026-05-28
|
||||||
|
|
||||||
|
**Minor — one new secret required before this version serves tag
|
||||||
|
suggestions; no schema migration; no behavior change for deployments
|
||||||
|
that do not bind the key.** Roadmap item #27: as a contributor fills in
|
||||||
|
the propose-RFC form, the backend asks Claude Haiku to recommend tags
|
||||||
|
drawn from the collection's existing tag set, surfaced as clickable
|
||||||
|
suggestion chips. Shipped from driver session 0025.0. This is the §9.1
|
||||||
|
"Slice 2" AI-suggested chips that the propose modal has carried a
|
||||||
|
deferred placeholder for since Slice 1.
|
||||||
|
|
||||||
|
- **`POST /api/rfcs/suggest-tags`** (contributor-gated, per-user
|
||||||
|
rate-limited). Takes the partial draft (`title`, `pitch`,
|
||||||
|
`use_case`) and returns `{ "suggestions": [{ "tag", "confidence" },
|
||||||
|
…] }`. The model is constrained to the corpus's existing tags — v1
|
||||||
|
tags are free-form chip input, so "the taxonomy" is the de-facto set
|
||||||
|
of distinct tags the existing RFCs carry. The model MUST NOT invent
|
||||||
|
tags; taxonomy extension is a deliberate out-of-scope follow-up.
|
||||||
|
- **Always Claude Haiku, for cost.** Tag suggestion uses Haiku
|
||||||
|
regardless of the `ENABLED_MODELS` chat-picker universe, via the new
|
||||||
|
`providers.construct_haiku()` factory. There is no RFC slug at propose
|
||||||
|
time, so the §6.7 per-RFC funder credential path does not apply — the
|
||||||
|
surface runs on the operator's own `ANTHROPIC_API_KEY`.
|
||||||
|
- **Degrades to silence, never error.** No key bound, an empty corpus,
|
||||||
|
an empty draft, a rate-limited caller, or an unparseable model reply
|
||||||
|
all yield an empty list; the propose modal hides its suggestion row,
|
||||||
|
and the rest of the app is unaffected. So a deployment that does not
|
||||||
|
set the key sees no change at all.
|
||||||
|
- **Privacy disclosure.** The modal carries an inline note that the
|
||||||
|
draft text is sent to Anthropic to generate the suggestions. This is
|
||||||
|
required for honesty (the text leaves the deployment before the RFC
|
||||||
|
is submitted); cookie-consent does not gate it because the user has
|
||||||
|
actively typed into a draft surface. The exact wording wants a
|
||||||
|
counsel pass before OHM's deploy, per the #22 drafting discipline.
|
||||||
|
- **Frontend.** `ProposeModal` debounce-posts the draft (700 ms, with a
|
||||||
|
stale-response guard); suggestion chips are clickable and add to the
|
||||||
|
tag list (nothing auto-applies). `api.suggestTags()` is forgiving —
|
||||||
|
any non-OK response resolves to `[]`.
|
||||||
|
|
||||||
|
Tests: 11 new (`test_tag_suggest_vertical.py`) — contributor-gating,
|
||||||
|
universe-constrained filtering, no-key / empty-corpus / rate-limit
|
||||||
|
paths, plus units for the universe gather, the tolerant reply parser,
|
||||||
|
the max cap, the empty-draft short-circuit, and provider-failure
|
||||||
|
fallback. Full suite 351 green.
|
||||||
|
|
||||||
|
Upgrade steps:
|
||||||
|
|
||||||
|
1. You **MUST** bind `ANTHROPIC_API_KEY` for the suggestion surface to
|
||||||
|
work. On OHM this is an operator gesture — the key never touches the
|
||||||
|
conversation. Set it from your own terminal via stdin (so the bytes
|
||||||
|
never land in shell history):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
printf '%s' "$(pbpaste)" | .venv/bin/ohm-rfc-app-flotilla \
|
||||||
|
secret set ohm-rfc-app ANTHROPIC_API_KEY
|
||||||
|
```
|
||||||
|
|
||||||
|
(The §18 chat stack already reads this same key, so a deployment
|
||||||
|
that has chat configured **MAY** already have it bound — confirm with
|
||||||
|
`flotilla secret list ohm-rfc-app`.) If the key is absent the app
|
||||||
|
still boots and serves normally; tag suggestions are simply
|
||||||
|
unavailable until it is set.
|
||||||
|
2. You **MAY** tune the per-user rate limit via `TAG_SUGGEST_RATE_MAX`
|
||||||
|
(default 30) and `TAG_SUGGEST_RATE_WINDOW_SECONDS` (default 60). The
|
||||||
|
defaults apply when unset.
|
||||||
|
3. You **SHOULD** have counsel review the modal's disclosure wording
|
||||||
|
before exposing the surface to users, per the #22 drafting
|
||||||
|
discipline — the copy is honest as written, but the legal review is
|
||||||
|
the right call for any "your text is sent to a third party" notice.
|
||||||
|
|
||||||
## 0.23.0 — 2026-05-28
|
## 0.23.0 — 2026-05-28
|
||||||
|
|
||||||
Roadmap item #29: signing in lands the user on their most recently
|
Roadmap item #29: signing in lands the user on their most recently
|
||||||
|
|||||||
@@ -38,10 +38,22 @@ GITEA_WEBHOOK_SECRET=change-me-to-a-shared-secret
|
|||||||
# Comma-separated list of provider keys to enable. Per the §19.2
|
# Comma-separated list of provider keys to enable. Per the §19.2
|
||||||
# per-RFC-model topic, this is app-wide until that topic lands.
|
# per-RFC-model topic, this is app-wide until that topic lands.
|
||||||
ENABLED_MODELS=claude
|
ENABLED_MODELS=claude
|
||||||
|
# ANTHROPIC_API_KEY also powers the §9.1 propose-RFC tag suggestions
|
||||||
|
# (roadmap #27) — that surface always uses Claude Haiku for cost,
|
||||||
|
# independent of ENABLED_MODELS. With no key set, tag suggestions are
|
||||||
|
# simply unavailable (the modal hides the row); the rest of the app is
|
||||||
|
# unaffected.
|
||||||
ANTHROPIC_API_KEY=
|
ANTHROPIC_API_KEY=
|
||||||
GOOGLE_API_KEY=
|
GOOGLE_API_KEY=
|
||||||
OPENAI_API_KEY=
|
OPENAI_API_KEY=
|
||||||
|
|
||||||
|
# --- Tag suggestions (§9.1 / roadmap #27) ---
|
||||||
|
# Per-user rate limit on the suggest-tags endpoint (cost backstop; the
|
||||||
|
# modal debounces and the endpoint is contributor-gated). Optional —
|
||||||
|
# these defaults apply when unset.
|
||||||
|
TAG_SUGGEST_RATE_MAX=30
|
||||||
|
TAG_SUGGEST_RATE_WINDOW_SECONDS=60
|
||||||
|
|
||||||
# --- Email (§15.4) ---
|
# --- Email (§15.4) ---
|
||||||
# Leave SMTP_HOST unset to use the stdout fallback — the integration
|
# Leave SMTP_HOST unset to use the stdout fallback — the integration
|
||||||
# tests rely on it, and a dev environment without a real SMTP provider
|
# tests rely on it, and a dev environment without a real SMTP provider
|
||||||
|
|||||||
@@ -39,6 +39,7 @@ from . import (
|
|||||||
notify,
|
notify,
|
||||||
philosophy,
|
philosophy,
|
||||||
providers as providers_mod,
|
providers as providers_mod,
|
||||||
|
tag_suggest,
|
||||||
)
|
)
|
||||||
from .bot import Bot
|
from .bot import Bot
|
||||||
from .config import Config
|
from .config import Config
|
||||||
@@ -58,6 +59,17 @@ class ProposeBody(BaseModel):
|
|||||||
proposed_use_case: str | None = Field(default=None, max_length=8000)
|
proposed_use_case: str | None = Field(default=None, max_length=8000)
|
||||||
|
|
||||||
|
|
||||||
|
class SuggestTagsBody(BaseModel):
|
||||||
|
# Roadmap #27: the partial propose-RFC draft, sent as the user types
|
||||||
|
# (debounced on the frontend). All fields optional — suggestions
|
||||||
|
# refine as the draft fills in. `pitch` is the "why is this needed"
|
||||||
|
# rationale; `use_case` is the #26 optional ground-truth field.
|
||||||
|
# Bounds mirror the propose body's free-text caps.
|
||||||
|
title: str = Field(default="", max_length=200)
|
||||||
|
pitch: str = Field(default="", max_length=8000)
|
||||||
|
use_case: str = Field(default="", max_length=8000)
|
||||||
|
|
||||||
|
|
||||||
class DeclineBody(BaseModel):
|
class DeclineBody(BaseModel):
|
||||||
comment: str = Field(min_length=1, max_length=4000)
|
comment: str = Field(min_length=1, max_length=4000)
|
||||||
|
|
||||||
@@ -843,6 +855,33 @@ def make_router(
|
|||||||
|
|
||||||
return {"pr_number": pr["number"], "slug": slug}
|
return {"pr_number": pr["number"], "slug": slug}
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------
|
||||||
|
# §9.1 Slice 2 (roadmap #27): Claude Haiku tag suggestions as the
|
||||||
|
# propose-RFC fields fill in. The modal debounce-posts the partial
|
||||||
|
# draft; we constrain Haiku to the corpus's existing tag set and
|
||||||
|
# return a short ranked list of clickable chips. Gated to
|
||||||
|
# contributors (same gate as propose) so the cost surface is bounded
|
||||||
|
# to people who can actually file an RFC; rate-limited per user as a
|
||||||
|
# backstop. Degrades to an empty list (no error) when no Anthropic
|
||||||
|
# key is bound, the corpus has no tags yet, or the draft is empty —
|
||||||
|
# so the modal simply shows nothing extra.
|
||||||
|
# ---------------------------------------------------------------
|
||||||
|
|
||||||
|
@router.post("/api/rfcs/suggest-tags")
|
||||||
|
async def suggest_rfc_tags(payload: SuggestTagsBody, request: Request) -> dict[str, Any]:
|
||||||
|
user = auth.require_contributor(request)
|
||||||
|
if not tag_suggest.rate_limit_ok(user.user_id):
|
||||||
|
raise HTTPException(429, "Too many tag-suggestion requests; please slow down.")
|
||||||
|
provider = tag_suggest.haiku_provider(config)
|
||||||
|
if provider is None:
|
||||||
|
return {"suggestions": []}
|
||||||
|
universe = tag_suggest.gather_tag_universe()
|
||||||
|
draft = tag_suggest.Draft(
|
||||||
|
title=payload.title, pitch=payload.pitch, use_case=payload.use_case
|
||||||
|
)
|
||||||
|
suggestions = tag_suggest.suggest(provider, draft, universe)
|
||||||
|
return {"suggestions": suggestions}
|
||||||
|
|
||||||
# ---------------------------------------------------------------
|
# ---------------------------------------------------------------
|
||||||
# §9.3: merge / decline / withdraw an idea PR
|
# §9.3: merge / decline / withdraw an idea PR
|
||||||
# ---------------------------------------------------------------
|
# ---------------------------------------------------------------
|
||||||
|
|||||||
@@ -40,7 +40,7 @@ from typing import Any
|
|||||||
from fastapi import APIRouter, HTTPException, Request
|
from fastapi import APIRouter, HTTPException, Request
|
||||||
from pydantic import BaseModel, Field
|
from pydantic import BaseModel, Field
|
||||||
|
|
||||||
from . import auth, chat as chat_layer, db
|
from . import auth, chat as chat_layer, db, rfc_links
|
||||||
|
|
||||||
log = logging.getLogger(__name__)
|
log = logging.getLogger(__name__)
|
||||||
|
|
||||||
@@ -171,9 +171,17 @@ def make_router() -> APIRouter:
|
|||||||
""",
|
""",
|
||||||
(thread_id,),
|
(thread_id,),
|
||||||
).fetchall()
|
).fetchall()
|
||||||
|
# Roadmap #28 Part 1: enrich each discussion comment with RFC
|
||||||
|
# auto-link segments scanned against the live accepted-RFC corpus
|
||||||
|
# (read-time; see rfc_links.py). exclude_slug suppresses self-links
|
||||||
|
# to this RFC inside its own discussion.
|
||||||
|
link_index = rfc_links.build_index(db.conn(), exclude_slug=slug)
|
||||||
|
messages = [_serialize_message(r) for r in rows]
|
||||||
|
for m in messages:
|
||||||
|
m["text_segments"] = link_index.segment(m["text"])
|
||||||
return {
|
return {
|
||||||
"thread": _serialize_thread(thread),
|
"thread": _serialize_thread(thread),
|
||||||
"messages": [_serialize_message(r) for r in rows],
|
"messages": messages,
|
||||||
}
|
}
|
||||||
|
|
||||||
# -------------------------------------------------------------------
|
# -------------------------------------------------------------------
|
||||||
|
|||||||
@@ -557,7 +557,26 @@ def make_router(config: Config) -> APIRouter:
|
|||||||
# stays unauthenticated for dev (the v1 contract).
|
# stays unauthenticated for dev (the v1 contract).
|
||||||
import os as _os
|
import os as _os
|
||||||
expected = _os.environ.get("WEBHOOK_EMAIL_BOUNCE_SECRET", "").strip()
|
expected = _os.environ.get("WEBHOOK_EMAIL_BOUNCE_SECRET", "").strip()
|
||||||
if expected:
|
# v0.25.0 (audit 0026 M5): fail closed. An unset secret used to
|
||||||
|
# leave this endpoint fully unauthenticated — anyone could suppress
|
||||||
|
# any user's mail by POSTing their address (email_opt_out_all flip
|
||||||
|
# below). Now an unset secret DISABLES the endpoint (503) instead
|
||||||
|
# of opening it. A dev that genuinely wants it open opts in
|
||||||
|
# explicitly with RFC_APP_INSECURE_BOUNCE_WEBHOOK=1, mirroring the
|
||||||
|
# RFC_APP_INSECURE_WEBHOOKS dev-bypass on the Gitea hook.
|
||||||
|
if not expected:
|
||||||
|
if _os.environ.get("RFC_APP_INSECURE_BOUNCE_WEBHOOK", "").strip() == "1":
|
||||||
|
log.warning(
|
||||||
|
"email-bounce webhook running UNAUTHENTICATED "
|
||||||
|
"(RFC_APP_INSECURE_BOUNCE_WEBHOOK=1) — never set this in production"
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
log.error(
|
||||||
|
"email-bounce webhook refused: WEBHOOK_EMAIL_BOUNCE_SECRET is unset "
|
||||||
|
"(set the secret to enable, or RFC_APP_INSECURE_BOUNCE_WEBHOOK=1 for dev)"
|
||||||
|
)
|
||||||
|
raise HTTPException(503, "Bounce webhook not configured")
|
||||||
|
else:
|
||||||
received = request.headers.get("X-Webhook-Secret", "")
|
received = request.headers.get("X-Webhook-Secret", "")
|
||||||
import hmac as _hmac
|
import hmac as _hmac
|
||||||
if not received or not _hmac.compare_digest(expected, received):
|
if not received or not _hmac.compare_digest(expected, received):
|
||||||
|
|||||||
+15
-1
@@ -23,7 +23,7 @@ from typing import Any
|
|||||||
from fastapi import APIRouter, HTTPException, Request
|
from fastapi import APIRouter, HTTPException, Request
|
||||||
from pydantic import BaseModel, Field
|
from pydantic import BaseModel, Field
|
||||||
|
|
||||||
from . import auth, cache, chat as chat_layer, db, entry as entry_mod, funder, models_resolver
|
from . import auth, cache, chat as chat_layer, db, entry as entry_mod, funder, models_resolver, rfc_links
|
||||||
from .bot import Bot
|
from .bot import Bot
|
||||||
from .config import Config
|
from .config import Config
|
||||||
from .gitea import Gitea, GiteaError
|
from .gitea import Gitea, GiteaError
|
||||||
@@ -213,6 +213,12 @@ def make_router(
|
|||||||
path = _file_path_for(rfc)
|
path = _file_path_for(rfc)
|
||||||
head_branch = pr_row["head_branch"]
|
head_branch = pr_row["head_branch"]
|
||||||
|
|
||||||
|
# Roadmap #28 Part 1: build the RFC auto-link index once for this
|
||||||
|
# PR view (read-time enrichment against the live accepted-RFC
|
||||||
|
# corpus; see rfc_links.py). exclude_slug suppresses self-links to
|
||||||
|
# this RFC inside its own PR.
|
||||||
|
link_index = rfc_links.build_index(db.conn(), exclude_slug=slug)
|
||||||
|
|
||||||
# §11.3: PRs are always public; no visibility check.
|
# §11.3: PRs are always public; no visibility check.
|
||||||
main_fetched = await gitea.read_file(owner, repo, path, ref="main")
|
main_fetched = await gitea.read_file(owner, repo, path, ref="main")
|
||||||
main_body = _extract_body(rfc, (main_fetched or ("", ""))[0])
|
main_body = _extract_body(rfc, (main_fetched or ("", ""))[0])
|
||||||
@@ -259,6 +265,13 @@ def make_router(
|
|||||||
for r in msg_rows:
|
for r in msg_rows:
|
||||||
messages_by_thread.setdefault(r["thread_id"], []).append(_serialize_message(r))
|
messages_by_thread.setdefault(r["thread_id"], []).append(_serialize_message(r))
|
||||||
|
|
||||||
|
# Roadmap #28 Part 1: enrich every comment with RFC auto-link
|
||||||
|
# segments (read-time; see rfc_links.py). The description is
|
||||||
|
# enriched alongside it in the return dict below.
|
||||||
|
for _msgs in messages_by_thread.values():
|
||||||
|
for _m in _msgs:
|
||||||
|
_m["text_segments"] = link_index.segment(_m["text"])
|
||||||
|
|
||||||
# Per-user seen cursor per §10.3. Anonymous viewers get no
|
# Per-user seen cursor per §10.3. Anonymous viewers get no
|
||||||
# cursor — they always see "everything new" but cannot advance
|
# cursor — they always see "everything new" but cannot advance
|
||||||
# the cursor (no row to write to).
|
# the cursor (no row to write to).
|
||||||
@@ -325,6 +338,7 @@ def make_router(
|
|||||||
"pr_number": pr_number,
|
"pr_number": pr_number,
|
||||||
"title": pr_row["title"],
|
"title": pr_row["title"],
|
||||||
"description": pr_row["description"],
|
"description": pr_row["description"],
|
||||||
|
"description_segments": link_index.segment(pr_row["description"]),
|
||||||
"proposed_use_case": _pr_use_case(pr_number),
|
"proposed_use_case": _pr_use_case(pr_number),
|
||||||
"state": pr_row["state"],
|
"state": pr_row["state"],
|
||||||
"opened_by": pr_row["opened_by"],
|
"opened_by": pr_row["opened_by"],
|
||||||
|
|||||||
+32
-21
@@ -97,6 +97,13 @@ class IssueOutcome:
|
|||||||
raw_token: str
|
raw_token: str
|
||||||
row_id: int
|
row_id: int
|
||||||
|
|
||||||
|
@property
|
||||||
|
def cookie_value(self) -> str:
|
||||||
|
"""The value to put in the `rfc_device_trust` cookie: the row-id
|
||||||
|
selector joined to the raw token (v0.25.0 / audit 0026 M1). The
|
||||||
|
selector lets `lookup` read one indexed row instead of scanning."""
|
||||||
|
return f"{self.row_id}.{self.raw_token}"
|
||||||
|
|
||||||
|
|
||||||
def _new_token() -> str:
|
def _new_token() -> str:
|
||||||
return secrets.token_urlsafe(TOKEN_BYTES)
|
return secrets.token_urlsafe(TOKEN_BYTES)
|
||||||
@@ -173,33 +180,37 @@ def lookup(raw_token: str) -> LookupOutcome:
|
|||||||
if not raw:
|
if not raw:
|
||||||
return LookupOutcome(ok=False, user=None, reason="invalid")
|
return LookupOutcome(ok=False, user=None, reason="invalid")
|
||||||
|
|
||||||
# The unique index on `device_token_hash` would let us SELECT by
|
# v0.25.0 (audit 0026 M1): the cookie is "<row_id>.<raw_token>". We
|
||||||
# hash if bcrypt were a stable hash, but bcrypt incorporates a
|
# parse the row-id selector and read exactly ONE row by its indexed
|
||||||
# per-row salt — equal tokens produce different hashes. We walk
|
# primary key, then bcrypt-check the token against that single row.
|
||||||
# the candidate set instead. In practice the set is small (a
|
|
||||||
# human has a handful of trusted devices) and bcrypt is cheap on
|
|
||||||
# the order of milliseconds; the walk is bounded by the user's
|
|
||||||
# active device count.
|
|
||||||
#
|
#
|
||||||
# We don't pre-filter by `revoked_at IS NULL` here so that a
|
# The previous shape read EVERY device_trust row (all users, including
|
||||||
# token presented for a recently-revoked row produces a
|
# revoked/expired) and bcrypt-checked each — an unauthenticated
|
||||||
# 'revoked' outcome (the endpoint surfaces a different shape).
|
# CPU-amplification DoS reachable at /auth/device-trust/start that
|
||||||
# Same for expired: we let the walk hit and classify after.
|
# grew without bound as the table accumulated. bcrypt's per-row salt
|
||||||
rows = db.conn().execute(
|
# is why we can't SELECT by hash; carrying the row-id in the cookie is
|
||||||
|
# the standard fix (the id is not secret; the token still is).
|
||||||
|
selector, sep, token = raw.partition(".")
|
||||||
|
if not sep or not selector.isdigit() or not token:
|
||||||
|
# Legacy bare-token cookies (pre-v0.25.0) and malformed values land
|
||||||
|
# here. We refuse rather than fall back to a full-table scan, so
|
||||||
|
# the amplification path is fully closed; affected users simply
|
||||||
|
# re-authenticate once via OTC/passcode and get a new cookie.
|
||||||
|
return LookupOutcome(ok=False, user=None, reason="invalid")
|
||||||
|
|
||||||
|
matched = db.conn().execute(
|
||||||
"""
|
"""
|
||||||
SELECT id, user_id, device_token_hash, expires_at, revoked_at
|
SELECT id, user_id, device_token_hash, expires_at, revoked_at
|
||||||
FROM device_trust
|
FROM device_trust
|
||||||
ORDER BY id DESC
|
WHERE id = ?
|
||||||
""",
|
""",
|
||||||
).fetchall()
|
(int(selector),),
|
||||||
|
).fetchone()
|
||||||
|
|
||||||
matched = None
|
# One bcrypt check, against the selected row only. A wrong/forged token
|
||||||
for row in rows:
|
# for a real id reads as 'unknown' (cookie cleared), same as a missing
|
||||||
if _check(raw, row["device_token_hash"]):
|
# row — a probing client can't distinguish the two.
|
||||||
matched = row
|
if matched is None or not _check(token, matched["device_token_hash"]):
|
||||||
break
|
|
||||||
|
|
||||||
if matched is None:
|
|
||||||
return LookupOutcome(ok=False, user=None, reason="unknown")
|
return LookupOutcome(ok=False, user=None, reason="unknown")
|
||||||
|
|
||||||
if matched["revoked_at"] is not None:
|
if matched["revoked_at"] is not None:
|
||||||
|
|||||||
+60
-18
@@ -7,6 +7,7 @@ no need for a separate worker.
|
|||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
import logging
|
import logging
|
||||||
|
import os
|
||||||
import secrets
|
import secrets
|
||||||
from contextlib import asynccontextmanager
|
from contextlib import asynccontextmanager
|
||||||
|
|
||||||
@@ -28,6 +29,7 @@ from . import (
|
|||||||
otc,
|
otc,
|
||||||
passcode as passcode_mod,
|
passcode as passcode_mod,
|
||||||
providers as providers_mod,
|
providers as providers_mod,
|
||||||
|
ratelimit,
|
||||||
turnstile,
|
turnstile,
|
||||||
webhooks,
|
webhooks,
|
||||||
)
|
)
|
||||||
@@ -142,12 +144,20 @@ def create_app() -> FastAPI:
|
|||||||
# eagerly via load_config(). Everything else waits for lifespan.
|
# eagerly via load_config(). Everything else waits for lifespan.
|
||||||
config = load_config()
|
config = load_config()
|
||||||
app = FastAPI(lifespan=lifespan)
|
app = FastAPI(lifespan=lifespan)
|
||||||
|
# v0.25.0 (audit 0026 M4): the session cookie is the primary 30-day
|
||||||
|
# auth credential and must carry `Secure` in production so it never
|
||||||
|
# travels cleartext. Default to Secure; a dev box serving over plain
|
||||||
|
# http opts out with SESSION_COOKIE_SECURE=false. Production (OHM is
|
||||||
|
# HTTPS-only with an HTTP->HTTPS 301) leaves this unset → Secure on.
|
||||||
|
session_secure = os.environ.get("SESSION_COOKIE_SECURE", "true").strip().lower() not in (
|
||||||
|
"0", "false", "no", "off",
|
||||||
|
)
|
||||||
app.add_middleware(
|
app.add_middleware(
|
||||||
SessionMiddleware,
|
SessionMiddleware,
|
||||||
secret_key=config.secret_key,
|
secret_key=config.secret_key,
|
||||||
session_cookie="rfc_session",
|
session_cookie="rfc_session",
|
||||||
max_age=60 * 60 * 24 * 30,
|
max_age=60 * 60 * 24 * 30,
|
||||||
https_only=False,
|
https_only=session_secure,
|
||||||
)
|
)
|
||||||
return app
|
return app
|
||||||
|
|
||||||
@@ -155,24 +165,25 @@ def create_app() -> FastAPI:
|
|||||||
app = create_app()
|
app = create_app()
|
||||||
|
|
||||||
|
|
||||||
def _set_device_trust_cookie(response: Response, raw_token: str) -> None:
|
def _set_device_trust_cookie(response: Response, cookie_value: str) -> None:
|
||||||
"""Attach the v0.11.0 device-trust cookie to the response.
|
"""Attach the v0.11.0 device-trust cookie to the response.
|
||||||
|
|
||||||
HttpOnly + Secure + SameSite=Lax + 30-day Max-Age + Path=/. The
|
HttpOnly + Secure + SameSite=Lax + 30-day Max-Age + Path=/. As of
|
||||||
cookie value is the raw token; server-side storage is the hash.
|
v0.25.0 (audit 0026 M1) the value is `IssueOutcome.cookie_value` —
|
||||||
The cookie is "essential" per the v0.13.0 cookie-consent contract
|
"<row_id>.<raw_token>" — so `device_trust.lookup` can read one indexed
|
||||||
(it is part of authentication), so we set it regardless of the
|
row instead of scanning; server-side storage remains the bcrypt hash
|
||||||
user's analytics / other-cookies choice.
|
of the token half only. The cookie is "essential" per the v0.13.0
|
||||||
|
cookie-consent contract (it is part of authentication), so we set it
|
||||||
|
regardless of the user's analytics / other-cookies choice.
|
||||||
|
|
||||||
Secure=True means the cookie is only ever sent over HTTPS. The
|
Secure=True means the cookie is only ever sent over HTTPS — the
|
||||||
SessionMiddleware in `create_app` keeps `https_only=False` for
|
device-trust cookie holds a 30-day credential and must never travel
|
||||||
dev parity, but the device-trust cookie holds a 30-day credential
|
cleartext. (The session cookie now also defaults to Secure; see M4 in
|
||||||
and must not travel cleartext — production deployments serve over
|
`create_app`.)
|
||||||
HTTPS, so Secure on the device-trust cookie is non-negotiable.
|
|
||||||
"""
|
"""
|
||||||
response.set_cookie(
|
response.set_cookie(
|
||||||
key=device_trust_mod.COOKIE_NAME,
|
key=device_trust_mod.COOKIE_NAME,
|
||||||
value=raw_token,
|
value=cookie_value,
|
||||||
max_age=device_trust_mod.COOKIE_MAX_AGE_SECONDS,
|
max_age=device_trust_mod.COOKIE_MAX_AGE_SECONDS,
|
||||||
path="/",
|
path="/",
|
||||||
secure=True,
|
secure=True,
|
||||||
@@ -245,6 +256,10 @@ def _oauth_router(config) -> APIRouter:
|
|||||||
|
|
||||||
@router.post("/auth/otc/request")
|
@router.post("/auth/otc/request")
|
||||||
async def otc_request(body: OtcRequestBody, request: Request):
|
async def otc_request(body: OtcRequestBody, request: Request):
|
||||||
|
# v0.25.0 (audit 0026 H1/L2): per-IP brake at the cheapest point,
|
||||||
|
# before the Turnstile network call or any bcrypt/SMTP work.
|
||||||
|
if not ratelimit.otc_request_limiter.allow(ratelimit.client_key(request)):
|
||||||
|
raise HTTPException(429, "Too many requests; please wait a few minutes")
|
||||||
# v0.12.0 / roadmap item #10: gate the request on a successful
|
# v0.12.0 / roadmap item #10: gate the request on a successful
|
||||||
# Turnstile siteverify before the bcrypt hash + SMTP send. The
|
# Turnstile siteverify before the bcrypt hash + SMTP send. The
|
||||||
# check runs first so a failed challenge spends no rate budget
|
# check runs first so a failed challenge spends no rate budget
|
||||||
@@ -278,9 +293,25 @@ def _oauth_router(config) -> APIRouter:
|
|||||||
|
|
||||||
@router.post("/auth/otc/verify")
|
@router.post("/auth/otc/verify")
|
||||||
async def otc_verify(body: OtcVerifyBody, request: Request, response: Response):
|
async def otc_verify(body: OtcVerifyBody, request: Request, response: Response):
|
||||||
|
# v0.25.0 (audit 0026 H1): per-IP brake against fan-out guessing,
|
||||||
|
# plus the per-email lockout enforced inside otc.verify_code.
|
||||||
|
ip = ratelimit.client_key(request)
|
||||||
|
if not ratelimit.verify_limiter.allow(ip):
|
||||||
|
raise HTTPException(429, "Too many attempts; please wait a few minutes")
|
||||||
result = otc.verify_code(body.email, body.code)
|
result = otc.verify_code(body.email, body.code)
|
||||||
|
if result.reason == "locked":
|
||||||
|
raise HTTPException(
|
||||||
|
423,
|
||||||
|
{
|
||||||
|
"detail": "Too many failed attempts; wait a few minutes or request a new code",
|
||||||
|
"locked_until": result.locked_until,
|
||||||
|
},
|
||||||
|
)
|
||||||
if not result.ok or result.user is None:
|
if not result.ok or result.user is None:
|
||||||
raise HTTPException(400, "Invalid or expired code")
|
raise HTTPException(400, "Invalid or expired code")
|
||||||
|
# Legit sign-in: clear this IP's window so a user who fat-fingered
|
||||||
|
# a couple of codes isn't left throttled.
|
||||||
|
ratelimit.verify_limiter.reset(ip)
|
||||||
auth.store_session(request, result.user)
|
auth.store_session(request, result.user)
|
||||||
# v0.8.0: surface `needs_profile` so the Login.jsx surface can
|
# v0.8.0: surface `needs_profile` so the Login.jsx surface can
|
||||||
# decide whether to advance to the first/last/why capture step
|
# decide whether to advance to the first/last/why capture step
|
||||||
@@ -313,7 +344,7 @@ def _oauth_router(config) -> APIRouter:
|
|||||||
if body.trust_device:
|
if body.trust_device:
|
||||||
ua = request.headers.get("user-agent", "")
|
ua = request.headers.get("user-agent", "")
|
||||||
outcome = device_trust_mod.issue(result.user.user_id, ua)
|
outcome = device_trust_mod.issue(result.user.user_id, ua)
|
||||||
_set_device_trust_cookie(response, outcome.raw_token)
|
_set_device_trust_cookie(response, outcome.cookie_value)
|
||||||
return {
|
return {
|
||||||
"ok": True,
|
"ok": True,
|
||||||
"user": {
|
"user": {
|
||||||
@@ -337,12 +368,17 @@ def _oauth_router(config) -> APIRouter:
|
|||||||
# ---------------------------------------------------------------
|
# ---------------------------------------------------------------
|
||||||
|
|
||||||
@router.get("/auth/passcode/check")
|
@router.get("/auth/passcode/check")
|
||||||
async def passcode_check(email: str = ""):
|
async def passcode_check(request: Request, email: str = ""):
|
||||||
"""Does this email have a passcode set? Anonymous endpoint —
|
"""Does this email have a passcode set? Anonymous endpoint —
|
||||||
the Login.jsx flow calls this after the user types their email
|
the Login.jsx flow calls this after the user types their email
|
||||||
to decide whether to render a passcode input or fall back to
|
to decide whether to render a passcode input or fall back to
|
||||||
OTC. We surface only the boolean; lockout state, the hash, and
|
OTC. We surface only the boolean; lockout state, the hash, and
|
||||||
the set-at stamp are not leaked here."""
|
the set-at stamp are not leaked here.
|
||||||
|
|
||||||
|
v0.25.0 (audit 0026 L3): per-IP rate limit so the has-passcode
|
||||||
|
boolean can't be bulk-harvested to enumerate accounts."""
|
||||||
|
if not ratelimit.check_limiter.allow(ratelimit.client_key(request)):
|
||||||
|
raise HTTPException(429, "Too many requests; please wait a few minutes")
|
||||||
status = passcode_mod.passcode_status(email)
|
status = passcode_mod.passcode_status(email)
|
||||||
return {"has_passcode": status.has_passcode}
|
return {"has_passcode": status.has_passcode}
|
||||||
|
|
||||||
@@ -376,6 +412,11 @@ def _oauth_router(config) -> APIRouter:
|
|||||||
v0.11.0: the body's `trust_device` flag, if true, mints a
|
v0.11.0: the body's `trust_device` flag, if true, mints a
|
||||||
fresh device-trust row and sets the long-lived cookie. Same
|
fresh device-trust row and sets the long-lived cookie. Same
|
||||||
opt-in contract as `/auth/otc/verify`."""
|
opt-in contract as `/auth/otc/verify`."""
|
||||||
|
# v0.25.0 (audit 0026 H1): per-IP brake in front of the per-account
|
||||||
|
# passcode lockout, so fan-out across emails is throttled too.
|
||||||
|
ip = ratelimit.client_key(request)
|
||||||
|
if not ratelimit.verify_limiter.allow(ip):
|
||||||
|
raise HTTPException(429, "Too many attempts; please wait a few minutes")
|
||||||
result = passcode_mod.verify_passcode(body.email, body.passcode)
|
result = passcode_mod.verify_passcode(body.email, body.passcode)
|
||||||
if result.reason == "locked":
|
if result.reason == "locked":
|
||||||
raise HTTPException(
|
raise HTTPException(
|
||||||
@@ -387,11 +428,12 @@ def _oauth_router(config) -> APIRouter:
|
|||||||
)
|
)
|
||||||
if not result.ok or result.user is None:
|
if not result.ok or result.user is None:
|
||||||
raise HTTPException(400, "Invalid passcode")
|
raise HTTPException(400, "Invalid passcode")
|
||||||
|
ratelimit.verify_limiter.reset(ip)
|
||||||
auth.store_session(request, result.user)
|
auth.store_session(request, result.user)
|
||||||
if body.trust_device:
|
if body.trust_device:
|
||||||
ua = request.headers.get("user-agent", "")
|
ua = request.headers.get("user-agent", "")
|
||||||
outcome = device_trust_mod.issue(result.user.user_id, ua)
|
outcome = device_trust_mod.issue(result.user.user_id, ua)
|
||||||
_set_device_trust_cookie(response, outcome.raw_token)
|
_set_device_trust_cookie(response, outcome.cookie_value)
|
||||||
return {
|
return {
|
||||||
"ok": True,
|
"ok": True,
|
||||||
"user": {
|
"user": {
|
||||||
@@ -459,7 +501,7 @@ def _oauth_router(config) -> APIRouter:
|
|||||||
if body.trust_device:
|
if body.trust_device:
|
||||||
ua = request.headers.get("user-agent", "")
|
ua = request.headers.get("user-agent", "")
|
||||||
outcome = device_trust_mod.issue(result.user.user_id, ua)
|
outcome = device_trust_mod.issue(result.user.user_id, ua)
|
||||||
_set_device_trust_cookie(response, outcome.raw_token)
|
_set_device_trust_cookie(response, outcome.cookie_value)
|
||||||
|
|
||||||
# Has the user already set a passcode? (Could only happen via
|
# Has the user already set a passcode? (Could only happen via
|
||||||
# an admin pre-population path that doesn't exist yet, but
|
# an admin pre-population path that doesn't exist yet, but
|
||||||
|
|||||||
@@ -85,6 +85,15 @@ def _cooldown_seconds() -> int:
|
|||||||
return 60
|
return 60
|
||||||
|
|
||||||
|
|
||||||
|
# v0.25.0 / security audit 0026 (H1): per-email OTC verify lockout,
|
||||||
|
# mirroring the passcode path (passcode.py). Five consecutive wrong codes
|
||||||
|
# for an email lock its OTC verify for 15 minutes. The per-IP limiter in
|
||||||
|
# ratelimit.py is the primary brute-force brake; this is the durable,
|
||||||
|
# passcode-parity layer.
|
||||||
|
LOCKOUT_AFTER_FAILED_ATTEMPTS = 5
|
||||||
|
LOCKOUT_DURATION_MINUTES = 15
|
||||||
|
|
||||||
|
|
||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
# Code generation + hashing
|
# Code generation + hashing
|
||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
@@ -206,6 +215,61 @@ class VerifyOutcome:
|
|||||||
ok: bool
|
ok: bool
|
||||||
user: SessionUser | None
|
user: SessionUser | None
|
||||||
reason: str
|
reason: str
|
||||||
|
# v0.25.0 (H1): ISO-8601 stamp when reason == 'locked'.
|
||||||
|
locked_until: str | None = None
|
||||||
|
|
||||||
|
|
||||||
|
def _verify_lockout_until(email: str) -> str | None:
|
||||||
|
"""Return the active lockout stamp for `email`, or None if not locked.
|
||||||
|
|
||||||
|
Clears an elapsed lockout (and resets the counter) as a side effect so
|
||||||
|
the next failure starts a fresh budget — mirrors passcode.verify_passcode.
|
||||||
|
"""
|
||||||
|
row = db.conn().execute(
|
||||||
|
"SELECT failed_attempts, locked_until FROM otc_verify_state WHERE email = ?",
|
||||||
|
(email,),
|
||||||
|
).fetchone()
|
||||||
|
if row is None or not row["locked_until"]:
|
||||||
|
return None
|
||||||
|
still_locked = db.conn().execute(
|
||||||
|
"SELECT datetime(?) > datetime('now') AS locked", (row["locked_until"],),
|
||||||
|
).fetchone()["locked"]
|
||||||
|
if still_locked:
|
||||||
|
return row["locked_until"]
|
||||||
|
db.conn().execute(
|
||||||
|
"UPDATE otc_verify_state SET failed_attempts = 0, locked_until = NULL WHERE email = ?",
|
||||||
|
(email,),
|
||||||
|
)
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def _record_verify_failure(email: str) -> None:
|
||||||
|
"""Increment the per-email failure counter; stamp a lockout once it
|
||||||
|
crosses the threshold. Mirrors the passcode lockout shape."""
|
||||||
|
db.conn().execute(
|
||||||
|
"""
|
||||||
|
INSERT INTO otc_verify_state (email, failed_attempts)
|
||||||
|
VALUES (?, 1)
|
||||||
|
ON CONFLICT(email) DO UPDATE SET failed_attempts = failed_attempts + 1
|
||||||
|
""",
|
||||||
|
(email,),
|
||||||
|
)
|
||||||
|
count = db.conn().execute(
|
||||||
|
"SELECT failed_attempts FROM otc_verify_state WHERE email = ?", (email,),
|
||||||
|
).fetchone()["failed_attempts"]
|
||||||
|
if count >= LOCKOUT_AFTER_FAILED_ATTEMPTS:
|
||||||
|
db.conn().execute(
|
||||||
|
f"""
|
||||||
|
UPDATE otc_verify_state
|
||||||
|
SET locked_until = datetime('now', '+{LOCKOUT_DURATION_MINUTES} minutes')
|
||||||
|
WHERE email = ?
|
||||||
|
""",
|
||||||
|
(email,),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _clear_verify_state(email: str) -> None:
|
||||||
|
db.conn().execute("DELETE FROM otc_verify_state WHERE email = ?", (email,))
|
||||||
|
|
||||||
|
|
||||||
def verify_code(email: str, code: str) -> VerifyOutcome:
|
def verify_code(email: str, code: str) -> VerifyOutcome:
|
||||||
@@ -214,6 +278,13 @@ def verify_code(email: str, code: str) -> VerifyOutcome:
|
|||||||
if not email or not code:
|
if not email or not code:
|
||||||
return VerifyOutcome(ok=False, user=None, reason="invalid")
|
return VerifyOutcome(ok=False, user=None, reason="invalid")
|
||||||
|
|
||||||
|
# v0.25.0 (H1): refuse before spending any bcrypt if this email is in
|
||||||
|
# its OTC-verify lockout window. The passcode path is unaffected — a
|
||||||
|
# locked-out OTC user can still set/use a passcode, and vice versa.
|
||||||
|
locked_until = _verify_lockout_until(email)
|
||||||
|
if locked_until:
|
||||||
|
return VerifyOutcome(ok=False, user=None, reason="locked", locked_until=locked_until)
|
||||||
|
|
||||||
rows = db.conn().execute(
|
rows = db.conn().execute(
|
||||||
"""
|
"""
|
||||||
SELECT id, code_hash, expires_at, consumed_at
|
SELECT id, code_hash, expires_at, consumed_at
|
||||||
@@ -237,6 +308,9 @@ def verify_code(email: str, code: str) -> VerifyOutcome:
|
|||||||
break
|
break
|
||||||
|
|
||||||
if matched is None:
|
if matched is None:
|
||||||
|
# A genuine wrong guess against this email — the brute-force
|
||||||
|
# signal. Count it toward the lockout threshold (H1).
|
||||||
|
_record_verify_failure(email)
|
||||||
return VerifyOutcome(ok=False, user=None, reason="wrong")
|
return VerifyOutcome(ok=False, user=None, reason="wrong")
|
||||||
|
|
||||||
if matched["consumed_at"] is not None:
|
if matched["consumed_at"] is not None:
|
||||||
@@ -255,6 +329,8 @@ def verify_code(email: str, code: str) -> VerifyOutcome:
|
|||||||
"UPDATE otc_codes SET consumed_at = datetime('now') WHERE id = ?",
|
"UPDATE otc_codes SET consumed_at = datetime('now') WHERE id = ?",
|
||||||
(matched["id"],),
|
(matched["id"],),
|
||||||
)
|
)
|
||||||
|
# Success wipes the per-email failure counter (H1).
|
||||||
|
_clear_verify_state(email)
|
||||||
user = provision_or_link_user(email)
|
user = provision_or_link_user(email)
|
||||||
return VerifyOutcome(ok=True, user=user, reason="ok")
|
return VerifyOutcome(ok=True, user=user, reason="ok")
|
||||||
|
|
||||||
|
|||||||
@@ -184,6 +184,21 @@ def load_providers(env: dict) -> dict[str, BaseProvider]:
|
|||||||
return providers
|
return providers
|
||||||
|
|
||||||
|
|
||||||
|
def construct_haiku(api_key: str) -> AnthropicProvider:
|
||||||
|
"""A dedicated Claude Haiku provider, independent of the
|
||||||
|
`ENABLED_MODELS` chat-picker universe.
|
||||||
|
|
||||||
|
The §9.1 tag-suggestion surface (roadmap #27) always wants the
|
||||||
|
cheap + fast model regardless of which models the operator surfaces
|
||||||
|
in the §8.12 picker, so it constructs Haiku directly from the
|
||||||
|
operator's Anthropic key rather than going through `load_providers`.
|
||||||
|
The model id is sourced from the same `_CLAUDE_VARIANTS` table the
|
||||||
|
picker uses, so a model-string bump lands in one place.
|
||||||
|
"""
|
||||||
|
model, name = _CLAUDE_VARIANTS["claude-haiku"]
|
||||||
|
return AnthropicProvider(api_key=api_key, model=model, display_name=name)
|
||||||
|
|
||||||
|
|
||||||
def load_from_config(config) -> dict[str, BaseProvider]:
|
def load_from_config(config) -> dict[str, BaseProvider]:
|
||||||
"""Convenience adapter so callers can pass our Config dataclass directly."""
|
"""Convenience adapter so callers can pass our Config dataclass directly."""
|
||||||
env = {
|
env = {
|
||||||
|
|||||||
@@ -0,0 +1,92 @@
|
|||||||
|
"""In-process per-IP sliding-window rate limiter (security audit 0026, H1).
|
||||||
|
|
||||||
|
The auth verify endpoints (`/auth/otc/verify`, `/auth/passcode/verify`)
|
||||||
|
had no per-IP brake, so an attacker could fan out guesses against a
|
||||||
|
target identity bounded only by bcrypt cost. This module is the brake.
|
||||||
|
|
||||||
|
It is deliberately tiny: §4.2 says the app is a single process with a
|
||||||
|
colocated SQLite file, so an in-memory dict of `key -> deque[timestamps]`
|
||||||
|
is sufficient and needs no shared store. State resets on restart, which
|
||||||
|
fails *open* for a brief window — acceptable because the per-email OTC
|
||||||
|
lockout (`otc_verify_state`) and the passcode lockout both persist in the
|
||||||
|
database and carry the durable guarantee; this limiter is the
|
||||||
|
anti-fan-out layer on top.
|
||||||
|
|
||||||
|
Chosen over a per-identity lockout *as the primary control* because a
|
||||||
|
per-IP window throttles the attacker without letting them grief a victim
|
||||||
|
by locking that victim's account (the known downside of identity
|
||||||
|
lockouts). Both layers run together.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import threading
|
||||||
|
import time
|
||||||
|
from collections import defaultdict, deque
|
||||||
|
|
||||||
|
|
||||||
|
class SlidingWindowLimiter:
|
||||||
|
"""Allow at most `max_events` per `window_seconds` per key.
|
||||||
|
|
||||||
|
`allow(key)` records an event and returns True if the key is still
|
||||||
|
within budget, False if it has exceeded it. Timestamps use a
|
||||||
|
monotonic clock so the limiter is immune to wall-clock jumps.
|
||||||
|
"""
|
||||||
|
|
||||||
|
def __init__(self, max_events: int, window_seconds: float) -> None:
|
||||||
|
self.max_events = max_events
|
||||||
|
self.window_seconds = window_seconds
|
||||||
|
self._events: dict[str, deque[float]] = defaultdict(deque)
|
||||||
|
self._lock = threading.Lock()
|
||||||
|
|
||||||
|
def allow(self, key: str) -> bool:
|
||||||
|
now = time.monotonic()
|
||||||
|
cutoff = now - self.window_seconds
|
||||||
|
with self._lock:
|
||||||
|
q = self._events[key]
|
||||||
|
while q and q[0] < cutoff:
|
||||||
|
q.popleft()
|
||||||
|
if len(q) >= self.max_events:
|
||||||
|
return False
|
||||||
|
q.append(now)
|
||||||
|
# Opportunistic cleanup so idle keys don't accumulate forever.
|
||||||
|
if not q:
|
||||||
|
self._events.pop(key, None)
|
||||||
|
return True
|
||||||
|
|
||||||
|
def reset(self, key: str) -> None:
|
||||||
|
"""Drop a key's window — e.g. after a successful sign-in so a
|
||||||
|
legitimate user who fat-fingered a few times isn't throttled."""
|
||||||
|
with self._lock:
|
||||||
|
self._events.pop(key, None)
|
||||||
|
|
||||||
|
|
||||||
|
# Module-level limiters shared across requests (one process, so module
|
||||||
|
# state is the natural home). Tunables are intentionally generous enough
|
||||||
|
# not to bother a human retyping a code, tight enough to kill fan-out:
|
||||||
|
# * verify: 10 attempts / 5 min / IP across the auth verify surfaces.
|
||||||
|
# * otc request: 5 sends / 5 min / IP (Turnstile is the primary gate;
|
||||||
|
# this is defense in depth against a solved-challenge replay loop).
|
||||||
|
verify_limiter = SlidingWindowLimiter(max_events=10, window_seconds=300)
|
||||||
|
otc_request_limiter = SlidingWindowLimiter(max_events=5, window_seconds=300)
|
||||||
|
# /auth/passcode/check is an anonymous has-passcode oracle (audit 0026 L3).
|
||||||
|
# It's a legitimate Login-flow affordance, so the budget is generous —
|
||||||
|
# enough for a human typing emails, tight enough to stop bulk scraping.
|
||||||
|
check_limiter = SlidingWindowLimiter(max_events=30, window_seconds=300)
|
||||||
|
|
||||||
|
|
||||||
|
def _reset_all_for_tests() -> None:
|
||||||
|
"""Clear every module-level limiter's window. Test support only — the
|
||||||
|
limiters are process-global singletons, so without a per-test reset
|
||||||
|
one test's requests bleed into the next and later tests trip the
|
||||||
|
budget (429). Not called in production."""
|
||||||
|
for lim in (verify_limiter, otc_request_limiter, check_limiter):
|
||||||
|
with lim._lock:
|
||||||
|
lim._events.clear()
|
||||||
|
|
||||||
|
|
||||||
|
def client_key(request) -> str:
|
||||||
|
"""Best-effort client identity for limiting. Behind nginx the app is
|
||||||
|
started with `--forwarded-allow-ips 127.0.0.1`, so `request.client.host`
|
||||||
|
reflects the real client IP via Uvicorn's ProxyHeaders handling."""
|
||||||
|
client = getattr(request, "client", None)
|
||||||
|
return client.host if client and client.host else "unknown"
|
||||||
@@ -0,0 +1,158 @@
|
|||||||
|
"""Roadmap #28 Part 1 — auto-link RFC references in submitted prose.
|
||||||
|
|
||||||
|
Scans plain-text PR descriptions and comment bodies for references to
|
||||||
|
existing **accepted** (state='active') RFCs and returns a structured list
|
||||||
|
of *segments* the frontend renders: plain-text runs interleaved with
|
||||||
|
``{"type": "rfc", ...}`` link segments. The backend never emits HTML —
|
||||||
|
the frontend maps link segments onto React anchors — so the surface is
|
||||||
|
XSS-safe by construction and independent of any HTML-sanitization layer.
|
||||||
|
|
||||||
|
**Read-time enrichment, not submit-time persistence.** The roadmap row
|
||||||
|
phrases the scan as happening "at submit/post time"; this module instead
|
||||||
|
enriches on read. The intent the roadmap actually names — "not as live
|
||||||
|
compose preview" — is honored (drafts are never scanned, only submitted
|
||||||
|
content on the read paths). Read-time was chosen for three reasons:
|
||||||
|
|
||||||
|
1. Correctness — links track the *live* active-RFC set. A newly-accepted
|
||||||
|
RFC starts linking in older comments; a withdrawn RFC stops linking
|
||||||
|
everywhere. Submit-time freezing would drift stale.
|
||||||
|
2. Zero migration — no derived data to store. (A concurrent session
|
||||||
|
already holds migration 023; staying migration-free keeps this slice
|
||||||
|
conflict-free as well as simpler.)
|
||||||
|
3. Cost — the active-RFC corpus is small and cache-resident, so building
|
||||||
|
the term index and scanning a ≤20k-char body per read is cheap.
|
||||||
|
|
||||||
|
**Matching is conservative by design.** Only references that are unlikely
|
||||||
|
to be coincidental link:
|
||||||
|
|
||||||
|
* ``rfc_id`` tokens (e.g. ``RFC-0001``) — inherently specific.
|
||||||
|
* Multi-word titles (containing whitespace, e.g. ``Open Human Model``).
|
||||||
|
* Hyphenated slugs (containing ``-``, e.g. ``open-human-model``).
|
||||||
|
|
||||||
|
Single common-word titles or slugs (e.g. a hypothetical RFC titled
|
||||||
|
"Human") are deliberately NOT auto-linked — they would turn every prose
|
||||||
|
"human" into a link. Surfacing those is the job of the roadmap's
|
||||||
|
"curated canonical-terms list", an explicit per-deployment opt-in left as
|
||||||
|
a future extension rather than guessed at here.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from typing import Any, Iterable
|
||||||
|
|
||||||
|
|
||||||
|
def _is_word_char(c: str) -> bool:
|
||||||
|
"""Word-boundary test. Hyphen and underscore count as word chars so a
|
||||||
|
match can't begin or end in the middle of a kebab/snake token."""
|
||||||
|
return c.isalnum() or c in ("-", "_")
|
||||||
|
|
||||||
|
|
||||||
|
def segment_text(text: str | None, terms: list[tuple[str, str, str]]) -> list[dict[str, Any]]:
|
||||||
|
"""Split ``text`` into text / rfc-link segments against ``terms``.
|
||||||
|
|
||||||
|
``terms`` is a list of ``(key_lower, slug, title)`` tuples; callers
|
||||||
|
pass it pre-sorted longest-first so the longest match wins at any
|
||||||
|
position (so "Open Human Model" wins over a bare "Open"). Matching is
|
||||||
|
case-insensitive and respects word boundaries on both ends. The
|
||||||
|
returned ``label`` preserves the source casing.
|
||||||
|
|
||||||
|
Always returns at least one segment; for empty/None input that is a
|
||||||
|
single empty text segment, so callers can render uniformly.
|
||||||
|
"""
|
||||||
|
if not text:
|
||||||
|
return [{"type": "text", "text": text or ""}]
|
||||||
|
|
||||||
|
out: list[dict[str, Any]] = []
|
||||||
|
buf: list[str] = []
|
||||||
|
low = text.lower()
|
||||||
|
n = len(text)
|
||||||
|
i = 0
|
||||||
|
while i < n:
|
||||||
|
match: tuple[str, str, str, int] | None = None
|
||||||
|
for key, slug, title in terms:
|
||||||
|
klen = len(key)
|
||||||
|
if klen == 0 or not low.startswith(key, i):
|
||||||
|
continue
|
||||||
|
before = text[i - 1] if i > 0 else ""
|
||||||
|
after = text[i + klen] if i + klen < n else ""
|
||||||
|
if _is_word_char(before) or _is_word_char(after):
|
||||||
|
continue
|
||||||
|
match = (key, slug, title, klen)
|
||||||
|
break
|
||||||
|
if match is not None:
|
||||||
|
_key, slug, title, klen = match
|
||||||
|
if buf:
|
||||||
|
out.append({"type": "text", "text": "".join(buf)})
|
||||||
|
buf = []
|
||||||
|
out.append({
|
||||||
|
"type": "rfc",
|
||||||
|
"slug": slug,
|
||||||
|
"label": text[i:i + klen],
|
||||||
|
"title": title,
|
||||||
|
})
|
||||||
|
i += klen
|
||||||
|
else:
|
||||||
|
buf.append(text[i])
|
||||||
|
i += 1
|
||||||
|
if buf:
|
||||||
|
out.append({"type": "text", "text": "".join(buf)})
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def _keys_for(slug: str, title: str, rfc_id: str | None) -> Iterable[str]:
|
||||||
|
"""The match keys an active RFC contributes. See the module docstring
|
||||||
|
for why each gate exists (conservative, false-positive-averse)."""
|
||||||
|
if rfc_id:
|
||||||
|
rid = rfc_id.strip()
|
||||||
|
if len(rid) >= 2:
|
||||||
|
yield rid.lower()
|
||||||
|
if title:
|
||||||
|
t = title.strip()
|
||||||
|
# Multi-word titles only — a single common word is too noisy.
|
||||||
|
if len(t) >= 2 and (" " in t or "\t" in t):
|
||||||
|
yield t.lower()
|
||||||
|
if slug:
|
||||||
|
s = slug.strip()
|
||||||
|
# Hyphenated slugs only — a single-token slug is a bare word.
|
||||||
|
if len(s) >= 2 and "-" in s:
|
||||||
|
yield s.lower()
|
||||||
|
|
||||||
|
|
||||||
|
class LinkIndex:
|
||||||
|
"""A reusable term index built once per request and applied to many
|
||||||
|
bodies (a PR's description plus every comment on it)."""
|
||||||
|
|
||||||
|
def __init__(self, terms: list[tuple[str, str, str]]):
|
||||||
|
# Longest key first so the longest reference wins at each position.
|
||||||
|
self._terms = sorted(terms, key=lambda t: len(t[0]), reverse=True)
|
||||||
|
|
||||||
|
def __bool__(self) -> bool:
|
||||||
|
return bool(self._terms)
|
||||||
|
|
||||||
|
def segment(self, text: str | None) -> list[dict[str, Any]]:
|
||||||
|
return segment_text(text, self._terms)
|
||||||
|
|
||||||
|
|
||||||
|
def build_index(conn, *, exclude_slug: str | None = None) -> LinkIndex:
|
||||||
|
"""Build a :class:`LinkIndex` from the accepted (active) RFC corpus.
|
||||||
|
|
||||||
|
``exclude_slug`` drops the RFC the surrounding surface is itself scoped
|
||||||
|
to, so an RFC's own title/id/slug don't self-link inside its own PR or
|
||||||
|
discussion. ``ORDER BY slug`` makes key de-duplication deterministic
|
||||||
|
when two RFCs would contribute the same key (first slug wins)."""
|
||||||
|
rows = conn.execute(
|
||||||
|
"SELECT slug, title, rfc_id FROM cached_rfcs WHERE state = 'active' ORDER BY slug"
|
||||||
|
).fetchall()
|
||||||
|
terms: list[tuple[str, str, str]] = []
|
||||||
|
seen: set[str] = set()
|
||||||
|
for r in rows:
|
||||||
|
slug = r["slug"]
|
||||||
|
if exclude_slug is not None and slug == exclude_slug:
|
||||||
|
continue
|
||||||
|
title = r["title"] or ""
|
||||||
|
rfc_id = r["rfc_id"] if "rfc_id" in r.keys() else None
|
||||||
|
for key in _keys_for(slug, title, rfc_id):
|
||||||
|
if key in seen:
|
||||||
|
continue
|
||||||
|
seen.add(key)
|
||||||
|
terms.append((key, slug, title))
|
||||||
|
return LinkIndex(terms)
|
||||||
@@ -0,0 +1,244 @@
|
|||||||
|
"""Roadmap #27 — Claude Haiku tag suggestions for the propose-RFC modal.
|
||||||
|
|
||||||
|
A cheap, fast assist: given the partial RFC draft a user is typing, ask
|
||||||
|
Claude Haiku to recommend tags drawn ONLY from the tags already in use
|
||||||
|
across the corpus. v1 has no curated tag list — tags are free-form chip
|
||||||
|
input (§9.1) — so "the taxonomy" is the de-facto set of distinct tags
|
||||||
|
the existing RFCs already carry. The model is constrained to that set
|
||||||
|
and MUST NOT invent new tags; taxonomy extension (letting the model
|
||||||
|
propose genuinely new tags) is a deferred follow-up per the roadmap.
|
||||||
|
|
||||||
|
Why Haiku specifically: cost. Picking a few reasonable tags from a known
|
||||||
|
set is well within Haiku's range, and the modal fires this repeatedly as
|
||||||
|
the user types, so the per-call price has to stay small.
|
||||||
|
|
||||||
|
The whole surface degrades to silence rather than error: no Anthropic
|
||||||
|
key, no provider, an empty corpus, a rate-limited caller, an empty
|
||||||
|
draft, or an unparseable model reply all yield an empty suggestion list.
|
||||||
|
The propose modal hides its suggestion row on an empty list, so the
|
||||||
|
fallback is simply the existing free-form chip input with nothing extra
|
||||||
|
shown.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
import logging
|
||||||
|
import os
|
||||||
|
import time
|
||||||
|
from dataclasses import dataclass
|
||||||
|
|
||||||
|
from . import db
|
||||||
|
from .providers import BaseProvider, construct_haiku
|
||||||
|
|
||||||
|
log = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
# Bound the universe handed to the model so a large corpus can't blow up
|
||||||
|
# the prompt size (and the cost). The most-common tags matter most.
|
||||||
|
_UNIVERSE_CAP = 200
|
||||||
|
|
||||||
|
# How many suggestions we ever return to the modal.
|
||||||
|
DEFAULT_MAX_SUGGESTIONS = 6
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class Draft:
|
||||||
|
"""The partial propose-RFC draft. `pitch` is the "why is this needed"
|
||||||
|
rationale; `use_case` is the #26 optional ground-truth field."""
|
||||||
|
|
||||||
|
title: str = ""
|
||||||
|
pitch: str = ""
|
||||||
|
use_case: str = ""
|
||||||
|
|
||||||
|
def is_empty(self) -> bool:
|
||||||
|
return not (self.title.strip() or self.pitch.strip() or self.use_case.strip())
|
||||||
|
|
||||||
|
|
||||||
|
def haiku_provider(config) -> BaseProvider | None:
|
||||||
|
"""Construct a dedicated Claude Haiku provider from the operator's
|
||||||
|
Anthropic key, or return None when no key is configured.
|
||||||
|
|
||||||
|
None means "suggestions unavailable" — the caller returns an empty
|
||||||
|
list and the modal shows nothing. This is the seam tests monkeypatch
|
||||||
|
to inject a fake provider without a real key. There is no RFC slug at
|
||||||
|
propose time, so the §6.7 per-RFC funder path deliberately does not
|
||||||
|
apply: tag suggestion always runs on the operator's own key.
|
||||||
|
"""
|
||||||
|
key = getattr(config, "anthropic_api_key", "") or ""
|
||||||
|
if not key:
|
||||||
|
return None
|
||||||
|
return construct_haiku(key)
|
||||||
|
|
||||||
|
|
||||||
|
def gather_tag_universe(cap: int = _UNIVERSE_CAP) -> list[str]:
|
||||||
|
"""Every distinct tag in use across the cached corpus, most-common
|
||||||
|
first (ties broken alphabetically for determinism), capped.
|
||||||
|
|
||||||
|
This is the universe the model is constrained to. An empty corpus
|
||||||
|
yields an empty list, which short-circuits suggestion to silence.
|
||||||
|
"""
|
||||||
|
rows = db.conn().execute("SELECT tags_json FROM cached_rfcs").fetchall()
|
||||||
|
counts: dict[str, int] = {}
|
||||||
|
for r in rows:
|
||||||
|
try:
|
||||||
|
tags = json.loads(r["tags_json"] or "[]")
|
||||||
|
except (ValueError, TypeError):
|
||||||
|
continue
|
||||||
|
if not isinstance(tags, list):
|
||||||
|
continue
|
||||||
|
for t in tags:
|
||||||
|
if not isinstance(t, str):
|
||||||
|
continue
|
||||||
|
tag = t.strip()
|
||||||
|
if tag:
|
||||||
|
counts[tag] = counts.get(tag, 0) + 1
|
||||||
|
ranked = sorted(counts, key=lambda t: (-counts[t], t.lower()))
|
||||||
|
return ranked[:cap]
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Rate limiting — in-process, per-user sliding window.
|
||||||
|
#
|
||||||
|
# Cost control, not security: the `require_contributor` gate already
|
||||||
|
# bounds callers to authenticated beta users, and the modal debounces.
|
||||||
|
# This is the backstop against a stuck/abusive client hammering the
|
||||||
|
# endpoint. In-memory is sufficient (single-process uvicorn on the VM)
|
||||||
|
# and resets on restart, which is fine for a cost guard.
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
_CALLS: dict[int, list[float]] = {}
|
||||||
|
|
||||||
|
|
||||||
|
def _rate_config() -> tuple[int, float]:
|
||||||
|
try:
|
||||||
|
max_calls = int(os.environ.get("TAG_SUGGEST_RATE_MAX", "30"))
|
||||||
|
except ValueError:
|
||||||
|
max_calls = 30
|
||||||
|
try:
|
||||||
|
window = float(os.environ.get("TAG_SUGGEST_RATE_WINDOW_SECONDS", "60"))
|
||||||
|
except ValueError:
|
||||||
|
window = 60.0
|
||||||
|
return max_calls, window
|
||||||
|
|
||||||
|
|
||||||
|
def rate_limit_ok(user_id: int, *, now: float | None = None) -> bool:
|
||||||
|
"""True if this call is within the per-user window; records the call.
|
||||||
|
`now` is injectable for tests (defaults to a monotonic clock)."""
|
||||||
|
max_calls, window = _rate_config()
|
||||||
|
t = time.monotonic() if now is None else now
|
||||||
|
calls = _CALLS.setdefault(user_id, [])
|
||||||
|
cutoff = t - window
|
||||||
|
calls[:] = [c for c in calls if c > cutoff]
|
||||||
|
if len(calls) >= max_calls:
|
||||||
|
return False
|
||||||
|
calls.append(t)
|
||||||
|
return True
|
||||||
|
|
||||||
|
|
||||||
|
def reset_rate_limits() -> None:
|
||||||
|
"""Test seam — clear the in-process window state."""
|
||||||
|
_CALLS.clear()
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Prompt + parse.
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
_SYSTEM = (
|
||||||
|
"You suggest topic tags for a draft RFC — a structured proposal "
|
||||||
|
"document in a collection. You are given the draft's title, its "
|
||||||
|
"rationale, an optional use case, and the exact set of tags already "
|
||||||
|
"in use across the collection. Choose the tags from that set that "
|
||||||
|
"best fit the draft.\n"
|
||||||
|
"Rules:\n"
|
||||||
|
"- Choose ONLY from the provided tag set. Never invent a tag.\n"
|
||||||
|
"- Order best-fit first. Omit weak fits rather than padding the list.\n"
|
||||||
|
"- Return at most {max} tags.\n"
|
||||||
|
"Respond with ONLY a JSON array, no prose, of the form:\n"
|
||||||
|
'[{{"tag": "<exact tag from the set>", "confidence": <number between 0 and 1>}}]\n'
|
||||||
|
"If no tag in the set fits the draft, return []."
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _build_messages(draft: Draft, universe: list[str], max_suggestions: int):
|
||||||
|
system = _SYSTEM.format(max=max_suggestions)
|
||||||
|
parts: list[str] = []
|
||||||
|
if draft.title.strip():
|
||||||
|
parts.append(f"Title: {draft.title.strip()[:300]}")
|
||||||
|
if draft.pitch.strip():
|
||||||
|
parts.append(f"Why this RFC is needed:\n{draft.pitch.strip()[:4000]}")
|
||||||
|
if draft.use_case.strip():
|
||||||
|
parts.append(f"What it will be used for:\n{draft.use_case.strip()[:4000]}")
|
||||||
|
parts.append(
|
||||||
|
"Tags already in use (choose only from these):\n" + ", ".join(universe)
|
||||||
|
)
|
||||||
|
return system, [{"role": "user", "content": "\n\n".join(parts)}]
|
||||||
|
|
||||||
|
|
||||||
|
def _extract_json_array(text: str):
|
||||||
|
start = text.find("[")
|
||||||
|
end = text.rfind("]")
|
||||||
|
if start == -1 or end == -1 or end < start:
|
||||||
|
return None
|
||||||
|
try:
|
||||||
|
return json.loads(text[start : end + 1])
|
||||||
|
except ValueError:
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def parse_reply(text: str, universe: list[str], max_suggestions: int) -> list[dict]:
|
||||||
|
"""Parse the model reply into a clean, universe-constrained list.
|
||||||
|
|
||||||
|
Tolerant of the model returning bare strings or objects, extra prose
|
||||||
|
around the JSON, unknown/invented tags (dropped), duplicate tags
|
||||||
|
(deduped), and missing/garbage confidences (defaulted + clamped).
|
||||||
|
"""
|
||||||
|
data = _extract_json_array(text or "")
|
||||||
|
if not isinstance(data, list):
|
||||||
|
return []
|
||||||
|
# Map back to the canonical spelling in the universe, case-insensitively,
|
||||||
|
# so a model that lowercases a tag still resolves to the real one.
|
||||||
|
canonical = {t.lower(): t for t in universe}
|
||||||
|
out: list[dict] = []
|
||||||
|
seen: set[str] = set()
|
||||||
|
for item in data:
|
||||||
|
if isinstance(item, dict):
|
||||||
|
raw = item.get("tag")
|
||||||
|
conf = item.get("confidence")
|
||||||
|
elif isinstance(item, str):
|
||||||
|
raw, conf = item, None
|
||||||
|
else:
|
||||||
|
continue
|
||||||
|
if not isinstance(raw, str):
|
||||||
|
continue
|
||||||
|
tag = canonical.get(raw.strip().lower())
|
||||||
|
if tag is None or tag in seen:
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
c = float(conf) if conf is not None else 0.5
|
||||||
|
except (ValueError, TypeError):
|
||||||
|
c = 0.5
|
||||||
|
c = max(0.0, min(1.0, c))
|
||||||
|
out.append({"tag": tag, "confidence": round(c, 3)})
|
||||||
|
seen.add(tag)
|
||||||
|
if len(out) >= max_suggestions:
|
||||||
|
break
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def suggest(
|
||||||
|
provider: BaseProvider,
|
||||||
|
draft: Draft,
|
||||||
|
universe: list[str],
|
||||||
|
max_suggestions: int = DEFAULT_MAX_SUGGESTIONS,
|
||||||
|
) -> list[dict]:
|
||||||
|
"""Orchestrate one suggestion call. Returns [] for an empty draft or
|
||||||
|
empty universe (no model call), and for any provider/parse failure."""
|
||||||
|
if draft.is_empty() or not universe:
|
||||||
|
return []
|
||||||
|
system, history = _build_messages(draft, universe, max_suggestions)
|
||||||
|
try:
|
||||||
|
text = provider.send(system, history)
|
||||||
|
except Exception as exc: # provider/network failure → silent empty
|
||||||
|
log.warning("tag-suggest provider failed: %s", exc)
|
||||||
|
return []
|
||||||
|
return parse_reply(text, universe, max_suggestions)
|
||||||
@@ -0,0 +1,18 @@
|
|||||||
|
-- v0.25.0 / security audit 0026, finding H1.
|
||||||
|
--
|
||||||
|
-- The OTC verify path had no attempt-limit or lockout, unlike the
|
||||||
|
-- passcode path (015_passcode.sql gave users.passcode_failed_attempts +
|
||||||
|
-- passcode_locked_until). This table gives the OTC verify endpoint the
|
||||||
|
-- same per-identity lockout shape. It is keyed by email rather than
|
||||||
|
-- user_id because an OTC sign-in may not have a users row yet — the row
|
||||||
|
-- is provisioned only on a *successful* verify, so the lockout state has
|
||||||
|
-- to survive independently of it.
|
||||||
|
--
|
||||||
|
-- The per-IP rate limiter (app/ratelimit.py) is the primary brute-force
|
||||||
|
-- defense; this table is the parity layer that mirrors the passcode
|
||||||
|
-- lockout and persists across restarts.
|
||||||
|
CREATE TABLE IF NOT EXISTS otc_verify_state (
|
||||||
|
email TEXT PRIMARY KEY,
|
||||||
|
failed_attempts INTEGER NOT NULL DEFAULT 0,
|
||||||
|
locked_until TEXT
|
||||||
|
);
|
||||||
@@ -0,0 +1,19 @@
|
|||||||
|
"""Shared pytest fixtures for the backend suite.
|
||||||
|
|
||||||
|
Added in v0.27.0 (security audit 0026) alongside the new per-IP rate
|
||||||
|
limiter. The limiters in `app.ratelimit` are process-global singletons,
|
||||||
|
so their state survives across tests within a run; without a reset, the
|
||||||
|
accumulated requests from earlier tests exhaust the budget and later
|
||||||
|
tests see spurious 429s. This autouse fixture gives every test a clean
|
||||||
|
limiter window.
|
||||||
|
"""
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from app import ratelimit
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture(autouse=True)
|
||||||
|
def _reset_rate_limiters():
|
||||||
|
ratelimit._reset_all_for_tests()
|
||||||
|
yield
|
||||||
|
ratelimit._reset_all_for_tests()
|
||||||
@@ -444,6 +444,18 @@ def tmp_env(monkeypatch):
|
|||||||
# the dev-bypass path monkeypatch `RFC_APP_INSECURE_WEBHOOKS=1`.
|
# the dev-bypass path monkeypatch `RFC_APP_INSECURE_WEBHOOKS=1`.
|
||||||
"GITEA_WEBHOOK_SECRET": "test-webhook-secret-for-signature-verification",
|
"GITEA_WEBHOOK_SECRET": "test-webhook-secret-for-signature-verification",
|
||||||
"ENABLED_MODELS": "claude",
|
"ENABLED_MODELS": "claude",
|
||||||
|
# v0.27.0 (audit 0026 M4): the session cookie now defaults to
|
||||||
|
# Secure. The TestClient talks plain http://testserver, so a
|
||||||
|
# Secure cookie is never sent back and every authenticated flow
|
||||||
|
# would fail. Tests opt out explicitly, exactly as a dev box on
|
||||||
|
# plain http does.
|
||||||
|
"SESSION_COOKIE_SECURE": "false",
|
||||||
|
# v0.27.0 (audit 0026 M5): the bounce webhook fails closed (503)
|
||||||
|
# when its secret is unset. Tests exercise the legacy behavioral
|
||||||
|
# path via the documented dev opt-in, mirroring the
|
||||||
|
# RFC_APP_INSECURE_WEBHOOKS bypass above. Tests that assert the
|
||||||
|
# fail-closed default delenv this key themselves.
|
||||||
|
"RFC_APP_INSECURE_BOUNCE_WEBHOOK": "1",
|
||||||
}
|
}
|
||||||
for k, v in env.items():
|
for k, v in env.items():
|
||||||
monkeypatch.setenv(k, v)
|
monkeypatch.setenv(k, v)
|
||||||
|
|||||||
@@ -0,0 +1,235 @@
|
|||||||
|
"""Roadmap #28 Part 1 — auto-link RFC references in PR text + comments.
|
||||||
|
|
||||||
|
Two layers:
|
||||||
|
|
||||||
|
* Unit tests over the pure scanner (`rfc_links.segment_text` /
|
||||||
|
`_keys_for` / `LinkIndex`) — the matching rules and their
|
||||||
|
false-positive guards, no DB.
|
||||||
|
* End-to-end tests that the PR description, PR review comments, and
|
||||||
|
PR-less discussion comments all surface `*_segments` enriched against
|
||||||
|
the live accepted-RFC corpus, with self-references suppressed.
|
||||||
|
|
||||||
|
Reuses the FakeGitea + session helpers from test_propose_vertical.py and
|
||||||
|
the active-RFC seed from test_rfc_view_vertical.py.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from app import rfc_links
|
||||||
|
|
||||||
|
from test_propose_vertical import ( # noqa: F401
|
||||||
|
FakeGitea,
|
||||||
|
app_with_fake_gitea,
|
||||||
|
grant_rfc_collaborator,
|
||||||
|
provision_user_row,
|
||||||
|
sign_in_as,
|
||||||
|
tmp_env,
|
||||||
|
)
|
||||||
|
from test_rfc_view_vertical import SEED_BODY, seed_active_rfc
|
||||||
|
from test_pr_flow_vertical import _cut_branch_and_accept_change
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Unit — the pure scanner
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def _idx(*terms):
|
||||||
|
"""Build a LinkIndex from raw (key, slug, title) tuples (keys lower)."""
|
||||||
|
return rfc_links.LinkIndex(list(terms))
|
||||||
|
|
||||||
|
|
||||||
|
def test_empty_text_is_single_empty_segment():
|
||||||
|
assert rfc_links.segment_text("", []) == [{"type": "text", "text": ""}]
|
||||||
|
assert rfc_links.segment_text(None, []) == [{"type": "text", "text": ""}]
|
||||||
|
|
||||||
|
|
||||||
|
def test_no_terms_returns_plain_text():
|
||||||
|
out = rfc_links.segment_text("hello world", [])
|
||||||
|
assert out == [{"type": "text", "text": "hello world"}]
|
||||||
|
|
||||||
|
|
||||||
|
def test_multiword_title_links_and_preserves_casing():
|
||||||
|
idx = _idx(("open human model", "open-human-model", "Open Human Model"))
|
||||||
|
out = idx.segment("See the Open Human Model for details.")
|
||||||
|
assert out == [
|
||||||
|
{"type": "text", "text": "See the "},
|
||||||
|
{"type": "rfc", "slug": "open-human-model", "label": "Open Human Model",
|
||||||
|
"title": "Open Human Model"},
|
||||||
|
{"type": "text", "text": " for details."},
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def test_match_is_case_insensitive():
|
||||||
|
idx = _idx(("open human model", "open-human-model", "Open Human Model"))
|
||||||
|
out = idx.segment("see the OPEN HUMAN MODEL")
|
||||||
|
assert out[-1] == {"type": "rfc", "slug": "open-human-model",
|
||||||
|
"label": "OPEN HUMAN MODEL", "title": "Open Human Model"}
|
||||||
|
|
||||||
|
|
||||||
|
def test_word_boundary_prevents_substring_match():
|
||||||
|
# "harm" must not match inside "charming" / "harmless".
|
||||||
|
idx = _idx(("rfc-0001", "open-human-model", "Open Human Model"))
|
||||||
|
out = idx.segment("a charming rfc-00012 not real")
|
||||||
|
# rfc-0001 is a prefix of rfc-00012 but the trailing '2' is a word char,
|
||||||
|
# so no match — the whole string stays plain text.
|
||||||
|
assert out == [{"type": "text", "text": "a charming rfc-00012 not real"}]
|
||||||
|
|
||||||
|
|
||||||
|
def test_rfc_id_token_links():
|
||||||
|
idx = _idx(("rfc-0001", "open-human-model", "Open Human Model"))
|
||||||
|
out = idx.segment("as established in RFC-0001.")
|
||||||
|
assert out[1] == {"type": "rfc", "slug": "open-human-model",
|
||||||
|
"label": "RFC-0001", "title": "Open Human Model"}
|
||||||
|
|
||||||
|
|
||||||
|
def test_longest_match_wins():
|
||||||
|
# A bare "Open" term and the full title both present; the full title
|
||||||
|
# (longer) must win at the position.
|
||||||
|
idx = _idx(
|
||||||
|
("open", "open", "Open"),
|
||||||
|
("open human model", "open-human-model", "Open Human Model"),
|
||||||
|
)
|
||||||
|
out = idx.segment("the Open Human Model")
|
||||||
|
assert out[-1]["slug"] == "open-human-model"
|
||||||
|
assert out[-1]["label"] == "Open Human Model"
|
||||||
|
|
||||||
|
|
||||||
|
def test_keys_for_gating():
|
||||||
|
keys = lambda **kw: set(rfc_links._keys_for(**kw))
|
||||||
|
# rfc_id always contributes.
|
||||||
|
assert "rfc-0001" in keys(slug="x", title="X", rfc_id="RFC-0001")
|
||||||
|
# multi-word title contributes; single common word does NOT.
|
||||||
|
assert "open human model" in keys(slug="ohm", title="Open Human Model", rfc_id=None)
|
||||||
|
assert keys(slug="human", title="Human", rfc_id=None) == set()
|
||||||
|
# hyphenated slug contributes; single-token slug does NOT.
|
||||||
|
assert "open-human-model" in keys(slug="open-human-model", title="X", rfc_id=None)
|
||||||
|
assert "ohm" not in keys(slug="ohm", title="OHM", rfc_id=None)
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# End-to-end — enrichment surfaces on the read paths
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def _open_pr_on(client, fake, *, host_slug: str, description: str):
|
||||||
|
"""Seed branch + accepted change on host_slug and open a PR. Returns
|
||||||
|
the pr_number."""
|
||||||
|
# `original` must exist verbatim in SEED_BODY or the accept is "stale".
|
||||||
|
branch, _ = _cut_branch_and_accept_change(
|
||||||
|
client, fake, slug=host_slug,
|
||||||
|
original="It defines consent, trait, and agency in compatible terms.",
|
||||||
|
proposed="It defines consent, trait, harm, and agency in compatible terms.",
|
||||||
|
)
|
||||||
|
r = client.post(
|
||||||
|
f"/api/rfcs/{host_slug}/branches/{branch}/open-pr",
|
||||||
|
json={"title": "A change", "description": description},
|
||||||
|
)
|
||||||
|
assert r.status_code == 200, r.text
|
||||||
|
return r.json()["pr_number"]
|
||||||
|
|
||||||
|
|
||||||
|
def _rfc_segments(segments):
|
||||||
|
return [s for s in segments if s["type"] == "rfc"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_pr_description_autolinks_other_rfc(app_with_fake_gitea):
|
||||||
|
from fastapi.testclient import TestClient
|
||||||
|
|
||||||
|
app, fake = app_with_fake_gitea
|
||||||
|
with TestClient(app) as client:
|
||||||
|
provision_user_row(user_id=2, login="alice", role="contributor")
|
||||||
|
# Two accepted RFCs: a host for the PR + a referenceable target.
|
||||||
|
seed_active_rfc(fake, slug="ohm", title="OHM", body=SEED_BODY)
|
||||||
|
seed_active_rfc(fake, slug="open-human-model", title="Open Human Model", body=SEED_BODY)
|
||||||
|
sign_in_as(client, user_id=2, gitea_login="alice", display_name="Alice", role="contributor")
|
||||||
|
|
||||||
|
pr_number = _open_pr_on(
|
||||||
|
client, fake, host_slug="ohm",
|
||||||
|
description="This builds on the Open Human Model definition.",
|
||||||
|
)
|
||||||
|
r = client.get(f"/api/rfcs/ohm/prs/{pr_number}")
|
||||||
|
assert r.status_code == 200, r.text
|
||||||
|
pr = r.json()
|
||||||
|
links = _rfc_segments(pr["description_segments"])
|
||||||
|
assert len(links) == 1
|
||||||
|
assert links[0]["slug"] == "open-human-model"
|
||||||
|
assert links[0]["label"] == "Open Human Model"
|
||||||
|
# The plain text is still present for non-segment callers (the bot
|
||||||
|
# appends a §6.5 On-behalf-of trailer, so this is a containment check).
|
||||||
|
assert "This builds on the Open Human Model definition." in pr["description"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_pr_review_comment_autolinked(app_with_fake_gitea):
|
||||||
|
from fastapi.testclient import TestClient
|
||||||
|
|
||||||
|
app, fake = app_with_fake_gitea
|
||||||
|
with TestClient(app) as client:
|
||||||
|
provision_user_row(user_id=2, login="alice", role="contributor")
|
||||||
|
seed_active_rfc(fake, slug="ohm", title="OHM", body=SEED_BODY)
|
||||||
|
seed_active_rfc(fake, slug="open-human-model", title="Open Human Model", body=SEED_BODY)
|
||||||
|
sign_in_as(client, user_id=2, gitea_login="alice", display_name="Alice", role="contributor")
|
||||||
|
|
||||||
|
pr_number = _open_pr_on(client, fake, host_slug="ohm", description="plain.")
|
||||||
|
r = client.post(
|
||||||
|
f"/api/rfcs/ohm/prs/{pr_number}/review",
|
||||||
|
json={"text": "See Open Human Model and RFC-0001.", "anchor_payload": {}},
|
||||||
|
)
|
||||||
|
assert r.status_code == 200, r.text
|
||||||
|
|
||||||
|
pr = client.get(f"/api/rfcs/ohm/prs/{pr_number}").json()
|
||||||
|
all_msgs = [m for msgs in pr["messages_by_thread"].values() for m in msgs]
|
||||||
|
review_msgs = [m for m in all_msgs if "Open Human Model" in (m["text"] or "")]
|
||||||
|
assert review_msgs, "review comment not found in payload"
|
||||||
|
links = _rfc_segments(review_msgs[0]["text_segments"])
|
||||||
|
# Both "Open Human Model" (title) and "RFC-0001" (id) point to the
|
||||||
|
# one referenceable RFC.
|
||||||
|
assert {s["slug"] for s in links} == {"open-human-model"}
|
||||||
|
assert {s["label"] for s in links} == {"Open Human Model", "RFC-0001"}
|
||||||
|
|
||||||
|
|
||||||
|
def test_discussion_comment_autolinked(app_with_fake_gitea):
|
||||||
|
from fastapi.testclient import TestClient
|
||||||
|
|
||||||
|
app, fake = app_with_fake_gitea
|
||||||
|
with TestClient(app) as client:
|
||||||
|
provision_user_row(user_id=2, login="alice", role="contributor")
|
||||||
|
seed_active_rfc(fake, slug="ohm", title="OHM", body=SEED_BODY)
|
||||||
|
seed_active_rfc(fake, slug="open-human-model", title="Open Human Model", body=SEED_BODY)
|
||||||
|
# alice is the seeded owner of ohm (owners=["alice"]); grant the
|
||||||
|
# per-RFC collaborator row explicitly so the #12 discuss gate passes.
|
||||||
|
grant_rfc_collaborator(user_id=2, rfc_slug="ohm", role_in_rfc="contributor")
|
||||||
|
sign_in_as(client, user_id=2, gitea_login="alice", display_name="Alice", role="contributor")
|
||||||
|
|
||||||
|
r = client.post(
|
||||||
|
"/api/rfcs/ohm/discussion/threads",
|
||||||
|
json={"message": "Compare with the Open Human Model."},
|
||||||
|
)
|
||||||
|
assert r.status_code == 200, r.text
|
||||||
|
thread_id = r.json()["thread_id"]
|
||||||
|
|
||||||
|
r = client.get(f"/api/rfcs/ohm/discussion/threads/{thread_id}/messages")
|
||||||
|
assert r.status_code == 200, r.text
|
||||||
|
msgs = r.json()["messages"]
|
||||||
|
assert msgs and "text_segments" in msgs[0]
|
||||||
|
links = _rfc_segments(msgs[0]["text_segments"])
|
||||||
|
assert len(links) == 1
|
||||||
|
assert links[0]["slug"] == "open-human-model"
|
||||||
|
|
||||||
|
|
||||||
|
def test_self_reference_not_linked(app_with_fake_gitea):
|
||||||
|
from fastapi.testclient import TestClient
|
||||||
|
|
||||||
|
app, fake = app_with_fake_gitea
|
||||||
|
with TestClient(app) as client:
|
||||||
|
provision_user_row(user_id=2, login="alice", role="contributor")
|
||||||
|
# The host RFC has a multi-word title, so absent exclude_slug it
|
||||||
|
# WOULD self-link. exclude_slug must suppress it.
|
||||||
|
seed_active_rfc(fake, slug="open-human-model", title="Open Human Model", body=SEED_BODY)
|
||||||
|
sign_in_as(client, user_id=2, gitea_login="alice", display_name="Alice", role="contributor")
|
||||||
|
|
||||||
|
pr_number = _open_pr_on(
|
||||||
|
client, fake, host_slug="open-human-model",
|
||||||
|
description="Refines the Open Human Model definition.",
|
||||||
|
)
|
||||||
|
pr = client.get(f"/api/rfcs/open-human-model/prs/{pr_number}").json()
|
||||||
|
assert _rfc_segments(pr["description_segments"]) == []
|
||||||
@@ -0,0 +1,264 @@
|
|||||||
|
"""Vertical + unit coverage for roadmap #27 (rfc-app v0.24.0): Claude
|
||||||
|
Haiku tag suggestions on the propose-RFC modal.
|
||||||
|
|
||||||
|
Reuses the FakeGitea + session helpers from test_propose_vertical.py.
|
||||||
|
The Anthropic call is never made for real — tests monkeypatch the
|
||||||
|
`tag_suggest.haiku_provider` seam with a stub provider whose `send`
|
||||||
|
returns canned text, so the HTTP contract is exercised without a key.
|
||||||
|
|
||||||
|
Proves:
|
||||||
|
(a) the endpoint is contributor-gated (anon → 401);
|
||||||
|
(b) a contributor gets suggestions, filtered to the corpus tag
|
||||||
|
universe, with invented tags dropped;
|
||||||
|
(c) no Anthropic key bound → empty list, not an error;
|
||||||
|
(d) an empty corpus → empty list (model is never even called);
|
||||||
|
(e) the per-user rate limit surfaces as a 429;
|
||||||
|
plus unit coverage of the universe gather, the reply parser's tolerance,
|
||||||
|
and the suggest() orchestration short-circuits.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from test_propose_vertical import ( # noqa: F401
|
||||||
|
FakeGitea,
|
||||||
|
app_with_fake_gitea,
|
||||||
|
provision_user_row,
|
||||||
|
sign_in_as,
|
||||||
|
tmp_env,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Helpers
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
class StubProvider:
|
||||||
|
"""A BaseProvider stand-in whose send() returns a fixed string (or
|
||||||
|
raises, to exercise the graceful-failure path)."""
|
||||||
|
|
||||||
|
def __init__(self, reply: str = "[]", *, raises: bool = False):
|
||||||
|
self.reply = reply
|
||||||
|
self.raises = raises
|
||||||
|
self.calls: list[tuple[str, list]] = []
|
||||||
|
|
||||||
|
def send(self, system, history):
|
||||||
|
self.calls.append((system, history))
|
||||||
|
if self.raises:
|
||||||
|
raise RuntimeError("boom")
|
||||||
|
return self.reply
|
||||||
|
|
||||||
|
|
||||||
|
def _seed_tags(slug: str, title: str, tags: list[str], state: str = "active") -> None:
|
||||||
|
from app import db
|
||||||
|
|
||||||
|
db.conn().execute(
|
||||||
|
"INSERT OR REPLACE INTO cached_rfcs (slug, title, state, tags_json) VALUES (?, ?, ?, ?)",
|
||||||
|
(slug, title, state, json.dumps(tags)),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _use_provider(monkeypatch, provider) -> None:
|
||||||
|
monkeypatch.setattr("app.tag_suggest.haiku_provider", lambda config: provider)
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture(autouse=True)
|
||||||
|
def _reset_rate_limits():
|
||||||
|
from app import tag_suggest
|
||||||
|
|
||||||
|
tag_suggest.reset_rate_limits()
|
||||||
|
yield
|
||||||
|
tag_suggest.reset_rate_limits()
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Endpoint (vertical)
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_anonymous_cannot_suggest(app_with_fake_gitea):
|
||||||
|
from fastapi.testclient import TestClient
|
||||||
|
|
||||||
|
app, _fake = app_with_fake_gitea
|
||||||
|
with TestClient(app) as client:
|
||||||
|
r = client.post("/api/rfcs/suggest-tags", json={"title": "X"})
|
||||||
|
assert r.status_code == 401
|
||||||
|
|
||||||
|
|
||||||
|
def test_contributor_gets_filtered_suggestions(app_with_fake_gitea, monkeypatch):
|
||||||
|
from fastapi.testclient import TestClient
|
||||||
|
|
||||||
|
app, _fake = app_with_fake_gitea
|
||||||
|
with TestClient(app) as client:
|
||||||
|
provision_user_row(user_id=2, login="alice", role="contributor")
|
||||||
|
sign_in_as(client, user_id=2, gitea_login="alice", display_name="Alice", role="contributor")
|
||||||
|
|
||||||
|
_seed_tags("ohm", "OHM", ["identity", "schema", "consent"])
|
||||||
|
_seed_tags("other", "Other", ["identity", "governance"])
|
||||||
|
|
||||||
|
# Model returns two real tags (one lowercased to test canonical
|
||||||
|
# mapping is exact-set anyway), plus one invented tag that MUST
|
||||||
|
# be dropped.
|
||||||
|
reply = json.dumps([
|
||||||
|
{"tag": "identity", "confidence": 0.9},
|
||||||
|
{"tag": "consent", "confidence": 0.7},
|
||||||
|
{"tag": "totally-invented", "confidence": 0.99},
|
||||||
|
])
|
||||||
|
stub = StubProvider(reply=reply)
|
||||||
|
_use_provider(monkeypatch, stub)
|
||||||
|
|
||||||
|
r = client.post("/api/rfcs/suggest-tags", json={
|
||||||
|
"title": "Consent and identity",
|
||||||
|
"pitch": "We need a shared definition of consent tied to identity.",
|
||||||
|
})
|
||||||
|
assert r.status_code == 200, r.text
|
||||||
|
tags = [s["tag"] for s in r.json()["suggestions"]]
|
||||||
|
assert tags == ["identity", "consent"]
|
||||||
|
# the model was actually invoked
|
||||||
|
assert len(stub.calls) == 1
|
||||||
|
# the universe (deduped distinct tags) was handed to the model
|
||||||
|
user_msg = stub.calls[0][1][0]["content"]
|
||||||
|
assert "identity" in user_msg and "governance" in user_msg
|
||||||
|
|
||||||
|
|
||||||
|
def test_no_api_key_returns_empty(app_with_fake_gitea, monkeypatch):
|
||||||
|
from fastapi.testclient import TestClient
|
||||||
|
|
||||||
|
app, _fake = app_with_fake_gitea
|
||||||
|
with TestClient(app) as client:
|
||||||
|
provision_user_row(user_id=2, login="alice", role="contributor")
|
||||||
|
sign_in_as(client, user_id=2, gitea_login="alice", display_name="Alice", role="contributor")
|
||||||
|
_seed_tags("ohm", "OHM", ["identity"])
|
||||||
|
|
||||||
|
# No key bound (test env has no ANTHROPIC_API_KEY) → provider None.
|
||||||
|
# (Explicitly assert the seam returns None given the test config.)
|
||||||
|
from app import tag_suggest
|
||||||
|
assert tag_suggest.haiku_provider(app.state.config) is None
|
||||||
|
|
||||||
|
r = client.post("/api/rfcs/suggest-tags", json={"title": "X", "pitch": "y"})
|
||||||
|
assert r.status_code == 200, r.text
|
||||||
|
assert r.json()["suggestions"] == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_empty_corpus_returns_empty_without_calling_model(app_with_fake_gitea, monkeypatch):
|
||||||
|
from fastapi.testclient import TestClient
|
||||||
|
|
||||||
|
app, _fake = app_with_fake_gitea
|
||||||
|
with TestClient(app) as client:
|
||||||
|
provision_user_row(user_id=2, login="alice", role="contributor")
|
||||||
|
sign_in_as(client, user_id=2, gitea_login="alice", display_name="Alice", role="contributor")
|
||||||
|
|
||||||
|
stub = StubProvider(reply=json.dumps([{"tag": "x", "confidence": 1}]))
|
||||||
|
_use_provider(monkeypatch, stub)
|
||||||
|
|
||||||
|
# No cached_rfcs rows → empty universe → suggest() short-circuits.
|
||||||
|
r = client.post("/api/rfcs/suggest-tags", json={"title": "X", "pitch": "y"})
|
||||||
|
assert r.status_code == 200, r.text
|
||||||
|
assert r.json()["suggestions"] == []
|
||||||
|
assert stub.calls == [] # model never invoked on an empty universe
|
||||||
|
|
||||||
|
|
||||||
|
def test_rate_limit_surfaces_429(app_with_fake_gitea, monkeypatch):
|
||||||
|
from fastapi.testclient import TestClient
|
||||||
|
|
||||||
|
monkeypatch.setenv("TAG_SUGGEST_RATE_MAX", "2")
|
||||||
|
monkeypatch.setenv("TAG_SUGGEST_RATE_WINDOW_SECONDS", "60")
|
||||||
|
|
||||||
|
app, _fake = app_with_fake_gitea
|
||||||
|
with TestClient(app) as client:
|
||||||
|
provision_user_row(user_id=2, login="alice", role="contributor")
|
||||||
|
sign_in_as(client, user_id=2, gitea_login="alice", display_name="Alice", role="contributor")
|
||||||
|
_seed_tags("ohm", "OHM", ["identity"])
|
||||||
|
_use_provider(monkeypatch, StubProvider(reply="[]"))
|
||||||
|
|
||||||
|
body = {"title": "X", "pitch": "y"}
|
||||||
|
assert client.post("/api/rfcs/suggest-tags", json=body).status_code == 200
|
||||||
|
assert client.post("/api/rfcs/suggest-tags", json=body).status_code == 200
|
||||||
|
# Third call inside the window trips the limit.
|
||||||
|
assert client.post("/api/rfcs/suggest-tags", json=body).status_code == 429
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Units
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
def test_gather_tag_universe_dedupes_and_ranks(app_with_fake_gitea):
|
||||||
|
from fastapi.testclient import TestClient
|
||||||
|
from app import tag_suggest
|
||||||
|
|
||||||
|
app, _fake = app_with_fake_gitea
|
||||||
|
# The db is initialized in the app's lifespan; enter the client
|
||||||
|
# context so cached_rfcs exists before we seed it directly.
|
||||||
|
with TestClient(app):
|
||||||
|
_seed_tags("a", "A", ["identity", "schema"])
|
||||||
|
_seed_tags("b", "B", ["identity", " schema ", "consent", ""]) # whitespace + empty
|
||||||
|
_seed_tags("c", "C", ["identity"])
|
||||||
|
|
||||||
|
universe = tag_suggest.gather_tag_universe()
|
||||||
|
# identity (3) > schema (2) > consent (1); whitespace trimmed/merged,
|
||||||
|
# empties dropped.
|
||||||
|
assert universe == ["identity", "schema", "consent"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_reply_tolerates_junk():
|
||||||
|
from app import tag_suggest
|
||||||
|
|
||||||
|
universe = ["identity", "schema", "consent"]
|
||||||
|
|
||||||
|
# Prose around the JSON, an invented tag, a bare string, a dup, a
|
||||||
|
# missing confidence, and a garbage confidence.
|
||||||
|
text = (
|
||||||
|
"Sure! Here are the tags:\n"
|
||||||
|
'[{"tag": "identity", "confidence": 0.9}, '
|
||||||
|
'{"tag": "invented", "confidence": 1}, '
|
||||||
|
'"schema", '
|
||||||
|
'{"tag": "identity", "confidence": 0.5}, '
|
||||||
|
'{"tag": "consent"}, '
|
||||||
|
'{"tag": "consent", "confidence": "high"}]\n'
|
||||||
|
"Hope that helps!"
|
||||||
|
)
|
||||||
|
out = tag_suggest.parse_reply(text, universe, max_suggestions=6)
|
||||||
|
tags = [s["tag"] for s in out]
|
||||||
|
assert tags == ["identity", "schema", "consent"] # invented dropped, deduped
|
||||||
|
by_tag = {s["tag"]: s["confidence"] for s in out}
|
||||||
|
assert by_tag["identity"] == 0.9
|
||||||
|
assert by_tag["schema"] == 0.5 # bare string defaults to 0.5
|
||||||
|
assert by_tag["consent"] == 0.5 # missing/garbage confidence → 0.5
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_reply_empty_on_unparseable():
|
||||||
|
from app import tag_suggest
|
||||||
|
|
||||||
|
assert tag_suggest.parse_reply("no json here", ["a"], 6) == []
|
||||||
|
assert tag_suggest.parse_reply("", ["a"], 6) == []
|
||||||
|
assert tag_suggest.parse_reply("[]", ["a"], 6) == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_reply_respects_max():
|
||||||
|
from app import tag_suggest
|
||||||
|
|
||||||
|
universe = ["a", "b", "c", "d", "e"]
|
||||||
|
text = json.dumps([{"tag": t, "confidence": 0.5} for t in universe])
|
||||||
|
out = tag_suggest.parse_reply(text, universe, max_suggestions=3)
|
||||||
|
assert [s["tag"] for s in out] == ["a", "b", "c"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_suggest_short_circuits_empty_draft():
|
||||||
|
from app import tag_suggest
|
||||||
|
|
||||||
|
stub = StubProvider(reply=json.dumps([{"tag": "a", "confidence": 1}]))
|
||||||
|
draft = tag_suggest.Draft(title=" ", pitch="", use_case="")
|
||||||
|
assert tag_suggest.suggest(stub, draft, ["a"]) == []
|
||||||
|
assert stub.calls == [] # never called for an empty draft
|
||||||
|
|
||||||
|
|
||||||
|
def test_suggest_returns_empty_on_provider_failure():
|
||||||
|
from app import tag_suggest
|
||||||
|
|
||||||
|
stub = StubProvider(raises=True)
|
||||||
|
draft = tag_suggest.Draft(title="Real title", pitch="a reason")
|
||||||
|
assert tag_suggest.suggest(stub, draft, ["identity"]) == []
|
||||||
@@ -18,6 +18,56 @@ server {
|
|||||||
listen [::]:80;
|
listen [::]:80;
|
||||||
server_name ohm.wiggleverse.org;
|
server_name ohm.wiggleverse.org;
|
||||||
|
|
||||||
|
# v0.25.0 security hardening (audit 0026 M2/L8)
|
||||||
|
#
|
||||||
|
# NOTE: certbot promotes THIS server block to the HTTPS listener
|
||||||
|
# (`listen 443 ssl`) and adds a separate port-80 → 443 redirect
|
||||||
|
# block (see the install comment above). These response headers
|
||||||
|
# therefore ride into the HTTPS server block on the VM. They use
|
||||||
|
# `add_header ... always` so they also apply to nginx-generated
|
||||||
|
# error responses (4xx/5xx), not just 200s.
|
||||||
|
#
|
||||||
|
# `server_tokens off` (L8) — suppress the nginx version in the
|
||||||
|
# Server header and on error pages so we don't advertise the
|
||||||
|
# build to scanners.
|
||||||
|
server_tokens off;
|
||||||
|
|
||||||
|
add_header Strict-Transport-Security "max-age=31536000; includeSubDomains" always;
|
||||||
|
add_header X-Frame-Options "DENY" always;
|
||||||
|
add_header X-Content-Type-Options "nosniff" always;
|
||||||
|
add_header Referrer-Policy "strict-origin-when-cross-origin" always;
|
||||||
|
|
||||||
|
# Content-Security-Policy (M2). Tuned to what the SPA actually loads:
|
||||||
|
# - default-src 'self': everything not called out below is same-origin.
|
||||||
|
# - script-src 'self' + challenges.cloudflare.com: the only external
|
||||||
|
# <script> tag the app injects is the CloudFlare Turnstile widget
|
||||||
|
# (frontend/src/components/TurnstileWidget.jsx). Amplitude and
|
||||||
|
# mermaid are BUNDLED (dynamic `import()` from node_modules, served
|
||||||
|
# from 'self'), so they need no extra script origin — *.amplitude.com
|
||||||
|
# is listed defensively in case a future SDK build script-injects.
|
||||||
|
# script-src DELIBERATELY OMITS 'unsafe-inline' — no inline <script>
|
||||||
|
# is used, so we keep XSS-via-inline-script blocked.
|
||||||
|
# - style-src 'unsafe-inline' IS REQUIRED by the current build: the
|
||||||
|
# JSX uses inline `style={...}` attributes throughout and mermaid
|
||||||
|
# injects <style> blocks at render time. Removing it would break
|
||||||
|
# layout; tightening this is a future build-side change (nonce/hash).
|
||||||
|
# - img-src 'self' data: https: — markdown/RFC bodies may embed remote
|
||||||
|
# images and data: URIs; svg/mermaid output uses data: too.
|
||||||
|
# - font-src 'self' data: — bundled fonts plus data: webfonts.
|
||||||
|
# - connect-src 'self' + *.amplitude.com + challenges.cloudflare.com:
|
||||||
|
# the app's API/auth/SSE are same-origin (nginx proxy); Amplitude
|
||||||
|
# Analytics + Session Replay (shipped at sampleRate 1) POST to
|
||||||
|
# *.amplitude.com; Turnstile verifies via challenges.cloudflare.com.
|
||||||
|
# - worker-src 'self' blob: — Amplitude Session Replay spins up a
|
||||||
|
# Web Worker from a blob: URL for capture/compression; without
|
||||||
|
# blob: here session replay breaks for every consenting user.
|
||||||
|
# - frame-src challenges.cloudflare.com — the Turnstile challenge
|
||||||
|
# renders in an iframe from that origin.
|
||||||
|
# - frame-ancestors 'none' — clickjacking defense, pairs with
|
||||||
|
# X-Frame-Options DENY for older agents.
|
||||||
|
# - base-uri 'self'; object-src 'none' — lock down <base>/<object>.
|
||||||
|
add_header Content-Security-Policy "default-src 'self'; script-src 'self' https://challenges.cloudflare.com https://*.amplitude.com; style-src 'self' 'unsafe-inline'; img-src 'self' data: https:; font-src 'self' data:; connect-src 'self' https://*.amplitude.com https://challenges.cloudflare.com; worker-src 'self' blob:; frame-src https://challenges.cloudflare.com; frame-ancestors 'none'; base-uri 'self'; object-src 'none'" always;
|
||||||
|
|
||||||
# Static SPA assets live in the Vite build output. The systemd unit
|
# Static SPA assets live in the Vite build output. The systemd unit
|
||||||
# runs as user `rfc-app`; make sure nginx (usually `www-data`) can
|
# runs as user `rfc-app`; make sure nginx (usually `www-data`) can
|
||||||
# read this path. Either group-add www-data into rfc-app's group, or
|
# read this path. Either group-add www-data into rfc-app's group, or
|
||||||
|
|||||||
@@ -42,5 +42,32 @@ ProtectHome=true
|
|||||||
PrivateTmp=true
|
PrivateTmp=true
|
||||||
ReadWritePaths=/opt/rfc-app/backend/data
|
ReadWritePaths=/opt/rfc-app/backend/data
|
||||||
|
|
||||||
|
# v0.25.0 security hardening (audit 0026 L4) — defense-in-depth.
|
||||||
|
# The service binds 127.0.0.1:8000 and runs plain CPython
|
||||||
|
# (FastAPI/uvicorn + sqlite + bcrypt + httpx), so it needs no
|
||||||
|
# capabilities and no exotic syscalls.
|
||||||
|
CapabilityBoundingSet=
|
||||||
|
AmbientCapabilities=
|
||||||
|
PrivateDevices=true
|
||||||
|
ProtectKernelTunables=true
|
||||||
|
ProtectKernelModules=true
|
||||||
|
ProtectKernelLogs=true
|
||||||
|
ProtectControlGroups=true
|
||||||
|
RestrictAddressFamilies=AF_INET AF_INET6 AF_UNIX
|
||||||
|
RestrictNamespaces=true
|
||||||
|
LockPersonality=true
|
||||||
|
# MemoryDenyWriteExecute=true blocks W^X memory — safe for stock
|
||||||
|
# CPython (no JIT) and the pure-Python/C-extension stack here, but
|
||||||
|
# would break a JIT or a C-ext that mmaps W+X. Watch the first
|
||||||
|
# restart's journal for a crash; if uvicorn fails to come up,
|
||||||
|
# comment this one line out and reload.
|
||||||
|
MemoryDenyWriteExecute=true
|
||||||
|
RestrictRealtime=true
|
||||||
|
RestrictSUIDSGID=true
|
||||||
|
SystemCallFilter=@system-service
|
||||||
|
SystemCallErrorNumber=EPERM
|
||||||
|
SystemCallArchitectures=native
|
||||||
|
UMask=0077
|
||||||
|
|
||||||
[Install]
|
[Install]
|
||||||
WantedBy=multi-user.target
|
WantedBy=multi-user.target
|
||||||
|
|||||||
Generated
+3
-2
@@ -1,12 +1,12 @@
|
|||||||
{
|
{
|
||||||
"name": "rfc-app-frontend",
|
"name": "rfc-app-frontend",
|
||||||
"version": "0.21.0",
|
"version": "0.24.0",
|
||||||
"lockfileVersion": 3,
|
"lockfileVersion": 3,
|
||||||
"requires": true,
|
"requires": true,
|
||||||
"packages": {
|
"packages": {
|
||||||
"": {
|
"": {
|
||||||
"name": "rfc-app-frontend",
|
"name": "rfc-app-frontend",
|
||||||
"version": "0.21.0",
|
"version": "0.24.0",
|
||||||
"dependencies": {
|
"dependencies": {
|
||||||
"@amplitude/unified": "^1.1.9",
|
"@amplitude/unified": "^1.1.9",
|
||||||
"@codemirror/commands": "^6.10.3",
|
"@codemirror/commands": "^6.10.3",
|
||||||
@@ -18,6 +18,7 @@
|
|||||||
"@tiptap/pm": "^3.5.0",
|
"@tiptap/pm": "^3.5.0",
|
||||||
"@tiptap/react": "^3.5.0",
|
"@tiptap/react": "^3.5.0",
|
||||||
"@tiptap/starter-kit": "^3.5.0",
|
"@tiptap/starter-kit": "^3.5.0",
|
||||||
|
"dompurify": "^3.2.4",
|
||||||
"marked": "^18.0.4",
|
"marked": "^18.0.4",
|
||||||
"mermaid": "^11.15.0",
|
"mermaid": "^11.15.0",
|
||||||
"react": "^19.2.6",
|
"react": "^19.2.6",
|
||||||
|
|||||||
@@ -1,7 +1,7 @@
|
|||||||
{
|
{
|
||||||
"name": "rfc-app-frontend",
|
"name": "rfc-app-frontend",
|
||||||
"private": true,
|
"private": true,
|
||||||
"version": "0.23.0",
|
"version": "0.27.0",
|
||||||
"type": "module",
|
"type": "module",
|
||||||
"scripts": {
|
"scripts": {
|
||||||
"dev": "vite",
|
"dev": "vite",
|
||||||
@@ -19,6 +19,7 @@
|
|||||||
"@tiptap/pm": "^3.5.0",
|
"@tiptap/pm": "^3.5.0",
|
||||||
"@tiptap/react": "^3.5.0",
|
"@tiptap/react": "^3.5.0",
|
||||||
"@tiptap/starter-kit": "^3.5.0",
|
"@tiptap/starter-kit": "^3.5.0",
|
||||||
|
"dompurify": "^3.2.4",
|
||||||
"marked": "^18.0.4",
|
"marked": "^18.0.4",
|
||||||
"mermaid": "^11.15.0",
|
"mermaid": "^11.15.0",
|
||||||
"react": "^19.2.6",
|
"react": "^19.2.6",
|
||||||
|
|||||||
@@ -1354,6 +1354,17 @@
|
|||||||
.pr-breadcrumb a:hover { text-decoration: underline; }
|
.pr-breadcrumb a:hover { text-decoration: underline; }
|
||||||
.pr-title { font-size: var(--text-xl); margin: 0 0 6px 0; }
|
.pr-title { font-size: var(--text-xl); margin: 0 0 6px 0; }
|
||||||
.pr-description { font-size: var(--text-base); color: var(--c-gray-600); margin: 0 0 6px 0; line-height: 1.55; }
|
.pr-description { font-size: var(--text-base); color: var(--c-gray-600); margin: 0 0 6px 0; line-height: 1.55; }
|
||||||
|
/* Roadmap #28 Part 1: inline auto-links to referenced RFCs inside PR text
|
||||||
|
and comments. Subtle accent + dotted underline so they read as enriched
|
||||||
|
references, not as primary navigation. */
|
||||||
|
.rfc-autolink {
|
||||||
|
color: var(--color-link);
|
||||||
|
text-decoration: underline;
|
||||||
|
text-decoration-style: dotted;
|
||||||
|
text-underline-offset: 2px;
|
||||||
|
font-weight: 500;
|
||||||
|
}
|
||||||
|
.rfc-autolink:hover { text-decoration-style: solid; }
|
||||||
.pr-header-edit { display: flex; flex-direction: column; gap: 8px; }
|
.pr-header-edit { display: flex; flex-direction: column; gap: 8px; }
|
||||||
.pr-header-right {
|
.pr-header-right {
|
||||||
display: flex; flex-direction: column; align-items: flex-end; gap: 8px;
|
display: flex; flex-direction: column; align-items: flex-end; gap: 8px;
|
||||||
|
|||||||
@@ -202,6 +202,30 @@ export async function proposeRFC({ title, slug, pitch, tags, proposedUseCase })
|
|||||||
return jsonOrThrow(res)
|
return jsonOrThrow(res)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Roadmap #27: Claude Haiku tag suggestions for the propose-RFC modal.
|
||||||
|
// Returns a (possibly empty) array of { tag, confidence }. Deliberately
|
||||||
|
// forgiving — any non-OK response (rate limit, transient error, no key
|
||||||
|
// configured server-side) resolves to [] so the modal just shows nothing
|
||||||
|
// rather than surfacing an error for what is a best-effort assist.
|
||||||
|
export async function suggestTags({ title, pitch, useCase }) {
|
||||||
|
try {
|
||||||
|
const res = await fetch('/api/rfcs/suggest-tags', {
|
||||||
|
method: 'POST',
|
||||||
|
headers: { 'Content-Type': 'application/json' },
|
||||||
|
body: JSON.stringify({
|
||||||
|
title: title || '',
|
||||||
|
pitch: pitch || '',
|
||||||
|
use_case: useCase || '',
|
||||||
|
}),
|
||||||
|
})
|
||||||
|
if (!res.ok) return []
|
||||||
|
const data = await res.json()
|
||||||
|
return Array.isArray(data.suggestions) ? data.suggestions : []
|
||||||
|
} catch {
|
||||||
|
return []
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
export async function mergeProposal(prNumber) {
|
export async function mergeProposal(prNumber) {
|
||||||
const res = await fetch(`/api/proposals/${prNumber}/merge`, { method: 'POST' })
|
const res = await fetch(`/api/proposals/${prNumber}/merge`, { method: 'POST' })
|
||||||
return jsonOrThrow(res)
|
return jsonOrThrow(res)
|
||||||
|
|||||||
@@ -17,7 +17,7 @@
|
|||||||
import { useEditor, EditorContent, Extension } from '@tiptap/react'
|
import { useEditor, EditorContent, Extension } from '@tiptap/react'
|
||||||
import StarterKit from '@tiptap/starter-kit'
|
import StarterKit from '@tiptap/starter-kit'
|
||||||
import { useEffect, useRef, useCallback } from 'react'
|
import { useEffect, useRef, useCallback } from 'react'
|
||||||
import { marked } from 'marked'
|
import { renderMarkdown } from '../lib/sanitizeHtml'
|
||||||
import { Plugin, PluginKey } from 'prosemirror-state'
|
import { Plugin, PluginKey } from 'prosemirror-state'
|
||||||
import { Decoration, DecorationSet } from 'prosemirror-view'
|
import { Decoration, DecorationSet } from 'prosemirror-view'
|
||||||
|
|
||||||
@@ -122,7 +122,7 @@ export default function Editor({
|
|||||||
|
|
||||||
useEffect(() => {
|
useEffect(() => {
|
||||||
if (!editor || content == null) return
|
if (!editor || content == null) return
|
||||||
const html = marked.parse(content)
|
const html = renderMarkdown(content)
|
||||||
editor.commands.setContent(html, false)
|
editor.commands.setContent(html, false)
|
||||||
}, [content, editor])
|
}, [content, editor])
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,38 @@
|
|||||||
|
// LinkedText.jsx — roadmap #28 Part 1.
|
||||||
|
//
|
||||||
|
// Renders a backend-provided list of text/rfc-link segments (see
|
||||||
|
// backend/app/rfc_links.py). RFC references in PR descriptions and
|
||||||
|
// comments arrive pre-scanned as structured segments — this component
|
||||||
|
// maps them onto plain text runs and anchor elements. It never renders
|
||||||
|
// HTML from the server (no dangerouslySetInnerHTML), so the surface is
|
||||||
|
// XSS-safe regardless of what a comment author typed.
|
||||||
|
//
|
||||||
|
// `segments` is the enriched array; `text` is the raw fallback used when
|
||||||
|
// the field is absent (an older cached response, or a caller that didn't
|
||||||
|
// pass segments). Either way the visible text is identical — only the
|
||||||
|
// links differ.
|
||||||
|
|
||||||
|
export default function LinkedText({ segments, text }) {
|
||||||
|
if (!Array.isArray(segments) || segments.length === 0) {
|
||||||
|
return <>{text ?? ''}</>
|
||||||
|
}
|
||||||
|
return (
|
||||||
|
<>
|
||||||
|
{segments.map((seg, i) => {
|
||||||
|
if (seg.type === 'rfc') {
|
||||||
|
return (
|
||||||
|
<a
|
||||||
|
key={i}
|
||||||
|
className="rfc-autolink"
|
||||||
|
href={`/rfc/${seg.slug}`}
|
||||||
|
title={seg.title ? `RFC: ${seg.title}` : undefined}
|
||||||
|
>
|
||||||
|
{seg.label}
|
||||||
|
</a>
|
||||||
|
)
|
||||||
|
}
|
||||||
|
return <span key={i}>{seg.text}</span>
|
||||||
|
})}
|
||||||
|
</>
|
||||||
|
)
|
||||||
|
}
|
||||||
@@ -21,6 +21,7 @@
|
|||||||
|
|
||||||
import { useEffect, useRef, useState, useCallback } from 'react'
|
import { useEffect, useRef, useState, useCallback } from 'react'
|
||||||
import { Marked } from 'marked'
|
import { Marked } from 'marked'
|
||||||
|
import { sanitizeHtml } from '../lib/sanitizeHtml'
|
||||||
import { decorateAcceptedChanges } from './trackedOverlay.js'
|
import { decorateAcceptedChanges } from './trackedOverlay.js'
|
||||||
import ChangeTooltip from './ChangeTooltip.jsx'
|
import ChangeTooltip from './ChangeTooltip.jsx'
|
||||||
|
|
||||||
@@ -98,7 +99,7 @@ export default function MarkdownPreview({
|
|||||||
// synchronously with the body itself — no flash of un-decorated text.
|
// synchronously with the body itself — no flash of un-decorated text.
|
||||||
useEffect(() => {
|
useEffect(() => {
|
||||||
if (!hostRef.current) return
|
if (!hostRef.current) return
|
||||||
const html = previewMarked.parse(content || '')
|
const html = sanitizeHtml(previewMarked.parse(content || ''))
|
||||||
hostRef.current.innerHTML = html
|
hostRef.current.innerHTML = html
|
||||||
const token = ++renderTokenRef.current
|
const token = ++renderTokenRef.current
|
||||||
// Reset memo so the new block set re-renders from scratch.
|
// Reset memo so the new block set re-renders from scratch.
|
||||||
|
|||||||
@@ -23,6 +23,7 @@ import {
|
|||||||
withdrawPR,
|
withdrawPR,
|
||||||
} from '../api'
|
} from '../api'
|
||||||
import { EVENTS, track } from '../lib/analytics'
|
import { EVENTS, track } from '../lib/analytics'
|
||||||
|
import LinkedText from './LinkedText'
|
||||||
|
|
||||||
export default function PRView({ viewer }) {
|
export default function PRView({ viewer }) {
|
||||||
const { slug, prNumber: prNumberParam } = useParams()
|
const { slug, prNumber: prNumberParam } = useParams()
|
||||||
@@ -218,7 +219,9 @@ export default function PRView({ viewer }) {
|
|||||||
<>
|
<>
|
||||||
<h1 className="pr-title">{pr.title}</h1>
|
<h1 className="pr-title">{pr.title}</h1>
|
||||||
{pr.description && (
|
{pr.description && (
|
||||||
<p className="pr-description">{pr.description}</p>
|
<p className="pr-description">
|
||||||
|
<LinkedText segments={pr.description_segments} text={pr.description} />
|
||||||
|
</p>
|
||||||
)}
|
)}
|
||||||
{/* #26: the optional ground-truth use case for this change,
|
{/* #26: the optional ground-truth use case for this change,
|
||||||
captured when the PR was opened. Muted "left blank"
|
captured when the PR was opened. Muted "left blank"
|
||||||
@@ -440,7 +443,9 @@ function PRConversation({ threads, messagesByThread, threadsByKind, seenMsgId })
|
|||||||
{isNew && <span className="chat-msg-new-pip" title="New since your last visit">●</span>}
|
{isNew && <span className="chat-msg-new-pip" title="New since your last visit">●</span>}
|
||||||
</div>
|
</div>
|
||||||
{m.quote && <pre className="chat-msg-quote">{m.quote}</pre>}
|
{m.quote && <pre className="chat-msg-quote">{m.quote}</pre>}
|
||||||
<div className="chat-msg-body">{m.text}</div>
|
<div className="chat-msg-body">
|
||||||
|
<LinkedText segments={m.text_segments} text={m.text} />
|
||||||
|
</div>
|
||||||
</li>
|
</li>
|
||||||
)
|
)
|
||||||
})}
|
})}
|
||||||
|
|||||||
@@ -10,7 +10,7 @@
|
|||||||
|
|
||||||
import { useEffect, useState } from 'react'
|
import { useEffect, useState } from 'react'
|
||||||
import { useParams, useNavigate } from 'react-router-dom'
|
import { useParams, useNavigate } from 'react-router-dom'
|
||||||
import { marked } from 'marked'
|
import { renderMarkdown } from '../lib/sanitizeHtml'
|
||||||
import { getProposal, mergeProposal, declineProposal, withdrawProposal } from '../api'
|
import { getProposal, mergeProposal, declineProposal, withdrawProposal } from '../api'
|
||||||
|
|
||||||
export default function ProposalView({ viewer, onChange }) {
|
export default function ProposalView({ viewer, onChange }) {
|
||||||
@@ -161,7 +161,7 @@ export default function ProposalView({ viewer, onChange }) {
|
|||||||
</h3>
|
</h3>
|
||||||
<div
|
<div
|
||||||
className="entry-body"
|
className="entry-body"
|
||||||
dangerouslySetInnerHTML={{ __html: marked.parse(data.entry?.body || '') }}
|
dangerouslySetInnerHTML={{ __html: renderMarkdown(data.entry?.body || '') }}
|
||||||
/>
|
/>
|
||||||
|
|
||||||
{/* #26: the optional ground-truth use case the proposer supplied. */}
|
{/* #26: the optional ground-truth use case the proposer supplied. */}
|
||||||
@@ -169,7 +169,7 @@ export default function ProposalView({ viewer, onChange }) {
|
|||||||
Intended use case
|
Intended use case
|
||||||
</h3>
|
</h3>
|
||||||
{data.proposed_use_case
|
{data.proposed_use_case
|
||||||
? <div className="entry-body" dangerouslySetInnerHTML={{ __html: marked.parse(data.proposed_use_case) }} />
|
? <div className="entry-body" dangerouslySetInnerHTML={{ __html: renderMarkdown(data.proposed_use_case) }} />
|
||||||
: <p style={{ color: '#999', fontStyle: 'italic' }}>Left blank by the proposer.</p>}
|
: <p style={{ color: '#999', fontStyle: 'italic' }}>Left blank by the proposer.</p>}
|
||||||
</article>
|
</article>
|
||||||
)
|
)
|
||||||
|
|||||||
@@ -2,17 +2,24 @@
|
|||||||
//
|
//
|
||||||
// Title (required) and pitch (required textarea), with a slug field
|
// Title (required) and pitch (required textarea), with a slug field
|
||||||
// that auto-fills from the title via the same deterministic kebab-case
|
// that auto-fills from the title via the same deterministic kebab-case
|
||||||
// the backend uses. Tags are chip-input (free-form for slice 1; the
|
// the backend uses. Tags are chip-input (free-form), with the §9.1
|
||||||
// AI-suggested chips of §9.1 are deferred to Slice 2 when the AI surface
|
// Slice 2 AI-suggested chips (roadmap #27) wired in: as the draft fills
|
||||||
// is wired up).
|
// in, the backend asks Claude Haiku for tags drawn from the corpus's
|
||||||
|
// existing tag set, surfaced as clickable suggestion chips. The assist
|
||||||
|
// is best-effort — it stays silent when unavailable.
|
||||||
//
|
//
|
||||||
// The submit button drives the §17 POST /api/rfcs/propose endpoint;
|
// The submit button drives the §17 POST /api/rfcs/propose endpoint;
|
||||||
// success navigates the proposer to the pending-idea view per §9.3.
|
// success navigates the proposer to the pending-idea view per §9.3.
|
||||||
|
|
||||||
import { useEffect, useState } from 'react'
|
import { useEffect, useRef, useState } from 'react'
|
||||||
import { proposeRFC } from '../api'
|
import { proposeRFC, suggestTags } from '../api'
|
||||||
import { EVENTS, track } from '../lib/analytics'
|
import { EVENTS, track } from '../lib/analytics'
|
||||||
|
|
||||||
|
// How long the draft must sit unchanged before we ask for suggestions —
|
||||||
|
// long enough to fire on typing pauses / field-blur, not on every
|
||||||
|
// keystroke (the backend is also per-user rate-limited as a backstop).
|
||||||
|
const SUGGEST_DEBOUNCE_MS = 700
|
||||||
|
|
||||||
function slugify(title) {
|
function slugify(title) {
|
||||||
return title
|
return title
|
||||||
.toLowerCase()
|
.toLowerCase()
|
||||||
@@ -32,17 +39,50 @@ export default function ProposeModal({ viewer, onClose, onSubmitted }) {
|
|||||||
const [tags, setTags] = useState([])
|
const [tags, setTags] = useState([])
|
||||||
const [submitting, setSubmitting] = useState(false)
|
const [submitting, setSubmitting] = useState(false)
|
||||||
const [error, setError] = useState(null)
|
const [error, setError] = useState(null)
|
||||||
|
// #27: Claude Haiku tag suggestions. `suggestions` is the latest
|
||||||
|
// ranked list from the backend ({ tag, confidence }); we render the
|
||||||
|
// subset not already chosen. `suggestedOnce` gates the disclosure +
|
||||||
|
// row so they only appear after the assist has actually run.
|
||||||
|
const [suggestions, setSuggestions] = useState([])
|
||||||
|
const [suggestedOnce, setSuggestedOnce] = useState(false)
|
||||||
|
|
||||||
useEffect(() => {
|
useEffect(() => {
|
||||||
if (!slugEdited) setSlug(slugify(title))
|
if (!slugEdited) setSlug(slugify(title))
|
||||||
}, [title, slugEdited])
|
}, [title, slugEdited])
|
||||||
|
|
||||||
|
// #27: debounced tag-suggestion fetch. Fires after the draft sits
|
||||||
|
// unchanged for SUGGEST_DEBOUNCE_MS, only once there's something to go
|
||||||
|
// on (a title). A stale-response guard keeps an earlier in-flight
|
||||||
|
// request from clobbering a newer one.
|
||||||
|
const suggestSeq = useRef(0)
|
||||||
|
useEffect(() => {
|
||||||
|
if (!title.trim()) {
|
||||||
|
setSuggestions([])
|
||||||
|
return
|
||||||
|
}
|
||||||
|
const handle = setTimeout(async () => {
|
||||||
|
const seq = ++suggestSeq.current
|
||||||
|
const result = await suggestTags({ title, pitch, useCase })
|
||||||
|
if (seq !== suggestSeq.current) return // a newer request superseded us
|
||||||
|
setSuggestions(Array.isArray(result) ? result : [])
|
||||||
|
if (result && result.length) setSuggestedOnce(true)
|
||||||
|
}, SUGGEST_DEBOUNCE_MS)
|
||||||
|
return () => clearTimeout(handle)
|
||||||
|
}, [title, pitch, useCase])
|
||||||
|
|
||||||
|
// Suggestions the user hasn't already added.
|
||||||
|
const freshSuggestions = suggestions.filter(s => !tags.includes(s.tag))
|
||||||
|
|
||||||
function addTag() {
|
function addTag() {
|
||||||
const t = tagInput.trim()
|
const t = tagInput.trim()
|
||||||
if (t && !tags.includes(t)) setTags([...tags, t])
|
if (t && !tags.includes(t)) setTags([...tags, t])
|
||||||
setTagInput('')
|
setTagInput('')
|
||||||
}
|
}
|
||||||
|
|
||||||
|
function addSuggested(tag) {
|
||||||
|
if (!tags.includes(tag)) setTags([...tags, tag])
|
||||||
|
}
|
||||||
|
|
||||||
async function handleSubmit(e) {
|
async function handleSubmit(e) {
|
||||||
e.preventDefault()
|
e.preventDefault()
|
||||||
if (!title.trim() || !slug || !pitch.trim()) return
|
if (!title.trim() || !slug || !pitch.trim()) return
|
||||||
@@ -152,6 +192,45 @@ export default function ProposeModal({ viewer, onClose, onSubmitted }) {
|
|||||||
</div>
|
</div>
|
||||||
)}
|
)}
|
||||||
|
|
||||||
|
{/* #27: Claude Haiku suggested tags. Clickable chips that add
|
||||||
|
to the tag list; nothing auto-applies. The disclosure is
|
||||||
|
required — the draft text is sent to Anthropic to generate
|
||||||
|
these — so it renders whenever the suggestion row does. */}
|
||||||
|
{(freshSuggestions.length > 0 || (suggestedOnce && suggestions.length > 0)) && (
|
||||||
|
<div style={{ marginTop: 4, marginBottom: 14 }}>
|
||||||
|
{freshSuggestions.length > 0 && (
|
||||||
|
<>
|
||||||
|
<p className="field-help" style={{ marginTop: 0, marginBottom: 4 }}>
|
||||||
|
Suggested tags — click to add:
|
||||||
|
</p>
|
||||||
|
<div>
|
||||||
|
{freshSuggestions.map(s => (
|
||||||
|
<button
|
||||||
|
key={s.tag}
|
||||||
|
type="button"
|
||||||
|
className="entry-tag"
|
||||||
|
onClick={() => addSuggested(s.tag)}
|
||||||
|
title={`Add "${s.tag}"`}
|
||||||
|
style={{
|
||||||
|
display: 'inline-block',
|
||||||
|
marginRight: 4,
|
||||||
|
marginBottom: 4,
|
||||||
|
border: '1px dashed var(--color-border, #ccc)',
|
||||||
|
background: 'none',
|
||||||
|
cursor: 'pointer',
|
||||||
|
}}
|
||||||
|
>+ {s.tag}</button>
|
||||||
|
))}
|
||||||
|
</div>
|
||||||
|
</>
|
||||||
|
)}
|
||||||
|
<p className="field-help" style={{ marginTop: 4, marginBottom: 0, fontStyle: 'italic' }}>
|
||||||
|
Suggestions are generated by Claude (Anthropic). The text you've
|
||||||
|
entered above is sent to Anthropic to produce them.
|
||||||
|
</p>
|
||||||
|
</div>
|
||||||
|
)}
|
||||||
|
|
||||||
{viewer && (
|
{viewer && (
|
||||||
<p className="field-help" style={{ marginTop: 14, marginBottom: 0 }}>
|
<p className="field-help" style={{ marginTop: 14, marginBottom: 0 }}>
|
||||||
Owner: <strong>{viewer.display_name || viewer.gitea_login}</strong> — you'll be the first owner of this super-draft. Additional owners can claim later (§13.1).
|
Owner: <strong>{viewer.display_name || viewer.gitea_login}</strong> — you'll be the first owner of this super-draft. Additional owners can claim later (§13.1).
|
||||||
|
|||||||
@@ -19,6 +19,7 @@ import {
|
|||||||
resolveDiscussionThread,
|
resolveDiscussionThread,
|
||||||
} from '../api'
|
} from '../api'
|
||||||
import { EVENTS, track } from '../lib/analytics'
|
import { EVENTS, track } from '../lib/analytics'
|
||||||
|
import LinkedText from './LinkedText'
|
||||||
|
|
||||||
export default function RFCDiscussionPanel({ slug, viewer }) {
|
export default function RFCDiscussionPanel({ slug, viewer }) {
|
||||||
const [threads, setThreads] = useState([])
|
const [threads, setThreads] = useState([])
|
||||||
@@ -279,7 +280,9 @@ function DiscussionMessage({ message }) {
|
|||||||
{message.quote && (
|
{message.quote && (
|
||||||
<div className="discussion-message-quote">"{message.quote}"</div>
|
<div className="discussion-message-quote">"{message.quote}"</div>
|
||||||
)}
|
)}
|
||||||
<div className="discussion-message-body">{message.text}</div>
|
<div className="discussion-message-body">
|
||||||
|
<LinkedText segments={message.text_segments} text={message.text} />
|
||||||
|
</div>
|
||||||
</div>
|
</div>
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,47 @@
|
|||||||
|
// sanitizeHtml.js — the single chokepoint for turning user-authored
|
||||||
|
// markdown into DOM-bound HTML.
|
||||||
|
//
|
||||||
|
// Security audit 0026 (finding C1, Critical): every `marked.parse(...)`
|
||||||
|
// result that reaches an `innerHTML` / `dangerouslySetInnerHTML` sink was
|
||||||
|
// previously written raw. `marked` passes through embedded HTML and
|
||||||
|
// `javascript:`/event-handler attributes verbatim, so any user-authored
|
||||||
|
// document (RFC body, proposal body, proposed_use_case, transcript) was a
|
||||||
|
// stored-XSS vector — a contributor's payload executed in the session of
|
||||||
|
// whoever viewed it, including an admin/owner during review.
|
||||||
|
//
|
||||||
|
// Fix: route EVERY markdown render through `renderMarkdown` (or, for
|
||||||
|
// already-rendered HTML, `sanitizeHtml`). DOMPurify's defaults already
|
||||||
|
// strip <script>, on* event handlers, and javascript:/unsafe-data: URIs;
|
||||||
|
// we add a hook so any link opening a new tab carries rel="noopener
|
||||||
|
// noreferrer". The html profile keeps the standard markdown tag set plus
|
||||||
|
// class + data-* attributes (the latter is what MarkdownPreview's mermaid
|
||||||
|
// placeholder relies on); mermaid renders its SVG into the DOM *after*
|
||||||
|
// sanitization and is itself locked down with securityLevel:'strict'.
|
||||||
|
|
||||||
|
import DOMPurify from 'dompurify'
|
||||||
|
import { marked } from 'marked'
|
||||||
|
|
||||||
|
let _hookInstalled = false
|
||||||
|
function ensureHook() {
|
||||||
|
if (_hookInstalled) return
|
||||||
|
DOMPurify.addHook('afterSanitizeAttributes', (node) => {
|
||||||
|
if (node.tagName === 'A' && node.getAttribute('target') === '_blank') {
|
||||||
|
node.setAttribute('rel', 'noopener noreferrer')
|
||||||
|
}
|
||||||
|
})
|
||||||
|
_hookInstalled = true
|
||||||
|
}
|
||||||
|
|
||||||
|
// Sanitize an already-rendered HTML string. Use when the HTML did not come
|
||||||
|
// from `marked` (rare) or when a caller parses markdown with a bespoke
|
||||||
|
// `Marked` instance and only needs the sanitize step.
|
||||||
|
export function sanitizeHtml(html) {
|
||||||
|
ensureHook()
|
||||||
|
return DOMPurify.sanitize(html || '', { USE_PROFILES: { html: true } })
|
||||||
|
}
|
||||||
|
|
||||||
|
// Parse markdown with the shared `marked` and sanitize the result. This is
|
||||||
|
// the drop-in replacement for `marked.parse(src)` at any HTML sink.
|
||||||
|
export function renderMarkdown(src) {
|
||||||
|
return sanitizeHtml(marked.parse(src || ''))
|
||||||
|
}
|
||||||
Reference in New Issue
Block a user