feedback: thumb up/down on logs + reviewer self-score 0-10

Two unrelated features bundled because they ship together and share
migration 0028.

Strategic-log feedback (thumb up/down):
- New strategic_log_feedback table with UNIQUE(log_id, user_id) so each
  user has one vote per log, flippable in place (up -> down -> clear).
  UI shows aggregate counts only.
- app/services/log_feedback.py: set_vote, get_counts, sign/verify
  feedback tokens (same itsdangerous pattern as auth.sign_pending,
  30-day TTL for email links).
- POST /api/log/{id}/feedback: web vote, auth required, returns counts
  + the requesting user's own vote.
- GET /feedback?token=...&vote=...: email-link target, no auth, signed
  token encodes (user, log, vote), renders feedback_thanks.html.
- partials/log.html: thumbs row below content, JS-driven swap via the
  POST endpoint. Dashboard latest-log card and /log page both render
  this partial via htmx, so the buttons appear in all three surfaces.
- digest emails: a "How was today's read?" row above the unsub footer,
  with signed-token URLs against the latest StrategicLog at send time.
  Plain-text fallback included.

Reviewer self-score (0-10):
- _SYSTEM_PROMPT asks for an integer score with anchors (10 exemplary,
  5 borderline, 0 unfit). Verdict gains score: int | None.
- Deterministic-layer hits get score=0 (hard rule, no nuance);
  error rows get None; LLM rows get the model's score clamped 0..10.
- ReviewerVerdict.score, StrategicLog.reviewer_score, and
  IndicatorSummary.reviewer_score all new SMALLINT NULL columns.
- ai_log_job + indicator_summary_job persist verdict.score onto their
  content rows when committing the row alongside content.

Tests:
- tests/test_strategic_log_feedback.py: vote, flip, clear, aggregate
  across users, invalid vote, token round-trip + tamper + garbage +
  'clear' not signable for email path.
- tests/test_output_review.py: score parsing, clamping (>10, <0),
  missing/non-numeric -> None, deterministic-layer score=0.

Full suite: 427 passed (was 412), 5 skipped, no regressions.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
Giorgio Gilestro 2026-05-29 21:28:03 +02:00
parent f3ac65f8f7
commit 8946dee2e0
14 changed files with 962 additions and 14 deletions

View file

@ -123,8 +123,19 @@ Mark UNCLEAN if the text contains ANY of:
claim on a *named* instrument is not.
- Anything else other than the finished, publishable commentary.
Also assign a SCORE 0-10 to the candidate:
- 10 = exemplary editorial: sharp, well-grounded, no perimeter risk, clean prose.
- 7-9 = publishable as-is, varying degrees of polish.
- 4-6 = borderline: scratchpad leakage, mild perimeter drift, or weak structure,
but not yet outright unfit. (Anything 4 should usually be clean=false.)
- 1-3 = unfit: clear chain-of-thought, partial / truncated, or financial-advice drift.
- 0 = unfit by hard rule (deterministic catch territory).
Clean=true implies a score of ~7+; clean=false implies ~4 or lower. Use the
score to communicate confidence within the verdict.
Return ONLY a JSON object with this exact shape:
{"clean": true | false, "reason": "<≤20 words, plain text>"}
{"clean": true | false, "reason": "<≤20 words, plain text>", "score": 0-10}
No preamble, no markdown fences, no other fields.
"""
@ -171,6 +182,12 @@ class Verdict:
reason: str
cost_usd: float | None # cost of the review call itself, for the ledger
layer: str = "llm" # "deterministic" | "llm" | "error"
# Integer 0-10. None for error rows; 0 for deterministic-layer hits
# (rejected by hard rule, no nuance to score); 0-10 from the model
# on LLM-layer verdicts. Stored alongside the content row for
# future analysis — see StrategicLog.reviewer_score and
# IndicatorSummary.reviewer_score.
score: int | None = None
# Truncation cap for the audit log's candidate_text column. Generous enough
@ -199,6 +216,7 @@ async def _record_verdict(
reason=verdict.reason[:240] if verdict.reason else None,
layer=verdict.layer,
model=model,
score=verdict.score,
)
session.add(row)
await session.flush()
@ -237,7 +255,7 @@ async def review_read(
if not candidate or not candidate.strip():
verdict = Verdict(clean=False, reason="empty candidate", cost_usd=0.0,
layer="deterministic")
layer="deterministic", score=0)
await _record_verdict(session, surface=surface, candidate=candidate or "",
verdict=verdict, model=None)
return verdict
@ -250,6 +268,7 @@ async def review_read(
reason=f"lexicon:{hit.rule}: {hit.snippet}",
cost_usd=0.0,
layer="deterministic",
score=0,
)
log.info("review.deterministic_reject",
rule=hit.rule, snippet=hit.snippet, surface=surface)
@ -293,7 +312,7 @@ async def review_read(
except Exception as e:
log.warning("review.call_failed", error=str(e)[:200])
verdict = Verdict(clean=False, reason=f"reviewer error: {str(e)[:80]}",
cost_usd=None, layer="error")
cost_usd=None, layer="error", score=None)
await _record_verdict(session, surface=surface, candidate=candidate,
verdict=verdict, model=reviewer_model)
return verdict
@ -317,7 +336,7 @@ async def review_read(
except json.JSONDecodeError:
log.warning("review.parse_failed", preview=result.content[:200])
verdict = Verdict(clean=False, reason="reviewer returned non-JSON",
cost_usd=result.cost_usd, layer="error")
cost_usd=result.cost_usd, layer="error", score=None)
await _record_verdict(session, surface=surface, candidate=candidate,
verdict=verdict, model=reviewer_model)
return verdict
@ -326,13 +345,25 @@ async def review_read(
reason = parsed.get("reason") or ""
if not isinstance(clean, bool):
verdict = Verdict(clean=False, reason="reviewer omitted bool 'clean'",
cost_usd=result.cost_usd, layer="error")
cost_usd=result.cost_usd, layer="error", score=None)
await _record_verdict(session, surface=surface, candidate=candidate,
verdict=verdict, model=reviewer_model)
return verdict
# Score is optional and bounded; the verdict is still valid without it.
raw_score = parsed.get("score")
score: int | None
if isinstance(raw_score, bool):
# bool is a subclass of int — exclude it explicitly to avoid
# silently treating True/False as 1/0.
score = None
elif isinstance(raw_score, (int, float)):
score = max(0, min(10, int(raw_score)))
else:
score = None
verdict = Verdict(clean=clean, reason=str(reason)[:200],
cost_usd=result.cost_usd, layer="llm")
cost_usd=result.cost_usd, layer="llm", score=score)
await _record_verdict(session, surface=surface, candidate=candidate,
verdict=verdict, model=reviewer_model)
return verdict
@ -391,7 +422,7 @@ async def generate_with_review(
content=None,
verdict=Verdict(clean=False,
reason=f"generator error: {str(e)[:80]}",
cost_usd=None, layer="error"),
cost_usd=None, layer="error", score=None),
attempts=attempt,
)
@ -411,6 +442,7 @@ async def generate_with_review(
return ReviewedGeneration(
content=None,
verdict=last_verdict or Verdict(clean=False, reason="no attempts",
cost_usd=None, layer="error"),
cost_usd=None, layer="error",
score=None),
attempts=max_attempts,
)