updating docs to match the implemented phase 3 pipeline

The pipeline docs still described pipeline.ts as pseudocode and the
validation module as unwritten. Both have been implemented and run.

- CLAUDE.md: replace the "no executable pipeline yet" description with
  the actual module flow, plus the two invariants worth preserving
  (resumability via headword diffing, raw responses saved before parsing)
- DATA_PIPELINE.md: mark the seven implemented modules, add a module
  responsibility map and the CLI flag table, drop the resolved warning
  about hardcoded prompt values
- roadmap.md: check off phase 3 tasks 3.2/3.3/3.4, record where the
  build diverged from the plan, note that validate.ts is stricter than
  its own spec
- STATUS.md: phase 3 is data work now, not code work

Also records two open issues: the systemic difficulty-ordering rejection
cause, and the hard-tier shortfall pending a full run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
lila 2026-08-20 12:10:48 +02:00
parent 37c978e230
commit da9cdbfa1b
4 changed files with 219 additions and 49 deletions

View file

@ -1,6 +1,6 @@
# Status — 2026-08-09
# Status — 2026-08-20
> Last updated: 2026-08-09. Update this file after every deploy or when switching tasks.
> Last updated: 2026-08-20. Update this file after every deploy or when switching tasks.
## What Works Today ✅
@ -12,7 +12,7 @@
## What's Broken / Blocked 🚧
- **Data quality** — Production still uses OpenWordNet/OMW translations. The replacement is the Gemini-only pipeline on branch `refactor/gemini-only-pipeline`: the sense-based schema (`words` → `senses` → `translations`) is migrated but empty, and the pipeline itself is not built yet. The earlier Kaikki/local-LLM pipeline was abandoned and removed (docs in `documentation/archive/`).
- **Data quality** — Production still uses OpenWordNet/OMW translations. The replacement is the Gemini-only pipeline on branch `refactor/gemini-only-pipeline`: the sense-based schema (`words` → `senses` → `translations`) is migrated but still empty in Postgres. The pipeline is **built and running**, staging into SQLite; the SQLite → Postgres import script (Phase 4) is what's missing. The earlier Kaikki/local-LLM pipeline was abandoned and removed (docs in `documentation/archive/`).
- **Guest play** — Auth is required for all game routes. No try-before-signup flow.
- **Game session store** — Still in-memory (`InMemoryGameSessionStore`). Valkey container exists in local dev but not wired up.
- **Rate limiting** — Partially implemented on auth endpoints; game endpoints not yet covered.
@ -21,7 +21,12 @@
## What I'm Working On Now 🔄
**Primary:** Phase 3 of the Gemini-only data pipeline (see [pipeline/roadmap.md](pipeline/roadmap.md)): template the Gemini prompt, write the validation module + `pipeline.ts`, and run the first full staging pass (~1,600–1,800 deduped nouns × 5 languages into SQLite).
**Primary:** Phase 3 of the Gemini-only data pipeline (see [pipeline/roadmap.md](pipeline/roadmap.md)). The pipeline code is complete and unit-tested — prompt templating, validation, Gemini structured output, and SQLite staging all work, and runs are resumable. Remaining Phase 3 work is data, not code:
1. Finish the full staging pass (~1,550–1,750 deduped nouns × 5 languages), German in progress.
2. Fix the one systemic rejection cause — sense difficulty tagged higher than its own easiest translation. It's derivable, so `validate.ts` should normalize rather than reject.
3. Fix the difficulty skew: `hard` translations are under 1% of staged rows, which leaves the "hard" game mode with too few rows to fill a round. Prompt calibration, needed before Phase 4 import.
4. Spot-check 50 entries by hand.
**Secondary:** Phase 7 hardening backlog items (see BACKLOG.md `next` section).