lila/documentation/STATUS.md
lila da9cdbfa1b updating docs to match the implemented phase 3 pipeline
The pipeline docs still described pipeline.ts as pseudocode and the
validation module as unwritten. Both have been implemented and run.

- CLAUDE.md: replace the "no executable pipeline yet" description with
  the actual module flow, plus the two invariants worth preserving
  (resumability via headword diffing, raw responses saved before parsing)
- DATA_PIPELINE.md: mark the seven implemented modules, add a module
  responsibility map and the CLI flag table, drop the resolved warning
  about hardcoded prompt values
- roadmap.md: check off phase 3 tasks 3.2/3.3/3.4, record where the
  build diverged from the plan, note that validate.ts is stricter than
  its own spec
- STATUS.md: phase 3 is data work now, not code work

Also records two open issues: the systemic difficulty-ordering rejection
cause, and the hard-tier shortfall pending a full run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 12:10:48 +02:00

3.6 KiB
Raw Blame History

Status — 2026-08-20

Last updated: 2026-08-20. Update this file after every deploy or when switching tasks.

What Works Today ✅

  • Singleplayer quiz — Duolingo-style, 5 language pairs (en↔it/de/es/fr), 3 or 10 rounds, POS + difficulty filters
  • Multiplayer — Create/join lobby by room code, 2–4 players, simultaneous answers, 15s server timer, live scoring, winner screen
  • Auth — Google + GitHub via Better Auth, cross-subdomain cookies, session middleware on protected routes
  • Deployment — Live at lilastudy.com, Hetzner VPS, Caddy HTTPS, Docker Compose, CI/CD via Forgejo Actions
  • Database — PostgreSQL with Drizzle ORM, daily backups, idempotent seeding

What's Broken / Blocked 🚧

  • Data quality — Production still uses OpenWordNet/OMW translations. The replacement is the Gemini-only pipeline on branch refactor/gemini-only-pipeline: the sense-based schema (words → senses → translations) is migrated but still empty in Postgres. The pipeline is built and running, staging into SQLite; the SQLite → Postgres import script (Phase 4) is what's missing. The earlier Kaikki/local-LLM pipeline was abandoned and removed (docs in documentation/archive/).
  • Guest play — Auth is required for all game routes. No try-before-signup flow.
  • Game session store — Still in-memory (InMemoryGameSessionStore). Valkey container exists in local dev but not wired up.
  • Rate limiting — Partially implemented on auth endpoints; game endpoints not yet covered.
  • React error boundaries — Not implemented; runtime crashes take down the whole app.
  • Monitoring — No uptime alerts or centralized logging on the VPS.

What I'm Working On Now 🔄

Primary: Phase 3 of the Gemini-only data pipeline (see pipeline/roadmap.md). The pipeline code is complete and unit-tested — prompt templating, validation, Gemini structured output, and SQLite staging all work, and runs are resumable. Remaining Phase 3 work is data, not code:

  1. Finish the full staging pass (~1,550–1,750 deduped nouns × 5 languages), German in progress.
  2. Fix the one systemic rejection cause — sense difficulty tagged higher than its own easiest translation. It's derivable, so validate.ts should normalize rather than reject.
  3. Fix the difficulty skew: hard translations are under 1% of staged rows, which leaves the "hard" game mode with too few rows to fill a round. Prompt calibration, needed before Phase 4 import.
  4. Spot-check 50 entries by hand.

Secondary: Phase 7 hardening backlog items (see BACKLOG.md next section).

Next 2-Week Goal 🎯

Finish pipeline Phase 3 (full noun run staged in SQLite, reject rate < 10%, 50 entries spot-checked) → Phase 4 import script (pipeline Postgres :5433, then dev :5432) → start rewriting getGameTerms/getDistractors against the new schema.

The Big Picture

Lila is a deployed, working vocabulary quiz app. The core loop (singleplayer + multiplayer) is solid. The next strategic milestone is media-based practice (learn vocab from a song/TV episode/book chapter), but that depends on:

  1. The Gemini data pipeline reaching production (fixes translation quality)
  2. A media ingestion prototype (subtitles/lyrics → text → vocab extraction → quiz)

Until then, the app is a generic vocabulary quiz — functional but not differentiated.