updating docs to match the implemented phase 3 pipeline

The pipeline docs still described pipeline.ts as pseudocode and the
validation module as unwritten. Both have been implemented and run.

- CLAUDE.md: replace the "no executable pipeline yet" description with
  the actual module flow, plus the two invariants worth preserving
  (resumability via headword diffing, raw responses saved before parsing)
- DATA_PIPELINE.md: mark the seven implemented modules, add a module
  responsibility map and the CLI flag table, drop the resolved warning
  about hardcoded prompt values
- roadmap.md: check off phase 3 tasks 3.2/3.3/3.4, record where the
  build diverged from the plan, note that validate.ts is stricter than
  its own spec
- STATUS.md: phase 3 is data work now, not code work

Also records two open issues: the systemic difficulty-ordering rejection
cause, and the hard-tier shortfall pending a full run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
lila 2026-08-20 12:10:48 +02:00
parent 37c978e230
commit da9cdbfa1b
4 changed files with 219 additions and 49 deletions

View file

@ -1,7 +1,7 @@
# Lila Data Pipeline
> How vocabulary data is generated and gets into PostgreSQL.
> Last updated: 2026-08-01 · Branch: `refactor/gemini-only-pipeline`
> Last updated: 2026-08-20 · Branch: `refactor/gemini-only-pipeline`
**Authoritative detail lives in two companion docs:**
@ -57,21 +57,55 @@ The app always reads from PostgreSQL. SQLite exists purely as a staging file so
## What exists on disk today
| Path | State |
| ---------------------------------------- | ------------------------------------------------------------------------- |
| `data-pipeline/source-data/{lang}/{pos}` | ✅ Noun lists for `de`, `en`, `es`, `fr`, `it` |
| `data-pipeline/prompt` | ✅ The Gemini system prompt (plain UTF-8 text, not a module) |
| `data-pipeline/db/schema.sql` | ✅ SQLite staging schema |
| `data-pipeline/db/staging.db` | ✅ Created, tables present, **0 rows** — gitignored |
| `data-pipeline/pipeline.ts` | 🚧 Design pseudocode in comments. No executable pipeline code yet. |
| Validation module | ❌ Not written (rules specced in design-doc §6.4) |
| SQLite → PostgreSQL import script | ❌ Not written |
| `data-pipeline/kaikki-source-files/` | ⚠️ Leftover JSONL dumps from the old pipeline; nothing reads them anymore |
| `data-pipeline/worddata/english/nouns/` | ⚠️ Empty leftover output directory from the old per-word-JSON design |
The pipeline is implemented and running. Every module below is executable code with
co-located unit tests in `data-pipeline/tests/`.
| Path | State |
| --------------------------------------------- | --------------------------------------------------------------------- |
| `data-pipeline/source-data/{lang}/{pos}` | ✅ Noun lists for `de`, `en`, `es`, `fr`, `it` |
| `data-pipeline/prompt` | ✅ Templated Gemini prompt with `{{PLACEHOLDER}}` substitutions |
| `data-pipeline/sourceLists.ts` | ✅ Wordlist discovery, trim/dedup normalization |
| `data-pipeline/promptTemplate.ts` | ✅ Placeholder rendering; throws on any unreplaced `{{...}}` |
| `data-pipeline/gemini.ts` | ✅ Structured-output API client with retry/backoff |
| `data-pipeline/validate.ts` | ✅ Per-entry validation (design-doc §6.4) |
| `data-pipeline/staging.ts` | ✅ SQLite writes, one transaction per word |
| `data-pipeline/pipeline.ts` | ✅ Orchestrator with CLI flags, resumable |
| `data-pipeline/db/schema.sql` | ✅ SQLite staging schema |
| `data-pipeline/db/staging.db` | ✅ Populated — gitignored |
| `data-pipeline/responses/` | ✅ Raw Gemini responses, one JSON per batch — gitignored |
| `data-pipeline/rejections/{lang}-{pos}.jsonl` | ✅ Rejection log, one JSON per failed entry — gitignored |
| SQLite → PostgreSQL import script | ❌ Not written (Phase 4) |
| `data-pipeline/kaikki-source-files/` | ⚠️ 5.7 GB of leftover JSONL from the old pipeline; nothing reads them |
Directory naming follows the language/POS codes used in `packages/shared/src/constants.ts` (`de/noun`, not `german/nouns`) so no name mapping is needed anywhere in the pipeline.
Note that `data-pipeline/vitest.config.ts` looks for tests in `tests/**/*.test.ts` — that directory does not exist yet.
## Module responsibilities
```
sourceLists.ts discoverSourceLists() → [{ sourceLanguage, pos, words, filePath }]
normalizeWords(): trim, drop empties, dedup preserving order
promptTemplate.ts renderPrompt(): substitutes SOURCE_LANGUAGE_NAME/_CODE, POS,
TARGET_LANGUAGE_CODES, TARGET_LANGUAGE_UNION, INPUT_WORDS
gemini.ts buildEntriesResponseSchema(): OpenAPI schema with enums narrowed
to this batch's source/POS/target languages
generateContent(): 5 attempts, retries 429/500/503, honours the
API's own retryDelay; rejects non-STOP finishReason
validate.ts validateEntry() → "valid" | "empty" | "invalid"
staging.ts openStaging(), getStagedHeadwords(), stageEntry(), countStagedRows()
pipeline.ts orchestration, batching, rate-limit delay, rejection logging
```
Two properties worth knowing:
- **Resumability is free.** `getStagedHeadwords()` diffs the input list against what is
already in `staging.db`, so an interrupted run picks up exactly where it stopped and
never re-spends quota on a staged word.
- **Raw responses are saved before parsing.** Validation-rule changes can be replayed
against `responses/` without calling the API again.
`"empty"` is a distinct outcome from `"invalid"`: the contract says a word that is not a
valid noun in that language comes back with `"senses": []`. Those are _skipped_ and
counted separately, not written to the rejection log.
---
@ -93,7 +127,7 @@ translations id, sense_id→senses, target_language_code, UNIQUE(sense_id, targe
## The prompt
`data-pipeline/prompt` is the current working prompt, checked in as a plain text file and edited by hand. It is currently pinned to a concrete sample run (Spanish nouns, 20 words inlined) rather than templated — source language, POS, target languages, and the word batch will need to become substitutions when `pipeline.ts` is implemented.
`data-pipeline/prompt` is the working prompt, checked in as a plain text file and edited by hand. It is fully templated: `promptTemplate.ts` substitutes source language (name and code), POS, target languages, and the word batch, then fails loudly if any `{{PLACEHOLDER}}` survives rendering — so a typo in a placeholder name can never silently reach the API.
What it enforces, beyond the JSON shape in design-doc §6.3:
@ -104,23 +138,59 @@ What it enforces, beyond the JSON shape in design-doc §6.3:
- Base dictionary form, no articles or determiners.
- 1–3 senses per word, most words 1; skip rare, archaic, and technical senses.
- Up to 2 translations per target language per sense, only genuine synonyms or difficulty variants.
- A translation's difficulty may never be lower than its sense's difficulty.
- A translation's difficulty may never be lower than its sense's difficulty, and a sense's
difficulty should equal the easiest translation difficulty in that sense.
- A word that isn't a valid noun in that language comes back with `"senses": []`.
⚠️ **Known inconsistencies in the current prompt file** — it was adapted from the English version and some hardcoded values were not updated: rules 2 and 3 still say `language` must be `"en"` and there is a stray "valid English noun" in rule 31, while the header correctly says Spanish. Rule 15 lists target languages `de, it, es, fr` while the header says `en, it, de, fr`. Fix these when templating the prompt.
The copy-paste bugs that existed in the pre-templating draft (hardcoded `"en"` in rules 2–3,
a contradictory target list in rule 15, "valid English noun" in rule 31) are fixed — those
values are now placeholders.
Beyond the prompt, `gemini.ts` constrains the output with a **response schema** sent on every
request, with enums narrowed to that batch's source language, POS, and target languages. The
shape of the JSON is therefore enforced by the API, and `validate.ts` is left to enforce the
things a schema cannot express: cross-field difficulty ordering, gender rules per target
language, sequential `sense_index`, translation coverage and caps, and that the headword was
actually in the input batch.
Validation is the safety net, not the prompt — every entry is checked before it reaches SQLite, and rejects go to a log for review rather than silently disappearing. Target reject rate is under 10%.
> **Known systemic rejection cause (open).** Effectively all current rejections are
> `translation difficulty lower than sense difficulty`. The model tends to tag a sense
> `medium` while correctly tagging some of its translations `easy`. Since the prompt already
> defines sense difficulty as the easiest translation difficulty in the sense, this value is
> derivable and the entry is otherwise good — normalizing it in `validate.ts` instead of
> rejecting would recover these words. Not yet implemented.
---
## Running it
```bash
docker compose up -d pipeline-database # dedicated PostgreSQL on :5433
pnpm --filter @lila/pipeline pipeline:run # tsx --env-file .env pipeline.ts (currently a no-op)
pnpm --filter @lila/pipeline pipeline:run # tsx --env-file .env pipeline.ts
pnpm --filter @lila/pipeline test
```
`pipeline.ts` takes CLI flags (pass them after `--`, which the script strips itself):
| Flag | Default | Meaning |
| --------------- | ------- | -------------------------------------------------------- |
| `--langs` | all | Comma-separated source languages, e.g. `de,es` |
| `--pos` | `noun` | Part of speech to process |
| `--batch-size` | `20` | Words per Gemini request |
| `--max-batches` | all | Cap batches per list — useful for smoke tests |
| `--delay-ms` | `6000` | Pause between requests, for rate limiting |
| `--dry-run` | `false` | Render prompts and print batches without calling the API |
```bash
# smoke test: one German batch, no API calls
pnpm --filter @lila/pipeline pipeline:run -- --langs de --max-batches 1 --dry-run
```
Because runs are resumable, interrupting with Ctrl-C is safe — already-staged words are
skipped on the next run.
The pipeline reads `.env` from the repo root: `GEMINI_API_KEY`, plus `PIPELINE_POSTGRES_USER` / `PIPELINE_POSTGRES_PASSWORD` / `PIPELINE_POSTGRES_DB` / `PIPELINE_DATABASE_URL`. The pipeline database is deliberately separate from the app database (`:5432`) so pipeline work can never damage dev data.
---
@ -133,12 +203,22 @@ Full breakdown in [pipeline/roadmap.md](pipeline/roadmap.md).
| ------------------------------------------------- | -------------------------------------------------------------- |
| 1 — Drizzle schema (words/senses/translations) | ✅ Complete |
| 2 — Preparation (wordlists, DBs, prompt) | ✅ Complete |
| 3 — Build the pipeline → SQLite | 🔄 **Current.** Validation + `pipeline.ts` + first run |
| 3 — Build the pipeline → SQLite | 🔄 **Current.** Code complete; full 5-language run in progress |
| 4 — Migration & SQLite → PostgreSQL import | ⬜ Not started |
| 5 — App integration (`getGameTerms`, distractors) | ⬜ Not started |
| 6 — Production deploy | ⬜ Not started |
| 7 — Extend to verbs, adjectives, adverbs | ⬜ Not started — new wordlists + prompt only, no schema change |
Target for Phase 3: ~1000 nouns × 5 languages in staging, reject rate under 10%, 50 entries spot-checked by hand.
All Phase 3 code is written and unit-tested. What remains is the data work: finish the
5-language run, resolve the systemic rejection cause noted above, and hand-check 50 entries.
Target for Phase 3: the full deduped noun lists (~1,550–1,750 unique words per language after
dedup) × 5 languages in staging, reject rate under 10%, 50 entries spot-checked by hand.
**Known risk for Phase 5 — the `hard` tier is nearly empty.** Generated difficulty skews
heavily easy, with `hard` translations well under 1% of the corpus. The design-doc §5.1 game
query filters translation difficulty as an _exact_ match, so a "hard" game currently has far
too few rows to fill one round plus distractors. This needs prompt calibration before import,
not after.
Phase 5 is where this becomes visible in the app: `packages/db/src/models/termModel.ts` still queries `vocabulary_entries`/`entry_translations` and must be rewritten against the sense-based schema. Until then, production runs on the old data.