The pipeline docs still described pipeline.ts as pseudocode and the validation module as unwritten. Both have been implemented and run. - CLAUDE.md: replace the "no executable pipeline yet" description with the actual module flow, plus the two invariants worth preserving (resumability via headword diffing, raw responses saved before parsing) - DATA_PIPELINE.md: mark the seven implemented modules, add a module responsibility map and the CLI flag table, drop the resolved warning about hardcoded prompt values - roadmap.md: check off phase 3 tasks 3.2/3.3/3.4, record where the build diverged from the plan, note that validate.ts is stricter than its own spec - STATUS.md: phase 3 is data work now, not code work Also records two open issues: the systemic difficulty-ordering rejection cause, and the hard-tier shortfall pending a full run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
14 KiB
Lila Data Pipeline
How vocabulary data is generated and gets into PostgreSQL. Last updated: 2026-08-20 · Branch:
refactor/gemini-only-pipeline
Authoritative detail lives in two companion docs:
| Doc | What's in it |
|---|---|
| pipeline/design-doc.md | Schema design, difficulty model, query patterns, Gemini JSON contract, indexes |
| pipeline/roadmap.md | Phase-by-phase plan with task checklists and acceptance criteria |
This file is the orientation layer: what the pipeline is, what exists on disk today, and what is not built yet.
The previous local-LLM pipeline (llama.cpp, adapter pattern, 10-model evaluation, CEFR voter ensemble, Kaikki gender lookup) has been removed from the codebase. Its documentation is preserved under archive/ and describes utils/ and config/ modules that no longer exist:
- archive/data-pipeline-local-llm.md — the old pipeline stages and file layout
- archive/llm-setup-local.md — llama.cpp / cloud provider configuration
- archive/model-strategy-cefr-voters.md — the multi-model voter architecture for sense-disambiguated CEFR assignment
What changed, and why
The old pipeline ran small local models and needed a deterministic Kaikki Wiktionary lookup to patch grammatical gender, because local models hallucinated it. The rewrite drops local inference entirely in favour of the Gemini API: one provider, no adapter layer, gender produced directly by the model and enforced by validation instead of by a second data source.
The data model changed with it. The live vocabulary_entries / entry_translations tables (one row per word sense, populated from Kaikki) are replaced by words → senses → translations, where translations hang off a sense, not off a flat entry. That is the whole point of the rewrite: a quiz question can now be tied to one specific meaning of a word.
Flow
source-data/{lang}/{pos} frequency wordlists, one word per line, UTF-8
│
▼
Gemini API batches of 20 words, one language at a time
│
▼
validation per-entry; invalid entries → rejection log, not the DB
│
▼
db/staging.db SQLite staging (words, senses, translations)
│
▼
import script SQLite → PostgreSQL via Drizzle, transaction per batch
│
▼
PostgreSQL (dev :5432, then prod)
Each language is processed independently so definitions and examples are written in that language — a German word gets a German definition, not a translation of an English one. Only the translations cross language boundaries.
The app always reads from PostgreSQL. SQLite exists purely as a staging file so re-runs, prompt tweaks, and spot-checks never touch a real database.
What exists on disk today
The pipeline is implemented and running. Every module below is executable code with
co-located unit tests in data-pipeline/tests/.
| Path | State |
|---|---|
data-pipeline/source-data/{lang}/{pos} |
✅ Noun lists for de, en, es, fr, it |
data-pipeline/prompt |
✅ Templated Gemini prompt with {{PLACEHOLDER}} substitutions |
data-pipeline/sourceLists.ts |
✅ Wordlist discovery, trim/dedup normalization |
data-pipeline/promptTemplate.ts |
✅ Placeholder rendering; throws on any unreplaced {{...}} |
data-pipeline/gemini.ts |
✅ Structured-output API client with retry/backoff |
data-pipeline/validate.ts |
✅ Per-entry validation (design-doc §6.4) |
data-pipeline/staging.ts |
✅ SQLite writes, one transaction per word |
data-pipeline/pipeline.ts |
✅ Orchestrator with CLI flags, resumable |
data-pipeline/db/schema.sql |
✅ SQLite staging schema |
data-pipeline/db/staging.db |
✅ Populated — gitignored |
data-pipeline/responses/ |
✅ Raw Gemini responses, one JSON per batch — gitignored |
data-pipeline/rejections/{lang}-{pos}.jsonl |
✅ Rejection log, one JSON per failed entry — gitignored |
| SQLite → PostgreSQL import script | ❌ Not written (Phase 4) |
data-pipeline/kaikki-source-files/ |
⚠️ 5.7 GB of leftover JSONL from the old pipeline; nothing reads them |
Directory naming follows the language/POS codes used in packages/shared/src/constants.ts (de/noun, not german/nouns) so no name mapping is needed anywhere in the pipeline.
Module responsibilities
sourceLists.ts discoverSourceLists() → [{ sourceLanguage, pos, words, filePath }]
normalizeWords(): trim, drop empties, dedup preserving order
promptTemplate.ts renderPrompt(): substitutes SOURCE_LANGUAGE_NAME/_CODE, POS,
TARGET_LANGUAGE_CODES, TARGET_LANGUAGE_UNION, INPUT_WORDS
gemini.ts buildEntriesResponseSchema(): OpenAPI schema with enums narrowed
to this batch's source/POS/target languages
generateContent(): 5 attempts, retries 429/500/503, honours the
API's own retryDelay; rejects non-STOP finishReason
validate.ts validateEntry() → "valid" | "empty" | "invalid"
staging.ts openStaging(), getStagedHeadwords(), stageEntry(), countStagedRows()
pipeline.ts orchestration, batching, rate-limit delay, rejection logging
Two properties worth knowing:
- Resumability is free.
getStagedHeadwords()diffs the input list against what is already instaging.db, so an interrupted run picks up exactly where it stopped and never re-spends quota on a staged word. - Raw responses are saved before parsing. Validation-rule changes can be replayed
against
responses/without calling the API again.
"empty" is a distinct outcome from "invalid": the contract says a word that is not a
valid noun in that language comes back with "senses": []. Those are skipped and
counted separately, not written to the rejection log.
Staging schema
data-pipeline/db/schema.sql mirrors the PostgreSQL schema with two SQLite concessions: IDs are TEXT (crypto.randomUUID()), and definitions / examples are JSON-encoded strings because SQLite has no array type. The import script parses them back into PostgreSQL TEXT[].
words id, headword, language_code, pos UNIQUE(headword, language_code, pos)
senses id, word_id→words, sense_index, UNIQUE(word_id, sense_index)
difficulty, definitions, examples
translations id, sense_id→senses, target_language_code, UNIQUE(sense_id, target_language_code, translation)
translation, gender, difficulty
difficulty is easy | medium | hard on both senses and translations, and they mean different things — sense difficulty is "is this meaning appropriate for the level", translation difficulty is "is this word an acceptable answer". design-doc §4 explains how queries use sense difficulty as a ceiling and translation difficulty as the target.
The prompt
data-pipeline/prompt is the working prompt, checked in as a plain text file and edited by hand. It is fully templated: promptTemplate.ts substitutes source language (name and code), POS, target languages, and the word batch, then fails loudly if any {{PLACEHOLDER}} survives rendering — so a typo in a placeholder name can never silently reach the API.
What it enforces, beyond the JSON shape in design-doc §6.3:
- Raw JSON only — no markdown fences, comments, or trailing commas; one object per input word, in input order.
- Definitions and examples in the source language.
- Gender required for
de(m/f/n) andit/es/fr(m/f); alwaysnullforen. - German translation nouns capitalized; Romance-language nouns lowercase unless proper nouns.
- Base dictionary form, no articles or determiners.
- 1–3 senses per word, most words 1; skip rare, archaic, and technical senses.
- Up to 2 translations per target language per sense, only genuine synonyms or difficulty variants.
- A translation's difficulty may never be lower than its sense's difficulty, and a sense's difficulty should equal the easiest translation difficulty in that sense.
- A word that isn't a valid noun in that language comes back with
"senses": [].
The copy-paste bugs that existed in the pre-templating draft (hardcoded "en" in rules 2–3,
a contradictory target list in rule 15, "valid English noun" in rule 31) are fixed — those
values are now placeholders.
Beyond the prompt, gemini.ts constrains the output with a response schema sent on every
request, with enums narrowed to that batch's source language, POS, and target languages. The
shape of the JSON is therefore enforced by the API, and validate.ts is left to enforce the
things a schema cannot express: cross-field difficulty ordering, gender rules per target
language, sequential sense_index, translation coverage and caps, and that the headword was
actually in the input batch.
Validation is the safety net, not the prompt — every entry is checked before it reaches SQLite, and rejects go to a log for review rather than silently disappearing. Target reject rate is under 10%.
Known systemic rejection cause (open). Effectively all current rejections are
translation difficulty lower than sense difficulty. The model tends to tag a sensemediumwhile correctly tagging some of its translationseasy. Since the prompt already defines sense difficulty as the easiest translation difficulty in the sense, this value is derivable and the entry is otherwise good — normalizing it invalidate.tsinstead of rejecting would recover these words. Not yet implemented.
Running it
docker compose up -d pipeline-database # dedicated PostgreSQL on :5433
pnpm --filter @lila/pipeline pipeline:run # tsx --env-file .env pipeline.ts
pnpm --filter @lila/pipeline test
pipeline.ts takes CLI flags (pass them after --, which the script strips itself):
| Flag | Default | Meaning |
|---|---|---|
--langs |
all | Comma-separated source languages, e.g. de,es |
--pos |
noun |
Part of speech to process |
--batch-size |
20 |
Words per Gemini request |
--max-batches |
all | Cap batches per list — useful for smoke tests |
--delay-ms |
6000 |
Pause between requests, for rate limiting |
--dry-run |
false |
Render prompts and print batches without calling the API |
# smoke test: one German batch, no API calls
pnpm --filter @lila/pipeline pipeline:run -- --langs de --max-batches 1 --dry-run
Because runs are resumable, interrupting with Ctrl-C is safe — already-staged words are skipped on the next run.
The pipeline reads .env from the repo root: GEMINI_API_KEY, plus PIPELINE_POSTGRES_USER / PIPELINE_POSTGRES_PASSWORD / PIPELINE_POSTGRES_DB / PIPELINE_DATABASE_URL. The pipeline database is deliberately separate from the app database (:5432) so pipeline work can never damage dev data.
Phase status
Full breakdown in pipeline/roadmap.md.
| Phase | State |
|---|---|
| 1 — Drizzle schema (words/senses/translations) | ✅ Complete |
| 2 — Preparation (wordlists, DBs, prompt) | ✅ Complete |
| 3 — Build the pipeline → SQLite | 🔄 Current. Code complete; full 5-language run in progress |
| 4 — Migration & SQLite → PostgreSQL import | ⬜ Not started |
5 — App integration (getGameTerms, distractors) |
⬜ Not started |
| 6 — Production deploy | ⬜ Not started |
| 7 — Extend to verbs, adjectives, adverbs | ⬜ Not started — new wordlists + prompt only, no schema change |
All Phase 3 code is written and unit-tested. What remains is the data work: finish the 5-language run, resolve the systemic rejection cause noted above, and hand-check 50 entries.
Target for Phase 3: the full deduped noun lists (~1,550–1,750 unique words per language after dedup) × 5 languages in staging, reject rate under 10%, 50 entries spot-checked by hand.
Known risk for Phase 5 — the hard tier is nearly empty. Generated difficulty skews
heavily easy, with hard translations well under 1% of the corpus. The design-doc §5.1 game
query filters translation difficulty as an exact match, so a "hard" game currently has far
too few rows to fill one round plus distractors. This needs prompt calibration before import,
not after.
Phase 5 is where this becomes visible in the app: packages/db/src/models/termModel.ts still queries vocabulary_entries/entry_translations and must be rewritten against the sense-based schema. Until then, production runs on the old data.