# Lila Data Pipeline > How vocabulary data is generated and gets into PostgreSQL. > Last updated: 2026-08-20 · Branch: `refactor/gemini-only-pipeline` **Authoritative detail lives in two companion docs:** | Doc | What's in it | | ------------------------------------------------ | ------------------------------------------------------------------------------ | | [pipeline/design-doc.md](pipeline/design-doc.md) | Schema design, difficulty model, query patterns, Gemini JSON contract, indexes | | [pipeline/roadmap.md](pipeline/roadmap.md) | Phase-by-phase plan with task checklists and acceptance criteria | This file is the orientation layer: what the pipeline is, what exists on disk today, and what is not built yet. The previous local-LLM pipeline (llama.cpp, adapter pattern, 10-model evaluation, CEFR voter ensemble, Kaikki gender lookup) has been removed from the codebase. Its documentation is preserved under `archive/` and describes `utils/` and `config/` modules that **no longer exist**: - [archive/data-pipeline-local-llm.md](archive/data-pipeline-local-llm.md) — the old pipeline stages and file layout - [archive/llm-setup-local.md](archive/llm-setup-local.md) — llama.cpp / cloud provider configuration - [archive/model-strategy-cefr-voters.md](archive/model-strategy-cefr-voters.md) — the multi-model voter architecture for sense-disambiguated CEFR assignment --- ## What changed, and why The old pipeline ran small local models and needed a deterministic Kaikki Wiktionary lookup to patch grammatical gender, because local models hallucinated it. The rewrite drops local inference entirely in favour of the Gemini API: one provider, no adapter layer, gender produced directly by the model and enforced by validation instead of by a second data source. The data model changed with it. The live `vocabulary_entries` / `entry_translations` tables (one row per word sense, populated from Kaikki) are replaced by `words` → `senses` → `translations`, where translations hang off a **sense**, not off a flat entry. That is the whole point of the rewrite: a quiz question can now be tied to one specific meaning of a word. --- ## Flow ``` source-data/{lang}/{pos} frequency wordlists, one word per line, UTF-8 │ ▼ Gemini API batches of 20 words, one language at a time │ ▼ validation per-entry; invalid entries → rejection log, not the DB │ ▼ db/staging.db SQLite staging (words, senses, translations) │ ▼ import script SQLite → PostgreSQL via Drizzle, transaction per batch │ ▼ PostgreSQL (dev :5432, then prod) ``` Each language is processed independently so definitions and examples are written **in that language** — a German word gets a German definition, not a translation of an English one. Only the translations cross language boundaries. The app always reads from PostgreSQL. SQLite exists purely as a staging file so re-runs, prompt tweaks, and spot-checks never touch a real database. --- ## What exists on disk today The pipeline is implemented and running. Every module below is executable code with co-located unit tests in `data-pipeline/tests/`. | Path | State | | --------------------------------------------- | --------------------------------------------------------------------- | | `data-pipeline/source-data/{lang}/{pos}` | ✅ Noun lists for `de`, `en`, `es`, `fr`, `it` | | `data-pipeline/prompt` | ✅ Templated Gemini prompt with `{{PLACEHOLDER}}` substitutions | | `data-pipeline/sourceLists.ts` | ✅ Wordlist discovery, trim/dedup normalization | | `data-pipeline/promptTemplate.ts` | ✅ Placeholder rendering; throws on any unreplaced `{{...}}` | | `data-pipeline/gemini.ts` | ✅ Structured-output API client with retry/backoff | | `data-pipeline/validate.ts` | ✅ Per-entry validation (design-doc §6.4) | | `data-pipeline/staging.ts` | ✅ SQLite writes, one transaction per word | | `data-pipeline/pipeline.ts` | ✅ Orchestrator with CLI flags, resumable | | `data-pipeline/replay.ts` | ✅ Re-validates saved responses without API calls | | `data-pipeline/db/schema.sql` | ✅ SQLite staging schema | | `data-pipeline/db/staging.db` | ✅ Populated — gitignored | | `data-pipeline/responses/` | ✅ Raw Gemini responses, one JSON per batch — gitignored | | `data-pipeline/rejections/{lang}-{pos}.jsonl` | ✅ Rejection log, one JSON per failed entry — gitignored | | SQLite → PostgreSQL import script | ❌ Not written (Phase 4) | | `data-pipeline/kaikki-source-files/` | ⚠️ 5.7 GB of leftover JSONL from the old pipeline; nothing reads them | Directory naming follows the language/POS codes used in `packages/shared/src/constants.ts` (`de/noun`, not `german/nouns`) so no name mapping is needed anywhere in the pipeline. ## Module responsibilities ``` sourceLists.ts discoverSourceLists() → [{ sourceLanguage, pos, words, filePath }] normalizeWords(): trim, drop empties, dedup preserving order promptTemplate.ts renderPrompt(): substitutes SOURCE_LANGUAGE_NAME/_CODE, POS, TARGET_LANGUAGE_CODES, TARGET_LANGUAGE_UNION, INPUT_WORDS gemini.ts buildEntriesResponseSchema(): OpenAPI schema with enums narrowed to this batch's source/POS/target languages generateContent(): 5 attempts, retries 429/500/503, honours the API's own retryDelay; rejects non-STOP finishReason validate.ts validateEntry() → "valid" | "empty" | "invalid" applySenseDifficultyFloor(): lowers a sense to its easiest translation, reported via the result's `normalizations` staging.ts openStaging(), getStagedHeadwords(), stageEntry(), countStagedRows() pipeline.ts orchestration, batching, rate-limit delay, rejection logging replay.ts re-validates responses/ with the current rules, no API calls ``` Two properties worth knowing: - **Resumability is free.** `getStagedHeadwords()` diffs the input list against what is already in `staging.db`, so an interrupted run picks up exactly where it stopped and never re-spends quota on a staged word. - **Raw responses are saved before parsing.** Validation-rule changes can be replayed against `responses/` without calling the API again. `"empty"` is a distinct outcome from `"invalid"`: the contract says a word that is not a valid noun in that language comes back with `"senses": []`. Those are _skipped_ and counted separately, not written to the rejection log. --- ## Staging schema `data-pipeline/db/schema.sql` mirrors the PostgreSQL schema with two SQLite concessions: IDs are `TEXT` (`crypto.randomUUID()`), and `definitions` / `examples` are JSON-encoded strings because SQLite has no array type. The import script parses them back into PostgreSQL `TEXT[]`. ``` words id, headword, language_code, pos UNIQUE(headword, language_code, pos) senses id, word_id→words, sense_index, UNIQUE(word_id, sense_index) difficulty, definitions, examples translations id, sense_id→senses, target_language_code, UNIQUE(sense_id, target_language_code, translation) translation, gender, difficulty ``` `difficulty` is `easy | medium | hard` on both `senses` and `translations`, and they mean different things — sense difficulty is "is this _meaning_ appropriate for the level", translation difficulty is "is this _word_ an acceptable answer". design-doc §4 explains how queries use sense difficulty as a ceiling and translation difficulty as the target. --- ## The prompt `data-pipeline/prompt` is the working prompt, checked in as a plain text file and edited by hand. It is fully templated: `promptTemplate.ts` substitutes source language (name and code), POS, target languages, and the word batch, then fails loudly if any `{{PLACEHOLDER}}` survives rendering — so a typo in a placeholder name can never silently reach the API. What it enforces, beyond the JSON shape in design-doc §6.3: - Raw JSON only — no markdown fences, comments, or trailing commas; one object per input word, in input order. - Definitions and examples in the **source** language. - Gender required for `de` (m/f/n) and `it`/`es`/`fr` (m/f); always `null` for `en`. - German translation nouns capitalized; Romance-language nouns lowercase unless proper nouns. - Base dictionary form, no articles or determiners. - 1–3 senses per word, most words 1; skip rare, archaic, and technical senses. - Up to 2 translations per target language per sense, only genuine synonyms or difficulty variants. - A translation's difficulty may never be lower than its sense's difficulty, and a sense's difficulty should equal the easiest translation difficulty in that sense. - A word that isn't a valid noun in that language comes back with `"senses": []`. The copy-paste bugs that existed in the pre-templating draft (hardcoded `"en"` in rules 2–3, a contradictory target list in rule 15, "valid English noun" in rule 31) are fixed — those values are now placeholders. Beyond the prompt, `gemini.ts` constrains the output with a **response schema** sent on every request, with enums narrowed to that batch's source language, POS, and target languages. The shape of the JSON is therefore enforced by the API, and `validate.ts` is left to enforce the things a schema cannot express: cross-field difficulty ordering, gender rules per target language, sequential `sense_index`, translation coverage and caps, and that the headword was actually in the input batch. Validation is the safety net, not the prompt — every entry is checked before it reaches SQLite, and rejects go to a log for review rather than silently disappearing. Target reject rate is under 10%. > **Systemic rejection cause — resolved.** Effectively every rejection used to be > `translation difficulty lower than sense difficulty`: the model tags a sense `medium` > while correctly tagging some of its translations `easy`. The prompt defines sense > difficulty as the easiest translation difficulty in the sense, which makes it a derived > value, so `validate.ts` now floors it (`applySenseDifficultyFloor`) instead of rejecting > the entry. The floor only ever lowers — raising a sense would gate a concept out of levels > it belongs in and collapse the concept-vs-word split of design-doc §4. Replaying every > saved response with the new rule took the reject rate from 20 entries to **zero**. --- ## Running it ```bash docker compose up -d pipeline-database # dedicated PostgreSQL on :5433 pnpm --filter @lila/pipeline pipeline:run # tsx --env-file .env pipeline.ts pnpm --filter @lila/pipeline test ``` `pipeline.ts` takes CLI flags (pass them after `--`, which the script strips itself): | Flag | Default | Meaning | | --------------- | ------- | -------------------------------------------------------- | | `--langs` | all | Comma-separated source languages, e.g. `de,es` | | `--pos` | `noun` | Part of speech to process | | `--batch-size` | `20` | Words per Gemini request | | `--max-batches` | all | Cap batches per list — useful for smoke tests | | `--delay-ms` | `6000` | Pause between requests, for rate limiting | | `--dry-run` | `false` | Render prompts and print batches without calling the API | ```bash # smoke test: one German batch, no API calls pnpm --filter @lila/pipeline pipeline:run -- --langs de --max-batches 1 --dry-run ``` Because runs are resumable, interrupting with Ctrl-C is safe — already-staged words are skipped on the next run. ### Replaying saved responses `replay.ts` re-validates everything in `responses/` **without calling the API**, which is how a validation-rule change is applied to data that has already been generated. ```bash pnpm --filter @lila/pipeline pipeline:replay # report only pnpm --filter @lila/pipeline pipeline:replay -- --write # stage recovered entries pnpm --filter @lila/pipeline pipeline:replay -- --verbose --langs de ``` It reports valid / normalized / empty / invalid counts and groups anything still invalid by error. Staging is idempotent, so `--write` is safe to repeat. One limitation: it can only _insert_. A word already in `staging.db` is left untouched, so a rule change cannot repair rows that were staged under the old rules — only recover ones that were rejected. ### API quota ⚠️ **The free tier allows far fewer requests than the batch count needs.** Observed 2026-08-20: `gemini-3.6-flash` returned `HTTP 429 … limit: 20` on metric `generate_content_free_tier_requests` after **17 successful requests in one day**, and did not recover across ~55 minutes of retrying — a daily window, not a per-minute one. The `retryDelay: ~59s` in the 429 payload is generic backoff advice and misleads the retry logic into grinding. Consequences to plan around: - What is rationed is **requests**, not words, so `--batch-size` is the cheap lever. The remaining ~7,300 words are ~365 requests at batch size 20, but only ~74 at batch size 100. - `gemini.ts` currently treats 429 as retryable and burns all 5 `MAX_ATTEMPTS` against a quota that will not clear for hours. It should distinguish rate-limiting from daily exhaustion and abort the run. - Replays cost nothing, which is why raw responses are persisted. The pipeline reads `.env` from the repo root: `GEMINI_API_KEY`, plus `PIPELINE_POSTGRES_USER` / `PIPELINE_POSTGRES_PASSWORD` / `PIPELINE_POSTGRES_DB` / `PIPELINE_DATABASE_URL`. The pipeline database is deliberately separate from the app database (`:5432`) so pipeline work can never damage dev data. --- ## Phase status Full breakdown in [pipeline/roadmap.md](pipeline/roadmap.md). | Phase | State | | ------------------------------------------------- | -------------------------------------------------------------- | | 1 — Drizzle schema (words/senses/translations) | ✅ Complete | | 2 — Preparation (wordlists, DBs, prompt) | ✅ Complete | | 3 — Build the pipeline → SQLite | 🔄 **Current.** Code complete; full 5-language run in progress | | 4 — Migration & SQLite → PostgreSQL import | ⬜ Not started | | 5 — App integration (`getGameTerms`, distractors) | ⬜ Not started | | 6 — Production deploy | ⬜ Not started | | 7 — Extend to verbs, adjectives, adverbs | ⬜ Not started — new wordlists + prompt only, no schema change | All Phase 3 code is written and unit-tested. What remains is the data work: finish the 5-language run, resolve the systemic rejection cause noted above, and hand-check 50 entries. Target for Phase 3: the full deduped noun lists (~1,550–1,750 unique words per language after dedup) × 5 languages in staging, reject rate under 10%, 50 entries spot-checked by hand. **Known risk for Phase 5 — the `hard` tier is nearly empty.** Generated difficulty skews heavily easy, with `hard` translations well under 1% of the corpus. The design-doc §5.1 game query filters translation difficulty as an _exact_ match, so a "hard" game currently has far too few rows to fill one round plus distractors. This needs prompt calibration before import, not after. Phase 5 is where this becomes visible in the app: `packages/db/src/models/termModel.ts` still queries `vocabulary_entries`/`entry_translations` and must be rewritten against the sense-based schema. Until then, production runs on the old data.