lila/documentation/DATA_PIPELINE.md
lila da9cdbfa1b updating docs to match the implemented phase 3 pipeline
The pipeline docs still described pipeline.ts as pseudocode and the
validation module as unwritten. Both have been implemented and run.

- CLAUDE.md: replace the "no executable pipeline yet" description with
  the actual module flow, plus the two invariants worth preserving
  (resumability via headword diffing, raw responses saved before parsing)
- DATA_PIPELINE.md: mark the seven implemented modules, add a module
  responsibility map and the CLI flag table, drop the resolved warning
  about hardcoded prompt values
- roadmap.md: check off phase 3 tasks 3.2/3.3/3.4, record where the
  build diverged from the plan, note that validate.ts is stricter than
  its own spec
- STATUS.md: phase 3 is data work now, not code work

Also records two open issues: the systemic difficulty-ordering rejection
cause, and the hard-tier shortfall pending a full run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 12:10:48 +02:00

14 KiB
Raw Blame History

Lila Data Pipeline

How vocabulary data is generated and gets into PostgreSQL. Last updated: 2026-08-20 · Branch: refactor/gemini-only-pipeline

Authoritative detail lives in two companion docs:

Doc What's in it
pipeline/design-doc.md Schema design, difficulty model, query patterns, Gemini JSON contract, indexes
pipeline/roadmap.md Phase-by-phase plan with task checklists and acceptance criteria

This file is the orientation layer: what the pipeline is, what exists on disk today, and what is not built yet.

The previous local-LLM pipeline (llama.cpp, adapter pattern, 10-model evaluation, CEFR voter ensemble, Kaikki gender lookup) has been removed from the codebase. Its documentation is preserved under archive/ and describes utils/ and config/ modules that no longer exist:


What changed, and why

The old pipeline ran small local models and needed a deterministic Kaikki Wiktionary lookup to patch grammatical gender, because local models hallucinated it. The rewrite drops local inference entirely in favour of the Gemini API: one provider, no adapter layer, gender produced directly by the model and enforced by validation instead of by a second data source.

The data model changed with it. The live vocabulary_entries / entry_translations tables (one row per word sense, populated from Kaikki) are replaced by words → senses → translations, where translations hang off a sense, not off a flat entry. That is the whole point of the rewrite: a quiz question can now be tied to one specific meaning of a word.


Flow

source-data/{lang}/{pos}   frequency wordlists, one word per line, UTF-8
        │
        ▼
Gemini API                 batches of 20 words, one language at a time
        │
        ▼
validation                 per-entry; invalid entries → rejection log, not the DB
        │
        ▼
db/staging.db              SQLite staging (words, senses, translations)
        │
        ▼
import script              SQLite → PostgreSQL via Drizzle, transaction per batch
        │
        ▼
PostgreSQL (dev :5432, then prod)

Each language is processed independently so definitions and examples are written in that language — a German word gets a German definition, not a translation of an English one. Only the translations cross language boundaries.

The app always reads from PostgreSQL. SQLite exists purely as a staging file so re-runs, prompt tweaks, and spot-checks never touch a real database.


What exists on disk today

The pipeline is implemented and running. Every module below is executable code with co-located unit tests in data-pipeline/tests/.

Path State
data-pipeline/source-data/{lang}/{pos} ✅ Noun lists for de, en, es, fr, it
data-pipeline/prompt ✅ Templated Gemini prompt with {{PLACEHOLDER}} substitutions
data-pipeline/sourceLists.ts ✅ Wordlist discovery, trim/dedup normalization
data-pipeline/promptTemplate.ts ✅ Placeholder rendering; throws on any unreplaced {{...}}
data-pipeline/gemini.ts ✅ Structured-output API client with retry/backoff
data-pipeline/validate.ts ✅ Per-entry validation (design-doc §6.4)
data-pipeline/staging.ts ✅ SQLite writes, one transaction per word
data-pipeline/pipeline.ts ✅ Orchestrator with CLI flags, resumable
data-pipeline/db/schema.sql ✅ SQLite staging schema
data-pipeline/db/staging.db ✅ Populated — gitignored
data-pipeline/responses/ ✅ Raw Gemini responses, one JSON per batch — gitignored
data-pipeline/rejections/{lang}-{pos}.jsonl ✅ Rejection log, one JSON per failed entry — gitignored
SQLite → PostgreSQL import script ❌ Not written (Phase 4)
data-pipeline/kaikki-source-files/ ⚠️ 5.7 GB of leftover JSONL from the old pipeline; nothing reads them

Directory naming follows the language/POS codes used in packages/shared/src/constants.ts (de/noun, not german/nouns) so no name mapping is needed anywhere in the pipeline.

Module responsibilities

sourceLists.ts     discoverSourceLists() → [{ sourceLanguage, pos, words, filePath }]
                   normalizeWords(): trim, drop empties, dedup preserving order
promptTemplate.ts  renderPrompt(): substitutes SOURCE_LANGUAGE_NAME/_CODE, POS,
                   TARGET_LANGUAGE_CODES, TARGET_LANGUAGE_UNION, INPUT_WORDS
gemini.ts          buildEntriesResponseSchema(): OpenAPI schema with enums narrowed
                   to this batch's source/POS/target languages
                   generateContent(): 5 attempts, retries 429/500/503, honours the
                   API's own retryDelay; rejects non-STOP finishReason
validate.ts        validateEntry() → "valid" | "empty" | "invalid"
staging.ts         openStaging(), getStagedHeadwords(), stageEntry(), countStagedRows()
pipeline.ts        orchestration, batching, rate-limit delay, rejection logging

Two properties worth knowing:

  • Resumability is free. getStagedHeadwords() diffs the input list against what is already in staging.db, so an interrupted run picks up exactly where it stopped and never re-spends quota on a staged word.
  • Raw responses are saved before parsing. Validation-rule changes can be replayed against responses/ without calling the API again.

"empty" is a distinct outcome from "invalid": the contract says a word that is not a valid noun in that language comes back with "senses": []. Those are skipped and counted separately, not written to the rejection log.


Staging schema

data-pipeline/db/schema.sql mirrors the PostgreSQL schema with two SQLite concessions: IDs are TEXT (crypto.randomUUID()), and definitions / examples are JSON-encoded strings because SQLite has no array type. The import script parses them back into PostgreSQL TEXT[].

words        id, headword, language_code, pos          UNIQUE(headword, language_code, pos)
senses       id, word_id→words, sense_index,           UNIQUE(word_id, sense_index)
             difficulty, definitions, examples
translations id, sense_id→senses, target_language_code, UNIQUE(sense_id, target_language_code, translation)
             translation, gender, difficulty

difficulty is easy | medium | hard on both senses and translations, and they mean different things — sense difficulty is "is this meaning appropriate for the level", translation difficulty is "is this word an acceptable answer". design-doc §4 explains how queries use sense difficulty as a ceiling and translation difficulty as the target.


The prompt

data-pipeline/prompt is the working prompt, checked in as a plain text file and edited by hand. It is fully templated: promptTemplate.ts substitutes source language (name and code), POS, target languages, and the word batch, then fails loudly if any {{PLACEHOLDER}} survives rendering — so a typo in a placeholder name can never silently reach the API.

What it enforces, beyond the JSON shape in design-doc §6.3:

  • Raw JSON only — no markdown fences, comments, or trailing commas; one object per input word, in input order.
  • Definitions and examples in the source language.
  • Gender required for de (m/f/n) and it/es/fr (m/f); always null for en.
  • German translation nouns capitalized; Romance-language nouns lowercase unless proper nouns.
  • Base dictionary form, no articles or determiners.
  • 1–3 senses per word, most words 1; skip rare, archaic, and technical senses.
  • Up to 2 translations per target language per sense, only genuine synonyms or difficulty variants.
  • A translation's difficulty may never be lower than its sense's difficulty, and a sense's difficulty should equal the easiest translation difficulty in that sense.
  • A word that isn't a valid noun in that language comes back with "senses": [].

The copy-paste bugs that existed in the pre-templating draft (hardcoded "en" in rules 2–3, a contradictory target list in rule 15, "valid English noun" in rule 31) are fixed — those values are now placeholders.

Beyond the prompt, gemini.ts constrains the output with a response schema sent on every request, with enums narrowed to that batch's source language, POS, and target languages. The shape of the JSON is therefore enforced by the API, and validate.ts is left to enforce the things a schema cannot express: cross-field difficulty ordering, gender rules per target language, sequential sense_index, translation coverage and caps, and that the headword was actually in the input batch.

Validation is the safety net, not the prompt — every entry is checked before it reaches SQLite, and rejects go to a log for review rather than silently disappearing. Target reject rate is under 10%.

Known systemic rejection cause (open). Effectively all current rejections are translation difficulty lower than sense difficulty. The model tends to tag a sense medium while correctly tagging some of its translations easy. Since the prompt already defines sense difficulty as the easiest translation difficulty in the sense, this value is derivable and the entry is otherwise good — normalizing it in validate.ts instead of rejecting would recover these words. Not yet implemented.


Running it

docker compose up -d pipeline-database      # dedicated PostgreSQL on :5433
pnpm --filter @lila/pipeline pipeline:run   # tsx --env-file .env pipeline.ts
pnpm --filter @lila/pipeline test

pipeline.ts takes CLI flags (pass them after --, which the script strips itself):

Flag Default Meaning
--langs all Comma-separated source languages, e.g. de,es
--pos noun Part of speech to process
--batch-size 20 Words per Gemini request
--max-batches all Cap batches per list — useful for smoke tests
--delay-ms 6000 Pause between requests, for rate limiting
--dry-run false Render prompts and print batches without calling the API
# smoke test: one German batch, no API calls
pnpm --filter @lila/pipeline pipeline:run -- --langs de --max-batches 1 --dry-run

Because runs are resumable, interrupting with Ctrl-C is safe — already-staged words are skipped on the next run.

The pipeline reads .env from the repo root: GEMINI_API_KEY, plus PIPELINE_POSTGRES_USER / PIPELINE_POSTGRES_PASSWORD / PIPELINE_POSTGRES_DB / PIPELINE_DATABASE_URL. The pipeline database is deliberately separate from the app database (:5432) so pipeline work can never damage dev data.


Phase status

Full breakdown in pipeline/roadmap.md.

Phase State
1 — Drizzle schema (words/senses/translations) ✅ Complete
2 — Preparation (wordlists, DBs, prompt) ✅ Complete
3 — Build the pipeline → SQLite 🔄 Current. Code complete; full 5-language run in progress
4 — Migration & SQLite → PostgreSQL import ⬜ Not started
5 — App integration (getGameTerms, distractors) ⬜ Not started
6 — Production deploy ⬜ Not started
7 — Extend to verbs, adjectives, adverbs ⬜ Not started — new wordlists + prompt only, no schema change

All Phase 3 code is written and unit-tested. What remains is the data work: finish the 5-language run, resolve the systemic rejection cause noted above, and hand-check 50 entries.

Target for Phase 3: the full deduped noun lists (~1,550–1,750 unique words per language after dedup) × 5 languages in staging, reject rate under 10%, 50 entries spot-checked by hand.

Known risk for Phase 5 — the hard tier is nearly empty. Generated difficulty skews heavily easy, with hard translations well under 1% of the corpus. The design-doc §5.1 game query filters translation difficulty as an exact match, so a "hard" game currently has far too few rows to fill one round plus distractors. This needs prompt calibration before import, not after.

Phase 5 is where this becomes visible in the app: packages/db/src/models/termModel.ts still queries vocabulary_entries/entry_translations and must be rewritten against the sense-based schema. Until then, production runs on the old data.