Every rejection in the German run was the same rule: a translation ranked below its own sense. The model tags a sense "medium" while correctly tagging some translations "easy" — the translations are right and the derived sense label is wrong, but the whole entry was discarded. The prompt defines sense difficulty as the easiest translation difficulty in that sense, so it is a derived value rather than an independent judgement. validate.ts now recomputes it via applySenseDifficultyFloor. The floor only ever lowers. Raising a sense to match its translations would gate a concept out of levels it belongs in and collapse the concept-vs-word distinction the two difficulty columns exist to express (design-doc section 4). - validate.ts: drop the cross-field rejection, add the floor; the valid result now carries "normalizations" so repairs are reported, not silent - pipeline.ts: count and print normalizations per batch and in the summary - replay.ts: new, re-validates responses/ with the current rules and no API calls; --write stages recovered entries, --langs and --verbose - tests: six cases covering the floor, replacing the old rejection test Replaying all 38 saved responses took the reject rate from 20 entries to zero. staging.db now holds 746 words / 774 senses / 3,436 translations with no sense ranked above its easiest translation. Docs also record the API quota ceiling found today: the free tier allows about 20 requests/day, not the 1,000 previously assumed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
264 lines
16 KiB
Markdown
264 lines
16 KiB
Markdown
# Lila Data Pipeline
|
||
|
||
> How vocabulary data is generated and gets into PostgreSQL.
|
||
> Last updated: 2026-08-20 · Branch: `refactor/gemini-only-pipeline`
|
||
|
||
**Authoritative detail lives in two companion docs:**
|
||
|
||
| Doc | What's in it |
|
||
| ------------------------------------------------ | ------------------------------------------------------------------------------ |
|
||
| [pipeline/design-doc.md](pipeline/design-doc.md) | Schema design, difficulty model, query patterns, Gemini JSON contract, indexes |
|
||
| [pipeline/roadmap.md](pipeline/roadmap.md) | Phase-by-phase plan with task checklists and acceptance criteria |
|
||
|
||
This file is the orientation layer: what the pipeline is, what exists on disk today, and what is not built yet.
|
||
|
||
The previous local-LLM pipeline (llama.cpp, adapter pattern, 10-model evaluation, CEFR voter ensemble, Kaikki gender lookup) has been removed from the codebase. Its documentation is preserved under `archive/` and describes `utils/` and `config/` modules that **no longer exist**:
|
||
|
||
- [archive/data-pipeline-local-llm.md](archive/data-pipeline-local-llm.md) — the old pipeline stages and file layout
|
||
- [archive/llm-setup-local.md](archive/llm-setup-local.md) — llama.cpp / cloud provider configuration
|
||
- [archive/model-strategy-cefr-voters.md](archive/model-strategy-cefr-voters.md) — the multi-model voter architecture for sense-disambiguated CEFR assignment
|
||
|
||
---
|
||
|
||
## What changed, and why
|
||
|
||
The old pipeline ran small local models and needed a deterministic Kaikki Wiktionary lookup to patch grammatical gender, because local models hallucinated it. The rewrite drops local inference entirely in favour of the Gemini API: one provider, no adapter layer, gender produced directly by the model and enforced by validation instead of by a second data source.
|
||
|
||
The data model changed with it. The live `vocabulary_entries` / `entry_translations` tables (one row per word sense, populated from Kaikki) are replaced by `words` → `senses` → `translations`, where translations hang off a **sense**, not off a flat entry. That is the whole point of the rewrite: a quiz question can now be tied to one specific meaning of a word.
|
||
|
||
---
|
||
|
||
## Flow
|
||
|
||
```
|
||
source-data/{lang}/{pos} frequency wordlists, one word per line, UTF-8
|
||
│
|
||
▼
|
||
Gemini API batches of 20 words, one language at a time
|
||
│
|
||
▼
|
||
validation per-entry; invalid entries → rejection log, not the DB
|
||
│
|
||
▼
|
||
db/staging.db SQLite staging (words, senses, translations)
|
||
│
|
||
▼
|
||
import script SQLite → PostgreSQL via Drizzle, transaction per batch
|
||
│
|
||
▼
|
||
PostgreSQL (dev :5432, then prod)
|
||
```
|
||
|
||
Each language is processed independently so definitions and examples are written **in that language** — a German word gets a German definition, not a translation of an English one. Only the translations cross language boundaries.
|
||
|
||
The app always reads from PostgreSQL. SQLite exists purely as a staging file so re-runs, prompt tweaks, and spot-checks never touch a real database.
|
||
|
||
---
|
||
|
||
## What exists on disk today
|
||
|
||
The pipeline is implemented and running. Every module below is executable code with
|
||
co-located unit tests in `data-pipeline/tests/`.
|
||
|
||
| Path | State |
|
||
| --------------------------------------------- | --------------------------------------------------------------------- |
|
||
| `data-pipeline/source-data/{lang}/{pos}` | ✅ Noun lists for `de`, `en`, `es`, `fr`, `it` |
|
||
| `data-pipeline/prompt` | ✅ Templated Gemini prompt with `{{PLACEHOLDER}}` substitutions |
|
||
| `data-pipeline/sourceLists.ts` | ✅ Wordlist discovery, trim/dedup normalization |
|
||
| `data-pipeline/promptTemplate.ts` | ✅ Placeholder rendering; throws on any unreplaced `{{...}}` |
|
||
| `data-pipeline/gemini.ts` | ✅ Structured-output API client with retry/backoff |
|
||
| `data-pipeline/validate.ts` | ✅ Per-entry validation (design-doc §6.4) |
|
||
| `data-pipeline/staging.ts` | ✅ SQLite writes, one transaction per word |
|
||
| `data-pipeline/pipeline.ts` | ✅ Orchestrator with CLI flags, resumable |
|
||
| `data-pipeline/replay.ts` | ✅ Re-validates saved responses without API calls |
|
||
| `data-pipeline/db/schema.sql` | ✅ SQLite staging schema |
|
||
| `data-pipeline/db/staging.db` | ✅ Populated — gitignored |
|
||
| `data-pipeline/responses/` | ✅ Raw Gemini responses, one JSON per batch — gitignored |
|
||
| `data-pipeline/rejections/{lang}-{pos}.jsonl` | ✅ Rejection log, one JSON per failed entry — gitignored |
|
||
| SQLite → PostgreSQL import script | ❌ Not written (Phase 4) |
|
||
| `data-pipeline/kaikki-source-files/` | ⚠️ 5.7 GB of leftover JSONL from the old pipeline; nothing reads them |
|
||
|
||
Directory naming follows the language/POS codes used in `packages/shared/src/constants.ts` (`de/noun`, not `german/nouns`) so no name mapping is needed anywhere in the pipeline.
|
||
|
||
## Module responsibilities
|
||
|
||
```
|
||
sourceLists.ts discoverSourceLists() → [{ sourceLanguage, pos, words, filePath }]
|
||
normalizeWords(): trim, drop empties, dedup preserving order
|
||
promptTemplate.ts renderPrompt(): substitutes SOURCE_LANGUAGE_NAME/_CODE, POS,
|
||
TARGET_LANGUAGE_CODES, TARGET_LANGUAGE_UNION, INPUT_WORDS
|
||
gemini.ts buildEntriesResponseSchema(): OpenAPI schema with enums narrowed
|
||
to this batch's source/POS/target languages
|
||
generateContent(): 5 attempts, retries 429/500/503, honours the
|
||
API's own retryDelay; rejects non-STOP finishReason
|
||
validate.ts validateEntry() → "valid" | "empty" | "invalid"
|
||
applySenseDifficultyFloor(): lowers a sense to its easiest
|
||
translation, reported via the result's `normalizations`
|
||
staging.ts openStaging(), getStagedHeadwords(), stageEntry(), countStagedRows()
|
||
pipeline.ts orchestration, batching, rate-limit delay, rejection logging
|
||
replay.ts re-validates responses/ with the current rules, no API calls
|
||
```
|
||
|
||
Two properties worth knowing:
|
||
|
||
- **Resumability is free.** `getStagedHeadwords()` diffs the input list against what is
|
||
already in `staging.db`, so an interrupted run picks up exactly where it stopped and
|
||
never re-spends quota on a staged word.
|
||
- **Raw responses are saved before parsing.** Validation-rule changes can be replayed
|
||
against `responses/` without calling the API again.
|
||
|
||
`"empty"` is a distinct outcome from `"invalid"`: the contract says a word that is not a
|
||
valid noun in that language comes back with `"senses": []`. Those are _skipped_ and
|
||
counted separately, not written to the rejection log.
|
||
|
||
---
|
||
|
||
## Staging schema
|
||
|
||
`data-pipeline/db/schema.sql` mirrors the PostgreSQL schema with two SQLite concessions: IDs are `TEXT` (`crypto.randomUUID()`), and `definitions` / `examples` are JSON-encoded strings because SQLite has no array type. The import script parses them back into PostgreSQL `TEXT[]`.
|
||
|
||
```
|
||
words id, headword, language_code, pos UNIQUE(headword, language_code, pos)
|
||
senses id, word_id→words, sense_index, UNIQUE(word_id, sense_index)
|
||
difficulty, definitions, examples
|
||
translations id, sense_id→senses, target_language_code, UNIQUE(sense_id, target_language_code, translation)
|
||
translation, gender, difficulty
|
||
```
|
||
|
||
`difficulty` is `easy | medium | hard` on both `senses` and `translations`, and they mean different things — sense difficulty is "is this _meaning_ appropriate for the level", translation difficulty is "is this _word_ an acceptable answer". design-doc §4 explains how queries use sense difficulty as a ceiling and translation difficulty as the target.
|
||
|
||
---
|
||
|
||
## The prompt
|
||
|
||
`data-pipeline/prompt` is the working prompt, checked in as a plain text file and edited by hand. It is fully templated: `promptTemplate.ts` substitutes source language (name and code), POS, target languages, and the word batch, then fails loudly if any `{{PLACEHOLDER}}` survives rendering — so a typo in a placeholder name can never silently reach the API.
|
||
|
||
What it enforces, beyond the JSON shape in design-doc §6.3:
|
||
|
||
- Raw JSON only — no markdown fences, comments, or trailing commas; one object per input word, in input order.
|
||
- Definitions and examples in the **source** language.
|
||
- Gender required for `de` (m/f/n) and `it`/`es`/`fr` (m/f); always `null` for `en`.
|
||
- German translation nouns capitalized; Romance-language nouns lowercase unless proper nouns.
|
||
- Base dictionary form, no articles or determiners.
|
||
- 1–3 senses per word, most words 1; skip rare, archaic, and technical senses.
|
||
- Up to 2 translations per target language per sense, only genuine synonyms or difficulty variants.
|
||
- A translation's difficulty may never be lower than its sense's difficulty, and a sense's
|
||
difficulty should equal the easiest translation difficulty in that sense.
|
||
- A word that isn't a valid noun in that language comes back with `"senses": []`.
|
||
|
||
The copy-paste bugs that existed in the pre-templating draft (hardcoded `"en"` in rules 2–3,
|
||
a contradictory target list in rule 15, "valid English noun" in rule 31) are fixed — those
|
||
values are now placeholders.
|
||
|
||
Beyond the prompt, `gemini.ts` constrains the output with a **response schema** sent on every
|
||
request, with enums narrowed to that batch's source language, POS, and target languages. The
|
||
shape of the JSON is therefore enforced by the API, and `validate.ts` is left to enforce the
|
||
things a schema cannot express: cross-field difficulty ordering, gender rules per target
|
||
language, sequential `sense_index`, translation coverage and caps, and that the headword was
|
||
actually in the input batch.
|
||
|
||
Validation is the safety net, not the prompt — every entry is checked before it reaches SQLite, and rejects go to a log for review rather than silently disappearing. Target reject rate is under 10%.
|
||
|
||
> **Systemic rejection cause — resolved.** Effectively every rejection used to be
|
||
> `translation difficulty lower than sense difficulty`: the model tags a sense `medium`
|
||
> while correctly tagging some of its translations `easy`. The prompt defines sense
|
||
> difficulty as the easiest translation difficulty in the sense, which makes it a derived
|
||
> value, so `validate.ts` now floors it (`applySenseDifficultyFloor`) instead of rejecting
|
||
> the entry. The floor only ever lowers — raising a sense would gate a concept out of levels
|
||
> it belongs in and collapse the concept-vs-word split of design-doc §4. Replaying every
|
||
> saved response with the new rule took the reject rate from 20 entries to **zero**.
|
||
|
||
---
|
||
|
||
## Running it
|
||
|
||
```bash
|
||
docker compose up -d pipeline-database # dedicated PostgreSQL on :5433
|
||
pnpm --filter @lila/pipeline pipeline:run # tsx --env-file .env pipeline.ts
|
||
pnpm --filter @lila/pipeline test
|
||
```
|
||
|
||
`pipeline.ts` takes CLI flags (pass them after `--`, which the script strips itself):
|
||
|
||
| Flag | Default | Meaning |
|
||
| --------------- | ------- | -------------------------------------------------------- |
|
||
| `--langs` | all | Comma-separated source languages, e.g. `de,es` |
|
||
| `--pos` | `noun` | Part of speech to process |
|
||
| `--batch-size` | `20` | Words per Gemini request |
|
||
| `--max-batches` | all | Cap batches per list — useful for smoke tests |
|
||
| `--delay-ms` | `6000` | Pause between requests, for rate limiting |
|
||
| `--dry-run` | `false` | Render prompts and print batches without calling the API |
|
||
|
||
```bash
|
||
# smoke test: one German batch, no API calls
|
||
pnpm --filter @lila/pipeline pipeline:run -- --langs de --max-batches 1 --dry-run
|
||
```
|
||
|
||
Because runs are resumable, interrupting with Ctrl-C is safe — already-staged words are
|
||
skipped on the next run.
|
||
|
||
### Replaying saved responses
|
||
|
||
`replay.ts` re-validates everything in `responses/` **without calling the API**, which is how
|
||
a validation-rule change is applied to data that has already been generated.
|
||
|
||
```bash
|
||
pnpm --filter @lila/pipeline pipeline:replay # report only
|
||
pnpm --filter @lila/pipeline pipeline:replay -- --write # stage recovered entries
|
||
pnpm --filter @lila/pipeline pipeline:replay -- --verbose --langs de
|
||
```
|
||
|
||
It reports valid / normalized / empty / invalid counts and groups anything still invalid by
|
||
error. Staging is idempotent, so `--write` is safe to repeat. One limitation: it can only
|
||
_insert_. A word already in `staging.db` is left untouched, so a rule change cannot repair
|
||
rows that were staged under the old rules — only recover ones that were rejected.
|
||
|
||
### API quota
|
||
|
||
⚠️ **The free tier allows far fewer requests than the batch count needs.** Observed
|
||
2026-08-20: `gemini-3.6-flash` returned `HTTP 429 … limit: 20` on metric
|
||
`generate_content_free_tier_requests` after **17 successful requests in one day**, and did
|
||
not recover across ~55 minutes of retrying — a daily window, not a per-minute one. The
|
||
`retryDelay: ~59s` in the 429 payload is generic backoff advice and misleads the retry logic
|
||
into grinding.
|
||
|
||
Consequences to plan around:
|
||
|
||
- What is rationed is **requests**, not words, so `--batch-size` is the cheap lever. The
|
||
remaining ~7,300 words are ~365 requests at batch size 20, but only ~74 at batch size 100.
|
||
- `gemini.ts` currently treats 429 as retryable and burns all 5 `MAX_ATTEMPTS` against a
|
||
quota that will not clear for hours. It should distinguish rate-limiting from daily
|
||
exhaustion and abort the run.
|
||
- Replays cost nothing, which is why raw responses are persisted.
|
||
|
||
The pipeline reads `.env` from the repo root: `GEMINI_API_KEY`, plus `PIPELINE_POSTGRES_USER` / `PIPELINE_POSTGRES_PASSWORD` / `PIPELINE_POSTGRES_DB` / `PIPELINE_DATABASE_URL`. The pipeline database is deliberately separate from the app database (`:5432`) so pipeline work can never damage dev data.
|
||
|
||
---
|
||
|
||
## Phase status
|
||
|
||
Full breakdown in [pipeline/roadmap.md](pipeline/roadmap.md).
|
||
|
||
| Phase | State |
|
||
| ------------------------------------------------- | -------------------------------------------------------------- |
|
||
| 1 — Drizzle schema (words/senses/translations) | ✅ Complete |
|
||
| 2 — Preparation (wordlists, DBs, prompt) | ✅ Complete |
|
||
| 3 — Build the pipeline → SQLite | 🔄 **Current.** Code complete; full 5-language run in progress |
|
||
| 4 — Migration & SQLite → PostgreSQL import | ⬜ Not started |
|
||
| 5 — App integration (`getGameTerms`, distractors) | ⬜ Not started |
|
||
| 6 — Production deploy | ⬜ Not started |
|
||
| 7 — Extend to verbs, adjectives, adverbs | ⬜ Not started — new wordlists + prompt only, no schema change |
|
||
|
||
All Phase 3 code is written and unit-tested. What remains is the data work: finish the
|
||
5-language run, resolve the systemic rejection cause noted above, and hand-check 50 entries.
|
||
|
||
Target for Phase 3: the full deduped noun lists (~1,550–1,750 unique words per language after
|
||
dedup) × 5 languages in staging, reject rate under 10%, 50 entries spot-checked by hand.
|
||
|
||
**Known risk for Phase 5 — the `hard` tier is nearly empty.** Generated difficulty skews
|
||
heavily easy, with `hard` translations well under 1% of the corpus. The design-doc §5.1 game
|
||
query filters translation difficulty as an _exact_ match, so a "hard" game currently has far
|
||
too few rows to fill one round plus distractors. This needs prompt calibration before import,
|
||
not after.
|
||
|
||
Phase 5 is where this becomes visible in the app: `packages/db/src/models/termModel.ts` still queries `vocabulary_entries`/`entry_translations` and must be rewritten against the sense-based schema. Until then, production runs on the old data.
|