updating pipeline roadmap and status to match branch reality

Roadmap: phase 4 re-scoped to import only (migration already applied),
phase 3 gains prompt templating, wordlist dedup, structured output,
raw-response persistence, and validation tests; idempotency decision
documented (skip already-staged words, one transaction per word).
Stale paths, postgres setup, and file structure corrected.

STATUS.md: refreshed from stale 2026-05-15 Kaikki state to current
gemini-only pipeline work.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
lila 2026-08-09 13:32:15 +02:00
parent e534b98bc5
commit 303bb9388c
2 changed files with 242 additions and 213 deletions

View file

@ -1,6 +1,6 @@
# Status — 2026-05-15 # Status — 2026-08-09
> Last updated: 2026-05-15. Update this file after every deploy or when switching tasks. > Last updated: 2026-08-09. Update this file after every deploy or when switching tasks.
## What Works Today ✅ ## What Works Today ✅
@ -12,7 +12,7 @@
## What's Broken / Blocked 🚧 ## What's Broken / Blocked 🚧
- **Data quality** — Production still uses OpenWordNet/OMW translations. Kaikki pipeline (sense-disambiguated) is in progress but not yet synced to production. - **Data quality** — Production still uses OpenWordNet/OMW translations. The replacement is the Gemini-only pipeline on branch `refactor/gemini-only-pipeline`: the sense-based schema (`words` → `senses` → `translations`) is migrated but empty, and the pipeline itself is not built yet. The earlier Kaikki/local-LLM pipeline was abandoned and removed (docs in `documentation/archive/`).
- **Guest play** — Auth is required for all game routes. No try-before-signup flow. - **Guest play** — Auth is required for all game routes. No try-before-signup flow.
- **Game session store** — Still in-memory (`InMemoryGameSessionStore`). Valkey container exists in local dev but not wired up. - **Game session store** — Still in-memory (`InMemoryGameSessionStore`). Valkey container exists in local dev but not wired up.
- **Rate limiting** — Partially implemented on auth endpoints; game endpoints not yet covered. - **Rate limiting** — Partially implemented on auth endpoints; game endpoints not yet covered.
@ -21,26 +21,26 @@
## What I'm Working On Now 🔄 ## What I'm Working On Now 🔄
**Primary:** Rewriting the Kaikki data pipeline enrich script for sub-stage architecture (round1_gloss → round1_example → round1_translations → round1_cefr). **Primary:** Phase 3 of the Gemini-only data pipeline (see [pipeline/roadmap.md](pipeline/roadmap.md)): template the Gemini prompt, write the validation module + `pipeline.ts`, and run the first full staging pass (~1,600–1,800 deduped nouns × 5 languages into SQLite).
**Secondary:** Phase 7 hardening backlog items (see BACKLOG.md `next` section). **Secondary:** Phase 7 hardening backlog items (see BACKLOG.md `next` section).
## Next 2-Week Goal 🎯 ## Next 2-Week Goal 🎯
Finish Kaikki Stage 3 (enrich) sub-stage rewrite → run full sample → compare quality → decide on production sync timeline. Finish pipeline Phase 3 (full noun run staged in SQLite, reject rate < 10%, 50 entries spot-checked) → Phase 4 import script (pipeline Postgres :5433, then dev :5432) → start rewriting `getGameTerms`/`getDistractors` against the new schema.
## The Big Picture ## The Big Picture
Lila is a **deployed, working vocabulary quiz app**. The core loop (singleplayer + multiplayer) is solid. The next strategic milestone is **media-based practice** (learn vocab from a song/TV episode/book chapter), but that depends on: Lila is a **deployed, working vocabulary quiz app**. The core loop (singleplayer + multiplayer) is solid. The next strategic milestone is **media-based practice** (learn vocab from a song/TV episode/book chapter), but that depends on:
1. Kaikki data pipeline reaching production (fixes translation quality) 1. The Gemini data pipeline reaching production (fixes translation quality)
2. A media ingestion prototype (subtitles/lyrics → text → vocab extraction → quiz) 2. A media ingestion prototype (subtitles/lyrics → text → vocab extraction → quiz)
Until then, the app is a generic vocabulary quiz — functional but not differentiated. Until then, the app is a generic vocabulary quiz — functional but not differentiated.
## Quick Links ## Quick Links
- [BACKLOG.md](BACKLOG.md) — Prioritized tasks - [BACKLOG.md](BACKLOG.md) — Prioritized tasks (`now` / `next` / `later`)
- [DATA_PIPELINE.md](DATA_PIPELINE.md) — Pipeline stages and current progress - [DATA_PIPELINE.md](DATA_PIPELINE.md) — Pipeline orientation and current progress
- [BACKLOG.md](BACKLOG.md) — `now` / `next` / `later` - [pipeline/roadmap.md](pipeline/roadmap.md) — Phase-by-phase pipeline plan
- [DEPLOYMENT.md](DEPLOYMENT.md) — Infrastructure ops - [DEPLOYMENT.md](DEPLOYMENT.md) — Infrastructure ops

View file

@ -4,9 +4,9 @@
> Gemini-powered pipeline that generates high-quality vocabulary data, > Gemini-powered pipeline that generates high-quality vocabulary data,
> backed by a normalized Postgres schema. > backed by a normalized Postgres schema.
> >
> **Author:** [Your Name] > **Author:** lila
> **Date:** July 2026 > **Date:** July 2026 · Last reviewed: 2026-08-09
> **Companion doc:** `docs/schema-design.md` > **Companion doc:** `documentation/pipeline/design-doc.md`
--- ---
@ -28,18 +28,25 @@ This roadmap has three zoom levels:
``` ```
Phase 1 Schema ✅ Phase 1 Schema ✅
New Drizzle schema (words, senses, translations) with New Drizzle schema (words, senses, translations) with
relations, constraints, and indexes. Committed. relations, constraints, and indexes. Committed and applied
(migration 0012_graceful_psynapse.sql) — tables exist, empty.
Phase 2 Preparation Phase 2 Preparation ✅
Get wordlists, set up tooling, finalize the Gemini prompt. Wordlists acquired, databases set up, prompt drafted and
tested. Two loose ends carried into Phase 3: the prompt is
not templated yet, and the wordlists still contain duplicates
(deduped at runtime, not in the files).
Phase 3 Data Pipeline Phase 3 Data Pipeline ← CURRENT
Build the Gemini → validate → SQLite pipeline. Build the Gemini → validate → SQLite pipeline.
Produce a clean dataset of ~1000 nouns × 5 languages. Produce a clean dataset: full deduped noun lists
(~1,600–1,800 words) × 5 languages.
Phase 4 Migration & Import Phase 4 Import
Generate and apply the Drizzle migration. (Migration already applied in Phase 1.)
Write the SQLite → Postgres import script. Write the SQLite → Postgres import script.
Test it against the pipeline Postgres (:5433) first,
then load the app dev database (:5432).
Phase 5 App Integration Phase 5 App Integration
Rewrite the game and distractor queries against the new schema. Rewrite the game and distractor queries against the new schema.
@ -47,8 +54,8 @@ Phase 5 App Integration
Test the full game flow in dev. Test the full game flow in dev.
Phase 6 Production Deploy Phase 6 Production Deploy
Run the Drizzle migration on prod. Import the dataset into prod (schema arrives via the normal
Import the dataset. Drizzle migration flow).
Verify the live app works end-to-end. Verify the live app works end-to-end.
Phase 7 Extend POS Phase 7 Extend POS
@ -63,7 +70,7 @@ Phase 8 Future Features (out of scope for now)
**Dependency chain:** **Dependency chain:**
``` ```
Phase 1 ✅ → Phase 2 → Phase 3 → Phase 4 → Phase 5 → Phase 6 → Phase 7 Phase 1 ✅ → Phase 2 ✅ → Phase 3 → Phase 4 → Phase 5 → Phase 6 → Phase 7
│ │
▼ ▼
Phase 8 Phase 8
@ -94,46 +101,66 @@ constraints, indexes, and relations.
- [x] Old tables (`vocabulary_entries`, `entry_translations`) untouched - [x] Old tables (`vocabulary_entries`, `entry_translations`) untouched
- [x] Auth and lobby tables untouched - [x] Auth and lobby tables untouched
- [x] Build passes, committed - [x] Build passes, committed
- [x] Migration generated, inspected, and applied to local Postgres
(`packages/db/drizzle/0012_graceful_psynapse.sql`)
--- ---
## Phase 2: Preparation ## Phase 2: Preparation ✅ COMPLETE
**Goal:** Have everything you need before writing pipeline code. **Goal:** Have everything you need before writing pipeline code.
**Tasks:** **Tasks:**
- [x] Acquire frequency-based noun lists for all 5 languages - [x] Acquire frequency-based noun lists for all 5 languages
- [x] Clean and format the lists (one word per line, UTF-8) - [x] Format the lists (one word per line, UTF-8) — **note:** the files
- [x] Set up local Postgres (Docker or native) still contain 110–147 duplicate words each; dedup happens at
runtime in Phase 3, not in the files
- [x] Set up local Postgres (docker compose: app DB :5432, dedicated
pipeline DB :5433)
- [x] Set up a SQLite database file for staging - [x] Set up a SQLite database file for staging
- [x] Write and test the Gemini prompt with 5 sample words (`data-pipeline/db/staging.db` from `db/schema.sql`)
- [x] Write and test the Gemini prompt with sample words
- [x] Refine the prompt until the JSON output matches the contract - [x] Refine the prompt until the JSON output matches the contract
defined in `docs/schema-design.md` §6.3 defined in `design-doc.md` §6.3 — **note:** the prompt works but
is pinned to a hardcoded Spanish sample and has known copy-paste
bugs; templating and fixes are Phase 3 tasks
**Dependencies:** Phase 1 complete. **Dependencies:** Phase 1 complete.
**Acceptance criteria:** **Acceptance criteria (met):**
- You have 5 wordlist files in `data-pipeline/source-data/` (one per - 5 wordlist files exist in `data-pipeline/source-data/{lang}/noun`
language). (language codes `de/en/es/fr/it`, matching `@lila/shared` constants).
- A Gemini call with 5 German nouns returns valid JSON matching the - A Gemini call with a batch of nouns returns valid JSON matching the
contract, including definitions, examples, translations with gender, contract, including definitions, examples, translations with gender,
and difficulty levels. and difficulty levels.
- Local Postgres is running and reachable from your app. - Local Postgres is running and reachable from the app.
--- ---
## Phase 3: Data Pipeline ## Phase 3: Data Pipeline ← CURRENT
**Goal:** A repeatable script that takes a wordlist, calls Gemini in **Goal:** A repeatable script that takes a wordlist, calls Gemini in
batches of 20, validates the output, and writes clean rows to SQLite. batches of 20, validates the output, and writes clean rows to SQLite.
Re-runs skip words that are already staged.
**Tasks:** **Tasks:**
- [ ] Write the validation module (see schema-design §6.4) - [x] Create the SQLite schema (`data-pipeline/db/schema.sql`, tables
- [ ] Write the pipeline script: - Read wordlist file - Split into batches of 20 - Call Gemini API per batch - Parse JSON response - Validate each entry - Write valid entries to SQLite - Log invalid entries to a rejection file created in `db/staging.db`)
- [ ] Create the SQLite schema (mirrors the Postgres schema) - [ ] Template the Gemini prompt: source language, POS, target
languages, and the word batch become substitutions; fix the known
copy-paste bugs while doing so (rules 2–3 hardcode `"en"`,
rule 15's target list contradicts the header, rule 31 says
"valid English noun")
- [ ] Write the wordlist normalization step: trim whitespace, drop
empty lines, dedup in memory
- [ ] Write the validation module (rules in design-doc §6.4) with unit
tests (`vitest.config.ts` expects `tests/**/*.test.ts` — the
directory doesn't exist yet)
- [ ] Write the pipeline script (`pipeline.ts`): - Read + normalize wordlist files - Skip words already present in staging.db (idempotency — see §3.3) - Split the remainder into batches of 20 - Call Gemini per batch using structured output
(`responseMimeType: "application/json"` + `responseSchema`) - Persist each raw response to disk before validating - Validate each entry - Write valid entries to SQLite, one transaction per word - Log invalid entries to a rejection file
- [ ] Run the pipeline for all 5 languages (nouns only) - [ ] Run the pipeline for all 5 languages (nouns only)
- [ ] Review the rejection log, fix prompt issues, re-run failed batches - [ ] Review the rejection log, fix prompt issues, re-run failed batches
- [ ] Spot-check 50 random entries for correctness - [ ] Spot-check 50 random entries for correctness
@ -142,39 +169,46 @@ batches of 20, validates the output, and writes clean rows to SQLite.
**Acceptance criteria:** **Acceptance criteria:**
- SQLite database contains ~1000 nouns × 5 languages with senses and - SQLite database contains the full deduped noun lists
translations. (~1,600–1,800 words × 5 languages) with senses and translations.
- Rejection rate is below 10%. - Rejection rate is below 10%.
- Spot-checked entries have correct definitions, plausible examples, - Spot-checked entries have correct definitions, plausible examples,
correct genders, and reasonable difficulty levels. correct genders, and reasonable difficulty levels.
- The pipeline is re-runnable (idempotent or with duplicate handling). - Re-running the pipeline skips already-staged words (no duplicate
rows, no repeated API calls for the same words).
- Raw Gemini responses are on disk, so validation-rule changes can be
re-applied without re-calling the API.
--- ---
## Phase 4: Migration & Import ## Phase 4: Import
**Goal:** The new schema exists in Postgres and the SQLite data is **Goal:** The SQLite data is imported into Postgres.
imported.
The Drizzle migration was already generated, inspected, and applied in
Phase 1 (`0012_graceful_psynapse.sql`) — the `words`/`senses`/
`translations` tables exist and are empty. What remains is the import
script.
**Tasks:** **Tasks:**
- [ ] Generate the Drizzle migration (`npx drizzle-kit generate`) - [ ] Write the import script (`data-pipeline/import-to-postgres.ts`): - Read all rows from SQLite - Insert into Postgres in dependency order:
- [ ] Inspect the generated SQL file — verify it creates the right words → senses → translations - Use batch inserts (not row-by-row) via Drizzle — add `@lila/db`
tables, constraints, and indexes as a workspace dependency (the pipeline currently has no
- [ ] Apply the migration to local Postgres (`npx drizzle-kit migrate`) Postgres client) - Wrap in transactions (per batch of 20 words) - Handle duplicates gracefully (`ON CONFLICT DO NOTHING`)
- [ ] Write the import script (SQLite → Postgres): - Read all rows from SQLite - Insert into Postgres in dependency order: - [ ] Run it against the **pipeline Postgres (:5433,
words → senses → translations - Use batch inserts (not row-by-row) - Wrap in transactions (per batch of 20 words) - Handle duplicates gracefully (skip or upsert) `PIPELINE_DATABASE_URL`)** first — this database exists so a bad
- [ ] Run the import script import can never damage dev data
- [ ] Verify row counts match between SQLite and Postgres - [ ] Verify row counts match between SQLite and Postgres
- [ ] Run 3–5 manual SQL queries against Postgres to sanity-check - [ ] Run 3–5 manual SQL queries against Postgres to sanity-check
the data the data
- [ ] Once trusted, run it against the app dev database (:5432)
**Dependencies:** Phase 3 complete (SQLite has data). **Dependencies:** Phase 3 complete (SQLite has data).
**Acceptance criteria:** **Acceptance criteria:**
- `npx drizzle-kit migrate` runs without errors on local Postgres. - Import script completes without errors on :5433 and then :5432.
- Import script completes without errors.
- Row counts in Postgres match SQLite (±rejection count). - Row counts in Postgres match SQLite (±rejection count).
- Manual query: "Give me 5 random German nouns with Spanish - Manual query: "Give me 5 random German nouns with Spanish
translations at easy difficulty" returns sensible results. translations at easy difficulty" returns sensible results.
@ -191,7 +225,8 @@ distractors work correctly for all language pairs.
- [ ] Rewrite `getGameTerms` query: - JOIN words → senses → translations - Filter: source language, pos, sense difficulty (ceiling), - [ ] Rewrite `getGameTerms` query: - JOIN words → senses → translations - Filter: source language, pos, sense difficulty (ceiling),
target language, translation difficulty (exact) - ORDER BY RANDOM(), LIMIT rounds target language, translation difficulty (exact) - ORDER BY RANDOM(), LIMIT rounds
- [ ] Rewrite `getDistractors` query: - Same JOINs and filters - Exclude: `sense_id != current`, `translation != correct` - ORDER BY RANDOM(), LIMIT 3 - [ ] Rewrite `getDistractors` query: - Same JOINs and filters - Exclude: `sense_id != current`, `translation != correct` - ORDER BY RANDOM(), LIMIT 3
- [ ] Update the exercise-generation logic: - Pick one random definition from the `definitions` array - Pick one random example from the `examples` array - Assemble the 4 answer options (1 correct + 3 distractors) - Shuffle the options - [ ] Update the exercise-generation logic: - Pick one random definition from the `definitions` array - Pick one random example from the `examples` array - Assemble the 4 answer options (1 correct + 3 distractors) - Shuffle the options - Keep the existing invariant: the correct answer is evaluated
server-side and is never sent to the client
- [ ] Remove or deprecate old schema references - [ ] Remove or deprecate old schema references
(old `vocabulary_entries`, `entry_translations` tables) (old `vocabulary_entries`, `entry_translations` tables)
- [ ] Test manually: - German → Spanish, nouns, easy, 10 rounds - Spanish → German, nouns, medium, 10 rounds - English → French, nouns, hard, 10 rounds - Italian → German, nouns, easy, 10 rounds - Verify: no duplicate answers, no same-sense synonyms as - [ ] Test manually: - German → Spanish, nouns, easy, 10 rounds - Spanish → German, nouns, medium, 10 rounds - English → French, nouns, hard, 10 rounds - Italian → German, nouns, easy, 10 rounds - Verify: no duplicate answers, no same-sense synonyms as
@ -278,105 +313,128 @@ Listed here for visibility. Not planned, not estimated.
--- ---
## Phase 2: Preparation — Detailed ## Phase 2: Preparation — Detailed ✅
### 2.1 Acquire wordlists ### 2.1 Wordlists (done)
- Search for frequency lists. Good starting points: - Frequency-based lists live in `data-pipeline/source-data/{lang}/noun`
- "Leipzig Corpora Collection" (academic, per-language) with `lang` ∈ `de/en/es/fr/it` — the same codes as
- Wiktionary frequency lists `packages/shared/src/constants.ts`, so no name mapping is needed
- GitHub repos: search "german noun frequency list", anywhere in the pipeline.
"spanish noun frequency list", etc.
- Tatoeba sentence counts as a proxy for word frequency
- Target: ~1000 nouns per language, sorted by frequency.
- Format: plain text, one word per line, UTF-8, no headers. - Format: plain text, one word per line, UTF-8, no headers.
- Save to `data-pipeline/source-data/{language}/nouns/`. - Sizes: ~1,700–1,900 lines per language. The files were **not**
- Clean the lists: remove duplicates, remove words with spaces deduplicated (110–147 duplicates each); the pipeline dedups at read
(multi-word expressions), remove proper nouns if desired. time (§3.3) rather than editing the source files.
### 2.2 Set up local Postgres ### 2.2 Databases (done)
- Option A (Docker): - `docker compose up -d` starts everything: app Postgres (:5432),
``` dedicated pipeline Postgres (:5433), Valkey (:6379).
docker run --name vocab-dev \ - Config lives in the single root `.env` (see `.env.example`):
-e POSTGRES_USER=dev \ `PIPELINE_POSTGRES_USER` / `PIPELINE_POSTGRES_PASSWORD` /
-e POSTGRES_PASSWORD=dev \ `PIPELINE_POSTGRES_DB` / `PIPELINE_DATABASE_URL`, plus
-e POSTGRES_DB=vocab \ `GEMINI_API_KEY` for the pipeline itself.
-p 5432:5432 \ - The pipeline Postgres is deliberately separate from the app database
-d postgres:16 so pipeline work can never damage dev data.
``` - SQLite staging: `data-pipeline/db/staging.db`, created from
- Option B (native install): install Postgres, create a `vocab` `data-pipeline/db/schema.sql`.
database.
- Update your `.env` / `.env.local` with the connection string.
- Verify: connect with `psql` or a GUI client (TablePlus, DBeaver,
pgAdmin). Run `SELECT 1;`.
### 2.3 Write and test the Gemini prompt ### 2.3 The Gemini prompt (drafted; templating is Phase 3)
- Start with 5 German nouns. - `data-pipeline/prompt` is a plain UTF-8 text file, edited by hand.
- The prompt should specify: It is currently pinned to a concrete sample run (Spanish nouns,
- The exact JSON structure expected (copy from schema-design §6.3) 20 words inlined).
- The prompt specifies:
- The exact JSON structure expected (from design-doc §6.3)
- That definitions and examples must be in the word's language - That definitions and examples must be in the word's language
- That translations are needed for all 4 other supported languages - That translations are needed for all 4 other supported languages
- That gender must be provided for de/fr/es/it, null for en - That gender must be provided for de/fr/es/it, null for en
- That difficulty must be one of: easy, medium, hard - That difficulty must be one of: easy, medium, hard
- That multiple senses should be included for polysemous words - That multiple senses should be included for polysemous words
- Call the API. Inspect the JSON. Common issues to fix: - **Known bugs to fix when templating (§3.2):** rules 2 and 3 still say
- Gemini wraps the JSON in markdown code fences → strip them `language` must be `"en"`; rule 15 lists target languages
- Gemini returns "intermediate" instead of "medium" → normalize `de, it, es, fr` while the header says `en, it, de, fr`; rule 31
- Gemini omits a language → re-prompt or reject says "valid English noun". All are leftovers from adapting the
- Gemini returns gender "common" for German → reject English version.
- Iterate until 3 consecutive batches of 20 return clean JSON.
--- ---
## Phase 3: Data Pipeline — Detailed ## Phase 3: Data Pipeline — Detailed
### 3.1 Create the SQLite schema ### 3.1 SQLite schema (done)
- Create a file `data-pipeline/schema.sql` or define it in your - `data-pipeline/db/schema.sql` mirrors the Postgres schema with two
pipeline script. SQLite concessions: IDs are `TEXT` (`crypto.randomUUID()`), and
- Tables mirror the Postgres schema: `definitions` / `examples` are JSON-encoded strings because SQLite
has no array type. The import script parses them back into Postgres
`TEXT[]`.
- The UNIQUE constraints match Postgres:
`words(headword, language_code, pos)`, `senses(word_id, sense_index)`,
`translations(sense_id, target_language_code, translation)`.
```sql ### 3.2 Template the prompt
CREATE TABLE words (
id TEXT PRIMARY KEY,
headword TEXT NOT NULL,
language_code TEXT NOT NULL,
pos TEXT NOT NULL,
created_at TEXT DEFAULT (datetime('now')),
UNIQUE(headword, language_code, pos)
);
CREATE TABLE senses ( - Turn `data-pipeline/prompt` into a template. Substitution slots:
id TEXT PRIMARY KEY, - source language (name + code)
word_id TEXT NOT NULL REFERENCES words(id), - POS
sense_index INTEGER NOT NULL DEFAULT 0, - target languages (the other 4 codes)
difficulty TEXT NOT NULL, - the word batch (20 words, one per line)
definitions TEXT NOT NULL DEFAULT '[]', -- JSON array as text - Fix the hardcoded leftovers listed in §2.3 as part of this — after
examples TEXT NOT NULL DEFAULT '[]', -- JSON array as text templating, the language codes in the rules must derive from the
created_at TEXT DEFAULT (datetime('now')), substitutions, so this class of bug can't recur.
UNIQUE(word_id, sense_index) - Request **structured output** from the API
); (`responseMimeType: "application/json"` with a `responseSchema`
matching design-doc §6.3) instead of relying on prompt instructions
alone. Keep fence-stripping as a defensive fallback only.
CREATE TABLE translations ( ### 3.3 Write the pipeline script
id TEXT PRIMARY KEY,
sense_id TEXT NOT NULL REFERENCES senses(id), - Entry point: `data-pipeline/pipeline.ts`
target_language_code TEXT NOT NULL, (`pnpm --filter @lila/pipeline pipeline:run`).
translation TEXT NOT NULL, - Pseudocode:
gender TEXT,
difficulty TEXT NOT NULL, ```
created_at TEXT DEFAULT (datetime('now')), for each language in [de, en, es, fr, it]:
UNIQUE(sense_id, target_language_code, translation) read wordlist file → trim, drop empties, dedup in memory
); query staging.db for existing (headword, language_code, pos)
skip words already staged ← idempotency
split the remainder into batches of 20
for each batch:
call Gemini API (structured output)
write the raw response to responses/{lang}-{pos}-{n}.json
parse JSON response
for each entry in response:
validate(entry)
if valid:
generate UUIDs for word, senses, translations
INSERT word + senses + translations in ONE transaction
else:
append to rejection log
log progress: "Batch 12/86 done. 238 words staged, 2 rejected."
sleep 1s (rate limiting); retry with backoff on API errors
``` ```
- Note: SQLite has no native TEXT[]. Store arrays as JSON text. - **Idempotency decision (resolves the pipeline.ts step 3 question):**
The import script will parse them into Postgres TEXT[] on import. wordlists are read fresh on every run; the staging DB is the record
of what's been processed. Words already present in `staging.db` are
skipped before batching, so re-runs cost no API calls for staged
words. Each validated entry (word + its senses + their translations)
is written in a single transaction, so a partially-written word can
never exist and no NOT NULL constraint needs relaxing. The UNIQUE
constraints remain as a backstop (`INSERT OR IGNORE`).
- Persisting raw responses means a validation-rule change re-validates
from disk instead of re-paying for ~430 API calls
(~8,700 words ÷ 20 per batch).
- Use `better-sqlite3` for SQLite access (synchronous, simple).
- Generate UUIDs with `crypto.randomUUID()`.
- Store definitions/examples as JSON strings in SQLite
(`JSON.stringify(arr)`).
### 3.2 Write the validation module ### 3.4 Validation module
- Create `data-pipeline/validate.ts`. - Create the validation module alongside the pipeline, with unit tests
(vitest is already configured; `data-pipeline/vitest.config.ts`
expects `tests/**/*.test.ts`).
- Input: one parsed Gemini entry (the JSON object for one word). - Input: one parsed Gemini entry (the JSON object for one word).
- Checks (return a list of errors, empty = valid): - Checks (return a list of errors, empty = valid):
- `headword` is a non-empty string - `headword` is a non-empty string
@ -394,43 +452,18 @@ Listed here for visibility. Not planned, not estimated.
- `target_language` != the word's own language - `target_language` != the word's own language
- `word` is a non-empty string - `word` is a non-empty string
- `gender` is valid for the target language - `gender` is valid for the target language
(de: masculine/feminine/neuter/null, (de: masculine/feminine/neuter,
fr/es/it: masculine/feminine/null, fr/es/it: masculine/feminine,
en: null) en: null)
- `difficulty` is in ['easy','medium','hard'] - `difficulty` is in ['easy','medium','hard']
- Output: `{ valid: boolean, errors: string[] }` - Output: `{ valid: boolean, errors: string[] }`
### 3.3 Write the pipeline script ### 3.5 Run and review
- Create `data-pipeline/run.ts`.
- Pseudocode:
```
for each language in [de, en, es, fr, it]:
read wordlist file → array of words
split into batches of 20
for each batch:
call Gemini API with the batch
parse JSON response (strip markdown fences if present)
for each entry in response:
validate(entry)
if valid:
generate UUIDs for word, senses, translations
INSERT into SQLite (words, senses, translations)
else:
append to rejection log
log progress: "Batch 12/50 done. 238 words imported, 2 rejected."
sleep 1s (rate limiting)
```
- Use `better-sqlite3` for SQLite access (synchronous, simple).
- Generate UUIDs with `crypto.randomUUID()`.
- Store definitions/examples as JSON strings in SQLite
(`JSON.stringify(arr)`).
### 3.4 Run and review
- Run the pipeline for all 5 languages. - Run the pipeline for all 5 languages.
- Check the rejection log. If rejection rate > 10%, fix the prompt - Check the rejection log. If rejection rate > 10%, fix the prompt
and re-run failed batches. and re-run — the skip-processed check means only rejected/missing
words are re-sent.
- Spot-check: open the SQLite DB, run: - Spot-check: open the SQLite DB, run:
```sql ```sql
SELECT w.headword, s.definitions, t.translation, t.gender SELECT w.headword, s.definitions, t.translation, t.gender
@ -446,50 +479,36 @@ Listed here for visibility. Not planned, not estimated.
--- ---
## Phase 4: Migration & Import — Detailed ## Phase 4: Import — Detailed
### 4.1 Generate and inspect the migration ### 4.1 Migration (done in Phase 1)
- Run: `npx drizzle-kit generate` `packages/db/drizzle/0012_graceful_psynapse.sql` creates the three
- Open the generated SQL file in `packages/db/drizzle/`. tables with all CHECK/UNIQUE constraints, the three indexes, and
- Read it. Verify: cascading FKs. It is applied locally; the tables exist and are empty.
- Three CREATE TABLE statements (words, senses, translations) Nothing to do here — prod gets the same migration in Phase 6.
- CHECK constraints match your schema
- UNIQUE constraints are present
- Three CREATE INDEX statements
- Foreign keys reference the correct tables with ON DELETE CASCADE
- If something looks wrong, fix the schema file and regenerate.
### 4.2 Apply the migration locally ### 4.2 Write the import script
- Run: `npx drizzle-kit migrate`
- Connect to local Postgres and verify:
```sql
\dt -- list tables
\d words -- describe words table
\d senses -- describe senses table
\d translations -- describe translations table
```
### 4.3 Write the import script
- Create `data-pipeline/import-to-postgres.ts`. - Create `data-pipeline/import-to-postgres.ts`.
- Add `@lila/db` as a workspace dependency for the Drizzle client and
schema (the pipeline currently ships only `better-sqlite3`).
- Pseudocode: - Pseudocode:
``` ```
open SQLite database (read-only) open SQLite database (read-only)
connect to Postgres via Drizzle connect to Postgres via Drizzle (PIPELINE_DATABASE_URL first)
read all words from SQLite read all words from SQLite
for each batch of 100 words: for each batch of words:
begin transaction begin transaction
INSERT words into Postgres INSERT words into Postgres
for each word: for each word:
read its senses from SQLite read its senses from SQLite
parse definitions/examples from JSON string → TEXT[]
INSERT senses into Postgres INSERT senses into Postgres
for each sense: for each sense:
read its translations from SQLite read its translations from SQLite
parse definitions/examples from JSON string → TEXT[]
INSERT translations into Postgres INSERT translations into Postgres
commit transaction commit transaction
log progress log progress
@ -501,9 +520,10 @@ Listed here for visibility. Not planned, not estimated.
into actual arrays (Postgres TEXT[]). into actual arrays (Postgres TEXT[]).
- Use `ON CONFLICT DO NOTHING` to handle re-runs gracefully. - Use `ON CONFLICT DO NOTHING` to handle re-runs gracefully.
### 4.4 Run and verify ### 4.3 Run and verify
- Run the import script. - Run against the pipeline Postgres (:5433) first; once the script is
trusted, point it at the app dev database (:5432).
- Compare counts: - Compare counts:
```sql ```sql
@ -574,6 +594,8 @@ Listed here for visibility. Not planned, not estimated.
- Combine 1 correct translation + 3 distractors. - Combine 1 correct translation + 3 distractors.
- Shuffle the 4 options (Fisher-Yates or similar). - Shuffle the 4 options (Fisher-Yates or similar).
- Attach gender to each option for display. - Attach gender to each option for display.
- Keep the server-side evaluation invariant: the correct answer is
never included in what is sent to the client.
### 5.4 Test matrix ### 5.4 Test matrix
@ -651,8 +673,9 @@ If something goes wrong:
Repeat the Phase 3 pipeline for each new POS: Repeat the Phase 3 pipeline for each new POS:
- Acquire wordlists (verbs, adjectives, adverbs) for all 5 languages. - Acquire wordlists (verbs, adjectives, adverbs) for all 5 languages,
- Adjust the Gemini prompt: saved as `source-data/{lang}/{pos}` with the shared POS codes.
- Adjust the Gemini prompt template:
- Verbs: ask for transitivity, common prepositions, or other - Verbs: ask for transitivity, common prepositions, or other
verb-specific metadata if needed for future exercises. verb-specific metadata if needed for future exercises.
- Adjectives: ask for the base form. Note that adjective metadata - Adjectives: ask for the base form. Note that adjective metadata
@ -673,41 +696,47 @@ English words. No migration needed.
# Risks & Mitigations # Risks & Mitigations
| Risk | Likelihood | Impact | Mitigation | | Risk | Likelihood | Impact | Mitigation |
| ------------------------------------------------------ | ----------------- | ------------------------- | ------------------------------------------------------------------------------------------------- | | ------------------------------------------------------ | ----------------- | ------------------------- | ----------------------------------------------------------------------------------------------------- |
| Gemini returns malformed JSON | Medium | Pipeline stalls | Strip markdown fences, wrap parsing in try/catch, log and skip bad batches | | Gemini returns malformed JSON | Low | Pipeline stalls | Structured output (responseSchema); fence-stripping + try/catch as fallback; log and skip bad batches |
| Validation rules change after a full run | Medium | Wasted API spend | Raw responses persisted per batch; re-validate from disk instead of re-calling the API |
| Gemini produces wrong genders or difficulties | Medium | Bad exercise data | Validation module rejects invalid entries. Spot-check 50+ entries per language | | Gemini produces wrong genders or difficulties | Medium | Bad exercise data | Validation module rejects invalid entries. Spot-check 50+ entries per language |
| Too few words at a given difficulty/pos/language combo | Medium | Game can't fill 4 options | Graceful fallback: reduce options or show a "not enough words" message. Log which combos are thin | | Too few words at a given difficulty/pos/language combo | Medium | Game can't fill 4 options | Graceful fallback: reduce options or show a "not enough words" message. Log which combos are thin |
| SQLite → Postgres import fails mid-way | Low | Partial data in Postgres | Transactions per batch. Re-run with ON CONFLICT DO NOTHING | | SQLite → Postgres import fails mid-way | Low | Partial data in Postgres | Transactions per batch. Re-run with ON CONFLICT DO NOTHING. Test on :5433 before touching :5432 |
| ORDER BY RANDOM() gets slow at scale | Low (not at 500k) | Slow game load | Add a TODO comment. Optimize with random() pre-filter when needed | | ORDER BY RANDOM() gets slow at scale | Low (not at 500k) | Slow game load | Add a TODO comment. Optimize with random() pre-filter when needed |
| Prod migration breaks the live app | Low | Downtime | Backup before migrating. Rollback plan documented. Deploy during low-traffic hours | | Prod migration breaks the live app | Low | Downtime | Backup before migrating. Rollback plan documented. Deploy during low-traffic hours |
--- ---
# File / Folder Structure (new files) # File / Folder Structure
``` ```
project/ project/
├── data-pipeline/ ├── data-pipeline/
│ ├── source-data/ │ ├── source-data/
│ │ ├── german/nouns/ ← wordlist files │ │ ├── de/noun ← wordlist files (shared lang/POS codes)
│ │ ├── english/nouns/ │ │ ├── en/noun
│ │ ├── spanish/nouns/ │ │ ├── es/noun
│ │ ├── french/nouns/ │ │ ├── fr/noun
│ │ └── italian/nouns/ │ │ └── it/noun
│ ├── rejections/ ← invalid Gemini entries (for review) │ ├── db/
│ ├── staging.db ← SQLite staging database │ │ ├── schema.sql ← SQLite schema definition ✅
│ ├── schema.sql ← SQLite schema definition │ │ └── staging.db ← SQLite staging database (gitignored) ✅
│ ├── validate.ts ← validation module │ ├── prompt ← Gemini prompt (plain UTF-8; to be templated)
│ ├── run.ts ← main pipeline script │ ├── responses/ ← raw Gemini responses, one file per batch (planned)
│ └── import-to-postgres.ts ← SQLite → Postgres import │ ├── rejections/ ← invalid Gemini entries for review (planned)
│ ├── pipeline.ts ← main pipeline script (pseudocode today)
│ ├── tests/ ← vitest unit tests, esp. validation (planned)
│ └── import-to-postgres.ts ← SQLite → Postgres import (planned)
├── packages/ ├── packages/
│ ├── db/ │ ├── db/
│ │ └── src/db/schema.ts ← Drizzle schema (updated ✅) │ │ ├── src/db/schema.ts ← Drizzle schema ✅
│ │ └── drizzle/0012_*.sql ← words/senses/translations migration ✅
│ └── shared/ │ └── shared/
│ └── src/constants.ts ← NOUN_GENDERS added ✅, "medium" ✅ │ └── src/constants.ts ← NOUN_GENDERS ✅, "medium" ✅
├── documentation/ ├── documentation/
│ ├── DATA_PIPELINE.md ← orientation layer
│ └── pipeline/ │ └── pipeline/
│ ├── design-doc.md ← schema design doc (updated ✅) │ ├── design-doc.md ← schema design doc ✅
│ └── roadmap.md ← this document (updated ✅) │ └── roadmap.md ← this document
└── ... └── ...
``` ```