# Vocabulary Trainer — Roadmap: Data Pipeline & Schema Migration > **Objective:** Replace the OpenWordNet/kaikki data source with a > Gemini-powered pipeline that generates high-quality vocabulary data, > backed by a normalized Postgres schema. > > **Author:** [Your Name] > **Date:** July 2026 > **Companion doc:** `docs/schema-design.md` --- ## How to Read This Document This roadmap has three zoom levels: - **Part 1 — Overview:** The phases at a glance. Read this to understand the full scope in 30 seconds. - **Part 2 — Phase Breakdown:** Goals, tasks, dependencies, and acceptance criteria per phase. Read this to plan your week. - **Part 3 — Detailed Tasks:** Step-by-step instructions within each phase. Read this when you sit down to code. --- # Part 1 — High-Level Overview ``` Phase 1 Schema ✅ New Drizzle schema (words, senses, translations) with relations, constraints, and indexes. Committed. Phase 2 Preparation Get wordlists, set up tooling, finalize the Gemini prompt. Phase 3 Data Pipeline Build the Gemini → validate → SQLite pipeline. Produce a clean dataset of ~1000 nouns × 5 languages. Phase 4 Migration & Import Generate and apply the Drizzle migration. Write the SQLite → Postgres import script. Phase 5 App Integration Rewrite the game and distractor queries against the new schema. Update the exercise-generation logic. Test the full game flow in dev. Phase 6 Production Deploy Run the Drizzle migration on prod. Import the dataset. Verify the live app works end-to-end. Phase 7 Extend POS Run the pipeline for verbs, adjectives, adverbs. No schema changes needed — new wordlists + adjusted prompts. Phase 8 Future Features (out of scope for now) Inflection tables, conjugation/declension exercises, gender exercises, spaced-repetition scheduling. ``` **Dependency chain:** ``` Phase 1 ✅ → Phase 2 → Phase 3 → Phase 4 → Phase 5 → Phase 6 → Phase 7 │ ▼ Phase 8 ``` --- # Part 2 — Phase Breakdown --- ## Phase 1: Schema ✅ COMPLETE **Goal:** New normalized schema exists in Drizzle with all tables, constraints, indexes, and relations. **Completed tasks:** - [x] `words` table: headword, language_code, pos, UNIQUE, CHECKs, index - [x] `senses` table: word_id FK (cascade), sense_index, difficulty, definitions TEXT[], examples TEXT[], UNIQUE, CHECK, index - [x] `translations` table: sense_id FK (cascade), target_language_code, translation, gender (nullable), difficulty, UNIQUE, CHECKs, index - [x] Relations: words→senses (many), senses→word (one) + translations (many), translations→sense (one) - [x] `NOUN_GENDERS` constant added to `@lila/shared` - [x] `DIFFICULTY_LEVELS` updated: "intermediate" → "medium" - [x] Old tables (`vocabulary_entries`, `entry_translations`) untouched - [x] Auth and lobby tables untouched - [x] Build passes, committed --- ## Phase 2: Preparation **Goal:** Have everything you need before writing pipeline code. **Tasks:** - [x] Acquire frequency-based noun lists for all 5 languages - [x] Clean and format the lists (one word per line, UTF-8) - [x] Set up local Postgres (Docker or native) - [ ] Set up a SQLite database file for staging - [ ] Write and test the Gemini prompt with 5 sample words - [ ] Refine the prompt until the JSON output matches the contract defined in `docs/schema-design.md` §6.3 **Dependencies:** Phase 1 complete. **Acceptance criteria:** - You have 5 wordlist files in `data-pipeline/source-data/` (one per language). - A Gemini call with 5 German nouns returns valid JSON matching the contract, including definitions, examples, translations with gender, and difficulty levels. - Local Postgres is running and reachable from your app. --- ## Phase 3: Data Pipeline **Goal:** A repeatable script that takes a wordlist, calls Gemini in batches of 20, validates the output, and writes clean rows to SQLite. **Tasks:** - [ ] Write the validation module (see schema-design §6.4) - [ ] Write the pipeline script: - Read wordlist file - Split into batches of 20 - Call Gemini API per batch - Parse JSON response - Validate each entry - Write valid entries to SQLite - Log invalid entries to a rejection file - [ ] Create the SQLite schema (mirrors the Postgres schema) - [ ] Run the pipeline for all 5 languages (nouns only) - [ ] Review the rejection log, fix prompt issues, re-run failed batches - [ ] Spot-check 50 random entries for correctness **Dependencies:** Phase 2 complete. **Acceptance criteria:** - SQLite database contains ~1000 nouns × 5 languages with senses and translations. - Rejection rate is below 10%. - Spot-checked entries have correct definitions, plausible examples, correct genders, and reasonable difficulty levels. - The pipeline is re-runnable (idempotent or with duplicate handling). --- ## Phase 4: Migration & Import **Goal:** The new schema exists in Postgres and the SQLite data is imported. **Tasks:** - [ ] Generate the Drizzle migration (`npx drizzle-kit generate`) - [ ] Inspect the generated SQL file — verify it creates the right tables, constraints, and indexes - [ ] Apply the migration to local Postgres (`npx drizzle-kit migrate`) - [ ] Write the import script (SQLite → Postgres): - Read all rows from SQLite - Insert into Postgres in dependency order: words → senses → translations - Use batch inserts (not row-by-row) - Wrap in transactions (per batch of 20 words) - Handle duplicates gracefully (skip or upsert) - [ ] Run the import script - [ ] Verify row counts match between SQLite and Postgres - [ ] Run 3–5 manual SQL queries against Postgres to sanity-check the data **Dependencies:** Phase 3 complete (SQLite has data). **Acceptance criteria:** - `npx drizzle-kit migrate` runs without errors on local Postgres. - Import script completes without errors. - Row counts in Postgres match SQLite (±rejection count). - Manual query: "Give me 5 random German nouns with Spanish translations at easy difficulty" returns sensible results. --- ## Phase 5: App Integration **Goal:** The running app uses the new schema. Game rounds and distractors work correctly for all language pairs. **Tasks:** - [ ] Rewrite `getGameTerms` query: - JOIN words → senses → translations - Filter: source language, pos, sense difficulty (ceiling), target language, translation difficulty (exact) - ORDER BY RANDOM(), LIMIT rounds - [ ] Rewrite `getDistractors` query: - Same JOINs and filters - Exclude: `sense_id != current`, `translation != correct` - ORDER BY RANDOM(), LIMIT 3 - [ ] Update the exercise-generation logic: - Pick one random definition from the `definitions` array - Pick one random example from the `examples` array - Assemble the 4 answer options (1 correct + 3 distractors) - Shuffle the options - [ ] Remove or deprecate old schema references (old `vocabulary_entries`, `entry_translations` tables) - [ ] Test manually: - German → Spanish, nouns, easy, 10 rounds - Spanish → German, nouns, medium, 10 rounds - English → French, nouns, hard, 10 rounds - Italian → German, nouns, easy, 10 rounds - Verify: no duplicate answers, no same-sense synonyms as distractors, definitions and examples are in the source language, genders display correctly - [ ] Test edge cases: - A difficulty/pos/language combo with very few words (does the app handle < 4 available words gracefully?) - A word with multiple senses (does the correct sense appear?) **Dependencies:** Phase 4 complete (Postgres has data, schema exists). **Acceptance criteria:** - Full game flow works in dev for at least 4 different language pairs. - No same-sense synonyms appear as distractors. - Definitions and examples are in the correct language. - Gender is displayed where applicable. - No console errors or unhandled query failures. --- ## Phase 6: Production Deploy **Goal:** The live app runs on the new schema with the new data. **Tasks:** - [ ] Back up the production database - [ ] Run the Drizzle migration on prod - [ ] Run the import script against prod Postgres - [ ] Verify row counts on prod - [ ] Test the live app: - Play 2 full games on the deployed app - Check different language pairs and difficulties - [ ] Monitor for errors (server logs, browser console) for 24h - [ ] Remove old tables from prod (after confirming everything works) **Dependencies:** Phase 5 complete (dev is fully working). **Acceptance criteria:** - Live app serves game rounds from the new schema. - No errors in server logs for 24 hours. - Old tables are dropped (or scheduled for removal). - Rollback plan exists: the prod backup can be restored if needed. --- ## Phase 7: Extend POS **Goal:** Verbs, adjectives, and adverbs are in the database and usable in the app. **Tasks:** - [ ] Acquire frequency lists for verbs, adjectives, adverbs (all 5 languages) - [ ] Adjust the Gemini prompt per POS: - Verbs: may need different metadata (transitivity, etc.) - Adjectives: may need base form info - Adverbs: typically simpler metadata - [ ] Run the pipeline for each POS - [ ] Validate, import to Postgres (dev → prod) - [ ] Test game flow with verbs, adjectives, adverbs - [ ] Verify the POS filter in the app UI works for all types **Dependencies:** Phase 6 complete. **Acceptance criteria:** - All 4 POS types are playable in the app. - Pipeline is repeatable for future data additions. --- ## Phase 8: Future Features (out of scope) Listed here for visibility. Not planned, not estimated. - [ ] `inflection_forms` table + conjugation exercises (verbs) - [ ] Adjective declension exercises (der grüne Mann, grüner Mann…) - [ ] Gender exercises (pick the correct article) - [ ] Spaced-repetition scheduling (track which words the user knows) - [ ] User accounts and progress persistence - [ ] Additional languages (if ever) --- # Part 3 — Detailed Task Breakdown --- ## Phase 2: Preparation — Detailed ### 2.1 Acquire wordlists - Search for frequency lists. Good starting points: - "Leipzig Corpora Collection" (academic, per-language) - Wiktionary frequency lists - GitHub repos: search "german noun frequency list", "spanish noun frequency list", etc. - Tatoeba sentence counts as a proxy for word frequency - Target: ~1000 nouns per language, sorted by frequency. - Format: plain text, one word per line, UTF-8, no headers. - Save to `data-pipeline/source-data/{language}/nouns/`. - Clean the lists: remove duplicates, remove words with spaces (multi-word expressions), remove proper nouns if desired. ### 2.2 Set up local Postgres - Option A (Docker): ``` docker run --name vocab-dev \ -e POSTGRES_USER=dev \ -e POSTGRES_PASSWORD=dev \ -e POSTGRES_DB=vocab \ -p 5432:5432 \ -d postgres:16 ``` - Option B (native install): install Postgres, create a `vocab` database. - Update your `.env` / `.env.local` with the connection string. - Verify: connect with `psql` or a GUI client (TablePlus, DBeaver, pgAdmin). Run `SELECT 1;`. ### 2.3 Write and test the Gemini prompt - Start with 5 German nouns. - The prompt should specify: - The exact JSON structure expected (copy from schema-design §6.3) - That definitions and examples must be in the word's language - That translations are needed for all 4 other supported languages - That gender must be provided for de/fr/es/it, null for en - That difficulty must be one of: easy, medium, hard - That multiple senses should be included for polysemous words - Call the API. Inspect the JSON. Common issues to fix: - Gemini wraps the JSON in markdown code fences → strip them - Gemini returns "intermediate" instead of "medium" → normalize - Gemini omits a language → re-prompt or reject - Gemini returns gender "common" for German → reject - Iterate until 3 consecutive batches of 20 return clean JSON. --- ## Phase 3: Data Pipeline — Detailed ### 3.1 Create the SQLite schema - Create a file `data-pipeline/schema.sql` or define it in your pipeline script. - Tables mirror the Postgres schema: ```sql CREATE TABLE words ( id TEXT PRIMARY KEY, headword TEXT NOT NULL, language_code TEXT NOT NULL, pos TEXT NOT NULL, created_at TEXT DEFAULT (datetime('now')), UNIQUE(headword, language_code, pos) ); CREATE TABLE senses ( id TEXT PRIMARY KEY, word_id TEXT NOT NULL REFERENCES words(id), sense_index INTEGER NOT NULL DEFAULT 0, difficulty TEXT NOT NULL, definitions TEXT NOT NULL DEFAULT '[]', -- JSON array as text examples TEXT NOT NULL DEFAULT '[]', -- JSON array as text created_at TEXT DEFAULT (datetime('now')), UNIQUE(word_id, sense_index) ); CREATE TABLE translations ( id TEXT PRIMARY KEY, sense_id TEXT NOT NULL REFERENCES senses(id), target_language_code TEXT NOT NULL, translation TEXT NOT NULL, gender TEXT, difficulty TEXT NOT NULL, created_at TEXT DEFAULT (datetime('now')), UNIQUE(sense_id, target_language_code, translation) ); ``` - Note: SQLite has no native TEXT[]. Store arrays as JSON text. The import script will parse them into Postgres TEXT[] on import. ### 3.2 Write the validation module - Create `data-pipeline/validate.ts`. - Input: one parsed Gemini entry (the JSON object for one word). - Checks (return a list of errors, empty = valid): - `headword` is a non-empty string - `language` is in ['en','de','it','fr','es'] - `pos` is in ['noun','verb','adjective','adverb'] - `senses` is a non-empty array - Each sense has: - `sense_index` is a non-negative integer - `difficulty` is in ['easy','medium','hard'] - `definitions` is a non-empty array of non-empty strings - `examples` is a non-empty array of non-empty strings - `translations` is a non-empty array - Each translation has: - `target_language` is in the supported list - `target_language` != the word's own language - `word` is a non-empty string - `gender` is valid for the target language (de: masculine/feminine/neuter/null, fr/es/it: masculine/feminine/null, en: null) - `difficulty` is in ['easy','medium','hard'] - Output: `{ valid: boolean, errors: string[] }` ### 3.3 Write the pipeline script - Create `data-pipeline/run.ts`. - Pseudocode: ``` for each language in [de, en, es, fr, it]: read wordlist file → array of words split into batches of 20 for each batch: call Gemini API with the batch parse JSON response (strip markdown fences if present) for each entry in response: validate(entry) if valid: generate UUIDs for word, senses, translations INSERT into SQLite (words, senses, translations) else: append to rejection log log progress: "Batch 12/50 done. 238 words imported, 2 rejected." sleep 1s (rate limiting) ``` - Use `better-sqlite3` for SQLite access (synchronous, simple). - Generate UUIDs with `crypto.randomUUID()`. - Store definitions/examples as JSON strings in SQLite (`JSON.stringify(arr)`). ### 3.4 Run and review - Run the pipeline for all 5 languages. - Check the rejection log. If rejection rate > 10%, fix the prompt and re-run failed batches. - Spot-check: open the SQLite DB, run: ```sql SELECT w.headword, s.definitions, t.translation, t.gender FROM words w JOIN senses s ON s.word_id = w.id JOIN translations t ON t.sense_id = s.id WHERE w.language_code = 'de' AND w.pos = 'noun' ORDER BY RANDOM() LIMIT 20; ``` - Read the definitions. Are they in German? Do they make sense? Are the genders correct? Are the difficulties reasonable? --- ## Phase 4: Migration & Import — Detailed ### 4.1 Generate and inspect the migration - Run: `npx drizzle-kit generate` - Open the generated SQL file in `packages/db/drizzle/`. - Read it. Verify: - Three CREATE TABLE statements (words, senses, translations) - CHECK constraints match your schema - UNIQUE constraints are present - Three CREATE INDEX statements - Foreign keys reference the correct tables with ON DELETE CASCADE - If something looks wrong, fix the schema file and regenerate. ### 4.2 Apply the migration locally - Run: `npx drizzle-kit migrate` - Connect to local Postgres and verify: ```sql \dt -- list tables \d words -- describe words table \d senses -- describe senses table \d translations -- describe translations table ``` ### 4.3 Write the import script - Create `data-pipeline/import-to-postgres.ts`. - Pseudocode: ``` open SQLite database (read-only) connect to Postgres via Drizzle read all words from SQLite for each batch of 100 words: begin transaction INSERT words into Postgres for each word: read its senses from SQLite INSERT senses into Postgres for each sense: read its translations from SQLite parse definitions/examples from JSON string → TEXT[] INSERT translations into Postgres commit transaction log progress log final counts: words, senses, translations ``` - Parse `definitions` and `examples` from JSON strings (SQLite) into actual arrays (Postgres TEXT[]). - Use `ON CONFLICT DO NOTHING` to handle re-runs gracefully. ### 4.4 Run and verify - Run the import script. - Compare counts: ```sql -- In SQLite SELECT COUNT(*) FROM words; SELECT COUNT(*) FROM senses; SELECT COUNT(*) FROM translations; -- In Postgres SELECT COUNT(*) FROM words; SELECT COUNT(*) FROM senses; SELECT COUNT(*) FROM translations; ``` - Counts should match (minus any rows that failed validation). - Run the game query manually in Postgres: ```sql SELECT w.headword, s.definitions, s.examples, t.translation, t.gender FROM words w JOIN senses s ON s.word_id = w.id JOIN translations t ON t.sense_id = s.id WHERE w.language_code = 'de' AND w.pos = 'noun' AND s.difficulty IN ('easy', 'medium') AND t.target_language_code = 'es' AND t.difficulty = 'medium' ORDER BY RANDOM() LIMIT 5; ``` - Verify the results make sense. --- ## Phase 5: App Integration — Detailed ### 5.1 Rewrite getGameTerms - Open `packages/db/src/models/termModel.ts`. - Replace the old query (2-table join on `vocabulary_entries` + `entry_translations`) with the new 3-table join (words → senses → translations). - Parameters: sourceLanguage, targetLanguage, pos, difficulty, rounds. - Difficulty filter: - `senses.difficulty IN (all levels up to and including selected)` — ceiling logic - `translations.difficulty = selected` — exact match - Return: word_id, headword, sense_id, definitions, examples, translation, gender. ### 5.2 Rewrite getDistractors - Same 3-table join. - Additional filters: - `t.sense_id != :currentSenseId` - `t.translation != :correctAnswer` - LIMIT 3. - If fewer than 3 distractors are found (small data pool), handle gracefully: reduce the number of options or log a warning. ### 5.3 Update exercise generation - In the function that assembles a game round: - Pick one random definition: `definitions[Math.floor(Math.random() * definitions.length)]` - Pick one random example: `examples[Math.floor(Math.random() * examples.length)]` - Combine 1 correct translation + 3 distractors. - Shuffle the 4 options (Fisher-Yates or similar). - Attach gender to each option for display. ### 5.4 Test matrix Run through this matrix manually in the dev app: | Source | Target | POS | Difficulty | Rounds | Pass? | | ------ | ------ | ---- | ---------- | ------ | ----- | | de | es | noun | easy | 10 | | | es | de | noun | medium | 10 | | | en | fr | noun | hard | 10 | | | it | de | noun | easy | 10 | | | fr | en | noun | medium | 10 | | For each: verify definitions are in the source language, translations are in the target language, genders are shown, no same-sense synonyms appear as distractors, no duplicate options. --- ## Phase 6: Production Deploy — Detailed ### 6.1 Pre-deploy checklist - [ ] All Phase 5 acceptance criteria pass - [ ] `git status` is clean, all changes committed - [ ] The Drizzle migration file is committed to the repo - [ ] You know your prod database connection string - [ ] You have a backup method for prod (pg_dump, hosting provider snapshot, etc.) ### 6.2 Deploy - Back up prod: ``` pg_dump -h -U -d > backup_$(date +%Y%m%d).sql ``` - Run migration on prod: ``` DATABASE_URL= npx drizzle-kit migrate ``` - Run import script against prod: ``` DATABASE_URL= npx tsx data-pipeline/import-to-postgres.ts ``` - Verify counts on prod. ### 6.3 Post-deploy verification - Open the live app. Play 2 full games with different settings. - Check server logs for errors. - Wait 24h. Check logs again. - If everything is clean, drop old tables: ```sql DROP TABLE IF EXISTS entry_translations; DROP TABLE IF EXISTS vocabulary_entries; ``` (Or keep them for another week if you want a safety net.) ### 6.4 Rollback plan If something goes wrong: - Restore the backup: ``` psql -h -U -d < backup_YYYYMMDD.sql ``` - Revert the code to the previous commit. - Redeploy. --- ## Phase 7: Extend POS — Detailed ### 7.1 Per POS Repeat the Phase 3 pipeline for each new POS: - Acquire wordlists (verbs, adjectives, adverbs) for all 5 languages. - Adjust the Gemini prompt: - Verbs: ask for transitivity, common prepositions, or other verb-specific metadata if needed for future exercises. - Adjectives: ask for the base form. Note that adjective metadata differs from noun metadata (no gender on the adjective itself in the same way — gender applies to the noun it modifies). - Adverbs: typically simpler. May not need gender at all. - Run pipeline → validate → SQLite → import → Postgres. - Test in the app with the POS filter set to the new type. ### 7.2 No schema changes The `pos` column already supports all four types. The `gender` column is nullable and simply won't be populated for adverbs or English words. No migration needed. --- # Risks & Mitigations | Risk | Likelihood | Impact | Mitigation | | ------------------------------------------------------ | ----------------- | ------------------------- | ------------------------------------------------------------------------------------------------- | | Gemini returns malformed JSON | Medium | Pipeline stalls | Strip markdown fences, wrap parsing in try/catch, log and skip bad batches | | Gemini produces wrong genders or difficulties | Medium | Bad exercise data | Validation module rejects invalid entries. Spot-check 50+ entries per language | | Too few words at a given difficulty/pos/language combo | Medium | Game can't fill 4 options | Graceful fallback: reduce options or show a "not enough words" message. Log which combos are thin | | SQLite → Postgres import fails mid-way | Low | Partial data in Postgres | Transactions per batch. Re-run with ON CONFLICT DO NOTHING | | ORDER BY RANDOM() gets slow at scale | Low (not at 500k) | Slow game load | Add a TODO comment. Optimize with random() pre-filter when needed | | Prod migration breaks the live app | Low | Downtime | Backup before migrating. Rollback plan documented. Deploy during low-traffic hours | --- # File / Folder Structure (new files) ``` project/ ├── data-pipeline/ │ ├── source-data/ │ │ ├── german/nouns/ ← wordlist files │ │ ├── english/nouns/ │ │ ├── spanish/nouns/ │ │ ├── french/nouns/ │ │ └── italian/nouns/ │ ├── rejections/ ← invalid Gemini entries (for review) │ ├── staging.db ← SQLite staging database │ ├── schema.sql ← SQLite schema definition │ ├── validate.ts ← validation module │ ├── run.ts ← main pipeline script │ └── import-to-postgres.ts ← SQLite → Postgres import ├── packages/ │ ├── db/ │ │ └── src/db/schema.ts ← Drizzle schema (updated ✅) │ └── shared/ │ └── src/constants.ts ← NOUN_GENDERS added ✅, "medium" ✅ ├── documentation/ │ └── pipeline/ │ ├── design-doc.md ← schema design doc (updated ✅) │ └── roadmap.md ← this document (updated ✅) └── ... ```