lila/documentation/pipeline/roadmap.md
2026-07-24 13:53:33 +02:00

25 KiB
Raw Blame History

Vocabulary Trainer — Roadmap: Data Pipeline & Schema Migration

Objective: Replace the OpenWordNet/kaikki data source with a Gemini-powered pipeline that generates high-quality vocabulary data, backed by a normalized Postgres schema.

Author: [Your Name] Date: July 2026 Companion doc: docs/schema-design.md


How to Read This Document

This roadmap has three zoom levels:

  • Part 1 — Overview: The phases at a glance. Read this to understand the full scope in 30 seconds.
  • Part 2 — Phase Breakdown: Goals, tasks, dependencies, and acceptance criteria per phase. Read this to plan your week.
  • Part 3 — Detailed Tasks: Step-by-step instructions within each phase. Read this when you sit down to code.

Part 1 — High-Level Overview

Phase 1  Schema ✅
         New Drizzle schema (words, senses, translations) with
         relations, constraints, and indexes. Committed.

Phase 2  Preparation
         Get wordlists, set up tooling, finalize the Gemini prompt.

Phase 3  Data Pipeline
         Build the Gemini → validate → SQLite pipeline.
         Produce a clean dataset of ~1000 nouns × 5 languages.

Phase 4  Migration & Import
         Generate and apply the Drizzle migration.
         Write the SQLite → Postgres import script.

Phase 5  App Integration
         Rewrite the game and distractor queries against the new schema.
         Update the exercise-generation logic.
         Test the full game flow in dev.

Phase 6  Production Deploy
         Run the Drizzle migration on prod.
         Import the dataset.
         Verify the live app works end-to-end.

Phase 7  Extend POS
         Run the pipeline for verbs, adjectives, adverbs.
         No schema changes needed — new wordlists + adjusted prompts.

Phase 8  Future Features (out of scope for now)
         Inflection tables, conjugation/declension exercises,
         gender exercises, spaced-repetition scheduling.

Dependency chain:

Phase 1 ✅ → Phase 2 → Phase 3 → Phase 4 → Phase 5 → Phase 6 → Phase 7
                                                                    │
                                                                    ▼
                                                                Phase 8

Part 2 — Phase Breakdown


Phase 1: Schema COMPLETE

Goal: New normalized schema exists in Drizzle with all tables, constraints, indexes, and relations.

Completed tasks:

  • words table: headword, language_code, pos, UNIQUE, CHECKs, index
  • senses table: word_id FK (cascade), sense_index, difficulty, definitions TEXT[], examples TEXT[], UNIQUE, CHECK, index
  • translations table: sense_id FK (cascade), target_language_code, translation, gender (nullable), difficulty, UNIQUE, CHECKs, index
  • Relations: words→senses (many), senses→word (one) + translations (many), translations→sense (one)
  • NOUN_GENDERS constant added to @lila/shared
  • DIFFICULTY_LEVELS updated: "intermediate" → "medium"
  • Old tables (vocabulary_entries, entry_translations) untouched
  • Auth and lobby tables untouched
  • Build passes, committed

Phase 2: Preparation

Goal: Have everything you need before writing pipeline code.

Tasks:

  • Acquire frequency-based noun lists for all 5 languages
  • Clean and format the lists (one word per line, UTF-8)
  • Set up local Postgres (Docker or native)
  • Set up a SQLite database file for staging
  • Write and test the Gemini prompt with 5 sample words
  • Refine the prompt until the JSON output matches the contract defined in docs/schema-design.md §6.3

Dependencies: Phase 1 complete.

Acceptance criteria:

  • You have 5 wordlist files in data-pipeline/source-data/ (one per language).
  • A Gemini call with 5 German nouns returns valid JSON matching the contract, including definitions, examples, translations with gender, and difficulty levels.
  • Local Postgres is running and reachable from your app.

Phase 3: Data Pipeline

Goal: A repeatable script that takes a wordlist, calls Gemini in batches of 20, validates the output, and writes clean rows to SQLite.

Tasks:

  • Write the validation module (see schema-design §6.4)
  • Write the pipeline script: - Read wordlist file - Split into batches of 20 - Call Gemini API per batch - Parse JSON response - Validate each entry - Write valid entries to SQLite - Log invalid entries to a rejection file
  • Create the SQLite schema (mirrors the Postgres schema)
  • Run the pipeline for all 5 languages (nouns only)
  • Review the rejection log, fix prompt issues, re-run failed batches
  • Spot-check 50 random entries for correctness

Dependencies: Phase 2 complete.

Acceptance criteria:

  • SQLite database contains ~1000 nouns × 5 languages with senses and translations.
  • Rejection rate is below 10%.
  • Spot-checked entries have correct definitions, plausible examples, correct genders, and reasonable difficulty levels.
  • The pipeline is re-runnable (idempotent or with duplicate handling).

Phase 4: Migration & Import

Goal: The new schema exists in Postgres and the SQLite data is imported.

Tasks:

  • Generate the Drizzle migration (npx drizzle-kit generate)
  • Inspect the generated SQL file — verify it creates the right tables, constraints, and indexes
  • Apply the migration to local Postgres (npx drizzle-kit migrate)
  • Write the import script (SQLite → Postgres): - Read all rows from SQLite - Insert into Postgres in dependency order: words → senses → translations - Use batch inserts (not row-by-row) - Wrap in transactions (per batch of 20 words) - Handle duplicates gracefully (skip or upsert)
  • Run the import script
  • Verify row counts match between SQLite and Postgres
  • Run 35 manual SQL queries against Postgres to sanity-check the data

Dependencies: Phase 3 complete (SQLite has data).

Acceptance criteria:

  • npx drizzle-kit migrate runs without errors on local Postgres.
  • Import script completes without errors.
  • Row counts in Postgres match SQLite (±rejection count).
  • Manual query: "Give me 5 random German nouns with Spanish translations at easy difficulty" returns sensible results.

Phase 5: App Integration

Goal: The running app uses the new schema. Game rounds and distractors work correctly for all language pairs.

Tasks:

  • Rewrite getGameTerms query: - JOIN words → senses → translations - Filter: source language, pos, sense difficulty (ceiling), target language, translation difficulty (exact) - ORDER BY RANDOM(), LIMIT rounds
  • Rewrite getDistractors query: - Same JOINs and filters - Exclude: sense_id != current, translation != correct - ORDER BY RANDOM(), LIMIT 3
  • Update the exercise-generation logic: - Pick one random definition from the definitions array - Pick one random example from the examples array - Assemble the 4 answer options (1 correct + 3 distractors) - Shuffle the options
  • Remove or deprecate old schema references (old vocabulary_entries, entry_translations tables)
  • Test manually: - German → Spanish, nouns, easy, 10 rounds - Spanish → German, nouns, medium, 10 rounds - English → French, nouns, hard, 10 rounds - Italian → German, nouns, easy, 10 rounds - Verify: no duplicate answers, no same-sense synonyms as distractors, definitions and examples are in the source language, genders display correctly
  • Test edge cases: - A difficulty/pos/language combo with very few words (does the app handle < 4 available words gracefully?) - A word with multiple senses (does the correct sense appear?)

Dependencies: Phase 4 complete (Postgres has data, schema exists).

Acceptance criteria:

  • Full game flow works in dev for at least 4 different language pairs.
  • No same-sense synonyms appear as distractors.
  • Definitions and examples are in the correct language.
  • Gender is displayed where applicable.
  • No console errors or unhandled query failures.

Phase 6: Production Deploy

Goal: The live app runs on the new schema with the new data.

Tasks:

  • Back up the production database
  • Run the Drizzle migration on prod
  • Run the import script against prod Postgres
  • Verify row counts on prod
  • Test the live app: - Play 2 full games on the deployed app - Check different language pairs and difficulties
  • Monitor for errors (server logs, browser console) for 24h
  • Remove old tables from prod (after confirming everything works)

Dependencies: Phase 5 complete (dev is fully working).

Acceptance criteria:

  • Live app serves game rounds from the new schema.
  • No errors in server logs for 24 hours.
  • Old tables are dropped (or scheduled for removal).
  • Rollback plan exists: the prod backup can be restored if needed.

Phase 7: Extend POS

Goal: Verbs, adjectives, and adverbs are in the database and usable in the app.

Tasks:

  • Acquire frequency lists for verbs, adjectives, adverbs (all 5 languages)
  • Adjust the Gemini prompt per POS: - Verbs: may need different metadata (transitivity, etc.) - Adjectives: may need base form info - Adverbs: typically simpler metadata
  • Run the pipeline for each POS
  • Validate, import to Postgres (dev → prod)
  • Test game flow with verbs, adjectives, adverbs
  • Verify the POS filter in the app UI works for all types

Dependencies: Phase 6 complete.

Acceptance criteria:

  • All 4 POS types are playable in the app.
  • Pipeline is repeatable for future data additions.

Phase 8: Future Features (out of scope)

Listed here for visibility. Not planned, not estimated.

  • inflection_forms table + conjugation exercises (verbs)
  • Adjective declension exercises (der grüne Mann, grüner Mann…)
  • Gender exercises (pick the correct article)
  • Spaced-repetition scheduling (track which words the user knows)
  • User accounts and progress persistence
  • Additional languages (if ever)

Part 3 — Detailed Task Breakdown


Phase 2: Preparation — Detailed

2.1 Acquire wordlists

  • Search for frequency lists. Good starting points:
    • "Leipzig Corpora Collection" (academic, per-language)
    • Wiktionary frequency lists
    • GitHub repos: search "german noun frequency list", "spanish noun frequency list", etc.
    • Tatoeba sentence counts as a proxy for word frequency
  • Target: ~1000 nouns per language, sorted by frequency.
  • Format: plain text, one word per line, UTF-8, no headers.
  • Save to data-pipeline/source-data/{language}/nouns/.
  • Clean the lists: remove duplicates, remove words with spaces (multi-word expressions), remove proper nouns if desired.

2.2 Set up local Postgres

  • Option A (Docker):
    docker run --name vocab-dev \
      -e POSTGRES_USER=dev \
      -e POSTGRES_PASSWORD=dev \
      -e POSTGRES_DB=vocab \
      -p 5432:5432 \
      -d postgres:16
    
  • Option B (native install): install Postgres, create a vocab database.
  • Update your .env / .env.local with the connection string.
  • Verify: connect with psql or a GUI client (TablePlus, DBeaver, pgAdmin). Run SELECT 1;.

2.3 Write and test the Gemini prompt

  • Start with 5 German nouns.
  • The prompt should specify:
    • The exact JSON structure expected (copy from schema-design §6.3)
    • That definitions and examples must be in the word's language
    • That translations are needed for all 4 other supported languages
    • That gender must be provided for de/fr/es/it, null for en
    • That difficulty must be one of: easy, medium, hard
    • That multiple senses should be included for polysemous words
  • Call the API. Inspect the JSON. Common issues to fix:
    • Gemini wraps the JSON in markdown code fences → strip them
    • Gemini returns "intermediate" instead of "medium" → normalize
    • Gemini omits a language → re-prompt or reject
    • Gemini returns gender "common" for German → reject
  • Iterate until 3 consecutive batches of 20 return clean JSON.

Phase 3: Data Pipeline — Detailed

3.1 Create the SQLite schema

  • Create a file data-pipeline/schema.sql or define it in your pipeline script.

  • Tables mirror the Postgres schema:

    CREATE TABLE words (
      id TEXT PRIMARY KEY,
      headword TEXT NOT NULL,
      language_code TEXT NOT NULL,
      pos TEXT NOT NULL,
      created_at TEXT DEFAULT (datetime('now')),
      UNIQUE(headword, language_code, pos)
    );
    
    CREATE TABLE senses (
      id TEXT PRIMARY KEY,
      word_id TEXT NOT NULL REFERENCES words(id),
      sense_index INTEGER NOT NULL DEFAULT 0,
      difficulty TEXT NOT NULL,
      definitions TEXT NOT NULL DEFAULT '[]',  -- JSON array as text
      examples TEXT NOT NULL DEFAULT '[]',     -- JSON array as text
      created_at TEXT DEFAULT (datetime('now')),
      UNIQUE(word_id, sense_index)
    );
    
    CREATE TABLE translations (
      id TEXT PRIMARY KEY,
      sense_id TEXT NOT NULL REFERENCES senses(id),
      target_language_code TEXT NOT NULL,
      translation TEXT NOT NULL,
      gender TEXT,
      difficulty TEXT NOT NULL,
      created_at TEXT DEFAULT (datetime('now')),
      UNIQUE(sense_id, target_language_code, translation)
    );
    
  • Note: SQLite has no native TEXT[]. Store arrays as JSON text. The import script will parse them into Postgres TEXT[] on import.

3.2 Write the validation module

  • Create data-pipeline/validate.ts.
  • Input: one parsed Gemini entry (the JSON object for one word).
  • Checks (return a list of errors, empty = valid):
    • headword is a non-empty string
    • language is in ['en','de','it','fr','es']
    • pos is in ['noun','verb','adjective','adverb']
    • senses is a non-empty array
    • Each sense has:
      • sense_index is a non-negative integer
      • difficulty is in ['easy','medium','hard']
      • definitions is a non-empty array of non-empty strings
      • examples is a non-empty array of non-empty strings
      • translations is a non-empty array
      • Each translation has:
        • target_language is in the supported list
        • target_language != the word's own language
        • word is a non-empty string
        • gender is valid for the target language (de: masculine/feminine/neuter/null, fr/es/it: masculine/feminine/null, en: null)
        • difficulty is in ['easy','medium','hard']
  • Output: { valid: boolean, errors: string[] }

3.3 Write the pipeline script

  • Create data-pipeline/run.ts.
  • Pseudocode:
    for each language in [de, en, es, fr, it]:
      read wordlist file → array of words
      split into batches of 20
      for each batch:
        call Gemini API with the batch
        parse JSON response (strip markdown fences if present)
        for each entry in response:
          validate(entry)
          if valid:
            generate UUIDs for word, senses, translations
            INSERT into SQLite (words, senses, translations)
          else:
            append to rejection log
        log progress: "Batch 12/50 done. 238 words imported, 2 rejected."
        sleep 1s (rate limiting)
    
  • Use better-sqlite3 for SQLite access (synchronous, simple).
  • Generate UUIDs with crypto.randomUUID().
  • Store definitions/examples as JSON strings in SQLite (JSON.stringify(arr)).

3.4 Run and review

  • Run the pipeline for all 5 languages.
  • Check the rejection log. If rejection rate > 10%, fix the prompt and re-run failed batches.
  • Spot-check: open the SQLite DB, run:
    SELECT w.headword, s.definitions, t.translation, t.gender
    FROM words w
    JOIN senses s ON s.word_id = w.id
    JOIN translations t ON t.sense_id = s.id
    WHERE w.language_code = 'de' AND w.pos = 'noun'
    ORDER BY RANDOM()
    LIMIT 20;
    
  • Read the definitions. Are they in German? Do they make sense? Are the genders correct? Are the difficulties reasonable?

Phase 4: Migration & Import — Detailed

4.1 Generate and inspect the migration

  • Run: npx drizzle-kit generate
  • Open the generated SQL file in packages/db/drizzle/.
  • Read it. Verify:
    • Three CREATE TABLE statements (words, senses, translations)
    • CHECK constraints match your schema
    • UNIQUE constraints are present
    • Three CREATE INDEX statements
    • Foreign keys reference the correct tables with ON DELETE CASCADE
  • If something looks wrong, fix the schema file and regenerate.

4.2 Apply the migration locally

  • Run: npx drizzle-kit migrate
  • Connect to local Postgres and verify:
    \dt                          -- list tables
    \d words                     -- describe words table
    \d senses                    -- describe senses table
    \d translations              -- describe translations table
    

4.3 Write the import script

  • Create data-pipeline/import-to-postgres.ts.

  • Pseudocode:

    open SQLite database (read-only)
    connect to Postgres via Drizzle
    
    read all words from SQLite
    for each batch of 100 words:
      begin transaction
      INSERT words into Postgres
      for each word:
        read its senses from SQLite
        INSERT senses into Postgres
        for each sense:
          read its translations from SQLite
          parse definitions/examples from JSON string → TEXT[]
          INSERT translations into Postgres
      commit transaction
      log progress
    
    log final counts: words, senses, translations
    
  • Parse definitions and examples from JSON strings (SQLite) into actual arrays (Postgres TEXT[]).

  • Use ON CONFLICT DO NOTHING to handle re-runs gracefully.

4.4 Run and verify

  • Run the import script.

  • Compare counts:

    -- In SQLite
    SELECT COUNT(*) FROM words;
    SELECT COUNT(*) FROM senses;
    SELECT COUNT(*) FROM translations;
    
    -- In Postgres
    SELECT COUNT(*) FROM words;
    SELECT COUNT(*) FROM senses;
    SELECT COUNT(*) FROM translations;
    
  • Counts should match (minus any rows that failed validation).

  • Run the game query manually in Postgres:

    SELECT w.headword, s.definitions, s.examples,
           t.translation, t.gender
    FROM words w
    JOIN senses s ON s.word_id = w.id
    JOIN translations t ON t.sense_id = s.id
    WHERE w.language_code = 'de'
      AND w.pos = 'noun'
      AND s.difficulty IN ('easy', 'medium')
      AND t.target_language_code = 'es'
      AND t.difficulty = 'medium'
    ORDER BY RANDOM()
    LIMIT 5;
    
  • Verify the results make sense.


Phase 5: App Integration — Detailed

5.1 Rewrite getGameTerms

  • Open packages/db/src/models/termModel.ts.
  • Replace the old query (2-table join on vocabulary_entries + entry_translations) with the new 3-table join (words → senses → translations).
  • Parameters: sourceLanguage, targetLanguage, pos, difficulty, rounds.
  • Difficulty filter:
    • senses.difficulty IN (all levels up to and including selected) — ceiling logic
    • translations.difficulty = selected — exact match
  • Return: word_id, headword, sense_id, definitions, examples, translation, gender.

5.2 Rewrite getDistractors

  • Same 3-table join.
  • Additional filters:
    • t.sense_id != :currentSenseId
    • t.translation != :correctAnswer
  • LIMIT 3.
  • If fewer than 3 distractors are found (small data pool), handle gracefully: reduce the number of options or log a warning.

5.3 Update exercise generation

  • In the function that assembles a game round:
    • Pick one random definition: definitions[Math.floor(Math.random() * definitions.length)]
    • Pick one random example: examples[Math.floor(Math.random() * examples.length)]
    • Combine 1 correct translation + 3 distractors.
    • Shuffle the 4 options (Fisher-Yates or similar).
    • Attach gender to each option for display.

5.4 Test matrix

Run through this matrix manually in the dev app:

Source Target POS Difficulty Rounds Pass?
de es noun easy 10
es de noun medium 10
en fr noun hard 10
it de noun easy 10
fr en noun medium 10

For each: verify definitions are in the source language, translations are in the target language, genders are shown, no same-sense synonyms appear as distractors, no duplicate options.


Phase 6: Production Deploy — Detailed

6.1 Pre-deploy checklist

  • All Phase 5 acceptance criteria pass
  • git status is clean, all changes committed
  • The Drizzle migration file is committed to the repo
  • You know your prod database connection string
  • You have a backup method for prod (pg_dump, hosting provider snapshot, etc.)

6.2 Deploy

  • Back up prod:
    pg_dump -h <host> -U <user> -d <db> > backup_$(date +%Y%m%d).sql
    
  • Run migration on prod:
    DATABASE_URL=<prod-url> npx drizzle-kit migrate
    
  • Run import script against prod:
    DATABASE_URL=<prod-url> npx tsx data-pipeline/import-to-postgres.ts
    
  • Verify counts on prod.

6.3 Post-deploy verification

  • Open the live app. Play 2 full games with different settings.
  • Check server logs for errors.
  • Wait 24h. Check logs again.
  • If everything is clean, drop old tables:
    DROP TABLE IF EXISTS entry_translations;
    DROP TABLE IF EXISTS vocabulary_entries;
    
    (Or keep them for another week if you want a safety net.)

6.4 Rollback plan

If something goes wrong:

  • Restore the backup:
    psql -h <host> -U <user> -d <db> < backup_YYYYMMDD.sql
    
  • Revert the code to the previous commit.
  • Redeploy.

Phase 7: Extend POS — Detailed

7.1 Per POS

Repeat the Phase 3 pipeline for each new POS:

  • Acquire wordlists (verbs, adjectives, adverbs) for all 5 languages.
  • Adjust the Gemini prompt:
    • Verbs: ask for transitivity, common prepositions, or other verb-specific metadata if needed for future exercises.
    • Adjectives: ask for the base form. Note that adjective metadata differs from noun metadata (no gender on the adjective itself in the same way — gender applies to the noun it modifies).
    • Adverbs: typically simpler. May not need gender at all.
  • Run pipeline → validate → SQLite → import → Postgres.
  • Test in the app with the POS filter set to the new type.

7.2 No schema changes

The pos column already supports all four types. The gender column is nullable and simply won't be populated for adverbs or English words. No migration needed.


Risks & Mitigations

Risk Likelihood Impact Mitigation
Gemini returns malformed JSON Medium Pipeline stalls Strip markdown fences, wrap parsing in try/catch, log and skip bad batches
Gemini produces wrong genders or difficulties Medium Bad exercise data Validation module rejects invalid entries. Spot-check 50+ entries per language
Too few words at a given difficulty/pos/language combo Medium Game can't fill 4 options Graceful fallback: reduce options or show a "not enough words" message. Log which combos are thin
SQLite → Postgres import fails mid-way Low Partial data in Postgres Transactions per batch. Re-run with ON CONFLICT DO NOTHING
ORDER BY RANDOM() gets slow at scale Low (not at 500k) Slow game load Add a TODO comment. Optimize with random() pre-filter when needed
Prod migration breaks the live app Low Downtime Backup before migrating. Rollback plan documented. Deploy during low-traffic hours

File / Folder Structure (new files)

project/
├── data-pipeline/
│   ├── source-data/
│   │   ├── german/nouns/         ← wordlist files
│   │   ├── english/nouns/
│   │   ├── spanish/nouns/
│   │   ├── french/nouns/
│   │   └── italian/nouns/
│   ├── rejections/               ← invalid Gemini entries (for review)
│   ├── staging.db                ← SQLite staging database
│   ├── schema.sql                ← SQLite schema definition
│   ├── validate.ts               ← validation module
│   ├── run.ts                    ← main pipeline script
│   └── import-to-postgres.ts     ← SQLite → Postgres import
├── packages/
│   ├── db/
│   │   └── src/db/schema.ts      ← Drizzle schema (updated ✅)
│   └── shared/
│       └── src/constants.ts      ← NOUN_GENDERS added ✅, "medium" ✅
├── documentation/
│   └── pipeline/
│       ├── design-doc.md         ← schema design doc (updated ✅)
│       └── roadmap.md            ← this document (updated ✅)
└── ...