From 303bb9388c45d6311a27e8277e4d9e3bbdf89a5f Mon Sep 17 00:00:00 2001 From: lila Date: Sun, 9 Aug 2026 13:32:15 +0200 Subject: [PATCH] updating pipeline roadmap and status to match branch reality Roadmap: phase 4 re-scoped to import only (migration already applied), phase 3 gains prompt templating, wordlist dedup, structured output, raw-response persistence, and validation tests; idempotency decision documented (skip already-staged words, one transaction per word). Stale paths, postgres setup, and file structure corrected. STATUS.md: refreshed from stale 2026-05-15 Kaikki state to current gemini-only pipeline work. Co-Authored-By: Claude Fable 5 --- documentation/STATUS.md | 18 +- documentation/pipeline/roadmap.md | 437 ++++++++++++++++-------------- 2 files changed, 242 insertions(+), 213 deletions(-) diff --git a/documentation/STATUS.md b/documentation/STATUS.md index 173ff91..816db2f 100644 --- a/documentation/STATUS.md +++ b/documentation/STATUS.md @@ -1,6 +1,6 @@ -# Status β€” 2026-05-15 +# Status β€” 2026-08-09 -> Last updated: 2026-05-15. Update this file after every deploy or when switching tasks. +> Last updated: 2026-08-09. Update this file after every deploy or when switching tasks. ## What Works Today βœ… @@ -12,7 +12,7 @@ ## What's Broken / Blocked 🚧 -- **Data quality** β€” Production still uses OpenWordNet/OMW translations. Kaikki pipeline (sense-disambiguated) is in progress but not yet synced to production. +- **Data quality** β€” Production still uses OpenWordNet/OMW translations. The replacement is the Gemini-only pipeline on branch `refactor/gemini-only-pipeline`: the sense-based schema (`words` β†’ `senses` β†’ `translations`) is migrated but empty, and the pipeline itself is not built yet. The earlier Kaikki/local-LLM pipeline was abandoned and removed (docs in `documentation/archive/`). - **Guest play** β€” Auth is required for all game routes. No try-before-signup flow. - **Game session store** β€” Still in-memory (`InMemoryGameSessionStore`). Valkey container exists in local dev but not wired up. - **Rate limiting** β€” Partially implemented on auth endpoints; game endpoints not yet covered. @@ -21,26 +21,26 @@ ## What I'm Working On Now πŸ”„ -**Primary:** Rewriting the Kaikki data pipeline enrich script for sub-stage architecture (round1_gloss β†’ round1_example β†’ round1_translations β†’ round1_cefr). +**Primary:** Phase 3 of the Gemini-only data pipeline (see [pipeline/roadmap.md](pipeline/roadmap.md)): template the Gemini prompt, write the validation module + `pipeline.ts`, and run the first full staging pass (~1,600–1,800 deduped nouns Γ— 5 languages into SQLite). **Secondary:** Phase 7 hardening backlog items (see BACKLOG.md `next` section). ## Next 2-Week Goal 🎯 -Finish Kaikki Stage 3 (enrich) sub-stage rewrite β†’ run full sample β†’ compare quality β†’ decide on production sync timeline. +Finish pipeline Phase 3 (full noun run staged in SQLite, reject rate < 10%, 50 entries spot-checked) β†’ Phase 4 import script (pipeline Postgres :5433, then dev :5432) β†’ start rewriting `getGameTerms`/`getDistractors` against the new schema. ## The Big Picture Lila is a **deployed, working vocabulary quiz app**. The core loop (singleplayer + multiplayer) is solid. The next strategic milestone is **media-based practice** (learn vocab from a song/TV episode/book chapter), but that depends on: -1. Kaikki data pipeline reaching production (fixes translation quality) +1. The Gemini data pipeline reaching production (fixes translation quality) 2. A media ingestion prototype (subtitles/lyrics β†’ text β†’ vocab extraction β†’ quiz) Until then, the app is a generic vocabulary quiz β€” functional but not differentiated. ## Quick Links -- [BACKLOG.md](BACKLOG.md) β€” Prioritized tasks -- [DATA_PIPELINE.md](DATA_PIPELINE.md) β€” Pipeline stages and current progress -- [BACKLOG.md](BACKLOG.md) β€” `now` / `next` / `later` +- [BACKLOG.md](BACKLOG.md) β€” Prioritized tasks (`now` / `next` / `later`) +- [DATA_PIPELINE.md](DATA_PIPELINE.md) β€” Pipeline orientation and current progress +- [pipeline/roadmap.md](pipeline/roadmap.md) β€” Phase-by-phase pipeline plan - [DEPLOYMENT.md](DEPLOYMENT.md) β€” Infrastructure ops diff --git a/documentation/pipeline/roadmap.md b/documentation/pipeline/roadmap.md index 722a0cd..9d4eb86 100644 --- a/documentation/pipeline/roadmap.md +++ b/documentation/pipeline/roadmap.md @@ -4,9 +4,9 @@ > Gemini-powered pipeline that generates high-quality vocabulary data, > backed by a normalized Postgres schema. > -> **Author:** [Your Name] -> **Date:** July 2026 -> **Companion doc:** `docs/schema-design.md` +> **Author:** lila +> **Date:** July 2026 Β· Last reviewed: 2026-08-09 +> **Companion doc:** `documentation/pipeline/design-doc.md` --- @@ -28,18 +28,25 @@ This roadmap has three zoom levels: ``` Phase 1 Schema βœ… New Drizzle schema (words, senses, translations) with - relations, constraints, and indexes. Committed. + relations, constraints, and indexes. Committed and applied + (migration 0012_graceful_psynapse.sql) β€” tables exist, empty. -Phase 2 Preparation - Get wordlists, set up tooling, finalize the Gemini prompt. +Phase 2 Preparation βœ… + Wordlists acquired, databases set up, prompt drafted and + tested. Two loose ends carried into Phase 3: the prompt is + not templated yet, and the wordlists still contain duplicates + (deduped at runtime, not in the files). -Phase 3 Data Pipeline +Phase 3 Data Pipeline ← CURRENT Build the Gemini β†’ validate β†’ SQLite pipeline. - Produce a clean dataset of ~1000 nouns Γ— 5 languages. + Produce a clean dataset: full deduped noun lists + (~1,600–1,800 words) Γ— 5 languages. -Phase 4 Migration & Import - Generate and apply the Drizzle migration. +Phase 4 Import + (Migration already applied in Phase 1.) Write the SQLite β†’ Postgres import script. + Test it against the pipeline Postgres (:5433) first, + then load the app dev database (:5432). Phase 5 App Integration Rewrite the game and distractor queries against the new schema. @@ -47,8 +54,8 @@ Phase 5 App Integration Test the full game flow in dev. Phase 6 Production Deploy - Run the Drizzle migration on prod. - Import the dataset. + Import the dataset into prod (schema arrives via the normal + Drizzle migration flow). Verify the live app works end-to-end. Phase 7 Extend POS @@ -63,10 +70,10 @@ Phase 8 Future Features (out of scope for now) **Dependency chain:** ``` -Phase 1 βœ… β†’ Phase 2 β†’ Phase 3 β†’ Phase 4 β†’ Phase 5 β†’ Phase 6 β†’ Phase 7 - β”‚ - β–Ό - Phase 8 +Phase 1 βœ… β†’ Phase 2 βœ… β†’ Phase 3 β†’ Phase 4 β†’ Phase 5 β†’ Phase 6 β†’ Phase 7 + β”‚ + β–Ό + Phase 8 ``` --- @@ -94,46 +101,66 @@ constraints, indexes, and relations. - [x] Old tables (`vocabulary_entries`, `entry_translations`) untouched - [x] Auth and lobby tables untouched - [x] Build passes, committed +- [x] Migration generated, inspected, and applied to local Postgres + (`packages/db/drizzle/0012_graceful_psynapse.sql`) --- -## Phase 2: Preparation +## Phase 2: Preparation βœ… COMPLETE **Goal:** Have everything you need before writing pipeline code. **Tasks:** - [x] Acquire frequency-based noun lists for all 5 languages -- [x] Clean and format the lists (one word per line, UTF-8) -- [x] Set up local Postgres (Docker or native) +- [x] Format the lists (one word per line, UTF-8) β€” **note:** the files + still contain 110–147 duplicate words each; dedup happens at + runtime in Phase 3, not in the files +- [x] Set up local Postgres (docker compose: app DB :5432, dedicated + pipeline DB :5433) - [x] Set up a SQLite database file for staging -- [x] Write and test the Gemini prompt with 5 sample words + (`data-pipeline/db/staging.db` from `db/schema.sql`) +- [x] Write and test the Gemini prompt with sample words - [x] Refine the prompt until the JSON output matches the contract - defined in `docs/schema-design.md` Β§6.3 + defined in `design-doc.md` Β§6.3 β€” **note:** the prompt works but + is pinned to a hardcoded Spanish sample and has known copy-paste + bugs; templating and fixes are Phase 3 tasks **Dependencies:** Phase 1 complete. -**Acceptance criteria:** +**Acceptance criteria (met):** -- You have 5 wordlist files in `data-pipeline/source-data/` (one per - language). -- A Gemini call with 5 German nouns returns valid JSON matching the +- 5 wordlist files exist in `data-pipeline/source-data/{lang}/noun` + (language codes `de/en/es/fr/it`, matching `@lila/shared` constants). +- A Gemini call with a batch of nouns returns valid JSON matching the contract, including definitions, examples, translations with gender, and difficulty levels. -- Local Postgres is running and reachable from your app. +- Local Postgres is running and reachable from the app. --- -## Phase 3: Data Pipeline +## Phase 3: Data Pipeline ← CURRENT **Goal:** A repeatable script that takes a wordlist, calls Gemini in batches of 20, validates the output, and writes clean rows to SQLite. +Re-runs skip words that are already staged. **Tasks:** -- [ ] Write the validation module (see schema-design Β§6.4) -- [ ] Write the pipeline script: - Read wordlist file - Split into batches of 20 - Call Gemini API per batch - Parse JSON response - Validate each entry - Write valid entries to SQLite - Log invalid entries to a rejection file -- [ ] Create the SQLite schema (mirrors the Postgres schema) +- [x] Create the SQLite schema (`data-pipeline/db/schema.sql`, tables + created in `db/staging.db`) +- [ ] Template the Gemini prompt: source language, POS, target + languages, and the word batch become substitutions; fix the known + copy-paste bugs while doing so (rules 2–3 hardcode `"en"`, + rule 15's target list contradicts the header, rule 31 says + "valid English noun") +- [ ] Write the wordlist normalization step: trim whitespace, drop + empty lines, dedup in memory +- [ ] Write the validation module (rules in design-doc Β§6.4) with unit + tests (`vitest.config.ts` expects `tests/**/*.test.ts` β€” the + directory doesn't exist yet) +- [ ] Write the pipeline script (`pipeline.ts`): - Read + normalize wordlist files - Skip words already present in staging.db (idempotency β€” see Β§3.3) - Split the remainder into batches of 20 - Call Gemini per batch using structured output + (`responseMimeType: "application/json"` + `responseSchema`) - Persist each raw response to disk before validating - Validate each entry - Write valid entries to SQLite, one transaction per word - Log invalid entries to a rejection file - [ ] Run the pipeline for all 5 languages (nouns only) - [ ] Review the rejection log, fix prompt issues, re-run failed batches - [ ] Spot-check 50 random entries for correctness @@ -142,39 +169,46 @@ batches of 20, validates the output, and writes clean rows to SQLite. **Acceptance criteria:** -- SQLite database contains ~1000 nouns Γ— 5 languages with senses and - translations. +- SQLite database contains the full deduped noun lists + (~1,600–1,800 words Γ— 5 languages) with senses and translations. - Rejection rate is below 10%. - Spot-checked entries have correct definitions, plausible examples, correct genders, and reasonable difficulty levels. -- The pipeline is re-runnable (idempotent or with duplicate handling). +- Re-running the pipeline skips already-staged words (no duplicate + rows, no repeated API calls for the same words). +- Raw Gemini responses are on disk, so validation-rule changes can be + re-applied without re-calling the API. --- -## Phase 4: Migration & Import +## Phase 4: Import -**Goal:** The new schema exists in Postgres and the SQLite data is -imported. +**Goal:** The SQLite data is imported into Postgres. + +The Drizzle migration was already generated, inspected, and applied in +Phase 1 (`0012_graceful_psynapse.sql`) β€” the `words`/`senses`/ +`translations` tables exist and are empty. What remains is the import +script. **Tasks:** -- [ ] Generate the Drizzle migration (`npx drizzle-kit generate`) -- [ ] Inspect the generated SQL file β€” verify it creates the right - tables, constraints, and indexes -- [ ] Apply the migration to local Postgres (`npx drizzle-kit migrate`) -- [ ] Write the import script (SQLite β†’ Postgres): - Read all rows from SQLite - Insert into Postgres in dependency order: - words β†’ senses β†’ translations - Use batch inserts (not row-by-row) - Wrap in transactions (per batch of 20 words) - Handle duplicates gracefully (skip or upsert) -- [ ] Run the import script +- [ ] Write the import script (`data-pipeline/import-to-postgres.ts`): - Read all rows from SQLite - Insert into Postgres in dependency order: + words β†’ senses β†’ translations - Use batch inserts (not row-by-row) via Drizzle β€” add `@lila/db` + as a workspace dependency (the pipeline currently has no + Postgres client) - Wrap in transactions (per batch of 20 words) - Handle duplicates gracefully (`ON CONFLICT DO NOTHING`) +- [ ] Run it against the **pipeline Postgres (:5433, + `PIPELINE_DATABASE_URL`)** first β€” this database exists so a bad + import can never damage dev data - [ ] Verify row counts match between SQLite and Postgres - [ ] Run 3–5 manual SQL queries against Postgres to sanity-check the data +- [ ] Once trusted, run it against the app dev database (:5432) **Dependencies:** Phase 3 complete (SQLite has data). **Acceptance criteria:** -- `npx drizzle-kit migrate` runs without errors on local Postgres. -- Import script completes without errors. +- Import script completes without errors on :5433 and then :5432. - Row counts in Postgres match SQLite (Β±rejection count). - Manual query: "Give me 5 random German nouns with Spanish translations at easy difficulty" returns sensible results. @@ -191,7 +225,8 @@ distractors work correctly for all language pairs. - [ ] Rewrite `getGameTerms` query: - JOIN words β†’ senses β†’ translations - Filter: source language, pos, sense difficulty (ceiling), target language, translation difficulty (exact) - ORDER BY RANDOM(), LIMIT rounds - [ ] Rewrite `getDistractors` query: - Same JOINs and filters - Exclude: `sense_id != current`, `translation != correct` - ORDER BY RANDOM(), LIMIT 3 -- [ ] Update the exercise-generation logic: - Pick one random definition from the `definitions` array - Pick one random example from the `examples` array - Assemble the 4 answer options (1 correct + 3 distractors) - Shuffle the options +- [ ] Update the exercise-generation logic: - Pick one random definition from the `definitions` array - Pick one random example from the `examples` array - Assemble the 4 answer options (1 correct + 3 distractors) - Shuffle the options - Keep the existing invariant: the correct answer is evaluated + server-side and is never sent to the client - [ ] Remove or deprecate old schema references (old `vocabulary_entries`, `entry_translations` tables) - [ ] Test manually: - German β†’ Spanish, nouns, easy, 10 rounds - Spanish β†’ German, nouns, medium, 10 rounds - English β†’ French, nouns, hard, 10 rounds - Italian β†’ German, nouns, easy, 10 rounds - Verify: no duplicate answers, no same-sense synonyms as @@ -278,105 +313,128 @@ Listed here for visibility. Not planned, not estimated. --- -## Phase 2: Preparation β€” Detailed +## Phase 2: Preparation β€” Detailed βœ… -### 2.1 Acquire wordlists +### 2.1 Wordlists (done) -- Search for frequency lists. Good starting points: - - "Leipzig Corpora Collection" (academic, per-language) - - Wiktionary frequency lists - - GitHub repos: search "german noun frequency list", - "spanish noun frequency list", etc. - - Tatoeba sentence counts as a proxy for word frequency -- Target: ~1000 nouns per language, sorted by frequency. +- Frequency-based lists live in `data-pipeline/source-data/{lang}/noun` + with `lang` ∈ `de/en/es/fr/it` β€” the same codes as + `packages/shared/src/constants.ts`, so no name mapping is needed + anywhere in the pipeline. - Format: plain text, one word per line, UTF-8, no headers. -- Save to `data-pipeline/source-data/{language}/nouns/`. -- Clean the lists: remove duplicates, remove words with spaces - (multi-word expressions), remove proper nouns if desired. +- Sizes: ~1,700–1,900 lines per language. The files were **not** + deduplicated (110–147 duplicates each); the pipeline dedups at read + time (Β§3.3) rather than editing the source files. -### 2.2 Set up local Postgres +### 2.2 Databases (done) -- Option A (Docker): - ``` - docker run --name vocab-dev \ - -e POSTGRES_USER=dev \ - -e POSTGRES_PASSWORD=dev \ - -e POSTGRES_DB=vocab \ - -p 5432:5432 \ - -d postgres:16 - ``` -- Option B (native install): install Postgres, create a `vocab` - database. -- Update your `.env` / `.env.local` with the connection string. -- Verify: connect with `psql` or a GUI client (TablePlus, DBeaver, - pgAdmin). Run `SELECT 1;`. +- `docker compose up -d` starts everything: app Postgres (:5432), + dedicated pipeline Postgres (:5433), Valkey (:6379). +- Config lives in the single root `.env` (see `.env.example`): + `PIPELINE_POSTGRES_USER` / `PIPELINE_POSTGRES_PASSWORD` / + `PIPELINE_POSTGRES_DB` / `PIPELINE_DATABASE_URL`, plus + `GEMINI_API_KEY` for the pipeline itself. +- The pipeline Postgres is deliberately separate from the app database + so pipeline work can never damage dev data. +- SQLite staging: `data-pipeline/db/staging.db`, created from + `data-pipeline/db/schema.sql`. -### 2.3 Write and test the Gemini prompt +### 2.3 The Gemini prompt (drafted; templating is Phase 3) -- Start with 5 German nouns. -- The prompt should specify: - - The exact JSON structure expected (copy from schema-design Β§6.3) +- `data-pipeline/prompt` is a plain UTF-8 text file, edited by hand. + It is currently pinned to a concrete sample run (Spanish nouns, + 20 words inlined). +- The prompt specifies: + - The exact JSON structure expected (from design-doc Β§6.3) - That definitions and examples must be in the word's language - That translations are needed for all 4 other supported languages - That gender must be provided for de/fr/es/it, null for en - That difficulty must be one of: easy, medium, hard - That multiple senses should be included for polysemous words -- Call the API. Inspect the JSON. Common issues to fix: - - Gemini wraps the JSON in markdown code fences β†’ strip them - - Gemini returns "intermediate" instead of "medium" β†’ normalize - - Gemini omits a language β†’ re-prompt or reject - - Gemini returns gender "common" for German β†’ reject -- Iterate until 3 consecutive batches of 20 return clean JSON. +- **Known bugs to fix when templating (Β§3.2):** rules 2 and 3 still say + `language` must be `"en"`; rule 15 lists target languages + `de, it, es, fr` while the header says `en, it, de, fr`; rule 31 + says "valid English noun". All are leftovers from adapting the + English version. --- ## Phase 3: Data Pipeline β€” Detailed -### 3.1 Create the SQLite schema +### 3.1 SQLite schema (done) -- Create a file `data-pipeline/schema.sql` or define it in your - pipeline script. -- Tables mirror the Postgres schema: +- `data-pipeline/db/schema.sql` mirrors the Postgres schema with two + SQLite concessions: IDs are `TEXT` (`crypto.randomUUID()`), and + `definitions` / `examples` are JSON-encoded strings because SQLite + has no array type. The import script parses them back into Postgres + `TEXT[]`. +- The UNIQUE constraints match Postgres: + `words(headword, language_code, pos)`, `senses(word_id, sense_index)`, + `translations(sense_id, target_language_code, translation)`. - ```sql - CREATE TABLE words ( - id TEXT PRIMARY KEY, - headword TEXT NOT NULL, - language_code TEXT NOT NULL, - pos TEXT NOT NULL, - created_at TEXT DEFAULT (datetime('now')), - UNIQUE(headword, language_code, pos) - ); +### 3.2 Template the prompt - CREATE TABLE senses ( - id TEXT PRIMARY KEY, - word_id TEXT NOT NULL REFERENCES words(id), - sense_index INTEGER NOT NULL DEFAULT 0, - difficulty TEXT NOT NULL, - definitions TEXT NOT NULL DEFAULT '[]', -- JSON array as text - examples TEXT NOT NULL DEFAULT '[]', -- JSON array as text - created_at TEXT DEFAULT (datetime('now')), - UNIQUE(word_id, sense_index) - ); +- Turn `data-pipeline/prompt` into a template. Substitution slots: + - source language (name + code) + - POS + - target languages (the other 4 codes) + - the word batch (20 words, one per line) +- Fix the hardcoded leftovers listed in Β§2.3 as part of this β€” after + templating, the language codes in the rules must derive from the + substitutions, so this class of bug can't recur. +- Request **structured output** from the API + (`responseMimeType: "application/json"` with a `responseSchema` + matching design-doc Β§6.3) instead of relying on prompt instructions + alone. Keep fence-stripping as a defensive fallback only. - CREATE TABLE translations ( - id TEXT PRIMARY KEY, - sense_id TEXT NOT NULL REFERENCES senses(id), - target_language_code TEXT NOT NULL, - translation TEXT NOT NULL, - gender TEXT, - difficulty TEXT NOT NULL, - created_at TEXT DEFAULT (datetime('now')), - UNIQUE(sense_id, target_language_code, translation) - ); +### 3.3 Write the pipeline script + +- Entry point: `data-pipeline/pipeline.ts` + (`pnpm --filter @lila/pipeline pipeline:run`). +- Pseudocode: + + ``` + for each language in [de, en, es, fr, it]: + read wordlist file β†’ trim, drop empties, dedup in memory + query staging.db for existing (headword, language_code, pos) + skip words already staged ← idempotency + split the remainder into batches of 20 + for each batch: + call Gemini API (structured output) + write the raw response to responses/{lang}-{pos}-{n}.json + parse JSON response + for each entry in response: + validate(entry) + if valid: + generate UUIDs for word, senses, translations + INSERT word + senses + translations in ONE transaction + else: + append to rejection log + log progress: "Batch 12/86 done. 238 words staged, 2 rejected." + sleep 1s (rate limiting); retry with backoff on API errors ``` -- Note: SQLite has no native TEXT[]. Store arrays as JSON text. - The import script will parse them into Postgres TEXT[] on import. +- **Idempotency decision (resolves the pipeline.ts step 3 question):** + wordlists are read fresh on every run; the staging DB is the record + of what's been processed. Words already present in `staging.db` are + skipped before batching, so re-runs cost no API calls for staged + words. Each validated entry (word + its senses + their translations) + is written in a single transaction, so a partially-written word can + never exist and no NOT NULL constraint needs relaxing. The UNIQUE + constraints remain as a backstop (`INSERT OR IGNORE`). +- Persisting raw responses means a validation-rule change re-validates + from disk instead of re-paying for ~430 API calls + (~8,700 words Γ· 20 per batch). +- Use `better-sqlite3` for SQLite access (synchronous, simple). +- Generate UUIDs with `crypto.randomUUID()`. +- Store definitions/examples as JSON strings in SQLite + (`JSON.stringify(arr)`). -### 3.2 Write the validation module +### 3.4 Validation module -- Create `data-pipeline/validate.ts`. +- Create the validation module alongside the pipeline, with unit tests + (vitest is already configured; `data-pipeline/vitest.config.ts` + expects `tests/**/*.test.ts`). - Input: one parsed Gemini entry (the JSON object for one word). - Checks (return a list of errors, empty = valid): - `headword` is a non-empty string @@ -394,43 +452,18 @@ Listed here for visibility. Not planned, not estimated. - `target_language` != the word's own language - `word` is a non-empty string - `gender` is valid for the target language - (de: masculine/feminine/neuter/null, - fr/es/it: masculine/feminine/null, + (de: masculine/feminine/neuter, + fr/es/it: masculine/feminine, en: null) - `difficulty` is in ['easy','medium','hard'] - Output: `{ valid: boolean, errors: string[] }` -### 3.3 Write the pipeline script - -- Create `data-pipeline/run.ts`. -- Pseudocode: - ``` - for each language in [de, en, es, fr, it]: - read wordlist file β†’ array of words - split into batches of 20 - for each batch: - call Gemini API with the batch - parse JSON response (strip markdown fences if present) - for each entry in response: - validate(entry) - if valid: - generate UUIDs for word, senses, translations - INSERT into SQLite (words, senses, translations) - else: - append to rejection log - log progress: "Batch 12/50 done. 238 words imported, 2 rejected." - sleep 1s (rate limiting) - ``` -- Use `better-sqlite3` for SQLite access (synchronous, simple). -- Generate UUIDs with `crypto.randomUUID()`. -- Store definitions/examples as JSON strings in SQLite - (`JSON.stringify(arr)`). - -### 3.4 Run and review +### 3.5 Run and review - Run the pipeline for all 5 languages. - Check the rejection log. If rejection rate > 10%, fix the prompt - and re-run failed batches. + and re-run β€” the skip-processed check means only rejected/missing + words are re-sent. - Spot-check: open the SQLite DB, run: ```sql SELECT w.headword, s.definitions, t.translation, t.gender @@ -446,50 +479,36 @@ Listed here for visibility. Not planned, not estimated. --- -## Phase 4: Migration & Import β€” Detailed +## Phase 4: Import β€” Detailed -### 4.1 Generate and inspect the migration +### 4.1 Migration (done in Phase 1) -- Run: `npx drizzle-kit generate` -- Open the generated SQL file in `packages/db/drizzle/`. -- Read it. Verify: - - Three CREATE TABLE statements (words, senses, translations) - - CHECK constraints match your schema - - UNIQUE constraints are present - - Three CREATE INDEX statements - - Foreign keys reference the correct tables with ON DELETE CASCADE -- If something looks wrong, fix the schema file and regenerate. +`packages/db/drizzle/0012_graceful_psynapse.sql` creates the three +tables with all CHECK/UNIQUE constraints, the three indexes, and +cascading FKs. It is applied locally; the tables exist and are empty. +Nothing to do here β€” prod gets the same migration in Phase 6. -### 4.2 Apply the migration locally - -- Run: `npx drizzle-kit migrate` -- Connect to local Postgres and verify: - ```sql - \dt -- list tables - \d words -- describe words table - \d senses -- describe senses table - \d translations -- describe translations table - ``` - -### 4.3 Write the import script +### 4.2 Write the import script - Create `data-pipeline/import-to-postgres.ts`. +- Add `@lila/db` as a workspace dependency for the Drizzle client and + schema (the pipeline currently ships only `better-sqlite3`). - Pseudocode: ``` open SQLite database (read-only) - connect to Postgres via Drizzle + connect to Postgres via Drizzle (PIPELINE_DATABASE_URL first) read all words from SQLite - for each batch of 100 words: + for each batch of words: begin transaction INSERT words into Postgres for each word: read its senses from SQLite + parse definitions/examples from JSON string β†’ TEXT[] INSERT senses into Postgres for each sense: read its translations from SQLite - parse definitions/examples from JSON string β†’ TEXT[] INSERT translations into Postgres commit transaction log progress @@ -501,9 +520,10 @@ Listed here for visibility. Not planned, not estimated. into actual arrays (Postgres TEXT[]). - Use `ON CONFLICT DO NOTHING` to handle re-runs gracefully. -### 4.4 Run and verify +### 4.3 Run and verify -- Run the import script. +- Run against the pipeline Postgres (:5433) first; once the script is + trusted, point it at the app dev database (:5432). - Compare counts: ```sql @@ -574,6 +594,8 @@ Listed here for visibility. Not planned, not estimated. - Combine 1 correct translation + 3 distractors. - Shuffle the 4 options (Fisher-Yates or similar). - Attach gender to each option for display. + - Keep the server-side evaluation invariant: the correct answer is + never included in what is sent to the client. ### 5.4 Test matrix @@ -651,8 +673,9 @@ If something goes wrong: Repeat the Phase 3 pipeline for each new POS: -- Acquire wordlists (verbs, adjectives, adverbs) for all 5 languages. -- Adjust the Gemini prompt: +- Acquire wordlists (verbs, adjectives, adverbs) for all 5 languages, + saved as `source-data/{lang}/{pos}` with the shared POS codes. +- Adjust the Gemini prompt template: - Verbs: ask for transitivity, common prepositions, or other verb-specific metadata if needed for future exercises. - Adjectives: ask for the base form. Note that adjective metadata @@ -672,42 +695,48 @@ English words. No migration needed. # Risks & Mitigations -| Risk | Likelihood | Impact | Mitigation | -| ------------------------------------------------------ | ----------------- | ------------------------- | ------------------------------------------------------------------------------------------------- | -| Gemini returns malformed JSON | Medium | Pipeline stalls | Strip markdown fences, wrap parsing in try/catch, log and skip bad batches | -| Gemini produces wrong genders or difficulties | Medium | Bad exercise data | Validation module rejects invalid entries. Spot-check 50+ entries per language | -| Too few words at a given difficulty/pos/language combo | Medium | Game can't fill 4 options | Graceful fallback: reduce options or show a "not enough words" message. Log which combos are thin | -| SQLite β†’ Postgres import fails mid-way | Low | Partial data in Postgres | Transactions per batch. Re-run with ON CONFLICT DO NOTHING | -| ORDER BY RANDOM() gets slow at scale | Low (not at 500k) | Slow game load | Add a TODO comment. Optimize with random() pre-filter when needed | -| Prod migration breaks the live app | Low | Downtime | Backup before migrating. Rollback plan documented. Deploy during low-traffic hours | +| Risk | Likelihood | Impact | Mitigation | +| ------------------------------------------------------ | ----------------- | ------------------------- | ----------------------------------------------------------------------------------------------------- | +| Gemini returns malformed JSON | Low | Pipeline stalls | Structured output (responseSchema); fence-stripping + try/catch as fallback; log and skip bad batches | +| Validation rules change after a full run | Medium | Wasted API spend | Raw responses persisted per batch; re-validate from disk instead of re-calling the API | +| Gemini produces wrong genders or difficulties | Medium | Bad exercise data | Validation module rejects invalid entries. Spot-check 50+ entries per language | +| Too few words at a given difficulty/pos/language combo | Medium | Game can't fill 4 options | Graceful fallback: reduce options or show a "not enough words" message. Log which combos are thin | +| SQLite β†’ Postgres import fails mid-way | Low | Partial data in Postgres | Transactions per batch. Re-run with ON CONFLICT DO NOTHING. Test on :5433 before touching :5432 | +| ORDER BY RANDOM() gets slow at scale | Low (not at 500k) | Slow game load | Add a TODO comment. Optimize with random() pre-filter when needed | +| Prod migration breaks the live app | Low | Downtime | Backup before migrating. Rollback plan documented. Deploy during low-traffic hours | --- -# File / Folder Structure (new files) +# File / Folder Structure ``` project/ β”œβ”€β”€ data-pipeline/ β”‚ β”œβ”€β”€ source-data/ -β”‚ β”‚ β”œβ”€β”€ german/nouns/ ← wordlist files -β”‚ β”‚ β”œβ”€β”€ english/nouns/ -β”‚ β”‚ β”œβ”€β”€ spanish/nouns/ -β”‚ β”‚ β”œβ”€β”€ french/nouns/ -β”‚ β”‚ └── italian/nouns/ -β”‚ β”œβ”€β”€ rejections/ ← invalid Gemini entries (for review) -β”‚ β”œβ”€β”€ staging.db ← SQLite staging database -β”‚ β”œβ”€β”€ schema.sql ← SQLite schema definition -β”‚ β”œβ”€β”€ validate.ts ← validation module -β”‚ β”œβ”€β”€ run.ts ← main pipeline script -β”‚ └── import-to-postgres.ts ← SQLite β†’ Postgres import +β”‚ β”‚ β”œβ”€β”€ de/noun ← wordlist files (shared lang/POS codes) +β”‚ β”‚ β”œβ”€β”€ en/noun +β”‚ β”‚ β”œβ”€β”€ es/noun +β”‚ β”‚ β”œβ”€β”€ fr/noun +β”‚ β”‚ └── it/noun +β”‚ β”œβ”€β”€ db/ +β”‚ β”‚ β”œβ”€β”€ schema.sql ← SQLite schema definition βœ… +β”‚ β”‚ └── staging.db ← SQLite staging database (gitignored) βœ… +β”‚ β”œβ”€β”€ prompt ← Gemini prompt (plain UTF-8; to be templated) +β”‚ β”œβ”€β”€ responses/ ← raw Gemini responses, one file per batch (planned) +β”‚ β”œβ”€β”€ rejections/ ← invalid Gemini entries for review (planned) +β”‚ β”œβ”€β”€ pipeline.ts ← main pipeline script (pseudocode today) +β”‚ β”œβ”€β”€ tests/ ← vitest unit tests, esp. validation (planned) +β”‚ └── import-to-postgres.ts ← SQLite β†’ Postgres import (planned) β”œβ”€β”€ packages/ β”‚ β”œβ”€β”€ db/ -β”‚ β”‚ └── src/db/schema.ts ← Drizzle schema (updated βœ…) +β”‚ β”‚ β”œβ”€β”€ src/db/schema.ts ← Drizzle schema βœ… +β”‚ β”‚ └── drizzle/0012_*.sql ← words/senses/translations migration βœ… β”‚ └── shared/ -β”‚ └── src/constants.ts ← NOUN_GENDERS added βœ…, "medium" βœ… +β”‚ └── src/constants.ts ← NOUN_GENDERS βœ…, "medium" βœ… β”œβ”€β”€ documentation/ +β”‚ β”œβ”€β”€ DATA_PIPELINE.md ← orientation layer β”‚ └── pipeline/ -β”‚ β”œβ”€β”€ design-doc.md ← schema design doc (updated βœ…) -β”‚ └── roadmap.md ← this document (updated βœ…) +β”‚ β”œβ”€β”€ design-doc.md ← schema design doc βœ… +β”‚ └── roadmap.md ← this document └── ... ```