# Vocabulary Trainer — Roadmap: Data Pipeline & Schema Migration > **Objective:** Replace the OpenWordNet/kaikki data source with a > Gemini-powered pipeline that generates high-quality vocabulary data, > backed by a normalized Postgres schema. > > **Author:** lila > **Date:** July 2026 · Last reviewed: 2026-08-09 > **Companion doc:** `documentation/pipeline/design-doc.md` --- ## How to Read This Document This roadmap has three zoom levels: - **Part 1 — Overview:** The phases at a glance. Read this to understand the full scope in 30 seconds. - **Part 2 — Phase Breakdown:** Goals, tasks, dependencies, and acceptance criteria per phase. Read this to plan your week. - **Part 3 — Detailed Tasks:** Step-by-step instructions within each phase. Read this when you sit down to code. --- # Part 1 — High-Level Overview ``` Phase 1 Schema ✅ New Drizzle schema (words, senses, translations) with relations, constraints, and indexes. Committed and applied (migration 0012_graceful_psynapse.sql) — tables exist, empty. Phase 2 Preparation ✅ Wordlists acquired, databases set up, prompt drafted and tested. Two loose ends carried into Phase 3: the prompt is not templated yet, and the wordlists still contain duplicates (deduped at runtime, not in the files). Phase 3 Data Pipeline ← CURRENT Build the Gemini → validate → SQLite pipeline. Produce a clean dataset: full deduped noun lists (~1,600–1,800 words) × 5 languages. Phase 4 Import (Migration already applied in Phase 1.) Write the SQLite → Postgres import script. Test it against the pipeline Postgres (:5433) first, then load the app dev database (:5432). Phase 5 App Integration Rewrite the game and distractor queries against the new schema. Update the exercise-generation logic. Test the full game flow in dev. Phase 6 Production Deploy Import the dataset into prod (schema arrives via the normal Drizzle migration flow). Verify the live app works end-to-end. Phase 7 Extend POS Run the pipeline for verbs, adjectives, adverbs. No schema changes needed — new wordlists + adjusted prompts. Phase 8 Future Features (out of scope for now) Inflection tables, conjugation/declension exercises, gender exercises, spaced-repetition scheduling. ``` **Dependency chain:** ``` Phase 1 ✅ → Phase 2 ✅ → Phase 3 → Phase 4 → Phase 5 → Phase 6 → Phase 7 │ ▼ Phase 8 ``` --- # Part 2 — Phase Breakdown --- ## Phase 1: Schema ✅ COMPLETE **Goal:** New normalized schema exists in Drizzle with all tables, constraints, indexes, and relations. **Completed tasks:** - [x] `words` table: headword, language_code, pos, UNIQUE, CHECKs, index - [x] `senses` table: word_id FK (cascade), sense_index, difficulty, definitions TEXT[], examples TEXT[], UNIQUE, CHECK, index - [x] `translations` table: sense_id FK (cascade), target_language_code, translation, gender (nullable), difficulty, UNIQUE, CHECKs, index - [x] Relations: words→senses (many), senses→word (one) + translations (many), translations→sense (one) - [x] `NOUN_GENDERS` constant added to `@lila/shared` - [x] `DIFFICULTY_LEVELS` updated: "intermediate" → "medium" - [x] Old tables (`vocabulary_entries`, `entry_translations`) untouched - [x] Auth and lobby tables untouched - [x] Build passes, committed - [x] Migration generated, inspected, and applied to local Postgres (`packages/db/drizzle/0012_graceful_psynapse.sql`) --- ## Phase 2: Preparation ✅ COMPLETE **Goal:** Have everything you need before writing pipeline code. **Tasks:** - [x] Acquire frequency-based noun lists for all 5 languages - [x] Format the lists (one word per line, UTF-8) — **note:** the files still contain 110–147 duplicate words each; dedup happens at runtime in Phase 3, not in the files - [x] Set up local Postgres (docker compose: app DB :5432, dedicated pipeline DB :5433) - [x] Set up a SQLite database file for staging (`data-pipeline/db/staging.db` from `db/schema.sql`) - [x] Write and test the Gemini prompt with sample words - [x] Refine the prompt until the JSON output matches the contract defined in `design-doc.md` §6.3 — **note:** the prompt works but is pinned to a hardcoded Spanish sample and has known copy-paste bugs; templating and fixes are Phase 3 tasks **Dependencies:** Phase 1 complete. **Acceptance criteria (met):** - 5 wordlist files exist in `data-pipeline/source-data/{lang}/noun` (language codes `de/en/es/fr/it`, matching `@lila/shared` constants). - A Gemini call with a batch of nouns returns valid JSON matching the contract, including definitions, examples, translations with gender, and difficulty levels. - Local Postgres is running and reachable from the app. --- ## Phase 3: Data Pipeline ← CURRENT **Goal:** A repeatable script that takes a wordlist, calls Gemini in batches of 20, validates the output, and writes clean rows to SQLite. Re-runs skip words that are already staged. **Tasks:** - [x] Create the SQLite schema (`data-pipeline/db/schema.sql`, tables created in `db/staging.db`) - [ ] Template the Gemini prompt: source language, POS, target languages, and the word batch become substitutions; fix the known copy-paste bugs while doing so (rules 2–3 hardcode `"en"`, rule 15's target list contradicts the header, rule 31 says "valid English noun") - [ ] Write the wordlist normalization step: trim whitespace, drop empty lines, dedup in memory - [ ] Write the validation module (rules in design-doc §6.4) with unit tests (`vitest.config.ts` expects `tests/**/*.test.ts` — the directory doesn't exist yet) - [ ] Write the pipeline script (`pipeline.ts`): - Read + normalize wordlist files - Skip words already present in staging.db (idempotency — see §3.3) - Split the remainder into batches of 20 - Call Gemini per batch using structured output (`responseMimeType: "application/json"` + `responseSchema`) - Persist each raw response to disk before validating - Validate each entry - Write valid entries to SQLite, one transaction per word - Log invalid entries to a rejection file - [ ] Run the pipeline for all 5 languages (nouns only) - [ ] Review the rejection log, fix prompt issues, re-run failed batches - [ ] Spot-check 50 random entries for correctness **Dependencies:** Phase 2 complete. **Acceptance criteria:** - SQLite database contains the full deduped noun lists (~1,600–1,800 words × 5 languages) with senses and translations. - Rejection rate is below 10%. - Spot-checked entries have correct definitions, plausible examples, correct genders, and reasonable difficulty levels. - Re-running the pipeline skips already-staged words (no duplicate rows, no repeated API calls for the same words). - Raw Gemini responses are on disk, so validation-rule changes can be re-applied without re-calling the API. --- ## Phase 4: Import **Goal:** The SQLite data is imported into Postgres. The Drizzle migration was already generated, inspected, and applied in Phase 1 (`0012_graceful_psynapse.sql`) — the `words`/`senses`/ `translations` tables exist and are empty. What remains is the import script. **Tasks:** - [ ] Write the import script (`data-pipeline/import-to-postgres.ts`): - Read all rows from SQLite - Insert into Postgres in dependency order: words → senses → translations - Use batch inserts (not row-by-row) via Drizzle — add `@lila/db` as a workspace dependency (the pipeline currently has no Postgres client) - Wrap in transactions (per batch of 20 words) - Handle duplicates gracefully (`ON CONFLICT DO NOTHING`) - [ ] Run it against the **pipeline Postgres (:5433, `PIPELINE_DATABASE_URL`)** first — this database exists so a bad import can never damage dev data - [ ] Verify row counts match between SQLite and Postgres - [ ] Run 3–5 manual SQL queries against Postgres to sanity-check the data - [ ] Once trusted, run it against the app dev database (:5432) **Dependencies:** Phase 3 complete (SQLite has data). **Acceptance criteria:** - Import script completes without errors on :5433 and then :5432. - Row counts in Postgres match SQLite (±rejection count). - Manual query: "Give me 5 random German nouns with Spanish translations at easy difficulty" returns sensible results. --- ## Phase 5: App Integration **Goal:** The running app uses the new schema. Game rounds and distractors work correctly for all language pairs. **Tasks:** - [ ] Rewrite `getGameTerms` query: - JOIN words → senses → translations - Filter: source language, pos, sense difficulty (ceiling), target language, translation difficulty (exact) - ORDER BY RANDOM(), LIMIT rounds - [ ] Rewrite `getDistractors` query: - Same JOINs and filters - Exclude: `sense_id != current`, `translation != correct` - ORDER BY RANDOM(), LIMIT 3 - [ ] Update the exercise-generation logic: - Pick one random definition from the `definitions` array - Pick one random example from the `examples` array - Assemble the 4 answer options (1 correct + 3 distractors) - Shuffle the options - Keep the existing invariant: the correct answer is evaluated server-side and is never sent to the client - [ ] Remove or deprecate old schema references (old `vocabulary_entries`, `entry_translations` tables) - [ ] Test manually: - German → Spanish, nouns, easy, 10 rounds - Spanish → German, nouns, medium, 10 rounds - English → French, nouns, hard, 10 rounds - Italian → German, nouns, easy, 10 rounds - Verify: no duplicate answers, no same-sense synonyms as distractors, definitions and examples are in the source language, genders display correctly - [ ] Test edge cases: - A difficulty/pos/language combo with very few words (does the app handle < 4 available words gracefully?) - A word with multiple senses (does the correct sense appear?) **Dependencies:** Phase 4 complete (Postgres has data, schema exists). **Acceptance criteria:** - Full game flow works in dev for at least 4 different language pairs. - No same-sense synonyms appear as distractors. - Definitions and examples are in the correct language. - Gender is displayed where applicable. - No console errors or unhandled query failures. --- ## Phase 6: Production Deploy **Goal:** The live app runs on the new schema with the new data. **Tasks:** - [ ] Back up the production database - [ ] Run the Drizzle migration on prod - [ ] Run the import script against prod Postgres - [ ] Verify row counts on prod - [ ] Test the live app: - Play 2 full games on the deployed app - Check different language pairs and difficulties - [ ] Monitor for errors (server logs, browser console) for 24h - [ ] Remove old tables from prod (after confirming everything works) **Dependencies:** Phase 5 complete (dev is fully working). **Acceptance criteria:** - Live app serves game rounds from the new schema. - No errors in server logs for 24 hours. - Old tables are dropped (or scheduled for removal). - Rollback plan exists: the prod backup can be restored if needed. --- ## Phase 7: Extend POS **Goal:** Verbs, adjectives, and adverbs are in the database and usable in the app. **Tasks:** - [ ] Acquire frequency lists for verbs, adjectives, adverbs (all 5 languages) - [ ] Adjust the Gemini prompt per POS: - Verbs: may need different metadata (transitivity, etc.) - Adjectives: may need base form info - Adverbs: typically simpler metadata - [ ] Run the pipeline for each POS - [ ] Validate, import to Postgres (dev → prod) - [ ] Test game flow with verbs, adjectives, adverbs - [ ] Verify the POS filter in the app UI works for all types **Dependencies:** Phase 6 complete. **Acceptance criteria:** - All 4 POS types are playable in the app. - Pipeline is repeatable for future data additions. --- ## Phase 8: Future Features (out of scope) Listed here for visibility. Not planned, not estimated. - [ ] `inflection_forms` table + conjugation exercises (verbs) - [ ] Adjective declension exercises (der grüne Mann, grüner Mann…) - [ ] Gender exercises (pick the correct article) - [ ] Spaced-repetition scheduling (track which words the user knows) - [ ] User accounts and progress persistence - [ ] Additional languages (if ever) --- # Part 3 — Detailed Task Breakdown --- ## Phase 2: Preparation — Detailed ✅ ### 2.1 Wordlists (done) - Frequency-based lists live in `data-pipeline/source-data/{lang}/noun` with `lang` ∈ `de/en/es/fr/it` — the same codes as `packages/shared/src/constants.ts`, so no name mapping is needed anywhere in the pipeline. - Format: plain text, one word per line, UTF-8, no headers. - Sizes: ~1,700–1,900 lines per language. The files were **not** deduplicated (110–147 duplicates each); the pipeline dedups at read time (§3.3) rather than editing the source files. ### 2.2 Databases (done) - `docker compose up -d` starts everything: app Postgres (:5432), dedicated pipeline Postgres (:5433), Valkey (:6379). - Config lives in the single root `.env` (see `.env.example`): `PIPELINE_POSTGRES_USER` / `PIPELINE_POSTGRES_PASSWORD` / `PIPELINE_POSTGRES_DB` / `PIPELINE_DATABASE_URL`, plus `GEMINI_API_KEY` for the pipeline itself. - The pipeline Postgres is deliberately separate from the app database so pipeline work can never damage dev data. - SQLite staging: `data-pipeline/db/staging.db`, created from `data-pipeline/db/schema.sql`. ### 2.3 The Gemini prompt (drafted; templating is Phase 3) - `data-pipeline/prompt` is a plain UTF-8 text file, edited by hand. It is currently pinned to a concrete sample run (Spanish nouns, 20 words inlined). - The prompt specifies: - The exact JSON structure expected (from design-doc §6.3) - That definitions and examples must be in the word's language - That translations are needed for all 4 other supported languages - That gender must be provided for de/fr/es/it, null for en - That difficulty must be one of: easy, medium, hard - That multiple senses should be included for polysemous words - **Known bugs to fix when templating (§3.2):** rules 2 and 3 still say `language` must be `"en"`; rule 15 lists target languages `de, it, es, fr` while the header says `en, it, de, fr`; rule 31 says "valid English noun". All are leftovers from adapting the English version. --- ## Phase 3: Data Pipeline — Detailed ### 3.1 SQLite schema (done) - `data-pipeline/db/schema.sql` mirrors the Postgres schema with two SQLite concessions: IDs are `TEXT` (`crypto.randomUUID()`), and `definitions` / `examples` are JSON-encoded strings because SQLite has no array type. The import script parses them back into Postgres `TEXT[]`. - The UNIQUE constraints match Postgres: `words(headword, language_code, pos)`, `senses(word_id, sense_index)`, `translations(sense_id, target_language_code, translation)`. ### 3.2 Template the prompt - Turn `data-pipeline/prompt` into a template. Substitution slots: - source language (name + code) - POS - target languages (the other 4 codes) - the word batch (20 words, one per line) - Fix the hardcoded leftovers listed in §2.3 as part of this — after templating, the language codes in the rules must derive from the substitutions, so this class of bug can't recur. - Request **structured output** from the API (`responseMimeType: "application/json"` with a `responseSchema` matching design-doc §6.3) instead of relying on prompt instructions alone. Keep fence-stripping as a defensive fallback only. ### 3.3 Write the pipeline script - Entry point: `data-pipeline/pipeline.ts` (`pnpm --filter @lila/pipeline pipeline:run`). - Pseudocode: ``` for each language in [de, en, es, fr, it]: read wordlist file → trim, drop empties, dedup in memory query staging.db for existing (headword, language_code, pos) skip words already staged ← idempotency split the remainder into batches of 20 for each batch: call Gemini API (structured output) write the raw response to responses/{lang}-{pos}-{n}.json parse JSON response for each entry in response: validate(entry) if valid: generate UUIDs for word, senses, translations INSERT word + senses + translations in ONE transaction else: append to rejection log log progress: "Batch 12/86 done. 238 words staged, 2 rejected." sleep 1s (rate limiting); retry with backoff on API errors ``` - **Idempotency decision (resolves the pipeline.ts step 3 question):** wordlists are read fresh on every run; the staging DB is the record of what's been processed. Words already present in `staging.db` are skipped before batching, so re-runs cost no API calls for staged words. Each validated entry (word + its senses + their translations) is written in a single transaction, so a partially-written word can never exist and no NOT NULL constraint needs relaxing. The UNIQUE constraints remain as a backstop (`INSERT OR IGNORE`). - Persisting raw responses means a validation-rule change re-validates from disk instead of re-paying for ~430 API calls (~8,700 words ÷ 20 per batch). - Use `better-sqlite3` for SQLite access (synchronous, simple). - Generate UUIDs with `crypto.randomUUID()`. - Store definitions/examples as JSON strings in SQLite (`JSON.stringify(arr)`). ### 3.4 Validation module - Create the validation module alongside the pipeline, with unit tests (vitest is already configured; `data-pipeline/vitest.config.ts` expects `tests/**/*.test.ts`). - Input: one parsed Gemini entry (the JSON object for one word). - Checks (return a list of errors, empty = valid): - `headword` is a non-empty string - `language` is in ['en','de','it','fr','es'] - `pos` is in ['noun','verb','adjective','adverb'] - `senses` is a non-empty array - Each sense has: - `sense_index` is a non-negative integer - `difficulty` is in ['easy','medium','hard'] - `definitions` is a non-empty array of non-empty strings - `examples` is a non-empty array of non-empty strings - `translations` is a non-empty array - Each translation has: - `target_language` is in the supported list - `target_language` != the word's own language - `word` is a non-empty string - `gender` is valid for the target language (de: masculine/feminine/neuter, fr/es/it: masculine/feminine, en: null) - `difficulty` is in ['easy','medium','hard'] - Output: `{ valid: boolean, errors: string[] }` ### 3.5 Run and review - Run the pipeline for all 5 languages. - Check the rejection log. If rejection rate > 10%, fix the prompt and re-run — the skip-processed check means only rejected/missing words are re-sent. - Spot-check: open the SQLite DB, run: ```sql SELECT w.headword, s.definitions, t.translation, t.gender FROM words w JOIN senses s ON s.word_id = w.id JOIN translations t ON t.sense_id = s.id WHERE w.language_code = 'de' AND w.pos = 'noun' ORDER BY RANDOM() LIMIT 20; ``` - Read the definitions. Are they in German? Do they make sense? Are the genders correct? Are the difficulties reasonable? --- ## Phase 4: Import — Detailed ### 4.1 Migration (done in Phase 1) `packages/db/drizzle/0012_graceful_psynapse.sql` creates the three tables with all CHECK/UNIQUE constraints, the three indexes, and cascading FKs. It is applied locally; the tables exist and are empty. Nothing to do here — prod gets the same migration in Phase 6. ### 4.2 Write the import script - Create `data-pipeline/import-to-postgres.ts`. - Add `@lila/db` as a workspace dependency for the Drizzle client and schema (the pipeline currently ships only `better-sqlite3`). - Pseudocode: ``` open SQLite database (read-only) connect to Postgres via Drizzle (PIPELINE_DATABASE_URL first) read all words from SQLite for each batch of words: begin transaction INSERT words into Postgres for each word: read its senses from SQLite parse definitions/examples from JSON string → TEXT[] INSERT senses into Postgres for each sense: read its translations from SQLite INSERT translations into Postgres commit transaction log progress log final counts: words, senses, translations ``` - Parse `definitions` and `examples` from JSON strings (SQLite) into actual arrays (Postgres TEXT[]). - Use `ON CONFLICT DO NOTHING` to handle re-runs gracefully. ### 4.3 Run and verify - Run against the pipeline Postgres (:5433) first; once the script is trusted, point it at the app dev database (:5432). - Compare counts: ```sql -- In SQLite SELECT COUNT(*) FROM words; SELECT COUNT(*) FROM senses; SELECT COUNT(*) FROM translations; -- In Postgres SELECT COUNT(*) FROM words; SELECT COUNT(*) FROM senses; SELECT COUNT(*) FROM translations; ``` - Counts should match (minus any rows that failed validation). - Run the game query manually in Postgres: ```sql SELECT w.headword, s.definitions, s.examples, t.translation, t.gender FROM words w JOIN senses s ON s.word_id = w.id JOIN translations t ON t.sense_id = s.id WHERE w.language_code = 'de' AND w.pos = 'noun' AND s.difficulty IN ('easy', 'medium') AND t.target_language_code = 'es' AND t.difficulty = 'medium' ORDER BY RANDOM() LIMIT 5; ``` - Verify the results make sense. --- ## Phase 5: App Integration — Detailed ### 5.1 Rewrite getGameTerms - Open `packages/db/src/models/termModel.ts`. - Replace the old query (2-table join on `vocabulary_entries` + `entry_translations`) with the new 3-table join (words → senses → translations). - Parameters: sourceLanguage, targetLanguage, pos, difficulty, rounds. - Difficulty filter: - `senses.difficulty IN (all levels up to and including selected)` — ceiling logic - `translations.difficulty = selected` — exact match - Return: word_id, headword, sense_id, definitions, examples, translation, gender. ### 5.2 Rewrite getDistractors - Same 3-table join. - Additional filters: - `t.sense_id != :currentSenseId` - `t.translation != :correctAnswer` - LIMIT 3. - If fewer than 3 distractors are found (small data pool), handle gracefully: reduce the number of options or log a warning. ### 5.3 Update exercise generation - In the function that assembles a game round: - Pick one random definition: `definitions[Math.floor(Math.random() * definitions.length)]` - Pick one random example: `examples[Math.floor(Math.random() * examples.length)]` - Combine 1 correct translation + 3 distractors. - Shuffle the 4 options (Fisher-Yates or similar). - Attach gender to each option for display. - Keep the server-side evaluation invariant: the correct answer is never included in what is sent to the client. ### 5.4 Test matrix Run through this matrix manually in the dev app: | Source | Target | POS | Difficulty | Rounds | Pass? | | ------ | ------ | ---- | ---------- | ------ | ----- | | de | es | noun | easy | 10 | | | es | de | noun | medium | 10 | | | en | fr | noun | hard | 10 | | | it | de | noun | easy | 10 | | | fr | en | noun | medium | 10 | | For each: verify definitions are in the source language, translations are in the target language, genders are shown, no same-sense synonyms appear as distractors, no duplicate options. --- ## Phase 6: Production Deploy — Detailed ### 6.1 Pre-deploy checklist - [ ] All Phase 5 acceptance criteria pass - [ ] `git status` is clean, all changes committed - [ ] The Drizzle migration file is committed to the repo - [ ] You know your prod database connection string - [ ] You have a backup method for prod (pg_dump, hosting provider snapshot, etc.) ### 6.2 Deploy - Back up prod: ``` pg_dump -h -U -d > backup_$(date +%Y%m%d).sql ``` - Run migration on prod: ``` DATABASE_URL= npx drizzle-kit migrate ``` - Run import script against prod: ``` DATABASE_URL= npx tsx data-pipeline/import-to-postgres.ts ``` - Verify counts on prod. ### 6.3 Post-deploy verification - Open the live app. Play 2 full games with different settings. - Check server logs for errors. - Wait 24h. Check logs again. - If everything is clean, drop old tables: ```sql DROP TABLE IF EXISTS entry_translations; DROP TABLE IF EXISTS vocabulary_entries; ``` (Or keep them for another week if you want a safety net.) ### 6.4 Rollback plan If something goes wrong: - Restore the backup: ``` psql -h -U -d < backup_YYYYMMDD.sql ``` - Revert the code to the previous commit. - Redeploy. --- ## Phase 7: Extend POS — Detailed ### 7.1 Per POS Repeat the Phase 3 pipeline for each new POS: - Acquire wordlists (verbs, adjectives, adverbs) for all 5 languages, saved as `source-data/{lang}/{pos}` with the shared POS codes. - Adjust the Gemini prompt template: - Verbs: ask for transitivity, common prepositions, or other verb-specific metadata if needed for future exercises. - Adjectives: ask for the base form. Note that adjective metadata differs from noun metadata (no gender on the adjective itself in the same way — gender applies to the noun it modifies). - Adverbs: typically simpler. May not need gender at all. - Run pipeline → validate → SQLite → import → Postgres. - Test in the app with the POS filter set to the new type. ### 7.2 No schema changes The `pos` column already supports all four types. The `gender` column is nullable and simply won't be populated for adverbs or English words. No migration needed. --- # Risks & Mitigations | Risk | Likelihood | Impact | Mitigation | | ------------------------------------------------------ | ----------------- | ------------------------- | ----------------------------------------------------------------------------------------------------- | | Gemini returns malformed JSON | Low | Pipeline stalls | Structured output (responseSchema); fence-stripping + try/catch as fallback; log and skip bad batches | | Validation rules change after a full run | Medium | Wasted API spend | Raw responses persisted per batch; re-validate from disk instead of re-calling the API | | Gemini produces wrong genders or difficulties | Medium | Bad exercise data | Validation module rejects invalid entries. Spot-check 50+ entries per language | | Too few words at a given difficulty/pos/language combo | Medium | Game can't fill 4 options | Graceful fallback: reduce options or show a "not enough words" message. Log which combos are thin | | SQLite → Postgres import fails mid-way | Low | Partial data in Postgres | Transactions per batch. Re-run with ON CONFLICT DO NOTHING. Test on :5433 before touching :5432 | | ORDER BY RANDOM() gets slow at scale | Low (not at 500k) | Slow game load | Add a TODO comment. Optimize with random() pre-filter when needed | | Prod migration breaks the live app | Low | Downtime | Backup before migrating. Rollback plan documented. Deploy during low-traffic hours | --- # File / Folder Structure ``` project/ ├── data-pipeline/ │ ├── source-data/ │ │ ├── de/noun ← wordlist files (shared lang/POS codes) │ │ ├── en/noun │ │ ├── es/noun │ │ ├── fr/noun │ │ └── it/noun │ ├── db/ │ │ ├── schema.sql ← SQLite schema definition ✅ │ │ └── staging.db ← SQLite staging database (gitignored) ✅ │ ├── prompt ← Gemini prompt (plain UTF-8; to be templated) │ ├── responses/ ← raw Gemini responses, one file per batch (planned) │ ├── rejections/ ← invalid Gemini entries for review (planned) │ ├── pipeline.ts ← main pipeline script (pseudocode today) │ ├── tests/ ← vitest unit tests, esp. validation (planned) │ └── import-to-postgres.ts ← SQLite → Postgres import (planned) ├── packages/ │ ├── db/ │ │ ├── src/db/schema.ts ← Drizzle schema ✅ │ │ └── drizzle/0012_*.sql ← words/senses/translations migration ✅ │ └── shared/ │ └── src/constants.ts ← NOUN_GENDERS ✅, "medium" ✅ ├── documentation/ │ ├── DATA_PIPELINE.md ← orientation layer │ └── pipeline/ │ ├── design-doc.md ← schema design doc ✅ │ └── roadmap.md ← this document └── ... ```