updating pipeline roadmap and status to match branch reality

Roadmap: phase 4 re-scoped to import only (migration already applied),
phase 3 gains prompt templating, wordlist dedup, structured output,
raw-response persistence, and validation tests; idempotency decision
documented (skip already-staged words, one transaction per word).
Stale paths, postgres setup, and file structure corrected.

STATUS.md: refreshed from stale 2026-05-15 Kaikki state to current
gemini-only pipeline work.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
lila 2026-08-09 13:32:15 +02:00
parent e534b98bc5
commit 303bb9388c
2 changed files with 242 additions and 213 deletions

View file

@ -4,9 +4,9 @@
> Gemini-powered pipeline that generates high-quality vocabulary data,
> backed by a normalized Postgres schema.
>
> **Author:** [Your Name]
> **Date:** July 2026
> **Companion doc:** `docs/schema-design.md`
> **Author:** lila
> **Date:** July 2026 · Last reviewed: 2026-08-09
> **Companion doc:** `documentation/pipeline/design-doc.md`
---
@ -28,18 +28,25 @@ This roadmap has three zoom levels:
```
Phase 1 Schema ✅
New Drizzle schema (words, senses, translations) with
relations, constraints, and indexes. Committed.
relations, constraints, and indexes. Committed and applied
(migration 0012_graceful_psynapse.sql) — tables exist, empty.
Phase 2 Preparation
Get wordlists, set up tooling, finalize the Gemini prompt.
Phase 2 Preparation ✅
Wordlists acquired, databases set up, prompt drafted and
tested. Two loose ends carried into Phase 3: the prompt is
not templated yet, and the wordlists still contain duplicates
(deduped at runtime, not in the files).
Phase 3 Data Pipeline
Phase 3 Data Pipeline ← CURRENT
Build the Gemini → validate → SQLite pipeline.
Produce a clean dataset of ~1000 nouns × 5 languages.
Produce a clean dataset: full deduped noun lists
(~1,600–1,800 words) × 5 languages.
Phase 4 Migration & Import
Generate and apply the Drizzle migration.
Phase 4 Import
(Migration already applied in Phase 1.)
Write the SQLite → Postgres import script.
Test it against the pipeline Postgres (:5433) first,
then load the app dev database (:5432).
Phase 5 App Integration
Rewrite the game and distractor queries against the new schema.
@ -47,8 +54,8 @@ Phase 5 App Integration
Test the full game flow in dev.
Phase 6 Production Deploy
Run the Drizzle migration on prod.
Import the dataset.
Import the dataset into prod (schema arrives via the normal
Drizzle migration flow).
Verify the live app works end-to-end.
Phase 7 Extend POS
@ -63,10 +70,10 @@ Phase 8 Future Features (out of scope for now)
**Dependency chain:**
```
Phase 1 ✅ → Phase 2 → Phase 3 → Phase 4 → Phase 5 → Phase 6 → Phase 7
│
▼
Phase 8
Phase 1 ✅ → Phase 2 ✅ → Phase 3 → Phase 4 → Phase 5 → Phase 6 → Phase 7
│
▼
Phase 8
```
---
@ -94,46 +101,66 @@ constraints, indexes, and relations.
- [x] Old tables (`vocabulary_entries`, `entry_translations`) untouched
- [x] Auth and lobby tables untouched
- [x] Build passes, committed
- [x] Migration generated, inspected, and applied to local Postgres
(`packages/db/drizzle/0012_graceful_psynapse.sql`)
---
## Phase 2: Preparation
## Phase 2: Preparation ✅ COMPLETE
**Goal:** Have everything you need before writing pipeline code.
**Tasks:**
- [x] Acquire frequency-based noun lists for all 5 languages
- [x] Clean and format the lists (one word per line, UTF-8)
- [x] Set up local Postgres (Docker or native)
- [x] Format the lists (one word per line, UTF-8) — **note:** the files
still contain 110–147 duplicate words each; dedup happens at
runtime in Phase 3, not in the files
- [x] Set up local Postgres (docker compose: app DB :5432, dedicated
pipeline DB :5433)
- [x] Set up a SQLite database file for staging
- [x] Write and test the Gemini prompt with 5 sample words
(`data-pipeline/db/staging.db` from `db/schema.sql`)
- [x] Write and test the Gemini prompt with sample words
- [x] Refine the prompt until the JSON output matches the contract
defined in `docs/schema-design.md` §6.3
defined in `design-doc.md` §6.3 — **note:** the prompt works but
is pinned to a hardcoded Spanish sample and has known copy-paste
bugs; templating and fixes are Phase 3 tasks
**Dependencies:** Phase 1 complete.
**Acceptance criteria:**
**Acceptance criteria (met):**
- You have 5 wordlist files in `data-pipeline/source-data/` (one per
language).
- A Gemini call with 5 German nouns returns valid JSON matching the
- 5 wordlist files exist in `data-pipeline/source-data/{lang}/noun`
(language codes `de/en/es/fr/it`, matching `@lila/shared` constants).
- A Gemini call with a batch of nouns returns valid JSON matching the
contract, including definitions, examples, translations with gender,
and difficulty levels.
- Local Postgres is running and reachable from your app.
- Local Postgres is running and reachable from the app.
---
## Phase 3: Data Pipeline
## Phase 3: Data Pipeline ← CURRENT
**Goal:** A repeatable script that takes a wordlist, calls Gemini in
batches of 20, validates the output, and writes clean rows to SQLite.
Re-runs skip words that are already staged.
**Tasks:**
- [ ] Write the validation module (see schema-design §6.4)
- [ ] Write the pipeline script: - Read wordlist file - Split into batches of 20 - Call Gemini API per batch - Parse JSON response - Validate each entry - Write valid entries to SQLite - Log invalid entries to a rejection file
- [ ] Create the SQLite schema (mirrors the Postgres schema)
- [x] Create the SQLite schema (`data-pipeline/db/schema.sql`, tables
created in `db/staging.db`)
- [ ] Template the Gemini prompt: source language, POS, target
languages, and the word batch become substitutions; fix the known
copy-paste bugs while doing so (rules 2–3 hardcode `"en"`,
rule 15's target list contradicts the header, rule 31 says
"valid English noun")
- [ ] Write the wordlist normalization step: trim whitespace, drop
empty lines, dedup in memory
- [ ] Write the validation module (rules in design-doc §6.4) with unit
tests (`vitest.config.ts` expects `tests/**/*.test.ts` — the
directory doesn't exist yet)
- [ ] Write the pipeline script (`pipeline.ts`): - Read + normalize wordlist files - Skip words already present in staging.db (idempotency — see §3.3) - Split the remainder into batches of 20 - Call Gemini per batch using structured output
(`responseMimeType: "application/json"` + `responseSchema`) - Persist each raw response to disk before validating - Validate each entry - Write valid entries to SQLite, one transaction per word - Log invalid entries to a rejection file
- [ ] Run the pipeline for all 5 languages (nouns only)
- [ ] Review the rejection log, fix prompt issues, re-run failed batches
- [ ] Spot-check 50 random entries for correctness
@ -142,39 +169,46 @@ batches of 20, validates the output, and writes clean rows to SQLite.
**Acceptance criteria:**
- SQLite database contains ~1000 nouns × 5 languages with senses and
translations.
- SQLite database contains the full deduped noun lists
(~1,600–1,800 words × 5 languages) with senses and translations.
- Rejection rate is below 10%.
- Spot-checked entries have correct definitions, plausible examples,
correct genders, and reasonable difficulty levels.
- The pipeline is re-runnable (idempotent or with duplicate handling).
- Re-running the pipeline skips already-staged words (no duplicate
rows, no repeated API calls for the same words).
- Raw Gemini responses are on disk, so validation-rule changes can be
re-applied without re-calling the API.
---
## Phase 4: Migration & Import
## Phase 4: Import
**Goal:** The new schema exists in Postgres and the SQLite data is
imported.
**Goal:** The SQLite data is imported into Postgres.
The Drizzle migration was already generated, inspected, and applied in
Phase 1 (`0012_graceful_psynapse.sql`) — the `words`/`senses`/
`translations` tables exist and are empty. What remains is the import
script.
**Tasks:**
- [ ] Generate the Drizzle migration (`npx drizzle-kit generate`)
- [ ] Inspect the generated SQL file — verify it creates the right
tables, constraints, and indexes
- [ ] Apply the migration to local Postgres (`npx drizzle-kit migrate`)
- [ ] Write the import script (SQLite → Postgres): - Read all rows from SQLite - Insert into Postgres in dependency order:
words → senses → translations - Use batch inserts (not row-by-row) - Wrap in transactions (per batch of 20 words) - Handle duplicates gracefully (skip or upsert)
- [ ] Run the import script
- [ ] Write the import script (`data-pipeline/import-to-postgres.ts`): - Read all rows from SQLite - Insert into Postgres in dependency order:
words → senses → translations - Use batch inserts (not row-by-row) via Drizzle — add `@lila/db`
as a workspace dependency (the pipeline currently has no
Postgres client) - Wrap in transactions (per batch of 20 words) - Handle duplicates gracefully (`ON CONFLICT DO NOTHING`)
- [ ] Run it against the **pipeline Postgres (:5433,
`PIPELINE_DATABASE_URL`)** first — this database exists so a bad
import can never damage dev data
- [ ] Verify row counts match between SQLite and Postgres
- [ ] Run 3–5 manual SQL queries against Postgres to sanity-check
the data
- [ ] Once trusted, run it against the app dev database (:5432)
**Dependencies:** Phase 3 complete (SQLite has data).
**Acceptance criteria:**
- `npx drizzle-kit migrate` runs without errors on local Postgres.
- Import script completes without errors.
- Import script completes without errors on :5433 and then :5432.
- Row counts in Postgres match SQLite (±rejection count).
- Manual query: "Give me 5 random German nouns with Spanish
translations at easy difficulty" returns sensible results.
@ -191,7 +225,8 @@ distractors work correctly for all language pairs.
- [ ] Rewrite `getGameTerms` query: - JOIN words → senses → translations - Filter: source language, pos, sense difficulty (ceiling),
target language, translation difficulty (exact) - ORDER BY RANDOM(), LIMIT rounds
- [ ] Rewrite `getDistractors` query: - Same JOINs and filters - Exclude: `sense_id != current`, `translation != correct` - ORDER BY RANDOM(), LIMIT 3
- [ ] Update the exercise-generation logic: - Pick one random definition from the `definitions` array - Pick one random example from the `examples` array - Assemble the 4 answer options (1 correct + 3 distractors) - Shuffle the options
- [ ] Update the exercise-generation logic: - Pick one random definition from the `definitions` array - Pick one random example from the `examples` array - Assemble the 4 answer options (1 correct + 3 distractors) - Shuffle the options - Keep the existing invariant: the correct answer is evaluated
server-side and is never sent to the client
- [ ] Remove or deprecate old schema references
(old `vocabulary_entries`, `entry_translations` tables)
- [ ] Test manually: - German → Spanish, nouns, easy, 10 rounds - Spanish → German, nouns, medium, 10 rounds - English → French, nouns, hard, 10 rounds - Italian → German, nouns, easy, 10 rounds - Verify: no duplicate answers, no same-sense synonyms as
@ -278,105 +313,128 @@ Listed here for visibility. Not planned, not estimated.
---
## Phase 2: Preparation — Detailed
## Phase 2: Preparation — Detailed ✅
### 2.1 Acquire wordlists
### 2.1 Wordlists (done)
- Search for frequency lists. Good starting points:
- "Leipzig Corpora Collection" (academic, per-language)
- Wiktionary frequency lists
- GitHub repos: search "german noun frequency list",
"spanish noun frequency list", etc.
- Tatoeba sentence counts as a proxy for word frequency
- Target: ~1000 nouns per language, sorted by frequency.
- Frequency-based lists live in `data-pipeline/source-data/{lang}/noun`
with `lang` ∈ `de/en/es/fr/it` — the same codes as
`packages/shared/src/constants.ts`, so no name mapping is needed
anywhere in the pipeline.
- Format: plain text, one word per line, UTF-8, no headers.
- Save to `data-pipeline/source-data/{language}/nouns/`.
- Clean the lists: remove duplicates, remove words with spaces
(multi-word expressions), remove proper nouns if desired.
- Sizes: ~1,700–1,900 lines per language. The files were **not**
deduplicated (110–147 duplicates each); the pipeline dedups at read
time (§3.3) rather than editing the source files.
### 2.2 Set up local Postgres
### 2.2 Databases (done)
- Option A (Docker):
```
docker run --name vocab-dev \
-e POSTGRES_USER=dev \
-e POSTGRES_PASSWORD=dev \
-e POSTGRES_DB=vocab \
-p 5432:5432 \
-d postgres:16
```
- Option B (native install): install Postgres, create a `vocab`
database.
- Update your `.env` / `.env.local` with the connection string.
- Verify: connect with `psql` or a GUI client (TablePlus, DBeaver,
pgAdmin). Run `SELECT 1;`.
- `docker compose up -d` starts everything: app Postgres (:5432),
dedicated pipeline Postgres (:5433), Valkey (:6379).
- Config lives in the single root `.env` (see `.env.example`):
`PIPELINE_POSTGRES_USER` / `PIPELINE_POSTGRES_PASSWORD` /
`PIPELINE_POSTGRES_DB` / `PIPELINE_DATABASE_URL`, plus
`GEMINI_API_KEY` for the pipeline itself.
- The pipeline Postgres is deliberately separate from the app database
so pipeline work can never damage dev data.
- SQLite staging: `data-pipeline/db/staging.db`, created from
`data-pipeline/db/schema.sql`.
### 2.3 Write and test the Gemini prompt
### 2.3 The Gemini prompt (drafted; templating is Phase 3)
- Start with 5 German nouns.
- The prompt should specify:
- The exact JSON structure expected (copy from schema-design §6.3)
- `data-pipeline/prompt` is a plain UTF-8 text file, edited by hand.
It is currently pinned to a concrete sample run (Spanish nouns,
20 words inlined).
- The prompt specifies:
- The exact JSON structure expected (from design-doc §6.3)
- That definitions and examples must be in the word's language
- That translations are needed for all 4 other supported languages
- That gender must be provided for de/fr/es/it, null for en
- That difficulty must be one of: easy, medium, hard
- That multiple senses should be included for polysemous words
- Call the API. Inspect the JSON. Common issues to fix:
- Gemini wraps the JSON in markdown code fences → strip them
- Gemini returns "intermediate" instead of "medium" → normalize
- Gemini omits a language → re-prompt or reject
- Gemini returns gender "common" for German → reject
- Iterate until 3 consecutive batches of 20 return clean JSON.
- **Known bugs to fix when templating (§3.2):** rules 2 and 3 still say
`language` must be `"en"`; rule 15 lists target languages
`de, it, es, fr` while the header says `en, it, de, fr`; rule 31
says "valid English noun". All are leftovers from adapting the
English version.
---
## Phase 3: Data Pipeline — Detailed
### 3.1 Create the SQLite schema
### 3.1 SQLite schema (done)
- Create a file `data-pipeline/schema.sql` or define it in your
pipeline script.
- Tables mirror the Postgres schema:
- `data-pipeline/db/schema.sql` mirrors the Postgres schema with two
SQLite concessions: IDs are `TEXT` (`crypto.randomUUID()`), and
`definitions` / `examples` are JSON-encoded strings because SQLite
has no array type. The import script parses them back into Postgres
`TEXT[]`.
- The UNIQUE constraints match Postgres:
`words(headword, language_code, pos)`, `senses(word_id, sense_index)`,
`translations(sense_id, target_language_code, translation)`.
```sql
CREATE TABLE words (
id TEXT PRIMARY KEY,
headword TEXT NOT NULL,
language_code TEXT NOT NULL,
pos TEXT NOT NULL,
created_at TEXT DEFAULT (datetime('now')),
UNIQUE(headword, language_code, pos)
);
### 3.2 Template the prompt
CREATE TABLE senses (
id TEXT PRIMARY KEY,
word_id TEXT NOT NULL REFERENCES words(id),
sense_index INTEGER NOT NULL DEFAULT 0,
difficulty TEXT NOT NULL,
definitions TEXT NOT NULL DEFAULT '[]', -- JSON array as text
examples TEXT NOT NULL DEFAULT '[]', -- JSON array as text
created_at TEXT DEFAULT (datetime('now')),
UNIQUE(word_id, sense_index)
);
- Turn `data-pipeline/prompt` into a template. Substitution slots:
- source language (name + code)
- POS
- target languages (the other 4 codes)
- the word batch (20 words, one per line)
- Fix the hardcoded leftovers listed in §2.3 as part of this — after
templating, the language codes in the rules must derive from the
substitutions, so this class of bug can't recur.
- Request **structured output** from the API
(`responseMimeType: "application/json"` with a `responseSchema`
matching design-doc §6.3) instead of relying on prompt instructions
alone. Keep fence-stripping as a defensive fallback only.
CREATE TABLE translations (
id TEXT PRIMARY KEY,
sense_id TEXT NOT NULL REFERENCES senses(id),
target_language_code TEXT NOT NULL,
translation TEXT NOT NULL,
gender TEXT,
difficulty TEXT NOT NULL,
created_at TEXT DEFAULT (datetime('now')),
UNIQUE(sense_id, target_language_code, translation)
);
### 3.3 Write the pipeline script
- Entry point: `data-pipeline/pipeline.ts`
(`pnpm --filter @lila/pipeline pipeline:run`).
- Pseudocode:
```
for each language in [de, en, es, fr, it]:
read wordlist file → trim, drop empties, dedup in memory
query staging.db for existing (headword, language_code, pos)
skip words already staged ← idempotency
split the remainder into batches of 20
for each batch:
call Gemini API (structured output)
write the raw response to responses/{lang}-{pos}-{n}.json
parse JSON response
for each entry in response:
validate(entry)
if valid:
generate UUIDs for word, senses, translations
INSERT word + senses + translations in ONE transaction
else:
append to rejection log
log progress: "Batch 12/86 done. 238 words staged, 2 rejected."
sleep 1s (rate limiting); retry with backoff on API errors
```
- Note: SQLite has no native TEXT[]. Store arrays as JSON text.
The import script will parse them into Postgres TEXT[] on import.
- **Idempotency decision (resolves the pipeline.ts step 3 question):**
wordlists are read fresh on every run; the staging DB is the record
of what's been processed. Words already present in `staging.db` are
skipped before batching, so re-runs cost no API calls for staged
words. Each validated entry (word + its senses + their translations)
is written in a single transaction, so a partially-written word can
never exist and no NOT NULL constraint needs relaxing. The UNIQUE
constraints remain as a backstop (`INSERT OR IGNORE`).
- Persisting raw responses means a validation-rule change re-validates
from disk instead of re-paying for ~430 API calls
(~8,700 words ÷ 20 per batch).
- Use `better-sqlite3` for SQLite access (synchronous, simple).
- Generate UUIDs with `crypto.randomUUID()`.
- Store definitions/examples as JSON strings in SQLite
(`JSON.stringify(arr)`).
### 3.2 Write the validation module
### 3.4 Validation module
- Create `data-pipeline/validate.ts`.
- Create the validation module alongside the pipeline, with unit tests
(vitest is already configured; `data-pipeline/vitest.config.ts`
expects `tests/**/*.test.ts`).
- Input: one parsed Gemini entry (the JSON object for one word).
- Checks (return a list of errors, empty = valid):
- `headword` is a non-empty string
@ -394,43 +452,18 @@ Listed here for visibility. Not planned, not estimated.
- `target_language` != the word's own language
- `word` is a non-empty string
- `gender` is valid for the target language
(de: masculine/feminine/neuter/null,
fr/es/it: masculine/feminine/null,
(de: masculine/feminine/neuter,
fr/es/it: masculine/feminine,
en: null)
- `difficulty` is in ['easy','medium','hard']
- Output: `{ valid: boolean, errors: string[] }`
### 3.3 Write the pipeline script
- Create `data-pipeline/run.ts`.
- Pseudocode:
```
for each language in [de, en, es, fr, it]:
read wordlist file → array of words
split into batches of 20
for each batch:
call Gemini API with the batch
parse JSON response (strip markdown fences if present)
for each entry in response:
validate(entry)
if valid:
generate UUIDs for word, senses, translations
INSERT into SQLite (words, senses, translations)
else:
append to rejection log
log progress: "Batch 12/50 done. 238 words imported, 2 rejected."
sleep 1s (rate limiting)
```
- Use `better-sqlite3` for SQLite access (synchronous, simple).
- Generate UUIDs with `crypto.randomUUID()`.
- Store definitions/examples as JSON strings in SQLite
(`JSON.stringify(arr)`).
### 3.4 Run and review
### 3.5 Run and review
- Run the pipeline for all 5 languages.
- Check the rejection log. If rejection rate > 10%, fix the prompt
and re-run failed batches.
and re-run — the skip-processed check means only rejected/missing
words are re-sent.
- Spot-check: open the SQLite DB, run:
```sql
SELECT w.headword, s.definitions, t.translation, t.gender
@ -446,50 +479,36 @@ Listed here for visibility. Not planned, not estimated.
---
## Phase 4: Migration & Import — Detailed
## Phase 4: Import — Detailed
### 4.1 Generate and inspect the migration
### 4.1 Migration (done in Phase 1)
- Run: `npx drizzle-kit generate`
- Open the generated SQL file in `packages/db/drizzle/`.
- Read it. Verify:
- Three CREATE TABLE statements (words, senses, translations)
- CHECK constraints match your schema
- UNIQUE constraints are present
- Three CREATE INDEX statements
- Foreign keys reference the correct tables with ON DELETE CASCADE
- If something looks wrong, fix the schema file and regenerate.
`packages/db/drizzle/0012_graceful_psynapse.sql` creates the three
tables with all CHECK/UNIQUE constraints, the three indexes, and
cascading FKs. It is applied locally; the tables exist and are empty.
Nothing to do here — prod gets the same migration in Phase 6.
### 4.2 Apply the migration locally
- Run: `npx drizzle-kit migrate`
- Connect to local Postgres and verify:
```sql
\dt -- list tables
\d words -- describe words table
\d senses -- describe senses table
\d translations -- describe translations table
```
### 4.3 Write the import script
### 4.2 Write the import script
- Create `data-pipeline/import-to-postgres.ts`.
- Add `@lila/db` as a workspace dependency for the Drizzle client and
schema (the pipeline currently ships only `better-sqlite3`).
- Pseudocode:
```
open SQLite database (read-only)
connect to Postgres via Drizzle
connect to Postgres via Drizzle (PIPELINE_DATABASE_URL first)
read all words from SQLite
for each batch of 100 words:
for each batch of words:
begin transaction
INSERT words into Postgres
for each word:
read its senses from SQLite
parse definitions/examples from JSON string → TEXT[]
INSERT senses into Postgres
for each sense:
read its translations from SQLite
parse definitions/examples from JSON string → TEXT[]
INSERT translations into Postgres
commit transaction
log progress
@ -501,9 +520,10 @@ Listed here for visibility. Not planned, not estimated.
into actual arrays (Postgres TEXT[]).
- Use `ON CONFLICT DO NOTHING` to handle re-runs gracefully.
### 4.4 Run and verify
### 4.3 Run and verify
- Run the import script.
- Run against the pipeline Postgres (:5433) first; once the script is
trusted, point it at the app dev database (:5432).
- Compare counts:
```sql
@ -574,6 +594,8 @@ Listed here for visibility. Not planned, not estimated.
- Combine 1 correct translation + 3 distractors.
- Shuffle the 4 options (Fisher-Yates or similar).
- Attach gender to each option for display.
- Keep the server-side evaluation invariant: the correct answer is
never included in what is sent to the client.
### 5.4 Test matrix
@ -651,8 +673,9 @@ If something goes wrong:
Repeat the Phase 3 pipeline for each new POS:
- Acquire wordlists (verbs, adjectives, adverbs) for all 5 languages.
- Adjust the Gemini prompt:
- Acquire wordlists (verbs, adjectives, adverbs) for all 5 languages,
saved as `source-data/{lang}/{pos}` with the shared POS codes.
- Adjust the Gemini prompt template:
- Verbs: ask for transitivity, common prepositions, or other
verb-specific metadata if needed for future exercises.
- Adjectives: ask for the base form. Note that adjective metadata
@ -672,42 +695,48 @@ English words. No migration needed.
# Risks & Mitigations
| Risk | Likelihood | Impact | Mitigation |
| ------------------------------------------------------ | ----------------- | ------------------------- | ------------------------------------------------------------------------------------------------- |
| Gemini returns malformed JSON | Medium | Pipeline stalls | Strip markdown fences, wrap parsing in try/catch, log and skip bad batches |
| Gemini produces wrong genders or difficulties | Medium | Bad exercise data | Validation module rejects invalid entries. Spot-check 50+ entries per language |
| Too few words at a given difficulty/pos/language combo | Medium | Game can't fill 4 options | Graceful fallback: reduce options or show a "not enough words" message. Log which combos are thin |
| SQLite → Postgres import fails mid-way | Low | Partial data in Postgres | Transactions per batch. Re-run with ON CONFLICT DO NOTHING |
| ORDER BY RANDOM() gets slow at scale | Low (not at 500k) | Slow game load | Add a TODO comment. Optimize with random() pre-filter when needed |
| Prod migration breaks the live app | Low | Downtime | Backup before migrating. Rollback plan documented. Deploy during low-traffic hours |
| Risk | Likelihood | Impact | Mitigation |
| ------------------------------------------------------ | ----------------- | ------------------------- | ----------------------------------------------------------------------------------------------------- |
| Gemini returns malformed JSON | Low | Pipeline stalls | Structured output (responseSchema); fence-stripping + try/catch as fallback; log and skip bad batches |
| Validation rules change after a full run | Medium | Wasted API spend | Raw responses persisted per batch; re-validate from disk instead of re-calling the API |
| Gemini produces wrong genders or difficulties | Medium | Bad exercise data | Validation module rejects invalid entries. Spot-check 50+ entries per language |
| Too few words at a given difficulty/pos/language combo | Medium | Game can't fill 4 options | Graceful fallback: reduce options or show a "not enough words" message. Log which combos are thin |
| SQLite → Postgres import fails mid-way | Low | Partial data in Postgres | Transactions per batch. Re-run with ON CONFLICT DO NOTHING. Test on :5433 before touching :5432 |
| ORDER BY RANDOM() gets slow at scale | Low (not at 500k) | Slow game load | Add a TODO comment. Optimize with random() pre-filter when needed |
| Prod migration breaks the live app | Low | Downtime | Backup before migrating. Rollback plan documented. Deploy during low-traffic hours |
---
# File / Folder Structure (new files)
# File / Folder Structure
```
project/
├── data-pipeline/
│ ├── source-data/
│ │ ├── german/nouns/ ← wordlist files
│ │ ├── english/nouns/
│ │ ├── spanish/nouns/
│ │ ├── french/nouns/
│ │ └── italian/nouns/
│ ├── rejections/ ← invalid Gemini entries (for review)
│ ├── staging.db ← SQLite staging database
│ ├── schema.sql ← SQLite schema definition
│ ├── validate.ts ← validation module
│ ├── run.ts ← main pipeline script
│ └── import-to-postgres.ts ← SQLite → Postgres import
│ │ ├── de/noun ← wordlist files (shared lang/POS codes)
│ │ ├── en/noun
│ │ ├── es/noun
│ │ ├── fr/noun
│ │ └── it/noun
│ ├── db/
│ │ ├── schema.sql ← SQLite schema definition ✅
│ │ └── staging.db ← SQLite staging database (gitignored) ✅
│ ├── prompt ← Gemini prompt (plain UTF-8; to be templated)
│ ├── responses/ ← raw Gemini responses, one file per batch (planned)
│ ├── rejections/ ← invalid Gemini entries for review (planned)
│ ├── pipeline.ts ← main pipeline script (pseudocode today)
│ ├── tests/ ← vitest unit tests, esp. validation (planned)
│ └── import-to-postgres.ts ← SQLite → Postgres import (planned)
├── packages/
│ ├── db/
│ │ └── src/db/schema.ts ← Drizzle schema (updated ✅)
│ │ ├── src/db/schema.ts ← Drizzle schema ✅
│ │ └── drizzle/0012_*.sql ← words/senses/translations migration ✅
│ └── shared/
│ └── src/constants.ts ← NOUN_GENDERS added ✅, "medium" ✅
│ └── src/constants.ts ← NOUN_GENDERS ✅, "medium" ✅
├── documentation/
│ ├── DATA_PIPELINE.md ← orientation layer
│ └── pipeline/
│ ├── design-doc.md ← schema design doc (updated ✅)
│ └── roadmap.md ← this document (updated ✅)
│ ├── design-doc.md ← schema design doc ✅
│ └── roadmap.md ← this document
└── ...
```