The pipeline docs still described pipeline.ts as pseudocode and the validation module as unwritten. Both have been implemented and run. - CLAUDE.md: replace the "no executable pipeline yet" description with the actual module flow, plus the two invariants worth preserving (resumability via headword diffing, raw responses saved before parsing) - DATA_PIPELINE.md: mark the seven implemented modules, add a module responsibility map and the CLI flag table, drop the resolved warning about hardcoded prompt values - roadmap.md: check off phase 3 tasks 3.2/3.3/3.4, record where the build diverged from the plan, note that validate.ts is stricter than its own spec - STATUS.md: phase 3 is data work now, not code work Also records two open issues: the systemic difficulty-ordering rejection cause, and the hard-tier shortfall pending a full run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
825 lines
33 KiB
Markdown
825 lines
33 KiB
Markdown
# Vocabulary Trainer — Roadmap: Data Pipeline & Schema Migration
|
||
|
||
> **Objective:** Replace the OpenWordNet/kaikki data source with a
|
||
> Gemini-powered pipeline that generates high-quality vocabulary data,
|
||
> backed by a normalized Postgres schema.
|
||
>
|
||
> **Author:** lila
|
||
> **Date:** July 2026 · Last reviewed: 2026-08-20
|
||
> **Companion doc:** `documentation/pipeline/design-doc.md`
|
||
|
||
---
|
||
|
||
## How to Read This Document
|
||
|
||
This roadmap has three zoom levels:
|
||
|
||
- **Part 1 — Overview:** The phases at a glance. Read this to
|
||
understand the full scope in 30 seconds.
|
||
- **Part 2 — Phase Breakdown:** Goals, tasks, dependencies, and
|
||
acceptance criteria per phase. Read this to plan your week.
|
||
- **Part 3 — Detailed Tasks:** Step-by-step instructions within each
|
||
phase. Read this when you sit down to code.
|
||
|
||
---
|
||
|
||
# Part 1 — High-Level Overview
|
||
|
||
```
|
||
Phase 1 Schema ✅
|
||
New Drizzle schema (words, senses, translations) with
|
||
relations, constraints, and indexes. Committed and applied
|
||
(migration 0012_graceful_psynapse.sql) — tables exist, empty.
|
||
|
||
Phase 2 Preparation ✅
|
||
Wordlists acquired, databases set up, prompt drafted and
|
||
tested. One loose end carried into Phase 3 and resolved
|
||
there: the prompt is now templated. The wordlist files still
|
||
contain duplicates, deduped at runtime rather than in the
|
||
files — harmless, and left as-is.
|
||
|
||
Phase 3 Data Pipeline ← CURRENT
|
||
Build the Gemini → validate → SQLite pipeline.
|
||
Produce a clean dataset: full deduped noun lists
|
||
(~1,550–1,750 unique words) × 5 languages.
|
||
Code is complete and unit-tested; the remaining work is the
|
||
data run, the rejection review, and the spot-check.
|
||
|
||
Phase 4 Import
|
||
(Migration already applied in Phase 1.)
|
||
Write the SQLite → Postgres import script.
|
||
Test it against the pipeline Postgres (:5433) first,
|
||
then load the app dev database (:5432).
|
||
|
||
Phase 5 App Integration
|
||
Rewrite the game and distractor queries against the new schema.
|
||
Update the exercise-generation logic.
|
||
Test the full game flow in dev.
|
||
|
||
Phase 6 Production Deploy
|
||
Import the dataset into prod (schema arrives via the normal
|
||
Drizzle migration flow).
|
||
Verify the live app works end-to-end.
|
||
|
||
Phase 7 Extend POS
|
||
Run the pipeline for verbs, adjectives, adverbs.
|
||
No schema changes needed — new wordlists + adjusted prompts.
|
||
|
||
Phase 8 Future Features (out of scope for now)
|
||
Inflection tables, conjugation/declension exercises,
|
||
gender exercises, spaced-repetition scheduling.
|
||
```
|
||
|
||
**Dependency chain:**
|
||
|
||
```
|
||
Phase 1 ✅ → Phase 2 ✅ → Phase 3 → Phase 4 → Phase 5 → Phase 6 → Phase 7
|
||
│
|
||
▼
|
||
Phase 8
|
||
```
|
||
|
||
---
|
||
|
||
# Part 2 — Phase Breakdown
|
||
|
||
---
|
||
|
||
## Phase 1: Schema ✅ COMPLETE
|
||
|
||
**Goal:** New normalized schema exists in Drizzle with all tables,
|
||
constraints, indexes, and relations.
|
||
|
||
**Completed tasks:**
|
||
|
||
- [x] `words` table: headword, language_code, pos, UNIQUE, CHECKs, index
|
||
- [x] `senses` table: word_id FK (cascade), sense_index, difficulty,
|
||
definitions TEXT[], examples TEXT[], UNIQUE, CHECK, index
|
||
- [x] `translations` table: sense_id FK (cascade), target_language_code,
|
||
translation, gender (nullable), difficulty, UNIQUE, CHECKs, index
|
||
- [x] Relations: words→senses (many), senses→word (one) + translations
|
||
(many), translations→sense (one)
|
||
- [x] `NOUN_GENDERS` constant added to `@lila/shared`
|
||
- [x] `DIFFICULTY_LEVELS` updated: "intermediate" → "medium"
|
||
- [x] Old tables (`vocabulary_entries`, `entry_translations`) untouched
|
||
- [x] Auth and lobby tables untouched
|
||
- [x] Build passes, committed
|
||
- [x] Migration generated, inspected, and applied to local Postgres
|
||
(`packages/db/drizzle/0012_graceful_psynapse.sql`)
|
||
|
||
---
|
||
|
||
## Phase 2: Preparation ✅ COMPLETE
|
||
|
||
**Goal:** Have everything you need before writing pipeline code.
|
||
|
||
**Tasks:**
|
||
|
||
- [x] Acquire frequency-based noun lists for all 5 languages
|
||
- [x] Format the lists (one word per line, UTF-8) — **note:** the files
|
||
still contain 110–147 duplicate words each; dedup happens at
|
||
runtime in Phase 3, not in the files
|
||
- [x] Set up local Postgres (docker compose: app DB :5432, dedicated
|
||
pipeline DB :5433)
|
||
- [x] Set up a SQLite database file for staging
|
||
(`data-pipeline/db/staging.db` from `db/schema.sql`)
|
||
- [x] Write and test the Gemini prompt with sample words
|
||
- [x] Refine the prompt until the JSON output matches the contract
|
||
defined in `design-doc.md` §6.3 — **note:** at the end of Phase 2 the
|
||
prompt worked but was pinned to a hardcoded Spanish sample and had
|
||
known copy-paste bugs. Both were resolved by the templating task in
|
||
Phase 3.
|
||
|
||
**Dependencies:** Phase 1 complete.
|
||
|
||
**Acceptance criteria (met):**
|
||
|
||
- 5 wordlist files exist in `data-pipeline/source-data/{lang}/noun`
|
||
(language codes `de/en/es/fr/it`, matching `@lila/shared` constants).
|
||
- A Gemini call with a batch of nouns returns valid JSON matching the
|
||
contract, including definitions, examples, translations with gender,
|
||
and difficulty levels.
|
||
- Local Postgres is running and reachable from the app.
|
||
|
||
---
|
||
|
||
## Phase 3: Data Pipeline ← CURRENT
|
||
|
||
**Goal:** A repeatable script that takes a wordlist, calls Gemini in
|
||
batches of 20, validates the output, and writes clean rows to SQLite.
|
||
Re-runs skip words that are already staged.
|
||
|
||
**Tasks:**
|
||
|
||
- [x] Create the SQLite schema (`data-pipeline/db/schema.sql`, tables
|
||
created in `db/staging.db`)
|
||
- [x] Template the Gemini prompt (`promptTemplate.ts` + placeholders in
|
||
`prompt`); the copy-paste bugs (rules 2–3 hardcoding `"en"`,
|
||
rule 15's contradictory target list, "valid English noun" in
|
||
rule 31) are fixed. `renderPrompt()` throws if any placeholder
|
||
survives rendering.
|
||
- [x] Write the wordlist normalization step (`sourceLists.ts`:
|
||
`normalizeWords()` trims, drops empty lines, dedups in memory
|
||
preserving order)
|
||
- [x] Write the validation module (`validate.ts`, rules in design-doc
|
||
§6.4) with unit tests — `tests/` now holds `validate.test.ts`,
|
||
`staging.test.ts`, `promptTemplate.test.ts`, `sourceLists.test.ts`
|
||
- [x] Write the pipeline script (`pipeline.ts`): reads + normalizes
|
||
wordlists, skips words already in `staging.db`, batches the
|
||
remainder, calls Gemini with structured output
|
||
(`responseMimeType: "application/json"` + `responseSchema` from
|
||
`gemini.ts`), persists each raw response to `responses/` before
|
||
validating, validates each entry, writes valid entries to SQLite
|
||
one transaction per word, and logs invalid entries to
|
||
`rejections/{lang}-{pos}.jsonl`. Adds CLI flags beyond the plan:
|
||
`--langs`, `--pos`, `--batch-size`, `--max-batches`, `--delay-ms`,
|
||
`--dry-run`.
|
||
- [ ] Run the pipeline for all 5 languages (nouns only) — **in progress**,
|
||
German first
|
||
- [ ] Review the rejection log, fix prompt issues, re-run failed batches
|
||
— reviewed; one systemic cause found (see below), fix not yet applied
|
||
- [ ] Spot-check 50 random entries for correctness
|
||
|
||
**Dependencies:** Phase 2 complete.
|
||
|
||
**Acceptance criteria:**
|
||
|
||
- SQLite database contains the full deduped noun lists
|
||
(~1,550–1,750 unique words × 5 languages) with senses and translations.
|
||
- Rejection rate is below 10%.
|
||
- Spot-checked entries have correct definitions, plausible examples,
|
||
correct genders, and reasonable difficulty levels.
|
||
- Re-running the pipeline skips already-staged words (no duplicate
|
||
rows, no repeated API calls for the same words).
|
||
- Raw Gemini responses are on disk, so validation-rule changes can be
|
||
re-applied without re-calling the API.
|
||
|
||
**Status against the criteria:** resumability is verified working (a live run
|
||
resumed correctly from previously staged words), raw responses are on disk,
|
||
and the reject rate is tracking well under 10%. The two open items are the
|
||
systemic rejection cause and the hard-tier shortfall below.
|
||
|
||
**Open issue — systemic rejection cause.** Effectively every rejection is
|
||
`translation difficulty lower than sense difficulty`: the model tags a sense
|
||
`medium` while correctly tagging some of its translations `easy`. Example:
|
||
`Ellbogen` with sense `medium` but `elbow` (en) and `codo` (es) as `easy` —
|
||
the translations are right and the sense label is wrong, yet the whole entry
|
||
is discarded. The prompt already defines sense difficulty as the easiest
|
||
translation difficulty in the sense, so the value is derivable. Normalizing it
|
||
in `validate.ts` rather than rejecting would recover these entries and can be
|
||
replayed against `responses/` without new API calls.
|
||
|
||
**Open issue — the `hard` tier is nearly empty.** Generated difficulty skews
|
||
heavily easy; `hard` translations are well under 1% of staged rows. Because
|
||
design-doc §5.1 filters translation difficulty as an _exact_ match, a "hard"
|
||
game currently resolves to a single-digit row count for a given language pair —
|
||
not enough for one round plus three distractors. This is prompt-calibration
|
||
work and belongs here in Phase 3, before the import in Phase 4, since fixing it
|
||
afterwards means re-importing.
|
||
|
||
---
|
||
|
||
## Phase 4: Import
|
||
|
||
**Goal:** The SQLite data is imported into Postgres.
|
||
|
||
The Drizzle migration was already generated, inspected, and applied in
|
||
Phase 1 (`0012_graceful_psynapse.sql`) — the `words`/`senses`/
|
||
`translations` tables exist and are empty. What remains is the import
|
||
script.
|
||
|
||
**Tasks:**
|
||
|
||
- [ ] Write the import script (`data-pipeline/import-to-postgres.ts`): - Read all rows from SQLite - Insert into Postgres in dependency order:
|
||
words → senses → translations - Use batch inserts (not row-by-row) via Drizzle — add `@lila/db`
|
||
as a workspace dependency (the pipeline currently has no
|
||
Postgres client) - Wrap in transactions (per batch of 20 words) - Handle duplicates gracefully (`ON CONFLICT DO NOTHING`)
|
||
- [ ] Run it against the **pipeline Postgres (:5433,
|
||
`PIPELINE_DATABASE_URL`)** first — this database exists so a bad
|
||
import can never damage dev data
|
||
- [ ] Verify row counts match between SQLite and Postgres
|
||
- [ ] Run 3–5 manual SQL queries against Postgres to sanity-check
|
||
the data
|
||
- [ ] Once trusted, run it against the app dev database (:5432)
|
||
|
||
**Dependencies:** Phase 3 complete (SQLite has data).
|
||
|
||
**Acceptance criteria:**
|
||
|
||
- Import script completes without errors on :5433 and then :5432.
|
||
- Row counts in Postgres match SQLite (±rejection count).
|
||
- Manual query: "Give me 5 random German nouns with Spanish
|
||
translations at easy difficulty" returns sensible results.
|
||
|
||
---
|
||
|
||
## Phase 5: App Integration
|
||
|
||
**Goal:** The running app uses the new schema. Game rounds and
|
||
distractors work correctly for all language pairs.
|
||
|
||
**Tasks:**
|
||
|
||
- [ ] Rewrite `getGameTerms` query: - JOIN words → senses → translations - Filter: source language, pos, sense difficulty (ceiling),
|
||
target language, translation difficulty (exact) - ORDER BY RANDOM(), LIMIT rounds
|
||
- [ ] Rewrite `getDistractors` query: - Same JOINs and filters - Exclude: `sense_id != current`, `translation != correct` - ORDER BY RANDOM(), LIMIT 3
|
||
- [ ] Update the exercise-generation logic: - Pick one random definition from the `definitions` array - Pick one random example from the `examples` array - Assemble the 4 answer options (1 correct + 3 distractors) - Shuffle the options - Keep the existing invariant: the correct answer is evaluated
|
||
server-side and is never sent to the client
|
||
- [ ] Remove or deprecate old schema references
|
||
(old `vocabulary_entries`, `entry_translations` tables)
|
||
- [ ] Test manually: - German → Spanish, nouns, easy, 10 rounds - Spanish → German, nouns, medium, 10 rounds - English → French, nouns, hard, 10 rounds - Italian → German, nouns, easy, 10 rounds - Verify: no duplicate answers, no same-sense synonyms as
|
||
distractors, definitions and examples are in the source
|
||
language, genders display correctly
|
||
- [ ] Test edge cases: - A difficulty/pos/language combo with very few words
|
||
(does the app handle < 4 available words gracefully?) - A word with multiple senses (does the correct sense appear?)
|
||
|
||
**Dependencies:** Phase 4 complete (Postgres has data, schema exists).
|
||
|
||
**Acceptance criteria:**
|
||
|
||
- Full game flow works in dev for at least 4 different language pairs.
|
||
- No same-sense synonyms appear as distractors.
|
||
- Definitions and examples are in the correct language.
|
||
- Gender is displayed where applicable.
|
||
- No console errors or unhandled query failures.
|
||
|
||
---
|
||
|
||
## Phase 6: Production Deploy
|
||
|
||
**Goal:** The live app runs on the new schema with the new data.
|
||
|
||
**Tasks:**
|
||
|
||
- [ ] Back up the production database
|
||
- [ ] Run the Drizzle migration on prod
|
||
- [ ] Run the import script against prod Postgres
|
||
- [ ] Verify row counts on prod
|
||
- [ ] Test the live app: - Play 2 full games on the deployed app - Check different language pairs and difficulties
|
||
- [ ] Monitor for errors (server logs, browser console) for 24h
|
||
- [ ] Remove old tables from prod (after confirming everything works)
|
||
|
||
**Dependencies:** Phase 5 complete (dev is fully working).
|
||
|
||
**Acceptance criteria:**
|
||
|
||
- Live app serves game rounds from the new schema.
|
||
- No errors in server logs for 24 hours.
|
||
- Old tables are dropped (or scheduled for removal).
|
||
- Rollback plan exists: the prod backup can be restored if needed.
|
||
|
||
---
|
||
|
||
## Phase 7: Extend POS
|
||
|
||
**Goal:** Verbs, adjectives, and adverbs are in the database and
|
||
usable in the app.
|
||
|
||
**Tasks:**
|
||
|
||
- [ ] Acquire frequency lists for verbs, adjectives, adverbs
|
||
(all 5 languages)
|
||
- [ ] Adjust the Gemini prompt per POS: - Verbs: may need different metadata (transitivity, etc.) - Adjectives: may need base form info - Adverbs: typically simpler metadata
|
||
- [ ] Run the pipeline for each POS
|
||
- [ ] Validate, import to Postgres (dev → prod)
|
||
- [ ] Test game flow with verbs, adjectives, adverbs
|
||
- [ ] Verify the POS filter in the app UI works for all types
|
||
|
||
**Dependencies:** Phase 6 complete.
|
||
|
||
**Acceptance criteria:**
|
||
|
||
- All 4 POS types are playable in the app.
|
||
- Pipeline is repeatable for future data additions.
|
||
|
||
---
|
||
|
||
## Phase 8: Future Features (out of scope)
|
||
|
||
Listed here for visibility. Not planned, not estimated.
|
||
|
||
- [ ] `inflection_forms` table + conjugation exercises (verbs)
|
||
- [ ] Adjective declension exercises (der grüne Mann, grüner Mann…)
|
||
- [ ] Gender exercises (pick the correct article)
|
||
- [ ] Spaced-repetition scheduling (track which words the user knows)
|
||
- [ ] User accounts and progress persistence
|
||
- [ ] Additional languages (if ever)
|
||
|
||
---
|
||
|
||
# Part 3 — Detailed Task Breakdown
|
||
|
||
---
|
||
|
||
## Phase 2: Preparation — Detailed ✅
|
||
|
||
### 2.1 Wordlists (done)
|
||
|
||
- Frequency-based lists live in `data-pipeline/source-data/{lang}/noun`
|
||
with `lang` ∈ `de/en/es/fr/it` — the same codes as
|
||
`packages/shared/src/constants.ts`, so no name mapping is needed
|
||
anywhere in the pipeline.
|
||
- Format: plain text, one word per line, UTF-8, no headers.
|
||
- Sizes: ~1,700–1,900 lines per language. The files were **not**
|
||
deduplicated (110–147 duplicates each); the pipeline dedups at read
|
||
time (§3.3) rather than editing the source files.
|
||
|
||
### 2.2 Databases (done)
|
||
|
||
- `docker compose up -d` starts everything: app Postgres (:5432),
|
||
dedicated pipeline Postgres (:5433), Valkey (:6379).
|
||
- Config lives in the single root `.env` (see `.env.example`):
|
||
`PIPELINE_POSTGRES_USER` / `PIPELINE_POSTGRES_PASSWORD` /
|
||
`PIPELINE_POSTGRES_DB` / `PIPELINE_DATABASE_URL`, plus
|
||
`GEMINI_API_KEY` for the pipeline itself.
|
||
- The pipeline Postgres is deliberately separate from the app database
|
||
so pipeline work can never damage dev data.
|
||
- SQLite staging: `data-pipeline/db/staging.db`, created from
|
||
`data-pipeline/db/schema.sql`.
|
||
|
||
### 2.3 The Gemini prompt (drafted; templating is Phase 3)
|
||
|
||
- `data-pipeline/prompt` is a plain UTF-8 text file, edited by hand.
|
||
It is currently pinned to a concrete sample run (Spanish nouns,
|
||
20 words inlined).
|
||
- The prompt specifies:
|
||
- The exact JSON structure expected (from design-doc §6.3)
|
||
- That definitions and examples must be in the word's language
|
||
- That translations are needed for all 4 other supported languages
|
||
- That gender must be provided for de/fr/es/it, null for en
|
||
- That difficulty must be one of: easy, medium, hard
|
||
- That multiple senses should be included for polysemous words
|
||
- **Known bugs to fix when templating (§3.2):** rules 2 and 3 still say
|
||
`language` must be `"en"`; rule 15 lists target languages
|
||
`de, it, es, fr` while the header says `en, it, de, fr`; rule 31
|
||
says "valid English noun". All are leftovers from adapting the
|
||
English version.
|
||
|
||
---
|
||
|
||
## Phase 3: Data Pipeline — Detailed
|
||
|
||
### 3.1 SQLite schema (done)
|
||
|
||
- `data-pipeline/db/schema.sql` mirrors the Postgres schema with two
|
||
SQLite concessions: IDs are `TEXT` (`crypto.randomUUID()`), and
|
||
`definitions` / `examples` are JSON-encoded strings because SQLite
|
||
has no array type. The import script parses them back into Postgres
|
||
`TEXT[]`.
|
||
- The UNIQUE constraints match Postgres:
|
||
`words(headword, language_code, pos)`, `senses(word_id, sense_index)`,
|
||
`translations(sense_id, target_language_code, translation)`.
|
||
|
||
### 3.2 Template the prompt (done)
|
||
|
||
Implemented in `promptTemplate.ts`. Structured output is implemented in
|
||
`gemini.ts` via `buildEntriesResponseSchema()`, which narrows the enums to
|
||
the batch's own source language, POS, and target languages rather than
|
||
allowing the full supported list. Fence-stripping is retained in
|
||
`parseEntries()` as the defensive fallback the plan called for.
|
||
|
||
Original spec:
|
||
|
||
- Turn `data-pipeline/prompt` into a template. Substitution slots:
|
||
- source language (name + code)
|
||
- POS
|
||
- target languages (the other 4 codes)
|
||
- the word batch (20 words, one per line)
|
||
- Fix the hardcoded leftovers listed in §2.3 as part of this — after
|
||
templating, the language codes in the rules must derive from the
|
||
substitutions, so this class of bug can't recur.
|
||
- Request **structured output** from the API
|
||
(`responseMimeType: "application/json"` with a `responseSchema`
|
||
matching design-doc §6.3) instead of relying on prompt instructions
|
||
alone. Keep fence-stripping as a defensive fallback only.
|
||
|
||
### 3.3 Write the pipeline script (done)
|
||
|
||
Implemented as specced. Divergences from the plan below: the rate-limit pause
|
||
defaults to 6s rather than 1s, batching/language selection is controlled by CLI
|
||
flags (`--langs`, `--pos`, `--batch-size`, `--max-batches`, `--delay-ms`,
|
||
`--dry-run`), and `stageEntry()` does an explicit existence check inside the
|
||
transaction instead of relying on `INSERT OR IGNORE`. Words the model omits
|
||
from a response entirely are logged to the rejection file as
|
||
`"missing from Gemini response"`, so a silent drop can't go unnoticed.
|
||
|
||
- Entry point: `data-pipeline/pipeline.ts`
|
||
(`pnpm --filter @lila/pipeline pipeline:run`).
|
||
- Pseudocode:
|
||
|
||
```
|
||
for each language in [de, en, es, fr, it]:
|
||
read wordlist file → trim, drop empties, dedup in memory
|
||
query staging.db for existing (headword, language_code, pos)
|
||
skip words already staged ← idempotency
|
||
split the remainder into batches of 20
|
||
for each batch:
|
||
call Gemini API (structured output)
|
||
write the raw response to responses/{lang}-{pos}-{n}.json
|
||
parse JSON response
|
||
for each entry in response:
|
||
validate(entry)
|
||
if valid:
|
||
generate UUIDs for word, senses, translations
|
||
INSERT word + senses + translations in ONE transaction
|
||
else:
|
||
append to rejection log
|
||
log progress: "Batch 12/86 done. 238 words staged, 2 rejected."
|
||
sleep 1s (rate limiting); retry with backoff on API errors
|
||
```
|
||
|
||
- **Idempotency decision (resolves the pipeline.ts step 3 question):**
|
||
wordlists are read fresh on every run; the staging DB is the record
|
||
of what's been processed. Words already present in `staging.db` are
|
||
skipped before batching, so re-runs cost no API calls for staged
|
||
words. Each validated entry (word + its senses + their translations)
|
||
is written in a single transaction, so a partially-written word can
|
||
never exist and no NOT NULL constraint needs relaxing. The UNIQUE
|
||
constraints remain as a backstop (`INSERT OR IGNORE`).
|
||
- Persisting raw responses means a validation-rule change re-validates
|
||
from disk instead of re-paying for ~430 API calls
|
||
(~8,700 words ÷ 20 per batch).
|
||
- Use `better-sqlite3` for SQLite access (synchronous, simple).
|
||
- Generate UUIDs with `crypto.randomUUID()`.
|
||
- Store definitions/examples as JSON strings in SQLite
|
||
(`JSON.stringify(arr)`).
|
||
|
||
### 3.4 Validation module (done)
|
||
|
||
Implemented in `validate.ts` with unit tests in `tests/validate.test.ts`.
|
||
|
||
**As built, it is stricter than this spec.** Two structural differences:
|
||
|
||
- The result is a three-way discriminated union, not `{ valid, errors }`.
|
||
`"empty"` (the contract's `"senses": []`, meaning "not a valid word of this
|
||
POS") is a distinct outcome from `"invalid"`, so genuinely-not-a-noun words
|
||
are skipped and counted separately rather than polluting the rejection log.
|
||
- Validation is context-aware. `ValidationContext` carries the batch's source
|
||
language, POS, target languages, and input words, so the checks below are
|
||
exact-match against the request rather than membership in the global
|
||
supported list.
|
||
|
||
Additional checks not in the original spec:
|
||
|
||
- `headword` must be one of the words actually sent in this batch.
|
||
- `language` must equal the batch's source language; `pos` must equal the
|
||
batch's POS.
|
||
- `sense_index` must equal the sense's position in the array (sequential
|
||
from 0), not merely be a non-negative integer.
|
||
- At most 3 senses per word.
|
||
- Every target language must have at least one translation, at most 2, with
|
||
no duplicate translation word within a target language.
|
||
- A translation's difficulty may not rank below its sense's difficulty —
|
||
this is the cross-field rule responsible for essentially all current
|
||
rejections (see the open issue in Phase 3 above).
|
||
|
||
Original spec:
|
||
|
||
- Create the validation module alongside the pipeline, with unit tests
|
||
(vitest is already configured; `data-pipeline/vitest.config.ts`
|
||
expects `tests/**/*.test.ts`).
|
||
- Input: one parsed Gemini entry (the JSON object for one word).
|
||
- Checks (return a list of errors, empty = valid):
|
||
- `headword` is a non-empty string
|
||
- `language` is in ['en','de','it','fr','es']
|
||
- `pos` is in ['noun','verb','adjective','adverb']
|
||
- `senses` is a non-empty array
|
||
- Each sense has:
|
||
- `sense_index` is a non-negative integer
|
||
- `difficulty` is in ['easy','medium','hard']
|
||
- `definitions` is a non-empty array of non-empty strings
|
||
- `examples` is a non-empty array of non-empty strings
|
||
- `translations` is a non-empty array
|
||
- Each translation has:
|
||
- `target_language` is in the supported list
|
||
- `target_language` != the word's own language
|
||
- `word` is a non-empty string
|
||
- `gender` is valid for the target language
|
||
(de: masculine/feminine/neuter,
|
||
fr/es/it: masculine/feminine,
|
||
en: null)
|
||
- `difficulty` is in ['easy','medium','hard']
|
||
- Output: `{ valid: boolean, errors: string[] }`
|
||
|
||
### 3.5 Run and review
|
||
|
||
- Run the pipeline for all 5 languages.
|
||
- Check the rejection log. If rejection rate > 10%, fix the prompt
|
||
and re-run — the skip-processed check means only rejected/missing
|
||
words are re-sent.
|
||
- Spot-check: open the SQLite DB, run:
|
||
```sql
|
||
SELECT w.headword, s.definitions, t.translation, t.gender
|
||
FROM words w
|
||
JOIN senses s ON s.word_id = w.id
|
||
JOIN translations t ON t.sense_id = s.id
|
||
WHERE w.language_code = 'de' AND w.pos = 'noun'
|
||
ORDER BY RANDOM()
|
||
LIMIT 20;
|
||
```
|
||
- Read the definitions. Are they in German? Do they make sense?
|
||
Are the genders correct? Are the difficulties reasonable?
|
||
|
||
---
|
||
|
||
## Phase 4: Import — Detailed
|
||
|
||
### 4.1 Migration (done in Phase 1)
|
||
|
||
`packages/db/drizzle/0012_graceful_psynapse.sql` creates the three
|
||
tables with all CHECK/UNIQUE constraints, the three indexes, and
|
||
cascading FKs. It is applied locally; the tables exist and are empty.
|
||
Nothing to do here — prod gets the same migration in Phase 6.
|
||
|
||
### 4.2 Write the import script
|
||
|
||
- Create `data-pipeline/import-to-postgres.ts`.
|
||
- Add `@lila/db` as a workspace dependency for the Drizzle client and
|
||
schema (the pipeline currently ships only `better-sqlite3`).
|
||
- Pseudocode:
|
||
|
||
```
|
||
open SQLite database (read-only)
|
||
connect to Postgres via Drizzle (PIPELINE_DATABASE_URL first)
|
||
|
||
read all words from SQLite
|
||
for each batch of words:
|
||
begin transaction
|
||
INSERT words into Postgres
|
||
for each word:
|
||
read its senses from SQLite
|
||
parse definitions/examples from JSON string → TEXT[]
|
||
INSERT senses into Postgres
|
||
for each sense:
|
||
read its translations from SQLite
|
||
INSERT translations into Postgres
|
||
commit transaction
|
||
log progress
|
||
|
||
log final counts: words, senses, translations
|
||
```
|
||
|
||
- Parse `definitions` and `examples` from JSON strings (SQLite)
|
||
into actual arrays (Postgres TEXT[]).
|
||
- Use `ON CONFLICT DO NOTHING` to handle re-runs gracefully.
|
||
|
||
### 4.3 Run and verify
|
||
|
||
- Run against the pipeline Postgres (:5433) first; once the script is
|
||
trusted, point it at the app dev database (:5432).
|
||
- Compare counts:
|
||
|
||
```sql
|
||
-- In SQLite
|
||
SELECT COUNT(*) FROM words;
|
||
SELECT COUNT(*) FROM senses;
|
||
SELECT COUNT(*) FROM translations;
|
||
|
||
-- In Postgres
|
||
SELECT COUNT(*) FROM words;
|
||
SELECT COUNT(*) FROM senses;
|
||
SELECT COUNT(*) FROM translations;
|
||
```
|
||
|
||
- Counts should match (minus any rows that failed validation).
|
||
- Run the game query manually in Postgres:
|
||
```sql
|
||
SELECT w.headword, s.definitions, s.examples,
|
||
t.translation, t.gender
|
||
FROM words w
|
||
JOIN senses s ON s.word_id = w.id
|
||
JOIN translations t ON t.sense_id = s.id
|
||
WHERE w.language_code = 'de'
|
||
AND w.pos = 'noun'
|
||
AND s.difficulty IN ('easy', 'medium')
|
||
AND t.target_language_code = 'es'
|
||
AND t.difficulty = 'medium'
|
||
ORDER BY RANDOM()
|
||
LIMIT 5;
|
||
```
|
||
- Verify the results make sense.
|
||
|
||
---
|
||
|
||
## Phase 5: App Integration — Detailed
|
||
|
||
### 5.1 Rewrite getGameTerms
|
||
|
||
- Open `packages/db/src/models/termModel.ts`.
|
||
- Replace the old query (2-table join on `vocabulary_entries` +
|
||
`entry_translations`) with the new 3-table join
|
||
(words → senses → translations).
|
||
- Parameters: sourceLanguage, targetLanguage, pos, difficulty, rounds.
|
||
- Difficulty filter:
|
||
- `senses.difficulty IN (all levels up to and including selected)`
|
||
— ceiling logic
|
||
- `translations.difficulty = selected` — exact match
|
||
- Return: word_id, headword, sense_id, definitions, examples,
|
||
translation, gender.
|
||
|
||
### 5.2 Rewrite getDistractors
|
||
|
||
- Same 3-table join.
|
||
- Additional filters:
|
||
- `t.sense_id != :currentSenseId`
|
||
- `t.translation != :correctAnswer`
|
||
- LIMIT 3.
|
||
- If fewer than 3 distractors are found (small data pool), handle
|
||
gracefully: reduce the number of options or log a warning.
|
||
|
||
### 5.3 Update exercise generation
|
||
|
||
- In the function that assembles a game round:
|
||
- Pick one random definition:
|
||
`definitions[Math.floor(Math.random() * definitions.length)]`
|
||
- Pick one random example:
|
||
`examples[Math.floor(Math.random() * examples.length)]`
|
||
- Combine 1 correct translation + 3 distractors.
|
||
- Shuffle the 4 options (Fisher-Yates or similar).
|
||
- Attach gender to each option for display.
|
||
- Keep the server-side evaluation invariant: the correct answer is
|
||
never included in what is sent to the client.
|
||
|
||
### 5.4 Test matrix
|
||
|
||
Run through this matrix manually in the dev app:
|
||
|
||
| Source | Target | POS | Difficulty | Rounds | Pass? |
|
||
| ------ | ------ | ---- | ---------- | ------ | ----- |
|
||
| de | es | noun | easy | 10 | |
|
||
| es | de | noun | medium | 10 | |
|
||
| en | fr | noun | hard | 10 | |
|
||
| it | de | noun | easy | 10 | |
|
||
| fr | en | noun | medium | 10 | |
|
||
|
||
For each: verify definitions are in the source language, translations
|
||
are in the target language, genders are shown, no same-sense synonyms
|
||
appear as distractors, no duplicate options.
|
||
|
||
---
|
||
|
||
## Phase 6: Production Deploy — Detailed
|
||
|
||
### 6.1 Pre-deploy checklist
|
||
|
||
- [ ] All Phase 5 acceptance criteria pass
|
||
- [ ] `git status` is clean, all changes committed
|
||
- [ ] The Drizzle migration file is committed to the repo
|
||
- [ ] You know your prod database connection string
|
||
- [ ] You have a backup method for prod (pg_dump, hosting provider
|
||
snapshot, etc.)
|
||
|
||
### 6.2 Deploy
|
||
|
||
- Back up prod:
|
||
```
|
||
pg_dump -h <host> -U <user> -d <db> > backup_$(date +%Y%m%d).sql
|
||
```
|
||
- Run migration on prod:
|
||
```
|
||
DATABASE_URL=<prod-url> npx drizzle-kit migrate
|
||
```
|
||
- Run import script against prod:
|
||
```
|
||
DATABASE_URL=<prod-url> npx tsx data-pipeline/import-to-postgres.ts
|
||
```
|
||
- Verify counts on prod.
|
||
|
||
### 6.3 Post-deploy verification
|
||
|
||
- Open the live app. Play 2 full games with different settings.
|
||
- Check server logs for errors.
|
||
- Wait 24h. Check logs again.
|
||
- If everything is clean, drop old tables:
|
||
```sql
|
||
DROP TABLE IF EXISTS entry_translations;
|
||
DROP TABLE IF EXISTS vocabulary_entries;
|
||
```
|
||
(Or keep them for another week if you want a safety net.)
|
||
|
||
### 6.4 Rollback plan
|
||
|
||
If something goes wrong:
|
||
|
||
- Restore the backup:
|
||
```
|
||
psql -h <host> -U <user> -d <db> < backup_YYYYMMDD.sql
|
||
```
|
||
- Revert the code to the previous commit.
|
||
- Redeploy.
|
||
|
||
---
|
||
|
||
## Phase 7: Extend POS — Detailed
|
||
|
||
### 7.1 Per POS
|
||
|
||
Repeat the Phase 3 pipeline for each new POS:
|
||
|
||
- Acquire wordlists (verbs, adjectives, adverbs) for all 5 languages,
|
||
saved as `source-data/{lang}/{pos}` with the shared POS codes.
|
||
- Adjust the Gemini prompt template:
|
||
- Verbs: ask for transitivity, common prepositions, or other
|
||
verb-specific metadata if needed for future exercises.
|
||
- Adjectives: ask for the base form. Note that adjective metadata
|
||
differs from noun metadata (no gender on the adjective itself
|
||
in the same way — gender applies to the noun it modifies).
|
||
- Adverbs: typically simpler. May not need gender at all.
|
||
- Run pipeline → validate → SQLite → import → Postgres.
|
||
- Test in the app with the POS filter set to the new type.
|
||
|
||
### 7.2 No schema changes
|
||
|
||
The `pos` column already supports all four types. The `gender`
|
||
column is nullable and simply won't be populated for adverbs or
|
||
English words. No migration needed.
|
||
|
||
---
|
||
|
||
# Risks & Mitigations
|
||
|
||
| Risk | Likelihood | Impact | Mitigation |
|
||
| ------------------------------------------------------ | ----------------- | ------------------------- | ----------------------------------------------------------------------------------------------------- |
|
||
| Gemini returns malformed JSON | Low | Pipeline stalls | Structured output (responseSchema); fence-stripping + try/catch as fallback; log and skip bad batches |
|
||
| Validation rules change after a full run | Medium | Wasted API spend | Raw responses persisted per batch; re-validate from disk instead of re-calling the API |
|
||
| Gemini produces wrong genders or difficulties | Medium | Bad exercise data | Validation module rejects invalid entries. Spot-check 50+ entries per language |
|
||
| Too few words at a given difficulty/pos/language combo | Medium | Game can't fill 4 options | Graceful fallback: reduce options or show a "not enough words" message. Log which combos are thin |
|
||
| SQLite → Postgres import fails mid-way | Low | Partial data in Postgres | Transactions per batch. Re-run with ON CONFLICT DO NOTHING. Test on :5433 before touching :5432 |
|
||
| ORDER BY RANDOM() gets slow at scale | Low (not at 500k) | Slow game load | Add a TODO comment. Optimize with random() pre-filter when needed |
|
||
| Prod migration breaks the live app | Low | Downtime | Backup before migrating. Rollback plan documented. Deploy during low-traffic hours |
|
||
|
||
---
|
||
|
||
# File / Folder Structure
|
||
|
||
```
|
||
project/
|
||
├── data-pipeline/
|
||
│ ├── source-data/
|
||
│ │ ├── de/noun ← wordlist files (shared lang/POS codes)
|
||
│ │ ├── en/noun
|
||
│ │ ├── es/noun
|
||
│ │ ├── fr/noun
|
||
│ │ └── it/noun
|
||
│ ├── db/
|
||
│ │ ├── schema.sql ← SQLite schema definition ✅
|
||
│ │ └── staging.db ← SQLite staging database (gitignored) ✅
|
||
│ ├── prompt ← Gemini prompt (plain UTF-8; to be templated)
|
||
│ ├── responses/ ← raw Gemini responses, one file per batch (planned)
|
||
│ ├── rejections/ ← invalid Gemini entries for review (planned)
|
||
│ ├── pipeline.ts ← main pipeline script (pseudocode today)
|
||
│ ├── tests/ ← vitest unit tests, esp. validation (planned)
|
||
│ └── import-to-postgres.ts ← SQLite → Postgres import (planned)
|
||
├── packages/
|
||
│ ├── db/
|
||
│ │ ├── src/db/schema.ts ← Drizzle schema ✅
|
||
│ │ └── drizzle/0012_*.sql ← words/senses/translations migration ✅
|
||
│ └── shared/
|
||
│ └── src/constants.ts ← NOUN_GENDERS ✅, "medium" ✅
|
||
├── documentation/
|
||
│ ├── DATA_PIPELINE.md ← orientation layer
|
||
│ └── pipeline/
|
||
│ ├── design-doc.md ← schema design doc ✅
|
||
│ └── roadmap.md ← this document
|
||
└── ...
|
||
```
|