Every rejection in the German run was the same rule: a translation ranked below its own sense. The model tags a sense "medium" while correctly tagging some translations "easy" — the translations are right and the derived sense label is wrong, but the whole entry was discarded. The prompt defines sense difficulty as the easiest translation difficulty in that sense, so it is a derived value rather than an independent judgement. validate.ts now recomputes it via applySenseDifficultyFloor. The floor only ever lowers. Raising a sense to match its translations would gate a concept out of levels it belongs in and collapse the concept-vs-word distinction the two difficulty columns exist to express (design-doc section 4). - validate.ts: drop the cross-field rejection, add the floor; the valid result now carries "normalizations" so repairs are reported, not silent - pipeline.ts: count and print normalizations per batch and in the summary - replay.ts: new, re-validates responses/ with the current rules and no API calls; --write stages recovered entries, --langs and --verbose - tests: six cases covering the floor, replacing the old rejection test Replaying all 38 saved responses took the reject rate from 20 entries to zero. staging.db now holds 746 words / 774 senses / 3,436 translations with no sense ranked above its easiest translation. Docs also record the API quota ceiling found today: the free tier allows about 20 requests/day, not the 1,000 previously assumed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
836 lines
34 KiB
Markdown
836 lines
34 KiB
Markdown
# Vocabulary Trainer — Roadmap: Data Pipeline & Schema Migration
|
||
|
||
> **Objective:** Replace the OpenWordNet/kaikki data source with a
|
||
> Gemini-powered pipeline that generates high-quality vocabulary data,
|
||
> backed by a normalized Postgres schema.
|
||
>
|
||
> **Author:** lila
|
||
> **Date:** July 2026 · Last reviewed: 2026-08-20
|
||
> **Companion doc:** `documentation/pipeline/design-doc.md`
|
||
|
||
---
|
||
|
||
## How to Read This Document
|
||
|
||
This roadmap has three zoom levels:
|
||
|
||
- **Part 1 — Overview:** The phases at a glance. Read this to
|
||
understand the full scope in 30 seconds.
|
||
- **Part 2 — Phase Breakdown:** Goals, tasks, dependencies, and
|
||
acceptance criteria per phase. Read this to plan your week.
|
||
- **Part 3 — Detailed Tasks:** Step-by-step instructions within each
|
||
phase. Read this when you sit down to code.
|
||
|
||
---
|
||
|
||
# Part 1 — High-Level Overview
|
||
|
||
```
|
||
Phase 1 Schema ✅
|
||
New Drizzle schema (words, senses, translations) with
|
||
relations, constraints, and indexes. Committed and applied
|
||
(migration 0012_graceful_psynapse.sql) — tables exist, empty.
|
||
|
||
Phase 2 Preparation ✅
|
||
Wordlists acquired, databases set up, prompt drafted and
|
||
tested. One loose end carried into Phase 3 and resolved
|
||
there: the prompt is now templated. The wordlist files still
|
||
contain duplicates, deduped at runtime rather than in the
|
||
files — harmless, and left as-is.
|
||
|
||
Phase 3 Data Pipeline ← CURRENT
|
||
Build the Gemini → validate → SQLite pipeline.
|
||
Produce a clean dataset: full deduped noun lists
|
||
(~1,550–1,750 unique words) × 5 languages.
|
||
Code is complete and unit-tested; the remaining work is the
|
||
data run, the rejection review, and the spot-check.
|
||
|
||
Phase 4 Import
|
||
(Migration already applied in Phase 1.)
|
||
Write the SQLite → Postgres import script.
|
||
Test it against the pipeline Postgres (:5433) first,
|
||
then load the app dev database (:5432).
|
||
|
||
Phase 5 App Integration
|
||
Rewrite the game and distractor queries against the new schema.
|
||
Update the exercise-generation logic.
|
||
Test the full game flow in dev.
|
||
|
||
Phase 6 Production Deploy
|
||
Import the dataset into prod (schema arrives via the normal
|
||
Drizzle migration flow).
|
||
Verify the live app works end-to-end.
|
||
|
||
Phase 7 Extend POS
|
||
Run the pipeline for verbs, adjectives, adverbs.
|
||
No schema changes needed — new wordlists + adjusted prompts.
|
||
|
||
Phase 8 Future Features (out of scope for now)
|
||
Inflection tables, conjugation/declension exercises,
|
||
gender exercises, spaced-repetition scheduling.
|
||
```
|
||
|
||
**Dependency chain:**
|
||
|
||
```
|
||
Phase 1 ✅ → Phase 2 ✅ → Phase 3 → Phase 4 → Phase 5 → Phase 6 → Phase 7
|
||
│
|
||
▼
|
||
Phase 8
|
||
```
|
||
|
||
---
|
||
|
||
# Part 2 — Phase Breakdown
|
||
|
||
---
|
||
|
||
## Phase 1: Schema ✅ COMPLETE
|
||
|
||
**Goal:** New normalized schema exists in Drizzle with all tables,
|
||
constraints, indexes, and relations.
|
||
|
||
**Completed tasks:**
|
||
|
||
- [x] `words` table: headword, language_code, pos, UNIQUE, CHECKs, index
|
||
- [x] `senses` table: word_id FK (cascade), sense_index, difficulty,
|
||
definitions TEXT[], examples TEXT[], UNIQUE, CHECK, index
|
||
- [x] `translations` table: sense_id FK (cascade), target_language_code,
|
||
translation, gender (nullable), difficulty, UNIQUE, CHECKs, index
|
||
- [x] Relations: words→senses (many), senses→word (one) + translations
|
||
(many), translations→sense (one)
|
||
- [x] `NOUN_GENDERS` constant added to `@lila/shared`
|
||
- [x] `DIFFICULTY_LEVELS` updated: "intermediate" → "medium"
|
||
- [x] Old tables (`vocabulary_entries`, `entry_translations`) untouched
|
||
- [x] Auth and lobby tables untouched
|
||
- [x] Build passes, committed
|
||
- [x] Migration generated, inspected, and applied to local Postgres
|
||
(`packages/db/drizzle/0012_graceful_psynapse.sql`)
|
||
|
||
---
|
||
|
||
## Phase 2: Preparation ✅ COMPLETE
|
||
|
||
**Goal:** Have everything you need before writing pipeline code.
|
||
|
||
**Tasks:**
|
||
|
||
- [x] Acquire frequency-based noun lists for all 5 languages
|
||
- [x] Format the lists (one word per line, UTF-8) — **note:** the files
|
||
still contain 110–147 duplicate words each; dedup happens at
|
||
runtime in Phase 3, not in the files
|
||
- [x] Set up local Postgres (docker compose: app DB :5432, dedicated
|
||
pipeline DB :5433)
|
||
- [x] Set up a SQLite database file for staging
|
||
(`data-pipeline/db/staging.db` from `db/schema.sql`)
|
||
- [x] Write and test the Gemini prompt with sample words
|
||
- [x] Refine the prompt until the JSON output matches the contract
|
||
defined in `design-doc.md` §6.3 — **note:** at the end of Phase 2 the
|
||
prompt worked but was pinned to a hardcoded Spanish sample and had
|
||
known copy-paste bugs. Both were resolved by the templating task in
|
||
Phase 3.
|
||
|
||
**Dependencies:** Phase 1 complete.
|
||
|
||
**Acceptance criteria (met):**
|
||
|
||
- 5 wordlist files exist in `data-pipeline/source-data/{lang}/noun`
|
||
(language codes `de/en/es/fr/it`, matching `@lila/shared` constants).
|
||
- A Gemini call with a batch of nouns returns valid JSON matching the
|
||
contract, including definitions, examples, translations with gender,
|
||
and difficulty levels.
|
||
- Local Postgres is running and reachable from the app.
|
||
|
||
---
|
||
|
||
## Phase 3: Data Pipeline ← CURRENT
|
||
|
||
**Goal:** A repeatable script that takes a wordlist, calls Gemini in
|
||
batches of 20, validates the output, and writes clean rows to SQLite.
|
||
Re-runs skip words that are already staged.
|
||
|
||
**Tasks:**
|
||
|
||
- [x] Create the SQLite schema (`data-pipeline/db/schema.sql`, tables
|
||
created in `db/staging.db`)
|
||
- [x] Template the Gemini prompt (`promptTemplate.ts` + placeholders in
|
||
`prompt`); the copy-paste bugs (rules 2–3 hardcoding `"en"`,
|
||
rule 15's contradictory target list, "valid English noun" in
|
||
rule 31) are fixed. `renderPrompt()` throws if any placeholder
|
||
survives rendering.
|
||
- [x] Write the wordlist normalization step (`sourceLists.ts`:
|
||
`normalizeWords()` trims, drops empty lines, dedups in memory
|
||
preserving order)
|
||
- [x] Write the validation module (`validate.ts`, rules in design-doc
|
||
§6.4) with unit tests — `tests/` now holds `validate.test.ts`,
|
||
`staging.test.ts`, `promptTemplate.test.ts`, `sourceLists.test.ts`
|
||
- [x] Write the pipeline script (`pipeline.ts`): reads + normalizes
|
||
wordlists, skips words already in `staging.db`, batches the
|
||
remainder, calls Gemini with structured output
|
||
(`responseMimeType: "application/json"` + `responseSchema` from
|
||
`gemini.ts`), persists each raw response to `responses/` before
|
||
validating, validates each entry, writes valid entries to SQLite
|
||
one transaction per word, and logs invalid entries to
|
||
`rejections/{lang}-{pos}.jsonl`. Adds CLI flags beyond the plan:
|
||
`--langs`, `--pos`, `--batch-size`, `--max-batches`, `--delay-ms`,
|
||
`--dry-run`.
|
||
- [ ] Run the pipeline for all 5 languages (nouns only) — **in progress**,
|
||
German first
|
||
- [x] Review the rejection log, fix prompt issues, re-run failed batches —
|
||
one systemic cause found and fixed in `validate.ts`; recovered from
|
||
`responses/` via `replay.ts`, no re-run needed. Reject rate now zero.
|
||
- [ ] Spot-check 50 random entries for correctness
|
||
|
||
**Dependencies:** Phase 2 complete.
|
||
|
||
**Acceptance criteria:**
|
||
|
||
- SQLite database contains the full deduped noun lists
|
||
(~1,550–1,750 unique words × 5 languages) with senses and translations.
|
||
- Rejection rate is below 10%.
|
||
- Spot-checked entries have correct definitions, plausible examples,
|
||
correct genders, and reasonable difficulty levels.
|
||
- Re-running the pipeline skips already-staged words (no duplicate
|
||
rows, no repeated API calls for the same words).
|
||
- Raw Gemini responses are on disk, so validation-rule changes can be
|
||
re-applied without re-calling the API.
|
||
|
||
**Status against the criteria:** resumability is verified working (a live run
|
||
resumed correctly from previously staged words), raw responses are on disk,
|
||
and the reject rate is tracking well under 10%. The two open items are the
|
||
systemic rejection cause and the hard-tier shortfall below.
|
||
|
||
**Resolved — systemic rejection cause.** Every rejection was
|
||
`translation difficulty lower than sense difficulty`: the model tags a sense
|
||
`medium` while correctly tagging some of its translations `easy`. Example:
|
||
`Ellbogen` with sense `medium` but `elbow` (en) and `codo` (es) as `easy` —
|
||
the translations are right and the sense label is wrong, yet the whole entry
|
||
was discarded. Since the prompt defines sense difficulty as the easiest
|
||
translation difficulty in the sense, the value is derived, so `validate.ts`
|
||
now floors it (`applySenseDifficultyFloor`) rather than rejecting. The floor
|
||
only lowers; raising a sense would gate a concept out of levels it belongs in.
|
||
Replaying all 38 saved responses took the reject rate from 20 entries to zero,
|
||
and `staging.db` now holds 746 words / 774 senses / 3,436 translations with no
|
||
sense ranked above its easiest translation.
|
||
|
||
**Open issue — API quota is the binding constraint.** The free tier allows
|
||
~20 requests/day for `gemini-3.6-flash` (observed 2026-08-20), not the ~1,000
|
||
previously assumed. At `--batch-size 20` the remaining ~7,300 words are ~365
|
||
requests, i.e. roughly 18 days of waiting. Requests are what is rationed, not
|
||
words, so raising `--batch-size` is the cheap lever (~74 requests at 100).
|
||
Paid billing is the alternative. Decide before scheduling the rest of the run.
|
||
|
||
**Open issue — the `hard` tier is nearly empty.** Generated difficulty skews
|
||
heavily easy; `hard` translations are well under 1% of staged rows. Because
|
||
design-doc §5.1 filters translation difficulty as an _exact_ match, a "hard"
|
||
game currently resolves to a single-digit row count for a given language pair —
|
||
not enough for one round plus three distractors. This is prompt-calibration
|
||
work and belongs here in Phase 3, before the import in Phase 4, since fixing it
|
||
afterwards means re-importing.
|
||
|
||
---
|
||
|
||
## Phase 4: Import
|
||
|
||
**Goal:** The SQLite data is imported into Postgres.
|
||
|
||
The Drizzle migration was already generated, inspected, and applied in
|
||
Phase 1 (`0012_graceful_psynapse.sql`) — the `words`/`senses`/
|
||
`translations` tables exist and are empty. What remains is the import
|
||
script.
|
||
|
||
**Tasks:**
|
||
|
||
- [ ] Write the import script (`data-pipeline/import-to-postgres.ts`): - Read all rows from SQLite - Insert into Postgres in dependency order:
|
||
words → senses → translations - Use batch inserts (not row-by-row) via Drizzle — add `@lila/db`
|
||
as a workspace dependency (the pipeline currently has no
|
||
Postgres client) - Wrap in transactions (per batch of 20 words) - Handle duplicates gracefully (`ON CONFLICT DO NOTHING`)
|
||
- [ ] Run it against the **pipeline Postgres (:5433,
|
||
`PIPELINE_DATABASE_URL`)** first — this database exists so a bad
|
||
import can never damage dev data
|
||
- [ ] Verify row counts match between SQLite and Postgres
|
||
- [ ] Run 3–5 manual SQL queries against Postgres to sanity-check
|
||
the data
|
||
- [ ] Once trusted, run it against the app dev database (:5432)
|
||
|
||
**Dependencies:** Phase 3 complete (SQLite has data).
|
||
|
||
**Acceptance criteria:**
|
||
|
||
- Import script completes without errors on :5433 and then :5432.
|
||
- Row counts in Postgres match SQLite (±rejection count).
|
||
- Manual query: "Give me 5 random German nouns with Spanish
|
||
translations at easy difficulty" returns sensible results.
|
||
|
||
---
|
||
|
||
## Phase 5: App Integration
|
||
|
||
**Goal:** The running app uses the new schema. Game rounds and
|
||
distractors work correctly for all language pairs.
|
||
|
||
**Tasks:**
|
||
|
||
- [ ] Rewrite `getGameTerms` query: - JOIN words → senses → translations - Filter: source language, pos, sense difficulty (ceiling),
|
||
target language, translation difficulty (exact) - ORDER BY RANDOM(), LIMIT rounds
|
||
- [ ] Rewrite `getDistractors` query: - Same JOINs and filters - Exclude: `sense_id != current`, `translation != correct` - ORDER BY RANDOM(), LIMIT 3
|
||
- [ ] Update the exercise-generation logic: - Pick one random definition from the `definitions` array - Pick one random example from the `examples` array - Assemble the 4 answer options (1 correct + 3 distractors) - Shuffle the options - Keep the existing invariant: the correct answer is evaluated
|
||
server-side and is never sent to the client
|
||
- [ ] Remove or deprecate old schema references
|
||
(old `vocabulary_entries`, `entry_translations` tables)
|
||
- [ ] Test manually: - German → Spanish, nouns, easy, 10 rounds - Spanish → German, nouns, medium, 10 rounds - English → French, nouns, hard, 10 rounds - Italian → German, nouns, easy, 10 rounds - Verify: no duplicate answers, no same-sense synonyms as
|
||
distractors, definitions and examples are in the source
|
||
language, genders display correctly
|
||
- [ ] Test edge cases: - A difficulty/pos/language combo with very few words
|
||
(does the app handle < 4 available words gracefully?) - A word with multiple senses (does the correct sense appear?)
|
||
|
||
**Dependencies:** Phase 4 complete (Postgres has data, schema exists).
|
||
|
||
**Acceptance criteria:**
|
||
|
||
- Full game flow works in dev for at least 4 different language pairs.
|
||
- No same-sense synonyms appear as distractors.
|
||
- Definitions and examples are in the correct language.
|
||
- Gender is displayed where applicable.
|
||
- No console errors or unhandled query failures.
|
||
|
||
---
|
||
|
||
## Phase 6: Production Deploy
|
||
|
||
**Goal:** The live app runs on the new schema with the new data.
|
||
|
||
**Tasks:**
|
||
|
||
- [ ] Back up the production database
|
||
- [ ] Run the Drizzle migration on prod
|
||
- [ ] Run the import script against prod Postgres
|
||
- [ ] Verify row counts on prod
|
||
- [ ] Test the live app: - Play 2 full games on the deployed app - Check different language pairs and difficulties
|
||
- [ ] Monitor for errors (server logs, browser console) for 24h
|
||
- [ ] Remove old tables from prod (after confirming everything works)
|
||
|
||
**Dependencies:** Phase 5 complete (dev is fully working).
|
||
|
||
**Acceptance criteria:**
|
||
|
||
- Live app serves game rounds from the new schema.
|
||
- No errors in server logs for 24 hours.
|
||
- Old tables are dropped (or scheduled for removal).
|
||
- Rollback plan exists: the prod backup can be restored if needed.
|
||
|
||
---
|
||
|
||
## Phase 7: Extend POS
|
||
|
||
**Goal:** Verbs, adjectives, and adverbs are in the database and
|
||
usable in the app.
|
||
|
||
**Tasks:**
|
||
|
||
- [ ] Acquire frequency lists for verbs, adjectives, adverbs
|
||
(all 5 languages)
|
||
- [ ] Adjust the Gemini prompt per POS: - Verbs: may need different metadata (transitivity, etc.) - Adjectives: may need base form info - Adverbs: typically simpler metadata
|
||
- [ ] Run the pipeline for each POS
|
||
- [ ] Validate, import to Postgres (dev → prod)
|
||
- [ ] Test game flow with verbs, adjectives, adverbs
|
||
- [ ] Verify the POS filter in the app UI works for all types
|
||
|
||
**Dependencies:** Phase 6 complete.
|
||
|
||
**Acceptance criteria:**
|
||
|
||
- All 4 POS types are playable in the app.
|
||
- Pipeline is repeatable for future data additions.
|
||
|
||
---
|
||
|
||
## Phase 8: Future Features (out of scope)
|
||
|
||
Listed here for visibility. Not planned, not estimated.
|
||
|
||
- [ ] `inflection_forms` table + conjugation exercises (verbs)
|
||
- [ ] Adjective declension exercises (der grüne Mann, grüner Mann…)
|
||
- [ ] Gender exercises (pick the correct article)
|
||
- [ ] Spaced-repetition scheduling (track which words the user knows)
|
||
- [ ] User accounts and progress persistence
|
||
- [ ] Additional languages (if ever)
|
||
|
||
---
|
||
|
||
# Part 3 — Detailed Task Breakdown
|
||
|
||
---
|
||
|
||
## Phase 2: Preparation — Detailed ✅
|
||
|
||
### 2.1 Wordlists (done)
|
||
|
||
- Frequency-based lists live in `data-pipeline/source-data/{lang}/noun`
|
||
with `lang` ∈ `de/en/es/fr/it` — the same codes as
|
||
`packages/shared/src/constants.ts`, so no name mapping is needed
|
||
anywhere in the pipeline.
|
||
- Format: plain text, one word per line, UTF-8, no headers.
|
||
- Sizes: ~1,700–1,900 lines per language. The files were **not**
|
||
deduplicated (110–147 duplicates each); the pipeline dedups at read
|
||
time (§3.3) rather than editing the source files.
|
||
|
||
### 2.2 Databases (done)
|
||
|
||
- `docker compose up -d` starts everything: app Postgres (:5432),
|
||
dedicated pipeline Postgres (:5433), Valkey (:6379).
|
||
- Config lives in the single root `.env` (see `.env.example`):
|
||
`PIPELINE_POSTGRES_USER` / `PIPELINE_POSTGRES_PASSWORD` /
|
||
`PIPELINE_POSTGRES_DB` / `PIPELINE_DATABASE_URL`, plus
|
||
`GEMINI_API_KEY` for the pipeline itself.
|
||
- The pipeline Postgres is deliberately separate from the app database
|
||
so pipeline work can never damage dev data.
|
||
- SQLite staging: `data-pipeline/db/staging.db`, created from
|
||
`data-pipeline/db/schema.sql`.
|
||
|
||
### 2.3 The Gemini prompt (drafted; templating is Phase 3)
|
||
|
||
- `data-pipeline/prompt` is a plain UTF-8 text file, edited by hand.
|
||
It is currently pinned to a concrete sample run (Spanish nouns,
|
||
20 words inlined).
|
||
- The prompt specifies:
|
||
- The exact JSON structure expected (from design-doc §6.3)
|
||
- That definitions and examples must be in the word's language
|
||
- That translations are needed for all 4 other supported languages
|
||
- That gender must be provided for de/fr/es/it, null for en
|
||
- That difficulty must be one of: easy, medium, hard
|
||
- That multiple senses should be included for polysemous words
|
||
- **Known bugs to fix when templating (§3.2):** rules 2 and 3 still say
|
||
`language` must be `"en"`; rule 15 lists target languages
|
||
`de, it, es, fr` while the header says `en, it, de, fr`; rule 31
|
||
says "valid English noun". All are leftovers from adapting the
|
||
English version.
|
||
|
||
---
|
||
|
||
## Phase 3: Data Pipeline — Detailed
|
||
|
||
### 3.1 SQLite schema (done)
|
||
|
||
- `data-pipeline/db/schema.sql` mirrors the Postgres schema with two
|
||
SQLite concessions: IDs are `TEXT` (`crypto.randomUUID()`), and
|
||
`definitions` / `examples` are JSON-encoded strings because SQLite
|
||
has no array type. The import script parses them back into Postgres
|
||
`TEXT[]`.
|
||
- The UNIQUE constraints match Postgres:
|
||
`words(headword, language_code, pos)`, `senses(word_id, sense_index)`,
|
||
`translations(sense_id, target_language_code, translation)`.
|
||
|
||
### 3.2 Template the prompt (done)
|
||
|
||
Implemented in `promptTemplate.ts`. Structured output is implemented in
|
||
`gemini.ts` via `buildEntriesResponseSchema()`, which narrows the enums to
|
||
the batch's own source language, POS, and target languages rather than
|
||
allowing the full supported list. Fence-stripping is retained in
|
||
`parseEntries()` as the defensive fallback the plan called for.
|
||
|
||
Original spec:
|
||
|
||
- Turn `data-pipeline/prompt` into a template. Substitution slots:
|
||
- source language (name + code)
|
||
- POS
|
||
- target languages (the other 4 codes)
|
||
- the word batch (20 words, one per line)
|
||
- Fix the hardcoded leftovers listed in §2.3 as part of this — after
|
||
templating, the language codes in the rules must derive from the
|
||
substitutions, so this class of bug can't recur.
|
||
- Request **structured output** from the API
|
||
(`responseMimeType: "application/json"` with a `responseSchema`
|
||
matching design-doc §6.3) instead of relying on prompt instructions
|
||
alone. Keep fence-stripping as a defensive fallback only.
|
||
|
||
### 3.3 Write the pipeline script (done)
|
||
|
||
Implemented as specced. Divergences from the plan below: the rate-limit pause
|
||
defaults to 6s rather than 1s, batching/language selection is controlled by CLI
|
||
flags (`--langs`, `--pos`, `--batch-size`, `--max-batches`, `--delay-ms`,
|
||
`--dry-run`), and `stageEntry()` does an explicit existence check inside the
|
||
transaction instead of relying on `INSERT OR IGNORE`. Words the model omits
|
||
from a response entirely are logged to the rejection file as
|
||
`"missing from Gemini response"`, so a silent drop can't go unnoticed.
|
||
|
||
- Entry point: `data-pipeline/pipeline.ts`
|
||
(`pnpm --filter @lila/pipeline pipeline:run`).
|
||
- Pseudocode:
|
||
|
||
```
|
||
for each language in [de, en, es, fr, it]:
|
||
read wordlist file → trim, drop empties, dedup in memory
|
||
query staging.db for existing (headword, language_code, pos)
|
||
skip words already staged ← idempotency
|
||
split the remainder into batches of 20
|
||
for each batch:
|
||
call Gemini API (structured output)
|
||
write the raw response to responses/{lang}-{pos}-{n}.json
|
||
parse JSON response
|
||
for each entry in response:
|
||
validate(entry)
|
||
if valid:
|
||
generate UUIDs for word, senses, translations
|
||
INSERT word + senses + translations in ONE transaction
|
||
else:
|
||
append to rejection log
|
||
log progress: "Batch 12/86 done. 238 words staged, 2 rejected."
|
||
sleep 1s (rate limiting); retry with backoff on API errors
|
||
```
|
||
|
||
- **Idempotency decision (resolves the pipeline.ts step 3 question):**
|
||
wordlists are read fresh on every run; the staging DB is the record
|
||
of what's been processed. Words already present in `staging.db` are
|
||
skipped before batching, so re-runs cost no API calls for staged
|
||
words. Each validated entry (word + its senses + their translations)
|
||
is written in a single transaction, so a partially-written word can
|
||
never exist and no NOT NULL constraint needs relaxing. The UNIQUE
|
||
constraints remain as a backstop (`INSERT OR IGNORE`).
|
||
- Persisting raw responses means a validation-rule change re-validates
|
||
from disk instead of re-paying for ~430 API calls
|
||
(~8,700 words ÷ 20 per batch).
|
||
- Use `better-sqlite3` for SQLite access (synchronous, simple).
|
||
- Generate UUIDs with `crypto.randomUUID()`.
|
||
- Store definitions/examples as JSON strings in SQLite
|
||
(`JSON.stringify(arr)`).
|
||
|
||
### 3.4 Validation module (done)
|
||
|
||
Implemented in `validate.ts` with unit tests in `tests/validate.test.ts`.
|
||
|
||
**As built, it is stricter than this spec.** Two structural differences:
|
||
|
||
- The result is a three-way discriminated union, not `{ valid, errors }`.
|
||
`"empty"` (the contract's `"senses": []`, meaning "not a valid word of this
|
||
POS") is a distinct outcome from `"invalid"`, so genuinely-not-a-noun words
|
||
are skipped and counted separately rather than polluting the rejection log.
|
||
- Validation is context-aware. `ValidationContext` carries the batch's source
|
||
language, POS, target languages, and input words, so the checks below are
|
||
exact-match against the request rather than membership in the global
|
||
supported list.
|
||
|
||
Additional checks not in the original spec:
|
||
|
||
- `headword` must be one of the words actually sent in this batch.
|
||
- `language` must equal the batch's source language; `pos` must equal the
|
||
batch's POS.
|
||
- `sense_index` must equal the sense's position in the array (sequential
|
||
from 0), not merely be a non-negative integer.
|
||
- At most 3 senses per word.
|
||
- Every target language must have at least one translation, at most 2, with
|
||
no duplicate translation word within a target language.
|
||
- A translation's difficulty may not rank below its sense's difficulty —
|
||
this is the cross-field rule responsible for essentially all current
|
||
rejections (see the open issue in Phase 3 above).
|
||
|
||
Original spec:
|
||
|
||
- Create the validation module alongside the pipeline, with unit tests
|
||
(vitest is already configured; `data-pipeline/vitest.config.ts`
|
||
expects `tests/**/*.test.ts`).
|
||
- Input: one parsed Gemini entry (the JSON object for one word).
|
||
- Checks (return a list of errors, empty = valid):
|
||
- `headword` is a non-empty string
|
||
- `language` is in ['en','de','it','fr','es']
|
||
- `pos` is in ['noun','verb','adjective','adverb']
|
||
- `senses` is a non-empty array
|
||
- Each sense has:
|
||
- `sense_index` is a non-negative integer
|
||
- `difficulty` is in ['easy','medium','hard']
|
||
- `definitions` is a non-empty array of non-empty strings
|
||
- `examples` is a non-empty array of non-empty strings
|
||
- `translations` is a non-empty array
|
||
- Each translation has:
|
||
- `target_language` is in the supported list
|
||
- `target_language` != the word's own language
|
||
- `word` is a non-empty string
|
||
- `gender` is valid for the target language
|
||
(de: masculine/feminine/neuter,
|
||
fr/es/it: masculine/feminine,
|
||
en: null)
|
||
- `difficulty` is in ['easy','medium','hard']
|
||
- Output: `{ valid: boolean, errors: string[] }`
|
||
|
||
### 3.5 Run and review
|
||
|
||
- Run the pipeline for all 5 languages.
|
||
- Check the rejection log. If rejection rate > 10%, fix the prompt
|
||
and re-run — the skip-processed check means only rejected/missing
|
||
words are re-sent.
|
||
- Spot-check: open the SQLite DB, run:
|
||
```sql
|
||
SELECT w.headword, s.definitions, t.translation, t.gender
|
||
FROM words w
|
||
JOIN senses s ON s.word_id = w.id
|
||
JOIN translations t ON t.sense_id = s.id
|
||
WHERE w.language_code = 'de' AND w.pos = 'noun'
|
||
ORDER BY RANDOM()
|
||
LIMIT 20;
|
||
```
|
||
- Read the definitions. Are they in German? Do they make sense?
|
||
Are the genders correct? Are the difficulties reasonable?
|
||
|
||
---
|
||
|
||
## Phase 4: Import — Detailed
|
||
|
||
### 4.1 Migration (done in Phase 1)
|
||
|
||
`packages/db/drizzle/0012_graceful_psynapse.sql` creates the three
|
||
tables with all CHECK/UNIQUE constraints, the three indexes, and
|
||
cascading FKs. It is applied locally; the tables exist and are empty.
|
||
Nothing to do here — prod gets the same migration in Phase 6.
|
||
|
||
### 4.2 Write the import script
|
||
|
||
- Create `data-pipeline/import-to-postgres.ts`.
|
||
- Add `@lila/db` as a workspace dependency for the Drizzle client and
|
||
schema (the pipeline currently ships only `better-sqlite3`).
|
||
- Pseudocode:
|
||
|
||
```
|
||
open SQLite database (read-only)
|
||
connect to Postgres via Drizzle (PIPELINE_DATABASE_URL first)
|
||
|
||
read all words from SQLite
|
||
for each batch of words:
|
||
begin transaction
|
||
INSERT words into Postgres
|
||
for each word:
|
||
read its senses from SQLite
|
||
parse definitions/examples from JSON string → TEXT[]
|
||
INSERT senses into Postgres
|
||
for each sense:
|
||
read its translations from SQLite
|
||
INSERT translations into Postgres
|
||
commit transaction
|
||
log progress
|
||
|
||
log final counts: words, senses, translations
|
||
```
|
||
|
||
- Parse `definitions` and `examples` from JSON strings (SQLite)
|
||
into actual arrays (Postgres TEXT[]).
|
||
- Use `ON CONFLICT DO NOTHING` to handle re-runs gracefully.
|
||
|
||
### 4.3 Run and verify
|
||
|
||
- Run against the pipeline Postgres (:5433) first; once the script is
|
||
trusted, point it at the app dev database (:5432).
|
||
- Compare counts:
|
||
|
||
```sql
|
||
-- In SQLite
|
||
SELECT COUNT(*) FROM words;
|
||
SELECT COUNT(*) FROM senses;
|
||
SELECT COUNT(*) FROM translations;
|
||
|
||
-- In Postgres
|
||
SELECT COUNT(*) FROM words;
|
||
SELECT COUNT(*) FROM senses;
|
||
SELECT COUNT(*) FROM translations;
|
||
```
|
||
|
||
- Counts should match (minus any rows that failed validation).
|
||
- Run the game query manually in Postgres:
|
||
```sql
|
||
SELECT w.headword, s.definitions, s.examples,
|
||
t.translation, t.gender
|
||
FROM words w
|
||
JOIN senses s ON s.word_id = w.id
|
||
JOIN translations t ON t.sense_id = s.id
|
||
WHERE w.language_code = 'de'
|
||
AND w.pos = 'noun'
|
||
AND s.difficulty IN ('easy', 'medium')
|
||
AND t.target_language_code = 'es'
|
||
AND t.difficulty = 'medium'
|
||
ORDER BY RANDOM()
|
||
LIMIT 5;
|
||
```
|
||
- Verify the results make sense.
|
||
|
||
---
|
||
|
||
## Phase 5: App Integration — Detailed
|
||
|
||
### 5.1 Rewrite getGameTerms
|
||
|
||
- Open `packages/db/src/models/termModel.ts`.
|
||
- Replace the old query (2-table join on `vocabulary_entries` +
|
||
`entry_translations`) with the new 3-table join
|
||
(words → senses → translations).
|
||
- Parameters: sourceLanguage, targetLanguage, pos, difficulty, rounds.
|
||
- Difficulty filter:
|
||
- `senses.difficulty IN (all levels up to and including selected)`
|
||
— ceiling logic
|
||
- `translations.difficulty = selected` — exact match
|
||
- Return: word_id, headword, sense_id, definitions, examples,
|
||
translation, gender.
|
||
|
||
### 5.2 Rewrite getDistractors
|
||
|
||
- Same 3-table join.
|
||
- Additional filters:
|
||
- `t.sense_id != :currentSenseId`
|
||
- `t.translation != :correctAnswer`
|
||
- LIMIT 3.
|
||
- If fewer than 3 distractors are found (small data pool), handle
|
||
gracefully: reduce the number of options or log a warning.
|
||
|
||
### 5.3 Update exercise generation
|
||
|
||
- In the function that assembles a game round:
|
||
- Pick one random definition:
|
||
`definitions[Math.floor(Math.random() * definitions.length)]`
|
||
- Pick one random example:
|
||
`examples[Math.floor(Math.random() * examples.length)]`
|
||
- Combine 1 correct translation + 3 distractors.
|
||
- Shuffle the 4 options (Fisher-Yates or similar).
|
||
- Attach gender to each option for display.
|
||
- Keep the server-side evaluation invariant: the correct answer is
|
||
never included in what is sent to the client.
|
||
|
||
### 5.4 Test matrix
|
||
|
||
Run through this matrix manually in the dev app:
|
||
|
||
| Source | Target | POS | Difficulty | Rounds | Pass? |
|
||
| ------ | ------ | ---- | ---------- | ------ | ----- |
|
||
| de | es | noun | easy | 10 | |
|
||
| es | de | noun | medium | 10 | |
|
||
| en | fr | noun | hard | 10 | |
|
||
| it | de | noun | easy | 10 | |
|
||
| fr | en | noun | medium | 10 | |
|
||
|
||
For each: verify definitions are in the source language, translations
|
||
are in the target language, genders are shown, no same-sense synonyms
|
||
appear as distractors, no duplicate options.
|
||
|
||
---
|
||
|
||
## Phase 6: Production Deploy — Detailed
|
||
|
||
### 6.1 Pre-deploy checklist
|
||
|
||
- [ ] All Phase 5 acceptance criteria pass
|
||
- [ ] `git status` is clean, all changes committed
|
||
- [ ] The Drizzle migration file is committed to the repo
|
||
- [ ] You know your prod database connection string
|
||
- [ ] You have a backup method for prod (pg_dump, hosting provider
|
||
snapshot, etc.)
|
||
|
||
### 6.2 Deploy
|
||
|
||
- Back up prod:
|
||
```
|
||
pg_dump -h <host> -U <user> -d <db> > backup_$(date +%Y%m%d).sql
|
||
```
|
||
- Run migration on prod:
|
||
```
|
||
DATABASE_URL=<prod-url> npx drizzle-kit migrate
|
||
```
|
||
- Run import script against prod:
|
||
```
|
||
DATABASE_URL=<prod-url> npx tsx data-pipeline/import-to-postgres.ts
|
||
```
|
||
- Verify counts on prod.
|
||
|
||
### 6.3 Post-deploy verification
|
||
|
||
- Open the live app. Play 2 full games with different settings.
|
||
- Check server logs for errors.
|
||
- Wait 24h. Check logs again.
|
||
- If everything is clean, drop old tables:
|
||
```sql
|
||
DROP TABLE IF EXISTS entry_translations;
|
||
DROP TABLE IF EXISTS vocabulary_entries;
|
||
```
|
||
(Or keep them for another week if you want a safety net.)
|
||
|
||
### 6.4 Rollback plan
|
||
|
||
If something goes wrong:
|
||
|
||
- Restore the backup:
|
||
```
|
||
psql -h <host> -U <user> -d <db> < backup_YYYYMMDD.sql
|
||
```
|
||
- Revert the code to the previous commit.
|
||
- Redeploy.
|
||
|
||
---
|
||
|
||
## Phase 7: Extend POS — Detailed
|
||
|
||
### 7.1 Per POS
|
||
|
||
Repeat the Phase 3 pipeline for each new POS:
|
||
|
||
- Acquire wordlists (verbs, adjectives, adverbs) for all 5 languages,
|
||
saved as `source-data/{lang}/{pos}` with the shared POS codes.
|
||
- Adjust the Gemini prompt template:
|
||
- Verbs: ask for transitivity, common prepositions, or other
|
||
verb-specific metadata if needed for future exercises.
|
||
- Adjectives: ask for the base form. Note that adjective metadata
|
||
differs from noun metadata (no gender on the adjective itself
|
||
in the same way — gender applies to the noun it modifies).
|
||
- Adverbs: typically simpler. May not need gender at all.
|
||
- Run pipeline → validate → SQLite → import → Postgres.
|
||
- Test in the app with the POS filter set to the new type.
|
||
|
||
### 7.2 No schema changes
|
||
|
||
The `pos` column already supports all four types. The `gender`
|
||
column is nullable and simply won't be populated for adverbs or
|
||
English words. No migration needed.
|
||
|
||
---
|
||
|
||
# Risks & Mitigations
|
||
|
||
| Risk | Likelihood | Impact | Mitigation |
|
||
| ------------------------------------------------------ | ----------------- | ------------------------- | ----------------------------------------------------------------------------------------------------- |
|
||
| Gemini returns malformed JSON | Low | Pipeline stalls | Structured output (responseSchema); fence-stripping + try/catch as fallback; log and skip bad batches |
|
||
| Validation rules change after a full run | Medium | Wasted API spend | Raw responses persisted per batch; re-validate from disk instead of re-calling the API |
|
||
| Gemini produces wrong genders or difficulties | Medium | Bad exercise data | Validation module rejects invalid entries. Spot-check 50+ entries per language |
|
||
| Too few words at a given difficulty/pos/language combo | Medium | Game can't fill 4 options | Graceful fallback: reduce options or show a "not enough words" message. Log which combos are thin |
|
||
| SQLite → Postgres import fails mid-way | Low | Partial data in Postgres | Transactions per batch. Re-run with ON CONFLICT DO NOTHING. Test on :5433 before touching :5432 |
|
||
| ORDER BY RANDOM() gets slow at scale | Low (not at 500k) | Slow game load | Add a TODO comment. Optimize with random() pre-filter when needed |
|
||
| Prod migration breaks the live app | Low | Downtime | Backup before migrating. Rollback plan documented. Deploy during low-traffic hours |
|
||
|
||
---
|
||
|
||
# File / Folder Structure
|
||
|
||
```
|
||
project/
|
||
├── data-pipeline/
|
||
│ ├── source-data/
|
||
│ │ ├── de/noun ← wordlist files (shared lang/POS codes)
|
||
│ │ ├── en/noun
|
||
│ │ ├── es/noun
|
||
│ │ ├── fr/noun
|
||
│ │ └── it/noun
|
||
│ ├── db/
|
||
│ │ ├── schema.sql ← SQLite schema definition ✅
|
||
│ │ └── staging.db ← SQLite staging database (gitignored) ✅
|
||
│ ├── prompt ← Gemini prompt (plain UTF-8; to be templated)
|
||
│ ├── responses/ ← raw Gemini responses, one file per batch (planned)
|
||
│ ├── rejections/ ← invalid Gemini entries for review (planned)
|
||
│ ├── pipeline.ts ← main pipeline script (pseudocode today)
|
||
│ ├── tests/ ← vitest unit tests, esp. validation (planned)
|
||
│ └── import-to-postgres.ts ← SQLite → Postgres import (planned)
|
||
├── packages/
|
||
│ ├── db/
|
||
│ │ ├── src/db/schema.ts ← Drizzle schema ✅
|
||
│ │ └── drizzle/0012_*.sql ← words/senses/translations migration ✅
|
||
│ └── shared/
|
||
│ └── src/constants.ts ← NOUN_GENDERS ✅, "medium" ✅
|
||
├── documentation/
|
||
│ ├── DATA_PIPELINE.md ← orientation layer
|
||
│ └── pipeline/
|
||
│ ├── design-doc.md ← schema design doc ✅
|
||
│ └── roadmap.md ← this document
|
||
└── ...
|
||
```
|