updating pipeline roadmap and status to match branch reality
Roadmap: phase 4 re-scoped to import only (migration already applied), phase 3 gains prompt templating, wordlist dedup, structured output, raw-response persistence, and validation tests; idempotency decision documented (skip already-staged words, one transaction per word). Stale paths, postgres setup, and file structure corrected. STATUS.md: refreshed from stale 2026-05-15 Kaikki state to current gemini-only pipeline work. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
parent
e534b98bc5
commit
303bb9388c
2 changed files with 242 additions and 213 deletions
|
|
@ -1,6 +1,6 @@
|
|||
# Status — 2026-05-15
|
||||
# Status — 2026-08-09
|
||||
|
||||
> Last updated: 2026-05-15. Update this file after every deploy or when switching tasks.
|
||||
> Last updated: 2026-08-09. Update this file after every deploy or when switching tasks.
|
||||
|
||||
## What Works Today ✅
|
||||
|
||||
|
|
@ -12,7 +12,7 @@
|
|||
|
||||
## What's Broken / Blocked 🚧
|
||||
|
||||
- **Data quality** — Production still uses OpenWordNet/OMW translations. Kaikki pipeline (sense-disambiguated) is in progress but not yet synced to production.
|
||||
- **Data quality** — Production still uses OpenWordNet/OMW translations. The replacement is the Gemini-only pipeline on branch `refactor/gemini-only-pipeline`: the sense-based schema (`words` → `senses` → `translations`) is migrated but empty, and the pipeline itself is not built yet. The earlier Kaikki/local-LLM pipeline was abandoned and removed (docs in `documentation/archive/`).
|
||||
- **Guest play** — Auth is required for all game routes. No try-before-signup flow.
|
||||
- **Game session store** — Still in-memory (`InMemoryGameSessionStore`). Valkey container exists in local dev but not wired up.
|
||||
- **Rate limiting** — Partially implemented on auth endpoints; game endpoints not yet covered.
|
||||
|
|
@ -21,26 +21,26 @@
|
|||
|
||||
## What I'm Working On Now 🔄
|
||||
|
||||
**Primary:** Rewriting the Kaikki data pipeline enrich script for sub-stage architecture (round1_gloss → round1_example → round1_translations → round1_cefr).
|
||||
**Primary:** Phase 3 of the Gemini-only data pipeline (see [pipeline/roadmap.md](pipeline/roadmap.md)): template the Gemini prompt, write the validation module + `pipeline.ts`, and run the first full staging pass (~1,600–1,800 deduped nouns × 5 languages into SQLite).
|
||||
|
||||
**Secondary:** Phase 7 hardening backlog items (see BACKLOG.md `next` section).
|
||||
|
||||
## Next 2-Week Goal 🎯
|
||||
|
||||
Finish Kaikki Stage 3 (enrich) sub-stage rewrite → run full sample → compare quality → decide on production sync timeline.
|
||||
Finish pipeline Phase 3 (full noun run staged in SQLite, reject rate < 10%, 50 entries spot-checked) → Phase 4 import script (pipeline Postgres :5433, then dev :5432) → start rewriting `getGameTerms`/`getDistractors` against the new schema.
|
||||
|
||||
## The Big Picture
|
||||
|
||||
Lila is a **deployed, working vocabulary quiz app**. The core loop (singleplayer + multiplayer) is solid. The next strategic milestone is **media-based practice** (learn vocab from a song/TV episode/book chapter), but that depends on:
|
||||
|
||||
1. Kaikki data pipeline reaching production (fixes translation quality)
|
||||
1. The Gemini data pipeline reaching production (fixes translation quality)
|
||||
2. A media ingestion prototype (subtitles/lyrics → text → vocab extraction → quiz)
|
||||
|
||||
Until then, the app is a generic vocabulary quiz — functional but not differentiated.
|
||||
|
||||
## Quick Links
|
||||
|
||||
- [BACKLOG.md](BACKLOG.md) — Prioritized tasks
|
||||
- [DATA_PIPELINE.md](DATA_PIPELINE.md) — Pipeline stages and current progress
|
||||
- [BACKLOG.md](BACKLOG.md) — `now` / `next` / `later`
|
||||
- [BACKLOG.md](BACKLOG.md) — Prioritized tasks (`now` / `next` / `later`)
|
||||
- [DATA_PIPELINE.md](DATA_PIPELINE.md) — Pipeline orientation and current progress
|
||||
- [pipeline/roadmap.md](pipeline/roadmap.md) — Phase-by-phase pipeline plan
|
||||
- [DEPLOYMENT.md](DEPLOYMENT.md) — Infrastructure ops
|
||||
|
|
|
|||
|
|
@ -4,9 +4,9 @@
|
|||
> Gemini-powered pipeline that generates high-quality vocabulary data,
|
||||
> backed by a normalized Postgres schema.
|
||||
>
|
||||
> **Author:** [Your Name]
|
||||
> **Date:** July 2026
|
||||
> **Companion doc:** `docs/schema-design.md`
|
||||
> **Author:** lila
|
||||
> **Date:** July 2026 · Last reviewed: 2026-08-09
|
||||
> **Companion doc:** `documentation/pipeline/design-doc.md`
|
||||
|
||||
---
|
||||
|
||||
|
|
@ -28,18 +28,25 @@ This roadmap has three zoom levels:
|
|||
```
|
||||
Phase 1 Schema ✅
|
||||
New Drizzle schema (words, senses, translations) with
|
||||
relations, constraints, and indexes. Committed.
|
||||
relations, constraints, and indexes. Committed and applied
|
||||
(migration 0012_graceful_psynapse.sql) — tables exist, empty.
|
||||
|
||||
Phase 2 Preparation
|
||||
Get wordlists, set up tooling, finalize the Gemini prompt.
|
||||
Phase 2 Preparation ✅
|
||||
Wordlists acquired, databases set up, prompt drafted and
|
||||
tested. Two loose ends carried into Phase 3: the prompt is
|
||||
not templated yet, and the wordlists still contain duplicates
|
||||
(deduped at runtime, not in the files).
|
||||
|
||||
Phase 3 Data Pipeline
|
||||
Phase 3 Data Pipeline ← CURRENT
|
||||
Build the Gemini → validate → SQLite pipeline.
|
||||
Produce a clean dataset of ~1000 nouns × 5 languages.
|
||||
Produce a clean dataset: full deduped noun lists
|
||||
(~1,600–1,800 words) × 5 languages.
|
||||
|
||||
Phase 4 Migration & Import
|
||||
Generate and apply the Drizzle migration.
|
||||
Phase 4 Import
|
||||
(Migration already applied in Phase 1.)
|
||||
Write the SQLite → Postgres import script.
|
||||
Test it against the pipeline Postgres (:5433) first,
|
||||
then load the app dev database (:5432).
|
||||
|
||||
Phase 5 App Integration
|
||||
Rewrite the game and distractor queries against the new schema.
|
||||
|
|
@ -47,8 +54,8 @@ Phase 5 App Integration
|
|||
Test the full game flow in dev.
|
||||
|
||||
Phase 6 Production Deploy
|
||||
Run the Drizzle migration on prod.
|
||||
Import the dataset.
|
||||
Import the dataset into prod (schema arrives via the normal
|
||||
Drizzle migration flow).
|
||||
Verify the live app works end-to-end.
|
||||
|
||||
Phase 7 Extend POS
|
||||
|
|
@ -63,7 +70,7 @@ Phase 8 Future Features (out of scope for now)
|
|||
**Dependency chain:**
|
||||
|
||||
```
|
||||
Phase 1 ✅ → Phase 2 → Phase 3 → Phase 4 → Phase 5 → Phase 6 → Phase 7
|
||||
Phase 1 ✅ → Phase 2 ✅ → Phase 3 → Phase 4 → Phase 5 → Phase 6 → Phase 7
|
||||
│
|
||||
▼
|
||||
Phase 8
|
||||
|
|
@ -94,46 +101,66 @@ constraints, indexes, and relations.
|
|||
- [x] Old tables (`vocabulary_entries`, `entry_translations`) untouched
|
||||
- [x] Auth and lobby tables untouched
|
||||
- [x] Build passes, committed
|
||||
- [x] Migration generated, inspected, and applied to local Postgres
|
||||
(`packages/db/drizzle/0012_graceful_psynapse.sql`)
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Preparation
|
||||
## Phase 2: Preparation ✅ COMPLETE
|
||||
|
||||
**Goal:** Have everything you need before writing pipeline code.
|
||||
|
||||
**Tasks:**
|
||||
|
||||
- [x] Acquire frequency-based noun lists for all 5 languages
|
||||
- [x] Clean and format the lists (one word per line, UTF-8)
|
||||
- [x] Set up local Postgres (Docker or native)
|
||||
- [x] Format the lists (one word per line, UTF-8) — **note:** the files
|
||||
still contain 110–147 duplicate words each; dedup happens at
|
||||
runtime in Phase 3, not in the files
|
||||
- [x] Set up local Postgres (docker compose: app DB :5432, dedicated
|
||||
pipeline DB :5433)
|
||||
- [x] Set up a SQLite database file for staging
|
||||
- [x] Write and test the Gemini prompt with 5 sample words
|
||||
(`data-pipeline/db/staging.db` from `db/schema.sql`)
|
||||
- [x] Write and test the Gemini prompt with sample words
|
||||
- [x] Refine the prompt until the JSON output matches the contract
|
||||
defined in `docs/schema-design.md` §6.3
|
||||
defined in `design-doc.md` §6.3 — **note:** the prompt works but
|
||||
is pinned to a hardcoded Spanish sample and has known copy-paste
|
||||
bugs; templating and fixes are Phase 3 tasks
|
||||
|
||||
**Dependencies:** Phase 1 complete.
|
||||
|
||||
**Acceptance criteria:**
|
||||
**Acceptance criteria (met):**
|
||||
|
||||
- You have 5 wordlist files in `data-pipeline/source-data/` (one per
|
||||
language).
|
||||
- A Gemini call with 5 German nouns returns valid JSON matching the
|
||||
- 5 wordlist files exist in `data-pipeline/source-data/{lang}/noun`
|
||||
(language codes `de/en/es/fr/it`, matching `@lila/shared` constants).
|
||||
- A Gemini call with a batch of nouns returns valid JSON matching the
|
||||
contract, including definitions, examples, translations with gender,
|
||||
and difficulty levels.
|
||||
- Local Postgres is running and reachable from your app.
|
||||
- Local Postgres is running and reachable from the app.
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: Data Pipeline
|
||||
## Phase 3: Data Pipeline ← CURRENT
|
||||
|
||||
**Goal:** A repeatable script that takes a wordlist, calls Gemini in
|
||||
batches of 20, validates the output, and writes clean rows to SQLite.
|
||||
Re-runs skip words that are already staged.
|
||||
|
||||
**Tasks:**
|
||||
|
||||
- [ ] Write the validation module (see schema-design §6.4)
|
||||
- [ ] Write the pipeline script: - Read wordlist file - Split into batches of 20 - Call Gemini API per batch - Parse JSON response - Validate each entry - Write valid entries to SQLite - Log invalid entries to a rejection file
|
||||
- [ ] Create the SQLite schema (mirrors the Postgres schema)
|
||||
- [x] Create the SQLite schema (`data-pipeline/db/schema.sql`, tables
|
||||
created in `db/staging.db`)
|
||||
- [ ] Template the Gemini prompt: source language, POS, target
|
||||
languages, and the word batch become substitutions; fix the known
|
||||
copy-paste bugs while doing so (rules 2–3 hardcode `"en"`,
|
||||
rule 15's target list contradicts the header, rule 31 says
|
||||
"valid English noun")
|
||||
- [ ] Write the wordlist normalization step: trim whitespace, drop
|
||||
empty lines, dedup in memory
|
||||
- [ ] Write the validation module (rules in design-doc §6.4) with unit
|
||||
tests (`vitest.config.ts` expects `tests/**/*.test.ts` — the
|
||||
directory doesn't exist yet)
|
||||
- [ ] Write the pipeline script (`pipeline.ts`): - Read + normalize wordlist files - Skip words already present in staging.db (idempotency — see §3.3) - Split the remainder into batches of 20 - Call Gemini per batch using structured output
|
||||
(`responseMimeType: "application/json"` + `responseSchema`) - Persist each raw response to disk before validating - Validate each entry - Write valid entries to SQLite, one transaction per word - Log invalid entries to a rejection file
|
||||
- [ ] Run the pipeline for all 5 languages (nouns only)
|
||||
- [ ] Review the rejection log, fix prompt issues, re-run failed batches
|
||||
- [ ] Spot-check 50 random entries for correctness
|
||||
|
|
@ -142,39 +169,46 @@ batches of 20, validates the output, and writes clean rows to SQLite.
|
|||
|
||||
**Acceptance criteria:**
|
||||
|
||||
- SQLite database contains ~1000 nouns × 5 languages with senses and
|
||||
translations.
|
||||
- SQLite database contains the full deduped noun lists
|
||||
(~1,600–1,800 words × 5 languages) with senses and translations.
|
||||
- Rejection rate is below 10%.
|
||||
- Spot-checked entries have correct definitions, plausible examples,
|
||||
correct genders, and reasonable difficulty levels.
|
||||
- The pipeline is re-runnable (idempotent or with duplicate handling).
|
||||
- Re-running the pipeline skips already-staged words (no duplicate
|
||||
rows, no repeated API calls for the same words).
|
||||
- Raw Gemini responses are on disk, so validation-rule changes can be
|
||||
re-applied without re-calling the API.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: Migration & Import
|
||||
## Phase 4: Import
|
||||
|
||||
**Goal:** The new schema exists in Postgres and the SQLite data is
|
||||
imported.
|
||||
**Goal:** The SQLite data is imported into Postgres.
|
||||
|
||||
The Drizzle migration was already generated, inspected, and applied in
|
||||
Phase 1 (`0012_graceful_psynapse.sql`) — the `words`/`senses`/
|
||||
`translations` tables exist and are empty. What remains is the import
|
||||
script.
|
||||
|
||||
**Tasks:**
|
||||
|
||||
- [ ] Generate the Drizzle migration (`npx drizzle-kit generate`)
|
||||
- [ ] Inspect the generated SQL file — verify it creates the right
|
||||
tables, constraints, and indexes
|
||||
- [ ] Apply the migration to local Postgres (`npx drizzle-kit migrate`)
|
||||
- [ ] Write the import script (SQLite → Postgres): - Read all rows from SQLite - Insert into Postgres in dependency order:
|
||||
words → senses → translations - Use batch inserts (not row-by-row) - Wrap in transactions (per batch of 20 words) - Handle duplicates gracefully (skip or upsert)
|
||||
- [ ] Run the import script
|
||||
- [ ] Write the import script (`data-pipeline/import-to-postgres.ts`): - Read all rows from SQLite - Insert into Postgres in dependency order:
|
||||
words → senses → translations - Use batch inserts (not row-by-row) via Drizzle — add `@lila/db`
|
||||
as a workspace dependency (the pipeline currently has no
|
||||
Postgres client) - Wrap in transactions (per batch of 20 words) - Handle duplicates gracefully (`ON CONFLICT DO NOTHING`)
|
||||
- [ ] Run it against the **pipeline Postgres (:5433,
|
||||
`PIPELINE_DATABASE_URL`)** first — this database exists so a bad
|
||||
import can never damage dev data
|
||||
- [ ] Verify row counts match between SQLite and Postgres
|
||||
- [ ] Run 3–5 manual SQL queries against Postgres to sanity-check
|
||||
the data
|
||||
- [ ] Once trusted, run it against the app dev database (:5432)
|
||||
|
||||
**Dependencies:** Phase 3 complete (SQLite has data).
|
||||
|
||||
**Acceptance criteria:**
|
||||
|
||||
- `npx drizzle-kit migrate` runs without errors on local Postgres.
|
||||
- Import script completes without errors.
|
||||
- Import script completes without errors on :5433 and then :5432.
|
||||
- Row counts in Postgres match SQLite (±rejection count).
|
||||
- Manual query: "Give me 5 random German nouns with Spanish
|
||||
translations at easy difficulty" returns sensible results.
|
||||
|
|
@ -191,7 +225,8 @@ distractors work correctly for all language pairs.
|
|||
- [ ] Rewrite `getGameTerms` query: - JOIN words → senses → translations - Filter: source language, pos, sense difficulty (ceiling),
|
||||
target language, translation difficulty (exact) - ORDER BY RANDOM(), LIMIT rounds
|
||||
- [ ] Rewrite `getDistractors` query: - Same JOINs and filters - Exclude: `sense_id != current`, `translation != correct` - ORDER BY RANDOM(), LIMIT 3
|
||||
- [ ] Update the exercise-generation logic: - Pick one random definition from the `definitions` array - Pick one random example from the `examples` array - Assemble the 4 answer options (1 correct + 3 distractors) - Shuffle the options
|
||||
- [ ] Update the exercise-generation logic: - Pick one random definition from the `definitions` array - Pick one random example from the `examples` array - Assemble the 4 answer options (1 correct + 3 distractors) - Shuffle the options - Keep the existing invariant: the correct answer is evaluated
|
||||
server-side and is never sent to the client
|
||||
- [ ] Remove or deprecate old schema references
|
||||
(old `vocabulary_entries`, `entry_translations` tables)
|
||||
- [ ] Test manually: - German → Spanish, nouns, easy, 10 rounds - Spanish → German, nouns, medium, 10 rounds - English → French, nouns, hard, 10 rounds - Italian → German, nouns, easy, 10 rounds - Verify: no duplicate answers, no same-sense synonyms as
|
||||
|
|
@ -278,105 +313,128 @@ Listed here for visibility. Not planned, not estimated.
|
|||
|
||||
---
|
||||
|
||||
## Phase 2: Preparation — Detailed
|
||||
## Phase 2: Preparation — Detailed ✅
|
||||
|
||||
### 2.1 Acquire wordlists
|
||||
### 2.1 Wordlists (done)
|
||||
|
||||
- Search for frequency lists. Good starting points:
|
||||
- "Leipzig Corpora Collection" (academic, per-language)
|
||||
- Wiktionary frequency lists
|
||||
- GitHub repos: search "german noun frequency list",
|
||||
"spanish noun frequency list", etc.
|
||||
- Tatoeba sentence counts as a proxy for word frequency
|
||||
- Target: ~1000 nouns per language, sorted by frequency.
|
||||
- Frequency-based lists live in `data-pipeline/source-data/{lang}/noun`
|
||||
with `lang` ∈ `de/en/es/fr/it` — the same codes as
|
||||
`packages/shared/src/constants.ts`, so no name mapping is needed
|
||||
anywhere in the pipeline.
|
||||
- Format: plain text, one word per line, UTF-8, no headers.
|
||||
- Save to `data-pipeline/source-data/{language}/nouns/`.
|
||||
- Clean the lists: remove duplicates, remove words with spaces
|
||||
(multi-word expressions), remove proper nouns if desired.
|
||||
- Sizes: ~1,700–1,900 lines per language. The files were **not**
|
||||
deduplicated (110–147 duplicates each); the pipeline dedups at read
|
||||
time (§3.3) rather than editing the source files.
|
||||
|
||||
### 2.2 Set up local Postgres
|
||||
### 2.2 Databases (done)
|
||||
|
||||
- Option A (Docker):
|
||||
```
|
||||
docker run --name vocab-dev \
|
||||
-e POSTGRES_USER=dev \
|
||||
-e POSTGRES_PASSWORD=dev \
|
||||
-e POSTGRES_DB=vocab \
|
||||
-p 5432:5432 \
|
||||
-d postgres:16
|
||||
```
|
||||
- Option B (native install): install Postgres, create a `vocab`
|
||||
database.
|
||||
- Update your `.env` / `.env.local` with the connection string.
|
||||
- Verify: connect with `psql` or a GUI client (TablePlus, DBeaver,
|
||||
pgAdmin). Run `SELECT 1;`.
|
||||
- `docker compose up -d` starts everything: app Postgres (:5432),
|
||||
dedicated pipeline Postgres (:5433), Valkey (:6379).
|
||||
- Config lives in the single root `.env` (see `.env.example`):
|
||||
`PIPELINE_POSTGRES_USER` / `PIPELINE_POSTGRES_PASSWORD` /
|
||||
`PIPELINE_POSTGRES_DB` / `PIPELINE_DATABASE_URL`, plus
|
||||
`GEMINI_API_KEY` for the pipeline itself.
|
||||
- The pipeline Postgres is deliberately separate from the app database
|
||||
so pipeline work can never damage dev data.
|
||||
- SQLite staging: `data-pipeline/db/staging.db`, created from
|
||||
`data-pipeline/db/schema.sql`.
|
||||
|
||||
### 2.3 Write and test the Gemini prompt
|
||||
### 2.3 The Gemini prompt (drafted; templating is Phase 3)
|
||||
|
||||
- Start with 5 German nouns.
|
||||
- The prompt should specify:
|
||||
- The exact JSON structure expected (copy from schema-design §6.3)
|
||||
- `data-pipeline/prompt` is a plain UTF-8 text file, edited by hand.
|
||||
It is currently pinned to a concrete sample run (Spanish nouns,
|
||||
20 words inlined).
|
||||
- The prompt specifies:
|
||||
- The exact JSON structure expected (from design-doc §6.3)
|
||||
- That definitions and examples must be in the word's language
|
||||
- That translations are needed for all 4 other supported languages
|
||||
- That gender must be provided for de/fr/es/it, null for en
|
||||
- That difficulty must be one of: easy, medium, hard
|
||||
- That multiple senses should be included for polysemous words
|
||||
- Call the API. Inspect the JSON. Common issues to fix:
|
||||
- Gemini wraps the JSON in markdown code fences → strip them
|
||||
- Gemini returns "intermediate" instead of "medium" → normalize
|
||||
- Gemini omits a language → re-prompt or reject
|
||||
- Gemini returns gender "common" for German → reject
|
||||
- Iterate until 3 consecutive batches of 20 return clean JSON.
|
||||
- **Known bugs to fix when templating (§3.2):** rules 2 and 3 still say
|
||||
`language` must be `"en"`; rule 15 lists target languages
|
||||
`de, it, es, fr` while the header says `en, it, de, fr`; rule 31
|
||||
says "valid English noun". All are leftovers from adapting the
|
||||
English version.
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: Data Pipeline — Detailed
|
||||
|
||||
### 3.1 Create the SQLite schema
|
||||
### 3.1 SQLite schema (done)
|
||||
|
||||
- Create a file `data-pipeline/schema.sql` or define it in your
|
||||
pipeline script.
|
||||
- Tables mirror the Postgres schema:
|
||||
- `data-pipeline/db/schema.sql` mirrors the Postgres schema with two
|
||||
SQLite concessions: IDs are `TEXT` (`crypto.randomUUID()`), and
|
||||
`definitions` / `examples` are JSON-encoded strings because SQLite
|
||||
has no array type. The import script parses them back into Postgres
|
||||
`TEXT[]`.
|
||||
- The UNIQUE constraints match Postgres:
|
||||
`words(headword, language_code, pos)`, `senses(word_id, sense_index)`,
|
||||
`translations(sense_id, target_language_code, translation)`.
|
||||
|
||||
```sql
|
||||
CREATE TABLE words (
|
||||
id TEXT PRIMARY KEY,
|
||||
headword TEXT NOT NULL,
|
||||
language_code TEXT NOT NULL,
|
||||
pos TEXT NOT NULL,
|
||||
created_at TEXT DEFAULT (datetime('now')),
|
||||
UNIQUE(headword, language_code, pos)
|
||||
);
|
||||
### 3.2 Template the prompt
|
||||
|
||||
CREATE TABLE senses (
|
||||
id TEXT PRIMARY KEY,
|
||||
word_id TEXT NOT NULL REFERENCES words(id),
|
||||
sense_index INTEGER NOT NULL DEFAULT 0,
|
||||
difficulty TEXT NOT NULL,
|
||||
definitions TEXT NOT NULL DEFAULT '[]', -- JSON array as text
|
||||
examples TEXT NOT NULL DEFAULT '[]', -- JSON array as text
|
||||
created_at TEXT DEFAULT (datetime('now')),
|
||||
UNIQUE(word_id, sense_index)
|
||||
);
|
||||
- Turn `data-pipeline/prompt` into a template. Substitution slots:
|
||||
- source language (name + code)
|
||||
- POS
|
||||
- target languages (the other 4 codes)
|
||||
- the word batch (20 words, one per line)
|
||||
- Fix the hardcoded leftovers listed in §2.3 as part of this — after
|
||||
templating, the language codes in the rules must derive from the
|
||||
substitutions, so this class of bug can't recur.
|
||||
- Request **structured output** from the API
|
||||
(`responseMimeType: "application/json"` with a `responseSchema`
|
||||
matching design-doc §6.3) instead of relying on prompt instructions
|
||||
alone. Keep fence-stripping as a defensive fallback only.
|
||||
|
||||
CREATE TABLE translations (
|
||||
id TEXT PRIMARY KEY,
|
||||
sense_id TEXT NOT NULL REFERENCES senses(id),
|
||||
target_language_code TEXT NOT NULL,
|
||||
translation TEXT NOT NULL,
|
||||
gender TEXT,
|
||||
difficulty TEXT NOT NULL,
|
||||
created_at TEXT DEFAULT (datetime('now')),
|
||||
UNIQUE(sense_id, target_language_code, translation)
|
||||
);
|
||||
### 3.3 Write the pipeline script
|
||||
|
||||
- Entry point: `data-pipeline/pipeline.ts`
|
||||
(`pnpm --filter @lila/pipeline pipeline:run`).
|
||||
- Pseudocode:
|
||||
|
||||
```
|
||||
for each language in [de, en, es, fr, it]:
|
||||
read wordlist file → trim, drop empties, dedup in memory
|
||||
query staging.db for existing (headword, language_code, pos)
|
||||
skip words already staged ← idempotency
|
||||
split the remainder into batches of 20
|
||||
for each batch:
|
||||
call Gemini API (structured output)
|
||||
write the raw response to responses/{lang}-{pos}-{n}.json
|
||||
parse JSON response
|
||||
for each entry in response:
|
||||
validate(entry)
|
||||
if valid:
|
||||
generate UUIDs for word, senses, translations
|
||||
INSERT word + senses + translations in ONE transaction
|
||||
else:
|
||||
append to rejection log
|
||||
log progress: "Batch 12/86 done. 238 words staged, 2 rejected."
|
||||
sleep 1s (rate limiting); retry with backoff on API errors
|
||||
```
|
||||
|
||||
- Note: SQLite has no native TEXT[]. Store arrays as JSON text.
|
||||
The import script will parse them into Postgres TEXT[] on import.
|
||||
- **Idempotency decision (resolves the pipeline.ts step 3 question):**
|
||||
wordlists are read fresh on every run; the staging DB is the record
|
||||
of what's been processed. Words already present in `staging.db` are
|
||||
skipped before batching, so re-runs cost no API calls for staged
|
||||
words. Each validated entry (word + its senses + their translations)
|
||||
is written in a single transaction, so a partially-written word can
|
||||
never exist and no NOT NULL constraint needs relaxing. The UNIQUE
|
||||
constraints remain as a backstop (`INSERT OR IGNORE`).
|
||||
- Persisting raw responses means a validation-rule change re-validates
|
||||
from disk instead of re-paying for ~430 API calls
|
||||
(~8,700 words ÷ 20 per batch).
|
||||
- Use `better-sqlite3` for SQLite access (synchronous, simple).
|
||||
- Generate UUIDs with `crypto.randomUUID()`.
|
||||
- Store definitions/examples as JSON strings in SQLite
|
||||
(`JSON.stringify(arr)`).
|
||||
|
||||
### 3.2 Write the validation module
|
||||
### 3.4 Validation module
|
||||
|
||||
- Create `data-pipeline/validate.ts`.
|
||||
- Create the validation module alongside the pipeline, with unit tests
|
||||
(vitest is already configured; `data-pipeline/vitest.config.ts`
|
||||
expects `tests/**/*.test.ts`).
|
||||
- Input: one parsed Gemini entry (the JSON object for one word).
|
||||
- Checks (return a list of errors, empty = valid):
|
||||
- `headword` is a non-empty string
|
||||
|
|
@ -394,43 +452,18 @@ Listed here for visibility. Not planned, not estimated.
|
|||
- `target_language` != the word's own language
|
||||
- `word` is a non-empty string
|
||||
- `gender` is valid for the target language
|
||||
(de: masculine/feminine/neuter/null,
|
||||
fr/es/it: masculine/feminine/null,
|
||||
(de: masculine/feminine/neuter,
|
||||
fr/es/it: masculine/feminine,
|
||||
en: null)
|
||||
- `difficulty` is in ['easy','medium','hard']
|
||||
- Output: `{ valid: boolean, errors: string[] }`
|
||||
|
||||
### 3.3 Write the pipeline script
|
||||
|
||||
- Create `data-pipeline/run.ts`.
|
||||
- Pseudocode:
|
||||
```
|
||||
for each language in [de, en, es, fr, it]:
|
||||
read wordlist file → array of words
|
||||
split into batches of 20
|
||||
for each batch:
|
||||
call Gemini API with the batch
|
||||
parse JSON response (strip markdown fences if present)
|
||||
for each entry in response:
|
||||
validate(entry)
|
||||
if valid:
|
||||
generate UUIDs for word, senses, translations
|
||||
INSERT into SQLite (words, senses, translations)
|
||||
else:
|
||||
append to rejection log
|
||||
log progress: "Batch 12/50 done. 238 words imported, 2 rejected."
|
||||
sleep 1s (rate limiting)
|
||||
```
|
||||
- Use `better-sqlite3` for SQLite access (synchronous, simple).
|
||||
- Generate UUIDs with `crypto.randomUUID()`.
|
||||
- Store definitions/examples as JSON strings in SQLite
|
||||
(`JSON.stringify(arr)`).
|
||||
|
||||
### 3.4 Run and review
|
||||
### 3.5 Run and review
|
||||
|
||||
- Run the pipeline for all 5 languages.
|
||||
- Check the rejection log. If rejection rate > 10%, fix the prompt
|
||||
and re-run failed batches.
|
||||
and re-run — the skip-processed check means only rejected/missing
|
||||
words are re-sent.
|
||||
- Spot-check: open the SQLite DB, run:
|
||||
```sql
|
||||
SELECT w.headword, s.definitions, t.translation, t.gender
|
||||
|
|
@ -446,50 +479,36 @@ Listed here for visibility. Not planned, not estimated.
|
|||
|
||||
---
|
||||
|
||||
## Phase 4: Migration & Import — Detailed
|
||||
## Phase 4: Import — Detailed
|
||||
|
||||
### 4.1 Generate and inspect the migration
|
||||
### 4.1 Migration (done in Phase 1)
|
||||
|
||||
- Run: `npx drizzle-kit generate`
|
||||
- Open the generated SQL file in `packages/db/drizzle/`.
|
||||
- Read it. Verify:
|
||||
- Three CREATE TABLE statements (words, senses, translations)
|
||||
- CHECK constraints match your schema
|
||||
- UNIQUE constraints are present
|
||||
- Three CREATE INDEX statements
|
||||
- Foreign keys reference the correct tables with ON DELETE CASCADE
|
||||
- If something looks wrong, fix the schema file and regenerate.
|
||||
`packages/db/drizzle/0012_graceful_psynapse.sql` creates the three
|
||||
tables with all CHECK/UNIQUE constraints, the three indexes, and
|
||||
cascading FKs. It is applied locally; the tables exist and are empty.
|
||||
Nothing to do here — prod gets the same migration in Phase 6.
|
||||
|
||||
### 4.2 Apply the migration locally
|
||||
|
||||
- Run: `npx drizzle-kit migrate`
|
||||
- Connect to local Postgres and verify:
|
||||
```sql
|
||||
\dt -- list tables
|
||||
\d words -- describe words table
|
||||
\d senses -- describe senses table
|
||||
\d translations -- describe translations table
|
||||
```
|
||||
|
||||
### 4.3 Write the import script
|
||||
### 4.2 Write the import script
|
||||
|
||||
- Create `data-pipeline/import-to-postgres.ts`.
|
||||
- Add `@lila/db` as a workspace dependency for the Drizzle client and
|
||||
schema (the pipeline currently ships only `better-sqlite3`).
|
||||
- Pseudocode:
|
||||
|
||||
```
|
||||
open SQLite database (read-only)
|
||||
connect to Postgres via Drizzle
|
||||
connect to Postgres via Drizzle (PIPELINE_DATABASE_URL first)
|
||||
|
||||
read all words from SQLite
|
||||
for each batch of 100 words:
|
||||
for each batch of words:
|
||||
begin transaction
|
||||
INSERT words into Postgres
|
||||
for each word:
|
||||
read its senses from SQLite
|
||||
parse definitions/examples from JSON string → TEXT[]
|
||||
INSERT senses into Postgres
|
||||
for each sense:
|
||||
read its translations from SQLite
|
||||
parse definitions/examples from JSON string → TEXT[]
|
||||
INSERT translations into Postgres
|
||||
commit transaction
|
||||
log progress
|
||||
|
|
@ -501,9 +520,10 @@ Listed here for visibility. Not planned, not estimated.
|
|||
into actual arrays (Postgres TEXT[]).
|
||||
- Use `ON CONFLICT DO NOTHING` to handle re-runs gracefully.
|
||||
|
||||
### 4.4 Run and verify
|
||||
### 4.3 Run and verify
|
||||
|
||||
- Run the import script.
|
||||
- Run against the pipeline Postgres (:5433) first; once the script is
|
||||
trusted, point it at the app dev database (:5432).
|
||||
- Compare counts:
|
||||
|
||||
```sql
|
||||
|
|
@ -574,6 +594,8 @@ Listed here for visibility. Not planned, not estimated.
|
|||
- Combine 1 correct translation + 3 distractors.
|
||||
- Shuffle the 4 options (Fisher-Yates or similar).
|
||||
- Attach gender to each option for display.
|
||||
- Keep the server-side evaluation invariant: the correct answer is
|
||||
never included in what is sent to the client.
|
||||
|
||||
### 5.4 Test matrix
|
||||
|
||||
|
|
@ -651,8 +673,9 @@ If something goes wrong:
|
|||
|
||||
Repeat the Phase 3 pipeline for each new POS:
|
||||
|
||||
- Acquire wordlists (verbs, adjectives, adverbs) for all 5 languages.
|
||||
- Adjust the Gemini prompt:
|
||||
- Acquire wordlists (verbs, adjectives, adverbs) for all 5 languages,
|
||||
saved as `source-data/{lang}/{pos}` with the shared POS codes.
|
||||
- Adjust the Gemini prompt template:
|
||||
- Verbs: ask for transitivity, common prepositions, or other
|
||||
verb-specific metadata if needed for future exercises.
|
||||
- Adjectives: ask for the base form. Note that adjective metadata
|
||||
|
|
@ -673,41 +696,47 @@ English words. No migration needed.
|
|||
# Risks & Mitigations
|
||||
|
||||
| Risk | Likelihood | Impact | Mitigation |
|
||||
| ------------------------------------------------------ | ----------------- | ------------------------- | ------------------------------------------------------------------------------------------------- |
|
||||
| Gemini returns malformed JSON | Medium | Pipeline stalls | Strip markdown fences, wrap parsing in try/catch, log and skip bad batches |
|
||||
| ------------------------------------------------------ | ----------------- | ------------------------- | ----------------------------------------------------------------------------------------------------- |
|
||||
| Gemini returns malformed JSON | Low | Pipeline stalls | Structured output (responseSchema); fence-stripping + try/catch as fallback; log and skip bad batches |
|
||||
| Validation rules change after a full run | Medium | Wasted API spend | Raw responses persisted per batch; re-validate from disk instead of re-calling the API |
|
||||
| Gemini produces wrong genders or difficulties | Medium | Bad exercise data | Validation module rejects invalid entries. Spot-check 50+ entries per language |
|
||||
| Too few words at a given difficulty/pos/language combo | Medium | Game can't fill 4 options | Graceful fallback: reduce options or show a "not enough words" message. Log which combos are thin |
|
||||
| SQLite → Postgres import fails mid-way | Low | Partial data in Postgres | Transactions per batch. Re-run with ON CONFLICT DO NOTHING |
|
||||
| SQLite → Postgres import fails mid-way | Low | Partial data in Postgres | Transactions per batch. Re-run with ON CONFLICT DO NOTHING. Test on :5433 before touching :5432 |
|
||||
| ORDER BY RANDOM() gets slow at scale | Low (not at 500k) | Slow game load | Add a TODO comment. Optimize with random() pre-filter when needed |
|
||||
| Prod migration breaks the live app | Low | Downtime | Backup before migrating. Rollback plan documented. Deploy during low-traffic hours |
|
||||
|
||||
---
|
||||
|
||||
# File / Folder Structure (new files)
|
||||
# File / Folder Structure
|
||||
|
||||
```
|
||||
project/
|
||||
├── data-pipeline/
|
||||
│ ├── source-data/
|
||||
│ │ ├── german/nouns/ ← wordlist files
|
||||
│ │ ├── english/nouns/
|
||||
│ │ ├── spanish/nouns/
|
||||
│ │ ├── french/nouns/
|
||||
│ │ └── italian/nouns/
|
||||
│ ├── rejections/ ← invalid Gemini entries (for review)
|
||||
│ ├── staging.db ← SQLite staging database
|
||||
│ ├── schema.sql ← SQLite schema definition
|
||||
│ ├── validate.ts ← validation module
|
||||
│ ├── run.ts ← main pipeline script
|
||||
│ └── import-to-postgres.ts ← SQLite → Postgres import
|
||||
│ │ ├── de/noun ← wordlist files (shared lang/POS codes)
|
||||
│ │ ├── en/noun
|
||||
│ │ ├── es/noun
|
||||
│ │ ├── fr/noun
|
||||
│ │ └── it/noun
|
||||
│ ├── db/
|
||||
│ │ ├── schema.sql ← SQLite schema definition ✅
|
||||
│ │ └── staging.db ← SQLite staging database (gitignored) ✅
|
||||
│ ├── prompt ← Gemini prompt (plain UTF-8; to be templated)
|
||||
│ ├── responses/ ← raw Gemini responses, one file per batch (planned)
|
||||
│ ├── rejections/ ← invalid Gemini entries for review (planned)
|
||||
│ ├── pipeline.ts ← main pipeline script (pseudocode today)
|
||||
│ ├── tests/ ← vitest unit tests, esp. validation (planned)
|
||||
│ └── import-to-postgres.ts ← SQLite → Postgres import (planned)
|
||||
├── packages/
|
||||
│ ├── db/
|
||||
│ │ └── src/db/schema.ts ← Drizzle schema (updated ✅)
|
||||
│ │ ├── src/db/schema.ts ← Drizzle schema ✅
|
||||
│ │ └── drizzle/0012_*.sql ← words/senses/translations migration ✅
|
||||
│ └── shared/
|
||||
│ └── src/constants.ts ← NOUN_GENDERS added ✅, "medium" ✅
|
||||
│ └── src/constants.ts ← NOUN_GENDERS ✅, "medium" ✅
|
||||
├── documentation/
|
||||
│ ├── DATA_PIPELINE.md ← orientation layer
|
||||
│ └── pipeline/
|
||||
│ ├── design-doc.md ← schema design doc (updated ✅)
|
||||
│ └── roadmap.md ← this document (updated ✅)
|
||||
│ ├── design-doc.md ← schema design doc ✅
|
||||
│ └── roadmap.md ← this document
|
||||
└── ...
|
||||
```
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue