The pipeline docs still described pipeline.ts as pseudocode and the
validation module as unwritten. Both have been implemented and run.
- CLAUDE.md: replace the "no executable pipeline yet" description with
the actual module flow, plus the two invariants worth preserving
(resumability via headword diffing, raw responses saved before parsing)
- DATA_PIPELINE.md: mark the seven implemented modules, add a module
responsibility map and the CLI flag table, drop the resolved warning
about hardcoded prompt values
- roadmap.md: check off phase 3 tasks 3.2/3.3/3.4, record where the
build diverged from the plan, note that validate.ts is stricter than
its own spec
- STATUS.md: phase 3 is data work now, not code work
Also records two open issues: the systemic difficulty-ordering rejection
cause, and the hard-tier shortfall pending a full run.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Prompt is now a template (fixes the hardcoded en/es leftovers in rules
2, 3, 15, 16, 26, 31). pipeline.ts replaces the pseudocode: wordlist
normalization, skip-already-staged idempotency, batches of 20 against
gemini-3.6-flash with responseSchema, raw responses persisted per batch,
per-entry validation with rejection log, one transaction per word into
db/staging.db. Flags: --langs --pos --max-batches --delay-ms --dry-run.
Smoke run: 40/40 words staged (de+es, one batch each), 0 rejections.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Roadmap: phase 4 re-scoped to import only (migration already applied),
phase 3 gains prompt templating, wordlist dedup, structured output,
raw-response persistence, and validation tests; idempotency decision
documented (skip already-staged words, one transaction per word).
Stale paths, postgres setup, and file structure corrected.
STATUS.md: refreshed from stale 2026-05-15 Kaikki state to current
gemini-only pipeline work.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Translation count test now adds reverse link count to expected total
- Non-English translations test now filters to kaikki source only
- Target language test now filters to kaikki source only — reverse links
to English are valid and expected
- Replace checkOmwExists with checkExtractedFilesExist
- Wire up importKaikki and reverseLink as real stage implementations
- Track reverse link completion via sentinel row in run_status
- Update report to use resolved_entry_cefr and entry counts
- Stages 3 onwards remain as stubs
- Extract.ts now processes all 5 language files, filters non-English
entries by lang_code, skips translation extraction for non-English
(no translations in source files)
- Import.ts now imports all 5 language output files, uses language
field from ExtractedSense instead of hardcoding en
- Sample limit hardcoded to 500 entries per language for development
- Add stage-1-extract/scripts/extract.ts — streams Kaikki JSONL,
filters to supported POS and languages, skips abbreviations and
senses with no translations in supported languages
- Rewrite db/import.ts for Kaikki flat model — tracks sense_index
offsets per headword+pos to handle duplicate JSONL entries
- Rewrite db/schema.sql for Kaikki model — entries, translations,
LLM vote tables, resolved tables
- Add extract and db:import scripts to package.json
- Sample mode hardcoded to 500 entries for development
- Replace terms/translations/term_glosses/term_examples with vocabulary_entries
and entry_translations
- Remove decks, topics and related tables (deferred)
- Add cefr_level and difficulty to entry_translations for game query filtering
- Update termModel.ts for new schema — getDistractors now takes sourceLanguage
- Update gameService.ts and multiplayerGameService.ts for entryId rename
- Update all test fixtures from termId to entryId
- Generate and apply migration 0011