diff --git a/CLAUDE.md b/CLAUDE.md index accb592..3ef472a 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -46,7 +46,9 @@ Monorepo: `apps/api` (Express + `ws`), `apps/web` (React 19, Vite, TanStack Rout ## Data pipeline -`data-pipeline/` is mid-rewrite on this branch (`refactor/gemini-only-pipeline`): the old local-LLM/adapter architecture is gone, and `pipeline.ts` currently holds the staged design as pseudocode comments — there is no executable pipeline yet. Current pieces: `source-data/{lang}/{pos}` wordlists, `prompt` (the Gemini system prompt, a plain UTF-8 file), `db/schema.sql` (SQLite staging: `words` → `senses` → `translations`), and a dedicated pipeline PostgreSQL on port 5433. +`data-pipeline/` was rewritten on this branch (`refactor/gemini-only-pipeline`): the old local-LLM/adapter architecture is gone, replaced by a Gemini-only pipeline that is implemented and running (Phase 3). Flow: `sourceLists.ts` (discover + dedup `source-data/{lang}/{pos}`) → `promptTemplate.ts` (render `prompt` placeholders) → `gemini.ts` (structured output with a per-batch `responseSchema`, retry/backoff) → `validate.ts` (per-entry, three-way `valid`/`empty`/`invalid`) → `staging.ts` (SQLite, one transaction per word), orchestrated by `pipeline.ts`. A dedicated pipeline PostgreSQL runs on port 5433; the SQLite → PostgreSQL import script is Phase 4 and does not exist yet. + +Two invariants to preserve when touching it: runs are **resumable** because already-staged headwords are diffed out before batching, and **raw responses are written to `responses/` before parsing**, so validation changes can be replayed without re-spending API quota. Rejected entries go to `rejections/{lang}-{pos}.jsonl`; `"senses": []` means "not a valid word of this POS" and is skipped, not rejected. Read `documentation/DATA_PIPELINE.md` for orientation and current phase status, `documentation/pipeline/design-doc.md` for the schema/difficulty model/Gemini JSON contract, and `documentation/pipeline/roadmap.md` for the phase plan. Everything in `documentation/archive/` is superseded — `data-pipeline-local-llm.md`, `llm-setup-local.md`, and `model-strategy-cefr-voters.md` describe the removed local-LLM / CEFR-voter pipeline and are historical only. diff --git a/documentation/DATA_PIPELINE.md b/documentation/DATA_PIPELINE.md index 80fde25..02b07eb 100644 --- a/documentation/DATA_PIPELINE.md +++ b/documentation/DATA_PIPELINE.md @@ -1,7 +1,7 @@ # Lila Data Pipeline > How vocabulary data is generated and gets into PostgreSQL. -> Last updated: 2026-08-01 · Branch: `refactor/gemini-only-pipeline` +> Last updated: 2026-08-20 · Branch: `refactor/gemini-only-pipeline` **Authoritative detail lives in two companion docs:** @@ -57,21 +57,55 @@ The app always reads from PostgreSQL. SQLite exists purely as a staging file so ## What exists on disk today -| Path | State | -| ---------------------------------------- | ------------------------------------------------------------------------- | -| `data-pipeline/source-data/{lang}/{pos}` | ✅ Noun lists for `de`, `en`, `es`, `fr`, `it` | -| `data-pipeline/prompt` | ✅ The Gemini system prompt (plain UTF-8 text, not a module) | -| `data-pipeline/db/schema.sql` | ✅ SQLite staging schema | -| `data-pipeline/db/staging.db` | ✅ Created, tables present, **0 rows** — gitignored | -| `data-pipeline/pipeline.ts` | 🚧 Design pseudocode in comments. No executable pipeline code yet. | -| Validation module | ❌ Not written (rules specced in design-doc §6.4) | -| SQLite → PostgreSQL import script | ❌ Not written | -| `data-pipeline/kaikki-source-files/` | ⚠️ Leftover JSONL dumps from the old pipeline; nothing reads them anymore | -| `data-pipeline/worddata/english/nouns/` | ⚠️ Empty leftover output directory from the old per-word-JSON design | +The pipeline is implemented and running. Every module below is executable code with +co-located unit tests in `data-pipeline/tests/`. + +| Path | State | +| --------------------------------------------- | --------------------------------------------------------------------- | +| `data-pipeline/source-data/{lang}/{pos}` | ✅ Noun lists for `de`, `en`, `es`, `fr`, `it` | +| `data-pipeline/prompt` | ✅ Templated Gemini prompt with `{{PLACEHOLDER}}` substitutions | +| `data-pipeline/sourceLists.ts` | ✅ Wordlist discovery, trim/dedup normalization | +| `data-pipeline/promptTemplate.ts` | ✅ Placeholder rendering; throws on any unreplaced `{{...}}` | +| `data-pipeline/gemini.ts` | ✅ Structured-output API client with retry/backoff | +| `data-pipeline/validate.ts` | ✅ Per-entry validation (design-doc §6.4) | +| `data-pipeline/staging.ts` | ✅ SQLite writes, one transaction per word | +| `data-pipeline/pipeline.ts` | ✅ Orchestrator with CLI flags, resumable | +| `data-pipeline/db/schema.sql` | ✅ SQLite staging schema | +| `data-pipeline/db/staging.db` | ✅ Populated — gitignored | +| `data-pipeline/responses/` | ✅ Raw Gemini responses, one JSON per batch — gitignored | +| `data-pipeline/rejections/{lang}-{pos}.jsonl` | ✅ Rejection log, one JSON per failed entry — gitignored | +| SQLite → PostgreSQL import script | ❌ Not written (Phase 4) | +| `data-pipeline/kaikki-source-files/` | ⚠️ 5.7 GB of leftover JSONL from the old pipeline; nothing reads them | Directory naming follows the language/POS codes used in `packages/shared/src/constants.ts` (`de/noun`, not `german/nouns`) so no name mapping is needed anywhere in the pipeline. -Note that `data-pipeline/vitest.config.ts` looks for tests in `tests/**/*.test.ts` — that directory does not exist yet. +## Module responsibilities + +``` +sourceLists.ts discoverSourceLists() → [{ sourceLanguage, pos, words, filePath }] + normalizeWords(): trim, drop empties, dedup preserving order +promptTemplate.ts renderPrompt(): substitutes SOURCE_LANGUAGE_NAME/_CODE, POS, + TARGET_LANGUAGE_CODES, TARGET_LANGUAGE_UNION, INPUT_WORDS +gemini.ts buildEntriesResponseSchema(): OpenAPI schema with enums narrowed + to this batch's source/POS/target languages + generateContent(): 5 attempts, retries 429/500/503, honours the + API's own retryDelay; rejects non-STOP finishReason +validate.ts validateEntry() → "valid" | "empty" | "invalid" +staging.ts openStaging(), getStagedHeadwords(), stageEntry(), countStagedRows() +pipeline.ts orchestration, batching, rate-limit delay, rejection logging +``` + +Two properties worth knowing: + +- **Resumability is free.** `getStagedHeadwords()` diffs the input list against what is + already in `staging.db`, so an interrupted run picks up exactly where it stopped and + never re-spends quota on a staged word. +- **Raw responses are saved before parsing.** Validation-rule changes can be replayed + against `responses/` without calling the API again. + +`"empty"` is a distinct outcome from `"invalid"`: the contract says a word that is not a +valid noun in that language comes back with `"senses": []`. Those are _skipped_ and +counted separately, not written to the rejection log. --- @@ -93,7 +127,7 @@ translations id, sense_id→senses, target_language_code, UNIQUE(sense_id, targe ## The prompt -`data-pipeline/prompt` is the current working prompt, checked in as a plain text file and edited by hand. It is currently pinned to a concrete sample run (Spanish nouns, 20 words inlined) rather than templated — source language, POS, target languages, and the word batch will need to become substitutions when `pipeline.ts` is implemented. +`data-pipeline/prompt` is the working prompt, checked in as a plain text file and edited by hand. It is fully templated: `promptTemplate.ts` substitutes source language (name and code), POS, target languages, and the word batch, then fails loudly if any `{{PLACEHOLDER}}` survives rendering — so a typo in a placeholder name can never silently reach the API. What it enforces, beyond the JSON shape in design-doc §6.3: @@ -104,23 +138,59 @@ What it enforces, beyond the JSON shape in design-doc §6.3: - Base dictionary form, no articles or determiners. - 1–3 senses per word, most words 1; skip rare, archaic, and technical senses. - Up to 2 translations per target language per sense, only genuine synonyms or difficulty variants. -- A translation's difficulty may never be lower than its sense's difficulty. +- A translation's difficulty may never be lower than its sense's difficulty, and a sense's + difficulty should equal the easiest translation difficulty in that sense. - A word that isn't a valid noun in that language comes back with `"senses": []`. -⚠️ **Known inconsistencies in the current prompt file** — it was adapted from the English version and some hardcoded values were not updated: rules 2 and 3 still say `language` must be `"en"` and there is a stray "valid English noun" in rule 31, while the header correctly says Spanish. Rule 15 lists target languages `de, it, es, fr` while the header says `en, it, de, fr`. Fix these when templating the prompt. +The copy-paste bugs that existed in the pre-templating draft (hardcoded `"en"` in rules 2–3, +a contradictory target list in rule 15, "valid English noun" in rule 31) are fixed — those +values are now placeholders. + +Beyond the prompt, `gemini.ts` constrains the output with a **response schema** sent on every +request, with enums narrowed to that batch's source language, POS, and target languages. The +shape of the JSON is therefore enforced by the API, and `validate.ts` is left to enforce the +things a schema cannot express: cross-field difficulty ordering, gender rules per target +language, sequential `sense_index`, translation coverage and caps, and that the headword was +actually in the input batch. Validation is the safety net, not the prompt — every entry is checked before it reaches SQLite, and rejects go to a log for review rather than silently disappearing. Target reject rate is under 10%. +> **Known systemic rejection cause (open).** Effectively all current rejections are +> `translation difficulty lower than sense difficulty`. The model tends to tag a sense +> `medium` while correctly tagging some of its translations `easy`. Since the prompt already +> defines sense difficulty as the easiest translation difficulty in the sense, this value is +> derivable and the entry is otherwise good — normalizing it in `validate.ts` instead of +> rejecting would recover these words. Not yet implemented. + --- ## Running it ```bash docker compose up -d pipeline-database # dedicated PostgreSQL on :5433 -pnpm --filter @lila/pipeline pipeline:run # tsx --env-file .env pipeline.ts (currently a no-op) +pnpm --filter @lila/pipeline pipeline:run # tsx --env-file .env pipeline.ts pnpm --filter @lila/pipeline test ``` +`pipeline.ts` takes CLI flags (pass them after `--`, which the script strips itself): + +| Flag | Default | Meaning | +| --------------- | ------- | -------------------------------------------------------- | +| `--langs` | all | Comma-separated source languages, e.g. `de,es` | +| `--pos` | `noun` | Part of speech to process | +| `--batch-size` | `20` | Words per Gemini request | +| `--max-batches` | all | Cap batches per list — useful for smoke tests | +| `--delay-ms` | `6000` | Pause between requests, for rate limiting | +| `--dry-run` | `false` | Render prompts and print batches without calling the API | + +```bash +# smoke test: one German batch, no API calls +pnpm --filter @lila/pipeline pipeline:run -- --langs de --max-batches 1 --dry-run +``` + +Because runs are resumable, interrupting with Ctrl-C is safe — already-staged words are +skipped on the next run. + The pipeline reads `.env` from the repo root: `GEMINI_API_KEY`, plus `PIPELINE_POSTGRES_USER` / `PIPELINE_POSTGRES_PASSWORD` / `PIPELINE_POSTGRES_DB` / `PIPELINE_DATABASE_URL`. The pipeline database is deliberately separate from the app database (`:5432`) so pipeline work can never damage dev data. --- @@ -133,12 +203,22 @@ Full breakdown in [pipeline/roadmap.md](pipeline/roadmap.md). | ------------------------------------------------- | -------------------------------------------------------------- | | 1 — Drizzle schema (words/senses/translations) | ✅ Complete | | 2 — Preparation (wordlists, DBs, prompt) | ✅ Complete | -| 3 — Build the pipeline → SQLite | 🔄 **Current.** Validation + `pipeline.ts` + first run | +| 3 — Build the pipeline → SQLite | 🔄 **Current.** Code complete; full 5-language run in progress | | 4 — Migration & SQLite → PostgreSQL import | ⬜ Not started | | 5 — App integration (`getGameTerms`, distractors) | ⬜ Not started | | 6 — Production deploy | ⬜ Not started | | 7 — Extend to verbs, adjectives, adverbs | ⬜ Not started — new wordlists + prompt only, no schema change | -Target for Phase 3: ~1000 nouns × 5 languages in staging, reject rate under 10%, 50 entries spot-checked by hand. +All Phase 3 code is written and unit-tested. What remains is the data work: finish the +5-language run, resolve the systemic rejection cause noted above, and hand-check 50 entries. + +Target for Phase 3: the full deduped noun lists (~1,550–1,750 unique words per language after +dedup) × 5 languages in staging, reject rate under 10%, 50 entries spot-checked by hand. + +**Known risk for Phase 5 — the `hard` tier is nearly empty.** Generated difficulty skews +heavily easy, with `hard` translations well under 1% of the corpus. The design-doc §5.1 game +query filters translation difficulty as an _exact_ match, so a "hard" game currently has far +too few rows to fill one round plus distractors. This needs prompt calibration before import, +not after. Phase 5 is where this becomes visible in the app: `packages/db/src/models/termModel.ts` still queries `vocabulary_entries`/`entry_translations` and must be rewritten against the sense-based schema. Until then, production runs on the old data. diff --git a/documentation/STATUS.md b/documentation/STATUS.md index 816db2f..75c0a7d 100644 --- a/documentation/STATUS.md +++ b/documentation/STATUS.md @@ -1,6 +1,6 @@ -# Status — 2026-08-09 +# Status — 2026-08-20 -> Last updated: 2026-08-09. Update this file after every deploy or when switching tasks. +> Last updated: 2026-08-20. Update this file after every deploy or when switching tasks. ## What Works Today ✅ @@ -12,7 +12,7 @@ ## What's Broken / Blocked 🚧 -- **Data quality** — Production still uses OpenWordNet/OMW translations. The replacement is the Gemini-only pipeline on branch `refactor/gemini-only-pipeline`: the sense-based schema (`words` → `senses` → `translations`) is migrated but empty, and the pipeline itself is not built yet. The earlier Kaikki/local-LLM pipeline was abandoned and removed (docs in `documentation/archive/`). +- **Data quality** — Production still uses OpenWordNet/OMW translations. The replacement is the Gemini-only pipeline on branch `refactor/gemini-only-pipeline`: the sense-based schema (`words` → `senses` → `translations`) is migrated but still empty in Postgres. The pipeline is **built and running**, staging into SQLite; the SQLite → Postgres import script (Phase 4) is what's missing. The earlier Kaikki/local-LLM pipeline was abandoned and removed (docs in `documentation/archive/`). - **Guest play** — Auth is required for all game routes. No try-before-signup flow. - **Game session store** — Still in-memory (`InMemoryGameSessionStore`). Valkey container exists in local dev but not wired up. - **Rate limiting** — Partially implemented on auth endpoints; game endpoints not yet covered. @@ -21,7 +21,12 @@ ## What I'm Working On Now 🔄 -**Primary:** Phase 3 of the Gemini-only data pipeline (see [pipeline/roadmap.md](pipeline/roadmap.md)): template the Gemini prompt, write the validation module + `pipeline.ts`, and run the first full staging pass (~1,600–1,800 deduped nouns × 5 languages into SQLite). +**Primary:** Phase 3 of the Gemini-only data pipeline (see [pipeline/roadmap.md](pipeline/roadmap.md)). The pipeline code is complete and unit-tested — prompt templating, validation, Gemini structured output, and SQLite staging all work, and runs are resumable. Remaining Phase 3 work is data, not code: + +1. Finish the full staging pass (~1,550–1,750 deduped nouns × 5 languages), German in progress. +2. Fix the one systemic rejection cause — sense difficulty tagged higher than its own easiest translation. It's derivable, so `validate.ts` should normalize rather than reject. +3. Fix the difficulty skew: `hard` translations are under 1% of staged rows, which leaves the "hard" game mode with too few rows to fill a round. Prompt calibration, needed before Phase 4 import. +4. Spot-check 50 entries by hand. **Secondary:** Phase 7 hardening backlog items (see BACKLOG.md `next` section). diff --git a/documentation/pipeline/roadmap.md b/documentation/pipeline/roadmap.md index 9d4eb86..43e4e38 100644 --- a/documentation/pipeline/roadmap.md +++ b/documentation/pipeline/roadmap.md @@ -5,7 +5,7 @@ > backed by a normalized Postgres schema. > > **Author:** lila -> **Date:** July 2026 · Last reviewed: 2026-08-09 +> **Date:** July 2026 · Last reviewed: 2026-08-20 > **Companion doc:** `documentation/pipeline/design-doc.md` --- @@ -33,14 +33,17 @@ Phase 1 Schema ✅ Phase 2 Preparation ✅ Wordlists acquired, databases set up, prompt drafted and - tested. Two loose ends carried into Phase 3: the prompt is - not templated yet, and the wordlists still contain duplicates - (deduped at runtime, not in the files). + tested. One loose end carried into Phase 3 and resolved + there: the prompt is now templated. The wordlist files still + contain duplicates, deduped at runtime rather than in the + files — harmless, and left as-is. Phase 3 Data Pipeline ← CURRENT Build the Gemini → validate → SQLite pipeline. Produce a clean dataset: full deduped noun lists - (~1,600–1,800 words) × 5 languages. + (~1,550–1,750 unique words) × 5 languages. + Code is complete and unit-tested; the remaining work is the + data run, the rejection review, and the spot-check. Phase 4 Import (Migration already applied in Phase 1.) @@ -122,9 +125,10 @@ constraints, indexes, and relations. (`data-pipeline/db/staging.db` from `db/schema.sql`) - [x] Write and test the Gemini prompt with sample words - [x] Refine the prompt until the JSON output matches the contract - defined in `design-doc.md` §6.3 — **note:** the prompt works but - is pinned to a hardcoded Spanish sample and has known copy-paste - bugs; templating and fixes are Phase 3 tasks + defined in `design-doc.md` §6.3 — **note:** at the end of Phase 2 the + prompt worked but was pinned to a hardcoded Spanish sample and had + known copy-paste bugs. Both were resolved by the templating task in + Phase 3. **Dependencies:** Phase 1 complete. @@ -149,20 +153,31 @@ Re-runs skip words that are already staged. - [x] Create the SQLite schema (`data-pipeline/db/schema.sql`, tables created in `db/staging.db`) -- [ ] Template the Gemini prompt: source language, POS, target - languages, and the word batch become substitutions; fix the known - copy-paste bugs while doing so (rules 2–3 hardcode `"en"`, - rule 15's target list contradicts the header, rule 31 says - "valid English noun") -- [ ] Write the wordlist normalization step: trim whitespace, drop - empty lines, dedup in memory -- [ ] Write the validation module (rules in design-doc §6.4) with unit - tests (`vitest.config.ts` expects `tests/**/*.test.ts` — the - directory doesn't exist yet) -- [ ] Write the pipeline script (`pipeline.ts`): - Read + normalize wordlist files - Skip words already present in staging.db (idempotency — see §3.3) - Split the remainder into batches of 20 - Call Gemini per batch using structured output - (`responseMimeType: "application/json"` + `responseSchema`) - Persist each raw response to disk before validating - Validate each entry - Write valid entries to SQLite, one transaction per word - Log invalid entries to a rejection file -- [ ] Run the pipeline for all 5 languages (nouns only) +- [x] Template the Gemini prompt (`promptTemplate.ts` + placeholders in + `prompt`); the copy-paste bugs (rules 2–3 hardcoding `"en"`, + rule 15's contradictory target list, "valid English noun" in + rule 31) are fixed. `renderPrompt()` throws if any placeholder + survives rendering. +- [x] Write the wordlist normalization step (`sourceLists.ts`: + `normalizeWords()` trims, drops empty lines, dedups in memory + preserving order) +- [x] Write the validation module (`validate.ts`, rules in design-doc + §6.4) with unit tests — `tests/` now holds `validate.test.ts`, + `staging.test.ts`, `promptTemplate.test.ts`, `sourceLists.test.ts` +- [x] Write the pipeline script (`pipeline.ts`): reads + normalizes + wordlists, skips words already in `staging.db`, batches the + remainder, calls Gemini with structured output + (`responseMimeType: "application/json"` + `responseSchema` from + `gemini.ts`), persists each raw response to `responses/` before + validating, validates each entry, writes valid entries to SQLite + one transaction per word, and logs invalid entries to + `rejections/{lang}-{pos}.jsonl`. Adds CLI flags beyond the plan: + `--langs`, `--pos`, `--batch-size`, `--max-batches`, `--delay-ms`, + `--dry-run`. +- [ ] Run the pipeline for all 5 languages (nouns only) — **in progress**, + German first - [ ] Review the rejection log, fix prompt issues, re-run failed batches + — reviewed; one systemic cause found (see below), fix not yet applied - [ ] Spot-check 50 random entries for correctness **Dependencies:** Phase 2 complete. @@ -170,7 +185,7 @@ Re-runs skip words that are already staged. **Acceptance criteria:** - SQLite database contains the full deduped noun lists - (~1,600–1,800 words × 5 languages) with senses and translations. + (~1,550–1,750 unique words × 5 languages) with senses and translations. - Rejection rate is below 10%. - Spot-checked entries have correct definitions, plausible examples, correct genders, and reasonable difficulty levels. @@ -179,6 +194,29 @@ Re-runs skip words that are already staged. - Raw Gemini responses are on disk, so validation-rule changes can be re-applied without re-calling the API. +**Status against the criteria:** resumability is verified working (a live run +resumed correctly from previously staged words), raw responses are on disk, +and the reject rate is tracking well under 10%. The two open items are the +systemic rejection cause and the hard-tier shortfall below. + +**Open issue — systemic rejection cause.** Effectively every rejection is +`translation difficulty lower than sense difficulty`: the model tags a sense +`medium` while correctly tagging some of its translations `easy`. Example: +`Ellbogen` with sense `medium` but `elbow` (en) and `codo` (es) as `easy` — +the translations are right and the sense label is wrong, yet the whole entry +is discarded. The prompt already defines sense difficulty as the easiest +translation difficulty in the sense, so the value is derivable. Normalizing it +in `validate.ts` rather than rejecting would recover these entries and can be +replayed against `responses/` without new API calls. + +**Open issue — the `hard` tier is nearly empty.** Generated difficulty skews +heavily easy; `hard` translations are well under 1% of staged rows. Because +design-doc §5.1 filters translation difficulty as an _exact_ match, a "hard" +game currently resolves to a single-digit row count for a given language pair — +not enough for one round plus three distractors. This is prompt-calibration +work and belongs here in Phase 3, before the import in Phase 4, since fixing it +afterwards means re-importing. + --- ## Phase 4: Import @@ -372,7 +410,15 @@ Listed here for visibility. Not planned, not estimated. `words(headword, language_code, pos)`, `senses(word_id, sense_index)`, `translations(sense_id, target_language_code, translation)`. -### 3.2 Template the prompt +### 3.2 Template the prompt (done) + +Implemented in `promptTemplate.ts`. Structured output is implemented in +`gemini.ts` via `buildEntriesResponseSchema()`, which narrows the enums to +the batch's own source language, POS, and target languages rather than +allowing the full supported list. Fence-stripping is retained in +`parseEntries()` as the defensive fallback the plan called for. + +Original spec: - Turn `data-pipeline/prompt` into a template. Substitution slots: - source language (name + code) @@ -387,7 +433,15 @@ Listed here for visibility. Not planned, not estimated. matching design-doc §6.3) instead of relying on prompt instructions alone. Keep fence-stripping as a defensive fallback only. -### 3.3 Write the pipeline script +### 3.3 Write the pipeline script (done) + +Implemented as specced. Divergences from the plan below: the rate-limit pause +defaults to 6s rather than 1s, batching/language selection is controlled by CLI +flags (`--langs`, `--pos`, `--batch-size`, `--max-batches`, `--delay-ms`, +`--dry-run`), and `stageEntry()` does an explicit existence check inside the +transaction instead of relying on `INSERT OR IGNORE`. Words the model omits +from a response entirely are logged to the rejection file as +`"missing from Gemini response"`, so a silent drop can't go unnoticed. - Entry point: `data-pipeline/pipeline.ts` (`pnpm --filter @lila/pipeline pipeline:run`). @@ -430,7 +484,36 @@ Listed here for visibility. Not planned, not estimated. - Store definitions/examples as JSON strings in SQLite (`JSON.stringify(arr)`). -### 3.4 Validation module +### 3.4 Validation module (done) + +Implemented in `validate.ts` with unit tests in `tests/validate.test.ts`. + +**As built, it is stricter than this spec.** Two structural differences: + +- The result is a three-way discriminated union, not `{ valid, errors }`. + `"empty"` (the contract's `"senses": []`, meaning "not a valid word of this + POS") is a distinct outcome from `"invalid"`, so genuinely-not-a-noun words + are skipped and counted separately rather than polluting the rejection log. +- Validation is context-aware. `ValidationContext` carries the batch's source + language, POS, target languages, and input words, so the checks below are + exact-match against the request rather than membership in the global + supported list. + +Additional checks not in the original spec: + +- `headword` must be one of the words actually sent in this batch. +- `language` must equal the batch's source language; `pos` must equal the + batch's POS. +- `sense_index` must equal the sense's position in the array (sequential + from 0), not merely be a non-negative integer. +- At most 3 senses per word. +- Every target language must have at least one translation, at most 2, with + no duplicate translation word within a target language. +- A translation's difficulty may not rank below its sense's difficulty — + this is the cross-field rule responsible for essentially all current + rejections (see the open issue in Phase 3 above). + +Original spec: - Create the validation module alongside the pipeline, with unit tests (vitest is already configured; `data-pipeline/vitest.config.ts`