updating docs to match the implemented phase 3 pipeline

The pipeline docs still described pipeline.ts as pseudocode and the
validation module as unwritten. Both have been implemented and run.

- CLAUDE.md: replace the "no executable pipeline yet" description with
  the actual module flow, plus the two invariants worth preserving
  (resumability via headword diffing, raw responses saved before parsing)
- DATA_PIPELINE.md: mark the seven implemented modules, add a module
  responsibility map and the CLI flag table, drop the resolved warning
  about hardcoded prompt values
- roadmap.md: check off phase 3 tasks 3.2/3.3/3.4, record where the
  build diverged from the plan, note that validate.ts is stricter than
  its own spec
- STATUS.md: phase 3 is data work now, not code work

Also records two open issues: the systemic difficulty-ordering rejection
cause, and the hard-tier shortfall pending a full run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
lila 2026-08-20 12:10:48 +02:00
parent 37c978e230
commit da9cdbfa1b
4 changed files with 219 additions and 49 deletions

View file

@ -46,7 +46,9 @@ Monorepo: `apps/api` (Express + `ws`), `apps/web` (React 19, Vite, TanStack Rout
## Data pipeline
`data-pipeline/` is mid-rewrite on this branch (`refactor/gemini-only-pipeline`): the old local-LLM/adapter architecture is gone, and `pipeline.ts` currently holds the staged design as pseudocode comments — there is no executable pipeline yet. Current pieces: `source-data/{lang}/{pos}` wordlists, `prompt` (the Gemini system prompt, a plain UTF-8 file), `db/schema.sql` (SQLite staging: `words` → `senses` → `translations`), and a dedicated pipeline PostgreSQL on port 5433.
`data-pipeline/` was rewritten on this branch (`refactor/gemini-only-pipeline`): the old local-LLM/adapter architecture is gone, replaced by a Gemini-only pipeline that is implemented and running (Phase 3). Flow: `sourceLists.ts` (discover + dedup `source-data/{lang}/{pos}`) → `promptTemplate.ts` (render `prompt` placeholders) → `gemini.ts` (structured output with a per-batch `responseSchema`, retry/backoff) → `validate.ts` (per-entry, three-way `valid`/`empty`/`invalid`) → `staging.ts` (SQLite, one transaction per word), orchestrated by `pipeline.ts`. A dedicated pipeline PostgreSQL runs on port 5433; the SQLite → PostgreSQL import script is Phase 4 and does not exist yet.
Two invariants to preserve when touching it: runs are **resumable** because already-staged headwords are diffed out before batching, and **raw responses are written to `responses/` before parsing**, so validation changes can be replayed without re-spending API quota. Rejected entries go to `rejections/{lang}-{pos}.jsonl`; `"senses": []` means "not a valid word of this POS" and is skipped, not rejected.
Read `documentation/DATA_PIPELINE.md` for orientation and current phase status, `documentation/pipeline/design-doc.md` for the schema/difficulty model/Gemini JSON contract, and `documentation/pipeline/roadmap.md` for the phase plan. Everything in `documentation/archive/` is superseded — `data-pipeline-local-llm.md`, `llm-setup-local.md`, and `model-strategy-cefr-voters.md` describe the removed local-LLM / CEFR-voter pipeline and are historical only.

View file

@ -1,7 +1,7 @@
# Lila Data Pipeline
> How vocabulary data is generated and gets into PostgreSQL.
> Last updated: 2026-08-01 · Branch: `refactor/gemini-only-pipeline`
> Last updated: 2026-08-20 · Branch: `refactor/gemini-only-pipeline`
**Authoritative detail lives in two companion docs:**
@ -57,21 +57,55 @@ The app always reads from PostgreSQL. SQLite exists purely as a staging file so
## What exists on disk today
| Path | State |
| ---------------------------------------- | ------------------------------------------------------------------------- |
| `data-pipeline/source-data/{lang}/{pos}` | ✅ Noun lists for `de`, `en`, `es`, `fr`, `it` |
| `data-pipeline/prompt` | ✅ The Gemini system prompt (plain UTF-8 text, not a module) |
| `data-pipeline/db/schema.sql` | ✅ SQLite staging schema |
| `data-pipeline/db/staging.db` | ✅ Created, tables present, **0 rows** — gitignored |
| `data-pipeline/pipeline.ts` | 🚧 Design pseudocode in comments. No executable pipeline code yet. |
| Validation module | ❌ Not written (rules specced in design-doc §6.4) |
| SQLite → PostgreSQL import script | ❌ Not written |
| `data-pipeline/kaikki-source-files/` | ⚠️ Leftover JSONL dumps from the old pipeline; nothing reads them anymore |
| `data-pipeline/worddata/english/nouns/` | ⚠️ Empty leftover output directory from the old per-word-JSON design |
The pipeline is implemented and running. Every module below is executable code with
co-located unit tests in `data-pipeline/tests/`.
| Path | State |
| --------------------------------------------- | --------------------------------------------------------------------- |
| `data-pipeline/source-data/{lang}/{pos}` | ✅ Noun lists for `de`, `en`, `es`, `fr`, `it` |
| `data-pipeline/prompt` | ✅ Templated Gemini prompt with `{{PLACEHOLDER}}` substitutions |
| `data-pipeline/sourceLists.ts` | ✅ Wordlist discovery, trim/dedup normalization |
| `data-pipeline/promptTemplate.ts` | ✅ Placeholder rendering; throws on any unreplaced `{{...}}` |
| `data-pipeline/gemini.ts` | ✅ Structured-output API client with retry/backoff |
| `data-pipeline/validate.ts` | ✅ Per-entry validation (design-doc §6.4) |
| `data-pipeline/staging.ts` | ✅ SQLite writes, one transaction per word |
| `data-pipeline/pipeline.ts` | ✅ Orchestrator with CLI flags, resumable |
| `data-pipeline/db/schema.sql` | ✅ SQLite staging schema |
| `data-pipeline/db/staging.db` | ✅ Populated — gitignored |
| `data-pipeline/responses/` | ✅ Raw Gemini responses, one JSON per batch — gitignored |
| `data-pipeline/rejections/{lang}-{pos}.jsonl` | ✅ Rejection log, one JSON per failed entry — gitignored |
| SQLite → PostgreSQL import script | ❌ Not written (Phase 4) |
| `data-pipeline/kaikki-source-files/` | ⚠️ 5.7 GB of leftover JSONL from the old pipeline; nothing reads them |
Directory naming follows the language/POS codes used in `packages/shared/src/constants.ts` (`de/noun`, not `german/nouns`) so no name mapping is needed anywhere in the pipeline.
Note that `data-pipeline/vitest.config.ts` looks for tests in `tests/**/*.test.ts` — that directory does not exist yet.
## Module responsibilities
```
sourceLists.ts discoverSourceLists() → [{ sourceLanguage, pos, words, filePath }]
normalizeWords(): trim, drop empties, dedup preserving order
promptTemplate.ts renderPrompt(): substitutes SOURCE_LANGUAGE_NAME/_CODE, POS,
TARGET_LANGUAGE_CODES, TARGET_LANGUAGE_UNION, INPUT_WORDS
gemini.ts buildEntriesResponseSchema(): OpenAPI schema with enums narrowed
to this batch's source/POS/target languages
generateContent(): 5 attempts, retries 429/500/503, honours the
API's own retryDelay; rejects non-STOP finishReason
validate.ts validateEntry() → "valid" | "empty" | "invalid"
staging.ts openStaging(), getStagedHeadwords(), stageEntry(), countStagedRows()
pipeline.ts orchestration, batching, rate-limit delay, rejection logging
```
Two properties worth knowing:
- **Resumability is free.** `getStagedHeadwords()` diffs the input list against what is
already in `staging.db`, so an interrupted run picks up exactly where it stopped and
never re-spends quota on a staged word.
- **Raw responses are saved before parsing.** Validation-rule changes can be replayed
against `responses/` without calling the API again.
`"empty"` is a distinct outcome from `"invalid"`: the contract says a word that is not a
valid noun in that language comes back with `"senses": []`. Those are _skipped_ and
counted separately, not written to the rejection log.
---
@ -93,7 +127,7 @@ translations id, sense_id→senses, target_language_code, UNIQUE(sense_id, targe
## The prompt
`data-pipeline/prompt` is the current working prompt, checked in as a plain text file and edited by hand. It is currently pinned to a concrete sample run (Spanish nouns, 20 words inlined) rather than templated — source language, POS, target languages, and the word batch will need to become substitutions when `pipeline.ts` is implemented.
`data-pipeline/prompt` is the working prompt, checked in as a plain text file and edited by hand. It is fully templated: `promptTemplate.ts` substitutes source language (name and code), POS, target languages, and the word batch, then fails loudly if any `{{PLACEHOLDER}}` survives rendering — so a typo in a placeholder name can never silently reach the API.
What it enforces, beyond the JSON shape in design-doc §6.3:
@ -104,23 +138,59 @@ What it enforces, beyond the JSON shape in design-doc §6.3:
- Base dictionary form, no articles or determiners.
- 1–3 senses per word, most words 1; skip rare, archaic, and technical senses.
- Up to 2 translations per target language per sense, only genuine synonyms or difficulty variants.
- A translation's difficulty may never be lower than its sense's difficulty.
- A translation's difficulty may never be lower than its sense's difficulty, and a sense's
difficulty should equal the easiest translation difficulty in that sense.
- A word that isn't a valid noun in that language comes back with `"senses": []`.
⚠️ **Known inconsistencies in the current prompt file** — it was adapted from the English version and some hardcoded values were not updated: rules 2 and 3 still say `language` must be `"en"` and there is a stray "valid English noun" in rule 31, while the header correctly says Spanish. Rule 15 lists target languages `de, it, es, fr` while the header says `en, it, de, fr`. Fix these when templating the prompt.
The copy-paste bugs that existed in the pre-templating draft (hardcoded `"en"` in rules 2–3,
a contradictory target list in rule 15, "valid English noun" in rule 31) are fixed — those
values are now placeholders.
Beyond the prompt, `gemini.ts` constrains the output with a **response schema** sent on every
request, with enums narrowed to that batch's source language, POS, and target languages. The
shape of the JSON is therefore enforced by the API, and `validate.ts` is left to enforce the
things a schema cannot express: cross-field difficulty ordering, gender rules per target
language, sequential `sense_index`, translation coverage and caps, and that the headword was
actually in the input batch.
Validation is the safety net, not the prompt — every entry is checked before it reaches SQLite, and rejects go to a log for review rather than silently disappearing. Target reject rate is under 10%.
> **Known systemic rejection cause (open).** Effectively all current rejections are
> `translation difficulty lower than sense difficulty`. The model tends to tag a sense
> `medium` while correctly tagging some of its translations `easy`. Since the prompt already
> defines sense difficulty as the easiest translation difficulty in the sense, this value is
> derivable and the entry is otherwise good — normalizing it in `validate.ts` instead of
> rejecting would recover these words. Not yet implemented.
---
## Running it
```bash
docker compose up -d pipeline-database # dedicated PostgreSQL on :5433
pnpm --filter @lila/pipeline pipeline:run # tsx --env-file .env pipeline.ts (currently a no-op)
pnpm --filter @lila/pipeline pipeline:run # tsx --env-file .env pipeline.ts
pnpm --filter @lila/pipeline test
```
`pipeline.ts` takes CLI flags (pass them after `--`, which the script strips itself):
| Flag | Default | Meaning |
| --------------- | ------- | -------------------------------------------------------- |
| `--langs` | all | Comma-separated source languages, e.g. `de,es` |
| `--pos` | `noun` | Part of speech to process |
| `--batch-size` | `20` | Words per Gemini request |
| `--max-batches` | all | Cap batches per list — useful for smoke tests |
| `--delay-ms` | `6000` | Pause between requests, for rate limiting |
| `--dry-run` | `false` | Render prompts and print batches without calling the API |
```bash
# smoke test: one German batch, no API calls
pnpm --filter @lila/pipeline pipeline:run -- --langs de --max-batches 1 --dry-run
```
Because runs are resumable, interrupting with Ctrl-C is safe — already-staged words are
skipped on the next run.
The pipeline reads `.env` from the repo root: `GEMINI_API_KEY`, plus `PIPELINE_POSTGRES_USER` / `PIPELINE_POSTGRES_PASSWORD` / `PIPELINE_POSTGRES_DB` / `PIPELINE_DATABASE_URL`. The pipeline database is deliberately separate from the app database (`:5432`) so pipeline work can never damage dev data.
---
@ -133,12 +203,22 @@ Full breakdown in [pipeline/roadmap.md](pipeline/roadmap.md).
| ------------------------------------------------- | -------------------------------------------------------------- |
| 1 — Drizzle schema (words/senses/translations) | ✅ Complete |
| 2 — Preparation (wordlists, DBs, prompt) | ✅ Complete |
| 3 — Build the pipeline → SQLite | 🔄 **Current.** Validation + `pipeline.ts` + first run |
| 3 — Build the pipeline → SQLite | 🔄 **Current.** Code complete; full 5-language run in progress |
| 4 — Migration & SQLite → PostgreSQL import | ⬜ Not started |
| 5 — App integration (`getGameTerms`, distractors) | ⬜ Not started |
| 6 — Production deploy | ⬜ Not started |
| 7 — Extend to verbs, adjectives, adverbs | ⬜ Not started — new wordlists + prompt only, no schema change |
Target for Phase 3: ~1000 nouns × 5 languages in staging, reject rate under 10%, 50 entries spot-checked by hand.
All Phase 3 code is written and unit-tested. What remains is the data work: finish the
5-language run, resolve the systemic rejection cause noted above, and hand-check 50 entries.
Target for Phase 3: the full deduped noun lists (~1,550–1,750 unique words per language after
dedup) × 5 languages in staging, reject rate under 10%, 50 entries spot-checked by hand.
**Known risk for Phase 5 — the `hard` tier is nearly empty.** Generated difficulty skews
heavily easy, with `hard` translations well under 1% of the corpus. The design-doc §5.1 game
query filters translation difficulty as an _exact_ match, so a "hard" game currently has far
too few rows to fill one round plus distractors. This needs prompt calibration before import,
not after.
Phase 5 is where this becomes visible in the app: `packages/db/src/models/termModel.ts` still queries `vocabulary_entries`/`entry_translations` and must be rewritten against the sense-based schema. Until then, production runs on the old data.

View file

@ -1,6 +1,6 @@
# Status — 2026-08-09
# Status — 2026-08-20
> Last updated: 2026-08-09. Update this file after every deploy or when switching tasks.
> Last updated: 2026-08-20. Update this file after every deploy or when switching tasks.
## What Works Today ✅
@ -12,7 +12,7 @@
## What's Broken / Blocked 🚧
- **Data quality** — Production still uses OpenWordNet/OMW translations. The replacement is the Gemini-only pipeline on branch `refactor/gemini-only-pipeline`: the sense-based schema (`words` → `senses` → `translations`) is migrated but empty, and the pipeline itself is not built yet. The earlier Kaikki/local-LLM pipeline was abandoned and removed (docs in `documentation/archive/`).
- **Data quality** — Production still uses OpenWordNet/OMW translations. The replacement is the Gemini-only pipeline on branch `refactor/gemini-only-pipeline`: the sense-based schema (`words` → `senses` → `translations`) is migrated but still empty in Postgres. The pipeline is **built and running**, staging into SQLite; the SQLite → Postgres import script (Phase 4) is what's missing. The earlier Kaikki/local-LLM pipeline was abandoned and removed (docs in `documentation/archive/`).
- **Guest play** — Auth is required for all game routes. No try-before-signup flow.
- **Game session store** — Still in-memory (`InMemoryGameSessionStore`). Valkey container exists in local dev but not wired up.
- **Rate limiting** — Partially implemented on auth endpoints; game endpoints not yet covered.
@ -21,7 +21,12 @@
## What I'm Working On Now 🔄
**Primary:** Phase 3 of the Gemini-only data pipeline (see [pipeline/roadmap.md](pipeline/roadmap.md)): template the Gemini prompt, write the validation module + `pipeline.ts`, and run the first full staging pass (~1,600–1,800 deduped nouns × 5 languages into SQLite).
**Primary:** Phase 3 of the Gemini-only data pipeline (see [pipeline/roadmap.md](pipeline/roadmap.md)). The pipeline code is complete and unit-tested — prompt templating, validation, Gemini structured output, and SQLite staging all work, and runs are resumable. Remaining Phase 3 work is data, not code:
1. Finish the full staging pass (~1,550–1,750 deduped nouns × 5 languages), German in progress.
2. Fix the one systemic rejection cause — sense difficulty tagged higher than its own easiest translation. It's derivable, so `validate.ts` should normalize rather than reject.
3. Fix the difficulty skew: `hard` translations are under 1% of staged rows, which leaves the "hard" game mode with too few rows to fill a round. Prompt calibration, needed before Phase 4 import.
4. Spot-check 50 entries by hand.
**Secondary:** Phase 7 hardening backlog items (see BACKLOG.md `next` section).

View file

@ -5,7 +5,7 @@
> backed by a normalized Postgres schema.
>
> **Author:** lila
> **Date:** July 2026 · Last reviewed: 2026-08-09
> **Date:** July 2026 · Last reviewed: 2026-08-20
> **Companion doc:** `documentation/pipeline/design-doc.md`
---
@ -33,14 +33,17 @@ Phase 1 Schema ✅
Phase 2 Preparation ✅
Wordlists acquired, databases set up, prompt drafted and
tested. Two loose ends carried into Phase 3: the prompt is
not templated yet, and the wordlists still contain duplicates
(deduped at runtime, not in the files).
tested. One loose end carried into Phase 3 and resolved
there: the prompt is now templated. The wordlist files still
contain duplicates, deduped at runtime rather than in the
files — harmless, and left as-is.
Phase 3 Data Pipeline ← CURRENT
Build the Gemini → validate → SQLite pipeline.
Produce a clean dataset: full deduped noun lists
(~1,600–1,800 words) × 5 languages.
(~1,550–1,750 unique words) × 5 languages.
Code is complete and unit-tested; the remaining work is the
data run, the rejection review, and the spot-check.
Phase 4 Import
(Migration already applied in Phase 1.)
@ -122,9 +125,10 @@ constraints, indexes, and relations.
(`data-pipeline/db/staging.db` from `db/schema.sql`)
- [x] Write and test the Gemini prompt with sample words
- [x] Refine the prompt until the JSON output matches the contract
defined in `design-doc.md` §6.3 — **note:** the prompt works but
is pinned to a hardcoded Spanish sample and has known copy-paste
bugs; templating and fixes are Phase 3 tasks
defined in `design-doc.md` §6.3 — **note:** at the end of Phase 2 the
prompt worked but was pinned to a hardcoded Spanish sample and had
known copy-paste bugs. Both were resolved by the templating task in
Phase 3.
**Dependencies:** Phase 1 complete.
@ -149,20 +153,31 @@ Re-runs skip words that are already staged.
- [x] Create the SQLite schema (`data-pipeline/db/schema.sql`, tables
created in `db/staging.db`)
- [ ] Template the Gemini prompt: source language, POS, target
languages, and the word batch become substitutions; fix the known
copy-paste bugs while doing so (rules 2–3 hardcode `"en"`,
rule 15's target list contradicts the header, rule 31 says
"valid English noun")
- [ ] Write the wordlist normalization step: trim whitespace, drop
empty lines, dedup in memory
- [ ] Write the validation module (rules in design-doc §6.4) with unit
tests (`vitest.config.ts` expects `tests/**/*.test.ts` — the
directory doesn't exist yet)
- [ ] Write the pipeline script (`pipeline.ts`): - Read + normalize wordlist files - Skip words already present in staging.db (idempotency — see §3.3) - Split the remainder into batches of 20 - Call Gemini per batch using structured output
(`responseMimeType: "application/json"` + `responseSchema`) - Persist each raw response to disk before validating - Validate each entry - Write valid entries to SQLite, one transaction per word - Log invalid entries to a rejection file
- [ ] Run the pipeline for all 5 languages (nouns only)
- [x] Template the Gemini prompt (`promptTemplate.ts` + placeholders in
`prompt`); the copy-paste bugs (rules 2–3 hardcoding `"en"`,
rule 15's contradictory target list, "valid English noun" in
rule 31) are fixed. `renderPrompt()` throws if any placeholder
survives rendering.
- [x] Write the wordlist normalization step (`sourceLists.ts`:
`normalizeWords()` trims, drops empty lines, dedups in memory
preserving order)
- [x] Write the validation module (`validate.ts`, rules in design-doc
§6.4) with unit tests — `tests/` now holds `validate.test.ts`,
`staging.test.ts`, `promptTemplate.test.ts`, `sourceLists.test.ts`
- [x] Write the pipeline script (`pipeline.ts`): reads + normalizes
wordlists, skips words already in `staging.db`, batches the
remainder, calls Gemini with structured output
(`responseMimeType: "application/json"` + `responseSchema` from
`gemini.ts`), persists each raw response to `responses/` before
validating, validates each entry, writes valid entries to SQLite
one transaction per word, and logs invalid entries to
`rejections/{lang}-{pos}.jsonl`. Adds CLI flags beyond the plan:
`--langs`, `--pos`, `--batch-size`, `--max-batches`, `--delay-ms`,
`--dry-run`.
- [ ] Run the pipeline for all 5 languages (nouns only) — **in progress**,
German first
- [ ] Review the rejection log, fix prompt issues, re-run failed batches
— reviewed; one systemic cause found (see below), fix not yet applied
- [ ] Spot-check 50 random entries for correctness
**Dependencies:** Phase 2 complete.
@ -170,7 +185,7 @@ Re-runs skip words that are already staged.
**Acceptance criteria:**
- SQLite database contains the full deduped noun lists
(~1,600–1,800 words × 5 languages) with senses and translations.
(~1,550–1,750 unique words × 5 languages) with senses and translations.
- Rejection rate is below 10%.
- Spot-checked entries have correct definitions, plausible examples,
correct genders, and reasonable difficulty levels.
@ -179,6 +194,29 @@ Re-runs skip words that are already staged.
- Raw Gemini responses are on disk, so validation-rule changes can be
re-applied without re-calling the API.
**Status against the criteria:** resumability is verified working (a live run
resumed correctly from previously staged words), raw responses are on disk,
and the reject rate is tracking well under 10%. The two open items are the
systemic rejection cause and the hard-tier shortfall below.
**Open issue — systemic rejection cause.** Effectively every rejection is
`translation difficulty lower than sense difficulty`: the model tags a sense
`medium` while correctly tagging some of its translations `easy`. Example:
`Ellbogen` with sense `medium` but `elbow` (en) and `codo` (es) as `easy` —
the translations are right and the sense label is wrong, yet the whole entry
is discarded. The prompt already defines sense difficulty as the easiest
translation difficulty in the sense, so the value is derivable. Normalizing it
in `validate.ts` rather than rejecting would recover these entries and can be
replayed against `responses/` without new API calls.
**Open issue — the `hard` tier is nearly empty.** Generated difficulty skews
heavily easy; `hard` translations are well under 1% of staged rows. Because
design-doc §5.1 filters translation difficulty as an _exact_ match, a "hard"
game currently resolves to a single-digit row count for a given language pair —
not enough for one round plus three distractors. This is prompt-calibration
work and belongs here in Phase 3, before the import in Phase 4, since fixing it
afterwards means re-importing.
---
## Phase 4: Import
@ -372,7 +410,15 @@ Listed here for visibility. Not planned, not estimated.
`words(headword, language_code, pos)`, `senses(word_id, sense_index)`,
`translations(sense_id, target_language_code, translation)`.
### 3.2 Template the prompt
### 3.2 Template the prompt (done)
Implemented in `promptTemplate.ts`. Structured output is implemented in
`gemini.ts` via `buildEntriesResponseSchema()`, which narrows the enums to
the batch's own source language, POS, and target languages rather than
allowing the full supported list. Fence-stripping is retained in
`parseEntries()` as the defensive fallback the plan called for.
Original spec:
- Turn `data-pipeline/prompt` into a template. Substitution slots:
- source language (name + code)
@ -387,7 +433,15 @@ Listed here for visibility. Not planned, not estimated.
matching design-doc §6.3) instead of relying on prompt instructions
alone. Keep fence-stripping as a defensive fallback only.
### 3.3 Write the pipeline script
### 3.3 Write the pipeline script (done)
Implemented as specced. Divergences from the plan below: the rate-limit pause
defaults to 6s rather than 1s, batching/language selection is controlled by CLI
flags (`--langs`, `--pos`, `--batch-size`, `--max-batches`, `--delay-ms`,
`--dry-run`), and `stageEntry()` does an explicit existence check inside the
transaction instead of relying on `INSERT OR IGNORE`. Words the model omits
from a response entirely are logged to the rejection file as
`"missing from Gemini response"`, so a silent drop can't go unnoticed.
- Entry point: `data-pipeline/pipeline.ts`
(`pnpm --filter @lila/pipeline pipeline:run`).
@ -430,7 +484,36 @@ Listed here for visibility. Not planned, not estimated.
- Store definitions/examples as JSON strings in SQLite
(`JSON.stringify(arr)`).
### 3.4 Validation module
### 3.4 Validation module (done)
Implemented in `validate.ts` with unit tests in `tests/validate.test.ts`.
**As built, it is stricter than this spec.** Two structural differences:
- The result is a three-way discriminated union, not `{ valid, errors }`.
`"empty"` (the contract's `"senses": []`, meaning "not a valid word of this
POS") is a distinct outcome from `"invalid"`, so genuinely-not-a-noun words
are skipped and counted separately rather than polluting the rejection log.
- Validation is context-aware. `ValidationContext` carries the batch's source
language, POS, target languages, and input words, so the checks below are
exact-match against the request rather than membership in the global
supported list.
Additional checks not in the original spec:
- `headword` must be one of the words actually sent in this batch.
- `language` must equal the batch's source language; `pos` must equal the
batch's POS.
- `sense_index` must equal the sense's position in the array (sequential
from 0), not merely be a non-negative integer.
- At most 3 senses per word.
- Every target language must have at least one translation, at most 2, with
no duplicate translation word within a target language.
- A translation's difficulty may not rank below its sense's difficulty —
this is the cross-field rule responsible for essentially all current
rejections (see the open issue in Phase 3 above).
Original spec:
- Create the validation module alongside the pipeline, with unit tests
(vitest is already configured; `data-pipeline/vitest.config.ts`