updating docs to match the implemented phase 3 pipeline

The pipeline docs still described pipeline.ts as pseudocode and the
validation module as unwritten. Both have been implemented and run.

- CLAUDE.md: replace the "no executable pipeline yet" description with
  the actual module flow, plus the two invariants worth preserving
  (resumability via headword diffing, raw responses saved before parsing)
- DATA_PIPELINE.md: mark the seven implemented modules, add a module
  responsibility map and the CLI flag table, drop the resolved warning
  about hardcoded prompt values
- roadmap.md: check off phase 3 tasks 3.2/3.3/3.4, record where the
  build diverged from the plan, note that validate.ts is stricter than
  its own spec
- STATUS.md: phase 3 is data work now, not code work

Also records two open issues: the systemic difficulty-ordering rejection
cause, and the hard-tier shortfall pending a full run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
lila 2026-08-20 12:10:48 +02:00
parent 37c978e230
commit da9cdbfa1b
4 changed files with 219 additions and 49 deletions

View file

@ -5,7 +5,7 @@
> backed by a normalized Postgres schema.
>
> **Author:** lila
> **Date:** July 2026 · Last reviewed: 2026-08-09
> **Date:** July 2026 · Last reviewed: 2026-08-20
> **Companion doc:** `documentation/pipeline/design-doc.md`
---
@ -33,14 +33,17 @@ Phase 1 Schema ✅
Phase 2 Preparation ✅
Wordlists acquired, databases set up, prompt drafted and
tested. Two loose ends carried into Phase 3: the prompt is
not templated yet, and the wordlists still contain duplicates
(deduped at runtime, not in the files).
tested. One loose end carried into Phase 3 and resolved
there: the prompt is now templated. The wordlist files still
contain duplicates, deduped at runtime rather than in the
files — harmless, and left as-is.
Phase 3 Data Pipeline ← CURRENT
Build the Gemini → validate → SQLite pipeline.
Produce a clean dataset: full deduped noun lists
(~1,600–1,800 words) × 5 languages.
(~1,550–1,750 unique words) × 5 languages.
Code is complete and unit-tested; the remaining work is the
data run, the rejection review, and the spot-check.
Phase 4 Import
(Migration already applied in Phase 1.)
@ -122,9 +125,10 @@ constraints, indexes, and relations.
(`data-pipeline/db/staging.db` from `db/schema.sql`)
- [x] Write and test the Gemini prompt with sample words
- [x] Refine the prompt until the JSON output matches the contract
defined in `design-doc.md` §6.3 — **note:** the prompt works but
is pinned to a hardcoded Spanish sample and has known copy-paste
bugs; templating and fixes are Phase 3 tasks
defined in `design-doc.md` §6.3 — **note:** at the end of Phase 2 the
prompt worked but was pinned to a hardcoded Spanish sample and had
known copy-paste bugs. Both were resolved by the templating task in
Phase 3.
**Dependencies:** Phase 1 complete.
@ -149,20 +153,31 @@ Re-runs skip words that are already staged.
- [x] Create the SQLite schema (`data-pipeline/db/schema.sql`, tables
created in `db/staging.db`)
- [ ] Template the Gemini prompt: source language, POS, target
languages, and the word batch become substitutions; fix the known
copy-paste bugs while doing so (rules 2–3 hardcode `"en"`,
rule 15's target list contradicts the header, rule 31 says
"valid English noun")
- [ ] Write the wordlist normalization step: trim whitespace, drop
empty lines, dedup in memory
- [ ] Write the validation module (rules in design-doc §6.4) with unit
tests (`vitest.config.ts` expects `tests/**/*.test.ts` — the
directory doesn't exist yet)
- [ ] Write the pipeline script (`pipeline.ts`): - Read + normalize wordlist files - Skip words already present in staging.db (idempotency — see §3.3) - Split the remainder into batches of 20 - Call Gemini per batch using structured output
(`responseMimeType: "application/json"` + `responseSchema`) - Persist each raw response to disk before validating - Validate each entry - Write valid entries to SQLite, one transaction per word - Log invalid entries to a rejection file
- [ ] Run the pipeline for all 5 languages (nouns only)
- [x] Template the Gemini prompt (`promptTemplate.ts` + placeholders in
`prompt`); the copy-paste bugs (rules 2–3 hardcoding `"en"`,
rule 15's contradictory target list, "valid English noun" in
rule 31) are fixed. `renderPrompt()` throws if any placeholder
survives rendering.
- [x] Write the wordlist normalization step (`sourceLists.ts`:
`normalizeWords()` trims, drops empty lines, dedups in memory
preserving order)
- [x] Write the validation module (`validate.ts`, rules in design-doc
§6.4) with unit tests — `tests/` now holds `validate.test.ts`,
`staging.test.ts`, `promptTemplate.test.ts`, `sourceLists.test.ts`
- [x] Write the pipeline script (`pipeline.ts`): reads + normalizes
wordlists, skips words already in `staging.db`, batches the
remainder, calls Gemini with structured output
(`responseMimeType: "application/json"` + `responseSchema` from
`gemini.ts`), persists each raw response to `responses/` before
validating, validates each entry, writes valid entries to SQLite
one transaction per word, and logs invalid entries to
`rejections/{lang}-{pos}.jsonl`. Adds CLI flags beyond the plan:
`--langs`, `--pos`, `--batch-size`, `--max-batches`, `--delay-ms`,
`--dry-run`.
- [ ] Run the pipeline for all 5 languages (nouns only) — **in progress**,
German first
- [ ] Review the rejection log, fix prompt issues, re-run failed batches
— reviewed; one systemic cause found (see below), fix not yet applied
- [ ] Spot-check 50 random entries for correctness
**Dependencies:** Phase 2 complete.
@ -170,7 +185,7 @@ Re-runs skip words that are already staged.
**Acceptance criteria:**
- SQLite database contains the full deduped noun lists
(~1,600–1,800 words × 5 languages) with senses and translations.
(~1,550–1,750 unique words × 5 languages) with senses and translations.
- Rejection rate is below 10%.
- Spot-checked entries have correct definitions, plausible examples,
correct genders, and reasonable difficulty levels.
@ -179,6 +194,29 @@ Re-runs skip words that are already staged.
- Raw Gemini responses are on disk, so validation-rule changes can be
re-applied without re-calling the API.
**Status against the criteria:** resumability is verified working (a live run
resumed correctly from previously staged words), raw responses are on disk,
and the reject rate is tracking well under 10%. The two open items are the
systemic rejection cause and the hard-tier shortfall below.
**Open issue — systemic rejection cause.** Effectively every rejection is
`translation difficulty lower than sense difficulty`: the model tags a sense
`medium` while correctly tagging some of its translations `easy`. Example:
`Ellbogen` with sense `medium` but `elbow` (en) and `codo` (es) as `easy` —
the translations are right and the sense label is wrong, yet the whole entry
is discarded. The prompt already defines sense difficulty as the easiest
translation difficulty in the sense, so the value is derivable. Normalizing it
in `validate.ts` rather than rejecting would recover these entries and can be
replayed against `responses/` without new API calls.
**Open issue — the `hard` tier is nearly empty.** Generated difficulty skews
heavily easy; `hard` translations are well under 1% of staged rows. Because
design-doc §5.1 filters translation difficulty as an _exact_ match, a "hard"
game currently resolves to a single-digit row count for a given language pair —
not enough for one round plus three distractors. This is prompt-calibration
work and belongs here in Phase 3, before the import in Phase 4, since fixing it
afterwards means re-importing.
---
## Phase 4: Import
@ -372,7 +410,15 @@ Listed here for visibility. Not planned, not estimated.
`words(headword, language_code, pos)`, `senses(word_id, sense_index)`,
`translations(sense_id, target_language_code, translation)`.
### 3.2 Template the prompt
### 3.2 Template the prompt (done)
Implemented in `promptTemplate.ts`. Structured output is implemented in
`gemini.ts` via `buildEntriesResponseSchema()`, which narrows the enums to
the batch's own source language, POS, and target languages rather than
allowing the full supported list. Fence-stripping is retained in
`parseEntries()` as the defensive fallback the plan called for.
Original spec:
- Turn `data-pipeline/prompt` into a template. Substitution slots:
- source language (name + code)
@ -387,7 +433,15 @@ Listed here for visibility. Not planned, not estimated.
matching design-doc §6.3) instead of relying on prompt instructions
alone. Keep fence-stripping as a defensive fallback only.
### 3.3 Write the pipeline script
### 3.3 Write the pipeline script (done)
Implemented as specced. Divergences from the plan below: the rate-limit pause
defaults to 6s rather than 1s, batching/language selection is controlled by CLI
flags (`--langs`, `--pos`, `--batch-size`, `--max-batches`, `--delay-ms`,
`--dry-run`), and `stageEntry()` does an explicit existence check inside the
transaction instead of relying on `INSERT OR IGNORE`. Words the model omits
from a response entirely are logged to the rejection file as
`"missing from Gemini response"`, so a silent drop can't go unnoticed.
- Entry point: `data-pipeline/pipeline.ts`
(`pnpm --filter @lila/pipeline pipeline:run`).
@ -430,7 +484,36 @@ Listed here for visibility. Not planned, not estimated.
- Store definitions/examples as JSON strings in SQLite
(`JSON.stringify(arr)`).
### 3.4 Validation module
### 3.4 Validation module (done)
Implemented in `validate.ts` with unit tests in `tests/validate.test.ts`.
**As built, it is stricter than this spec.** Two structural differences:
- The result is a three-way discriminated union, not `{ valid, errors }`.
`"empty"` (the contract's `"senses": []`, meaning "not a valid word of this
POS") is a distinct outcome from `"invalid"`, so genuinely-not-a-noun words
are skipped and counted separately rather than polluting the rejection log.
- Validation is context-aware. `ValidationContext` carries the batch's source
language, POS, target languages, and input words, so the checks below are
exact-match against the request rather than membership in the global
supported list.
Additional checks not in the original spec:
- `headword` must be one of the words actually sent in this batch.
- `language` must equal the batch's source language; `pos` must equal the
batch's POS.
- `sense_index` must equal the sense's position in the array (sequential
from 0), not merely be a non-negative integer.
- At most 3 senses per word.
- Every target language must have at least one translation, at most 2, with
no duplicate translation word within a target language.
- A translation's difficulty may not rank below its sense's difficulty —
this is the cross-field rule responsible for essentially all current
rejections (see the open issue in Phase 3 above).
Original spec:
- Create the validation module alongside the pipeline, with unit tests
(vitest is already configured; `data-pipeline/vitest.config.ts`