updating docs to match the implemented phase 3 pipeline
The pipeline docs still described pipeline.ts as pseudocode and the validation module as unwritten. Both have been implemented and run. - CLAUDE.md: replace the "no executable pipeline yet" description with the actual module flow, plus the two invariants worth preserving (resumability via headword diffing, raw responses saved before parsing) - DATA_PIPELINE.md: mark the seven implemented modules, add a module responsibility map and the CLI flag table, drop the resolved warning about hardcoded prompt values - roadmap.md: check off phase 3 tasks 3.2/3.3/3.4, record where the build diverged from the plan, note that validate.ts is stricter than its own spec - STATUS.md: phase 3 is data work now, not code work Also records two open issues: the systemic difficulty-ordering rejection cause, and the hard-tier shortfall pending a full run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
37c978e230
commit
da9cdbfa1b
4 changed files with 219 additions and 49 deletions
|
|
@ -5,7 +5,7 @@
|
|||
> backed by a normalized Postgres schema.
|
||||
>
|
||||
> **Author:** lila
|
||||
> **Date:** July 2026 · Last reviewed: 2026-08-09
|
||||
> **Date:** July 2026 · Last reviewed: 2026-08-20
|
||||
> **Companion doc:** `documentation/pipeline/design-doc.md`
|
||||
|
||||
---
|
||||
|
|
@ -33,14 +33,17 @@ Phase 1 Schema ✅
|
|||
|
||||
Phase 2 Preparation ✅
|
||||
Wordlists acquired, databases set up, prompt drafted and
|
||||
tested. Two loose ends carried into Phase 3: the prompt is
|
||||
not templated yet, and the wordlists still contain duplicates
|
||||
(deduped at runtime, not in the files).
|
||||
tested. One loose end carried into Phase 3 and resolved
|
||||
there: the prompt is now templated. The wordlist files still
|
||||
contain duplicates, deduped at runtime rather than in the
|
||||
files — harmless, and left as-is.
|
||||
|
||||
Phase 3 Data Pipeline ← CURRENT
|
||||
Build the Gemini → validate → SQLite pipeline.
|
||||
Produce a clean dataset: full deduped noun lists
|
||||
(~1,600–1,800 words) × 5 languages.
|
||||
(~1,550–1,750 unique words) × 5 languages.
|
||||
Code is complete and unit-tested; the remaining work is the
|
||||
data run, the rejection review, and the spot-check.
|
||||
|
||||
Phase 4 Import
|
||||
(Migration already applied in Phase 1.)
|
||||
|
|
@ -122,9 +125,10 @@ constraints, indexes, and relations.
|
|||
(`data-pipeline/db/staging.db` from `db/schema.sql`)
|
||||
- [x] Write and test the Gemini prompt with sample words
|
||||
- [x] Refine the prompt until the JSON output matches the contract
|
||||
defined in `design-doc.md` §6.3 — **note:** the prompt works but
|
||||
is pinned to a hardcoded Spanish sample and has known copy-paste
|
||||
bugs; templating and fixes are Phase 3 tasks
|
||||
defined in `design-doc.md` §6.3 — **note:** at the end of Phase 2 the
|
||||
prompt worked but was pinned to a hardcoded Spanish sample and had
|
||||
known copy-paste bugs. Both were resolved by the templating task in
|
||||
Phase 3.
|
||||
|
||||
**Dependencies:** Phase 1 complete.
|
||||
|
||||
|
|
@ -149,20 +153,31 @@ Re-runs skip words that are already staged.
|
|||
|
||||
- [x] Create the SQLite schema (`data-pipeline/db/schema.sql`, tables
|
||||
created in `db/staging.db`)
|
||||
- [ ] Template the Gemini prompt: source language, POS, target
|
||||
languages, and the word batch become substitutions; fix the known
|
||||
copy-paste bugs while doing so (rules 2–3 hardcode `"en"`,
|
||||
rule 15's target list contradicts the header, rule 31 says
|
||||
"valid English noun")
|
||||
- [ ] Write the wordlist normalization step: trim whitespace, drop
|
||||
empty lines, dedup in memory
|
||||
- [ ] Write the validation module (rules in design-doc §6.4) with unit
|
||||
tests (`vitest.config.ts` expects `tests/**/*.test.ts` — the
|
||||
directory doesn't exist yet)
|
||||
- [ ] Write the pipeline script (`pipeline.ts`): - Read + normalize wordlist files - Skip words already present in staging.db (idempotency — see §3.3) - Split the remainder into batches of 20 - Call Gemini per batch using structured output
|
||||
(`responseMimeType: "application/json"` + `responseSchema`) - Persist each raw response to disk before validating - Validate each entry - Write valid entries to SQLite, one transaction per word - Log invalid entries to a rejection file
|
||||
- [ ] Run the pipeline for all 5 languages (nouns only)
|
||||
- [x] Template the Gemini prompt (`promptTemplate.ts` + placeholders in
|
||||
`prompt`); the copy-paste bugs (rules 2–3 hardcoding `"en"`,
|
||||
rule 15's contradictory target list, "valid English noun" in
|
||||
rule 31) are fixed. `renderPrompt()` throws if any placeholder
|
||||
survives rendering.
|
||||
- [x] Write the wordlist normalization step (`sourceLists.ts`:
|
||||
`normalizeWords()` trims, drops empty lines, dedups in memory
|
||||
preserving order)
|
||||
- [x] Write the validation module (`validate.ts`, rules in design-doc
|
||||
§6.4) with unit tests — `tests/` now holds `validate.test.ts`,
|
||||
`staging.test.ts`, `promptTemplate.test.ts`, `sourceLists.test.ts`
|
||||
- [x] Write the pipeline script (`pipeline.ts`): reads + normalizes
|
||||
wordlists, skips words already in `staging.db`, batches the
|
||||
remainder, calls Gemini with structured output
|
||||
(`responseMimeType: "application/json"` + `responseSchema` from
|
||||
`gemini.ts`), persists each raw response to `responses/` before
|
||||
validating, validates each entry, writes valid entries to SQLite
|
||||
one transaction per word, and logs invalid entries to
|
||||
`rejections/{lang}-{pos}.jsonl`. Adds CLI flags beyond the plan:
|
||||
`--langs`, `--pos`, `--batch-size`, `--max-batches`, `--delay-ms`,
|
||||
`--dry-run`.
|
||||
- [ ] Run the pipeline for all 5 languages (nouns only) — **in progress**,
|
||||
German first
|
||||
- [ ] Review the rejection log, fix prompt issues, re-run failed batches
|
||||
— reviewed; one systemic cause found (see below), fix not yet applied
|
||||
- [ ] Spot-check 50 random entries for correctness
|
||||
|
||||
**Dependencies:** Phase 2 complete.
|
||||
|
|
@ -170,7 +185,7 @@ Re-runs skip words that are already staged.
|
|||
**Acceptance criteria:**
|
||||
|
||||
- SQLite database contains the full deduped noun lists
|
||||
(~1,600–1,800 words × 5 languages) with senses and translations.
|
||||
(~1,550–1,750 unique words × 5 languages) with senses and translations.
|
||||
- Rejection rate is below 10%.
|
||||
- Spot-checked entries have correct definitions, plausible examples,
|
||||
correct genders, and reasonable difficulty levels.
|
||||
|
|
@ -179,6 +194,29 @@ Re-runs skip words that are already staged.
|
|||
- Raw Gemini responses are on disk, so validation-rule changes can be
|
||||
re-applied without re-calling the API.
|
||||
|
||||
**Status against the criteria:** resumability is verified working (a live run
|
||||
resumed correctly from previously staged words), raw responses are on disk,
|
||||
and the reject rate is tracking well under 10%. The two open items are the
|
||||
systemic rejection cause and the hard-tier shortfall below.
|
||||
|
||||
**Open issue — systemic rejection cause.** Effectively every rejection is
|
||||
`translation difficulty lower than sense difficulty`: the model tags a sense
|
||||
`medium` while correctly tagging some of its translations `easy`. Example:
|
||||
`Ellbogen` with sense `medium` but `elbow` (en) and `codo` (es) as `easy` —
|
||||
the translations are right and the sense label is wrong, yet the whole entry
|
||||
is discarded. The prompt already defines sense difficulty as the easiest
|
||||
translation difficulty in the sense, so the value is derivable. Normalizing it
|
||||
in `validate.ts` rather than rejecting would recover these entries and can be
|
||||
replayed against `responses/` without new API calls.
|
||||
|
||||
**Open issue — the `hard` tier is nearly empty.** Generated difficulty skews
|
||||
heavily easy; `hard` translations are well under 1% of staged rows. Because
|
||||
design-doc §5.1 filters translation difficulty as an _exact_ match, a "hard"
|
||||
game currently resolves to a single-digit row count for a given language pair —
|
||||
not enough for one round plus three distractors. This is prompt-calibration
|
||||
work and belongs here in Phase 3, before the import in Phase 4, since fixing it
|
||||
afterwards means re-importing.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: Import
|
||||
|
|
@ -372,7 +410,15 @@ Listed here for visibility. Not planned, not estimated.
|
|||
`words(headword, language_code, pos)`, `senses(word_id, sense_index)`,
|
||||
`translations(sense_id, target_language_code, translation)`.
|
||||
|
||||
### 3.2 Template the prompt
|
||||
### 3.2 Template the prompt (done)
|
||||
|
||||
Implemented in `promptTemplate.ts`. Structured output is implemented in
|
||||
`gemini.ts` via `buildEntriesResponseSchema()`, which narrows the enums to
|
||||
the batch's own source language, POS, and target languages rather than
|
||||
allowing the full supported list. Fence-stripping is retained in
|
||||
`parseEntries()` as the defensive fallback the plan called for.
|
||||
|
||||
Original spec:
|
||||
|
||||
- Turn `data-pipeline/prompt` into a template. Substitution slots:
|
||||
- source language (name + code)
|
||||
|
|
@ -387,7 +433,15 @@ Listed here for visibility. Not planned, not estimated.
|
|||
matching design-doc §6.3) instead of relying on prompt instructions
|
||||
alone. Keep fence-stripping as a defensive fallback only.
|
||||
|
||||
### 3.3 Write the pipeline script
|
||||
### 3.3 Write the pipeline script (done)
|
||||
|
||||
Implemented as specced. Divergences from the plan below: the rate-limit pause
|
||||
defaults to 6s rather than 1s, batching/language selection is controlled by CLI
|
||||
flags (`--langs`, `--pos`, `--batch-size`, `--max-batches`, `--delay-ms`,
|
||||
`--dry-run`), and `stageEntry()` does an explicit existence check inside the
|
||||
transaction instead of relying on `INSERT OR IGNORE`. Words the model omits
|
||||
from a response entirely are logged to the rejection file as
|
||||
`"missing from Gemini response"`, so a silent drop can't go unnoticed.
|
||||
|
||||
- Entry point: `data-pipeline/pipeline.ts`
|
||||
(`pnpm --filter @lila/pipeline pipeline:run`).
|
||||
|
|
@ -430,7 +484,36 @@ Listed here for visibility. Not planned, not estimated.
|
|||
- Store definitions/examples as JSON strings in SQLite
|
||||
(`JSON.stringify(arr)`).
|
||||
|
||||
### 3.4 Validation module
|
||||
### 3.4 Validation module (done)
|
||||
|
||||
Implemented in `validate.ts` with unit tests in `tests/validate.test.ts`.
|
||||
|
||||
**As built, it is stricter than this spec.** Two structural differences:
|
||||
|
||||
- The result is a three-way discriminated union, not `{ valid, errors }`.
|
||||
`"empty"` (the contract's `"senses": []`, meaning "not a valid word of this
|
||||
POS") is a distinct outcome from `"invalid"`, so genuinely-not-a-noun words
|
||||
are skipped and counted separately rather than polluting the rejection log.
|
||||
- Validation is context-aware. `ValidationContext` carries the batch's source
|
||||
language, POS, target languages, and input words, so the checks below are
|
||||
exact-match against the request rather than membership in the global
|
||||
supported list.
|
||||
|
||||
Additional checks not in the original spec:
|
||||
|
||||
- `headword` must be one of the words actually sent in this batch.
|
||||
- `language` must equal the batch's source language; `pos` must equal the
|
||||
batch's POS.
|
||||
- `sense_index` must equal the sense's position in the array (sequential
|
||||
from 0), not merely be a non-negative integer.
|
||||
- At most 3 senses per word.
|
||||
- Every target language must have at least one translation, at most 2, with
|
||||
no duplicate translation word within a target language.
|
||||
- A translation's difficulty may not rank below its sense's difficulty —
|
||||
this is the cross-field rule responsible for essentially all current
|
||||
rejections (see the open issue in Phase 3 above).
|
||||
|
||||
Original spec:
|
||||
|
||||
- Create the validation module alongside the pipeline, with unit tests
|
||||
(vitest is already configured; `data-pipeline/vitest.config.ts`
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue