Every rejection in the German run was the same rule: a translation ranked
below its own sense. The model tags a sense "medium" while correctly
tagging some translations "easy" — the translations are right and the
derived sense label is wrong, but the whole entry was discarded.
The prompt defines sense difficulty as the easiest translation difficulty
in that sense, so it is a derived value rather than an independent
judgement. validate.ts now recomputes it via applySenseDifficultyFloor.
The floor only ever lowers. Raising a sense to match its translations
would gate a concept out of levels it belongs in and collapse the
concept-vs-word distinction the two difficulty columns exist to express
(design-doc section 4).
- validate.ts: drop the cross-field rejection, add the floor; the valid
result now carries "normalizations" so repairs are reported, not silent
- pipeline.ts: count and print normalizations per batch and in the summary
- replay.ts: new, re-validates responses/ with the current rules and no
API calls; --write stages recovered entries, --langs and --verbose
- tests: six cases covering the floor, replacing the old rejection test
Replaying all 38 saved responses took the reject rate from 20 entries to
zero. staging.db now holds 746 words / 774 senses / 3,436 translations
with no sense ranked above its easiest translation.
Docs also record the API quota ceiling found today: the free tier allows
about 20 requests/day, not the 1,000 previously assumed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The pipeline docs still described pipeline.ts as pseudocode and the
validation module as unwritten. Both have been implemented and run.
- CLAUDE.md: replace the "no executable pipeline yet" description with
the actual module flow, plus the two invariants worth preserving
(resumability via headword diffing, raw responses saved before parsing)
- DATA_PIPELINE.md: mark the seven implemented modules, add a module
responsibility map and the CLI flag table, drop the resolved warning
about hardcoded prompt values
- roadmap.md: check off phase 3 tasks 3.2/3.3/3.4, record where the
build diverged from the plan, note that validate.ts is stricter than
its own spec
- STATUS.md: phase 3 is data work now, not code work
Also records two open issues: the systemic difficulty-ordering rejection
cause, and the hard-tier shortfall pending a full run.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Roadmap: phase 4 re-scoped to import only (migration already applied),
phase 3 gains prompt templating, wordlist dedup, structured output,
raw-response persistence, and validation tests; idempotency decision
documented (skip already-staged words, one transaction per word).
Stale paths, postgres setup, and file structure corrected.
STATUS.md: refreshed from stale 2026-05-15 Kaikki state to current
gemini-only pipeline work.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>