flooring sense difficulty instead of rejecting the entry
Every rejection in the German run was the same rule: a translation ranked below its own sense. The model tags a sense "medium" while correctly tagging some translations "easy" — the translations are right and the derived sense label is wrong, but the whole entry was discarded. The prompt defines sense difficulty as the easiest translation difficulty in that sense, so it is a derived value rather than an independent judgement. validate.ts now recomputes it via applySenseDifficultyFloor. The floor only ever lowers. Raising a sense to match its translations would gate a concept out of levels it belongs in and collapse the concept-vs-word distinction the two difficulty columns exist to express (design-doc section 4). - validate.ts: drop the cross-field rejection, add the floor; the valid result now carries "normalizations" so repairs are reported, not silent - pipeline.ts: count and print normalizations per batch and in the summary - replay.ts: new, re-validates responses/ with the current rules and no API calls; --write stages recovered entries, --langs and --verbose - tests: six cases covering the floor, replacing the old rejection test Replaying all 38 saved responses took the reject rate from 20 entries to zero. staging.db now holds 746 words / 774 senses / 3,436 translations with no sense ranked above its easiest translation. Docs also record the API quota ceiling found today: the free tier allows about 20 requests/day, not the 1,000 previously assumed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
da9cdbfa1b
commit
5ee594334c
7 changed files with 454 additions and 50 deletions
|
|
@ -176,8 +176,9 @@ Re-runs skip words that are already staged.
|
|||
`--dry-run`.
|
||||
- [ ] Run the pipeline for all 5 languages (nouns only) — **in progress**,
|
||||
German first
|
||||
- [ ] Review the rejection log, fix prompt issues, re-run failed batches
|
||||
— reviewed; one systemic cause found (see below), fix not yet applied
|
||||
- [x] Review the rejection log, fix prompt issues, re-run failed batches —
|
||||
one systemic cause found and fixed in `validate.ts`; recovered from
|
||||
`responses/` via `replay.ts`, no re-run needed. Reject rate now zero.
|
||||
- [ ] Spot-check 50 random entries for correctness
|
||||
|
||||
**Dependencies:** Phase 2 complete.
|
||||
|
|
@ -199,15 +200,25 @@ resumed correctly from previously staged words), raw responses are on disk,
|
|||
and the reject rate is tracking well under 10%. The two open items are the
|
||||
systemic rejection cause and the hard-tier shortfall below.
|
||||
|
||||
**Open issue — systemic rejection cause.** Effectively every rejection is
|
||||
**Resolved — systemic rejection cause.** Every rejection was
|
||||
`translation difficulty lower than sense difficulty`: the model tags a sense
|
||||
`medium` while correctly tagging some of its translations `easy`. Example:
|
||||
`Ellbogen` with sense `medium` but `elbow` (en) and `codo` (es) as `easy` —
|
||||
the translations are right and the sense label is wrong, yet the whole entry
|
||||
is discarded. The prompt already defines sense difficulty as the easiest
|
||||
translation difficulty in the sense, so the value is derivable. Normalizing it
|
||||
in `validate.ts` rather than rejecting would recover these entries and can be
|
||||
replayed against `responses/` without new API calls.
|
||||
was discarded. Since the prompt defines sense difficulty as the easiest
|
||||
translation difficulty in the sense, the value is derived, so `validate.ts`
|
||||
now floors it (`applySenseDifficultyFloor`) rather than rejecting. The floor
|
||||
only lowers; raising a sense would gate a concept out of levels it belongs in.
|
||||
Replaying all 38 saved responses took the reject rate from 20 entries to zero,
|
||||
and `staging.db` now holds 746 words / 774 senses / 3,436 translations with no
|
||||
sense ranked above its easiest translation.
|
||||
|
||||
**Open issue — API quota is the binding constraint.** The free tier allows
|
||||
~20 requests/day for `gemini-3.6-flash` (observed 2026-08-20), not the ~1,000
|
||||
previously assumed. At `--batch-size 20` the remaining ~7,300 words are ~365
|
||||
requests, i.e. roughly 18 days of waiting. Requests are what is rationed, not
|
||||
words, so raising `--batch-size` is the cheap lever (~74 requests at 100).
|
||||
Paid billing is the alternative. Decide before scheduling the rest of the run.
|
||||
|
||||
**Open issue — the `hard` tier is nearly empty.** Generated difficulty skews
|
||||
heavily easy; `hard` translations are well under 1% of staged rows. Because
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue