flooring sense difficulty instead of rejecting the entry

Every rejection in the German run was the same rule: a translation ranked
below its own sense. The model tags a sense "medium" while correctly
tagging some translations "easy" — the translations are right and the
derived sense label is wrong, but the whole entry was discarded.

The prompt defines sense difficulty as the easiest translation difficulty
in that sense, so it is a derived value rather than an independent
judgement. validate.ts now recomputes it via applySenseDifficultyFloor.

The floor only ever lowers. Raising a sense to match its translations
would gate a concept out of levels it belongs in and collapse the
concept-vs-word distinction the two difficulty columns exist to express
(design-doc section 4).

- validate.ts: drop the cross-field rejection, add the floor; the valid
  result now carries "normalizations" so repairs are reported, not silent
- pipeline.ts: count and print normalizations per batch and in the summary
- replay.ts: new, re-validates responses/ with the current rules and no
  API calls; --write stages recovered entries, --langs and --verbose
- tests: six cases covering the floor, replacing the old rejection test

Replaying all 38 saved responses took the reject rate from 20 entries to
zero. staging.db now holds 746 words / 774 senses / 3,436 translations
with no sense ranked above its easiest translation.

Docs also record the API quota ceiling found today: the free tier allows
about 20 requests/day, not the 1,000 previously assumed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
lila 2026-08-20 12:34:37 +02:00
parent da9cdbfa1b
commit 5ee594334c
7 changed files with 454 additions and 50 deletions

View file

@ -70,6 +70,7 @@ co-located unit tests in `data-pipeline/tests/`.
| `data-pipeline/validate.ts` | ✅ Per-entry validation (design-doc §6.4) |
| `data-pipeline/staging.ts` | ✅ SQLite writes, one transaction per word |
| `data-pipeline/pipeline.ts` | ✅ Orchestrator with CLI flags, resumable |
| `data-pipeline/replay.ts` | ✅ Re-validates saved responses without API calls |
| `data-pipeline/db/schema.sql` | ✅ SQLite staging schema |
| `data-pipeline/db/staging.db` | ✅ Populated — gitignored |
| `data-pipeline/responses/` | ✅ Raw Gemini responses, one JSON per batch — gitignored |
@ -91,8 +92,11 @@ gemini.ts buildEntriesResponseSchema(): OpenAPI schema with enums narro
generateContent(): 5 attempts, retries 429/500/503, honours the
API's own retryDelay; rejects non-STOP finishReason
validate.ts validateEntry() → "valid" | "empty" | "invalid"
applySenseDifficultyFloor(): lowers a sense to its easiest
translation, reported via the result's `normalizations`
staging.ts openStaging(), getStagedHeadwords(), stageEntry(), countStagedRows()
pipeline.ts orchestration, batching, rate-limit delay, rejection logging
replay.ts re-validates responses/ with the current rules, no API calls
```
Two properties worth knowing:
@ -155,12 +159,14 @@ actually in the input batch.
Validation is the safety net, not the prompt — every entry is checked before it reaches SQLite, and rejects go to a log for review rather than silently disappearing. Target reject rate is under 10%.
> **Known systemic rejection cause (open).** Effectively all current rejections are
> `translation difficulty lower than sense difficulty`. The model tends to tag a sense
> `medium` while correctly tagging some of its translations `easy`. Since the prompt already
> defines sense difficulty as the easiest translation difficulty in the sense, this value is
> derivable and the entry is otherwise good — normalizing it in `validate.ts` instead of
> rejecting would recover these words. Not yet implemented.
> **Systemic rejection cause — resolved.** Effectively every rejection used to be
> `translation difficulty lower than sense difficulty`: the model tags a sense `medium`
> while correctly tagging some of its translations `easy`. The prompt defines sense
> difficulty as the easiest translation difficulty in the sense, which makes it a derived
> value, so `validate.ts` now floors it (`applySenseDifficultyFloor`) instead of rejecting
> the entry. The floor only ever lowers — raising a sense would gate a concept out of levels
> it belongs in and collapse the concept-vs-word split of design-doc §4. Replaying every
> saved response with the new rule took the reject rate from 20 entries to **zero**.
---
@ -191,6 +197,40 @@ pnpm --filter @lila/pipeline pipeline:run -- --langs de --max-batches 1 --dry-ru
Because runs are resumable, interrupting with Ctrl-C is safe — already-staged words are
skipped on the next run.
### Replaying saved responses
`replay.ts` re-validates everything in `responses/` **without calling the API**, which is how
a validation-rule change is applied to data that has already been generated.
```bash
pnpm --filter @lila/pipeline pipeline:replay # report only
pnpm --filter @lila/pipeline pipeline:replay -- --write # stage recovered entries
pnpm --filter @lila/pipeline pipeline:replay -- --verbose --langs de
```
It reports valid / normalized / empty / invalid counts and groups anything still invalid by
error. Staging is idempotent, so `--write` is safe to repeat. One limitation: it can only
_insert_. A word already in `staging.db` is left untouched, so a rule change cannot repair
rows that were staged under the old rules — only recover ones that were rejected.
### API quota
⚠️ **The free tier allows far fewer requests than the batch count needs.** Observed
2026-08-20: `gemini-3.6-flash` returned `HTTP 429 … limit: 20` on metric
`generate_content_free_tier_requests` after **17 successful requests in one day**, and did
not recover across ~55 minutes of retrying — a daily window, not a per-minute one. The
`retryDelay: ~59s` in the 429 payload is generic backoff advice and misleads the retry logic
into grinding.
Consequences to plan around:
- What is rationed is **requests**, not words, so `--batch-size` is the cheap lever. The
remaining ~7,300 words are ~365 requests at batch size 20, but only ~74 at batch size 100.
- `gemini.ts` currently treats 429 as retryable and burns all 5 `MAX_ATTEMPTS` against a
quota that will not clear for hours. It should distinguish rate-limiting from daily
exhaustion and abort the run.
- Replays cost nothing, which is why raw responses are persisted.
The pipeline reads `.env` from the repo root: `GEMINI_API_KEY`, plus `PIPELINE_POSTGRES_USER` / `PIPELINE_POSTGRES_PASSWORD` / `PIPELINE_POSTGRES_DB` / `PIPELINE_DATABASE_URL`. The pipeline database is deliberately separate from the app database (`:5432`) so pipeline work can never damage dev data.
---