implementing phase 3 pipeline: gemini structured output, validation, sqlite staging

Prompt is now a template (fixes the hardcoded en/es leftovers in rules
2, 3, 15, 16, 26, 31). pipeline.ts replaces the pseudocode: wordlist
normalization, skip-already-staged idempotency, batches of 20 against
gemini-3.6-flash with responseSchema, raw responses persisted per batch,
per-entry validation with rejection log, one transaction per word into
db/staging.db. Flags: --langs --pos --max-batches --delay-ms --dry-run.

Smoke run: 40/40 words staged (de+es, one batch each), 0 rejections.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
lila 2026-08-09 19:04:04 +02:00
parent 303bb9388c
commit 37c978e230
12 changed files with 1477 additions and 66 deletions

View file

@ -1,31 +1,11 @@
You are a multilingual lexicographer generating vocabulary data for a language-learning app.
Source language: spanish
Part of speech: noun
Target languages: en, it, de, fr
Source language: {{SOURCE_LANGUAGE_NAME}} ("{{SOURCE_LANGUAGE_CODE}}")
Part of speech: {{POS}}
Target languages: {{TARGET_LANGUAGE_CODES}}
Input words:
suelo
pared
techo
portón
valla
esquina
centro
borde
superficie
centro
suburbio
medioambiente
oportunidad
ventaja
decisión
paciencia
comportamiento
industria
conocimiento
solución
{{INPUT_WORDS}}
Return ONLY valid JSON.
Do not include markdown fences.
@ -39,8 +19,8 @@ Each word object must have this shape:
{
"headword": string,
"language": "es",
"pos": "noun",
"language": "{{SOURCE_LANGUAGE_CODE}}",
"pos": "{{POS}}",
"senses": [
{
"sense_index": number,
@ -49,7 +29,7 @@ Each word object must have this shape:
"examples": string[],
"translations": [
{
"target_language": "de" | "it" | "en" | "fr",
"target_language": {{TARGET_LANGUAGE_UNION}},
"word": string,
"gender": "masculine" | "feminine" | "neuter" | null,
"difficulty": "easy" | "medium" | "hard"
@ -62,36 +42,36 @@ Each word object must have this shape:
Rules:
1. The headword must be exactly one of the input words.
2. language must be "en".
3. pos must be "noun".
2. language must be "{{SOURCE_LANGUAGE_CODE}}".
3. pos must be "{{POS}}".
4. Include only common, learner-relevant senses.
5. Most words should have 1 sense.
6. Polysemous words may have 2 or 3 senses.
7. Do not include rare, archaic, highly technical, or literary senses unless they are common.
8. sense_index must start at 0 and increase by 1.
9. definitions must be written in English.
10. examples must be written in English.
9. definitions must be written in {{SOURCE_LANGUAGE_NAME}}.
10. examples must be written in {{SOURCE_LANGUAGE_NAME}}.
11. Include 1 or 2 definitions per sense.
12. Include 1 or 2 example sentences per sense.
13. Each definition must be student-friendly and at most 15 words.
14. Each example should naturally contain the headword or a clear form of it.
15. Every sense must have translations for all target languages: de, it, es, fr.
16. Do not include English as a target_language.
15. Every sense must have translations for all target languages: {{TARGET_LANGUAGE_CODES}}.
16. Do not include {{SOURCE_LANGUAGE_NAME}} ("{{SOURCE_LANGUAGE_CODE}}") as a target_language.
17. You may include up to 2 translations per target language per sense if they are genuinely common synonyms or difficulty variants.
18. Do not include more than 2 translations per target language per sense.
19. Do not duplicate the same translation word for the same target language within one sense.
20. Use the base dictionary form of the translated noun.
20. Use the base dictionary form of the translated {{POS}}.
21. Do not include articles or determiners in translations.
22. German translation nouns must be capitalized.
23. Spanish, French, and Italian translation nouns should be lowercase unless they are proper nouns.
24. For target_language "de", gender must be "masculine", "feminine", or "neuter".
25. For target_language "it", "es", or "fr", gender must be "masculine" or "feminine".
26. For this prompt, gender must never be null.
26. gender must be null if and only if target_language is "en".
27. difficulty must be one of: "easy", "medium", "hard".
28. senses.difficulty describes how common or advanced the meaning is.
29. translations.difficulty describes how difficult the specific target-language word is for a learner.
30. A common translation like "Bank" may be easy, while a formal synonym like "Geldinstitut" may be medium.
31. If a word cannot be treated as a valid English noun, return it with "senses": [].
31. If a word cannot be treated as a valid {{SOURCE_LANGUAGE_NAME}} {{POS}}, return it with "senses": [].
Difficulty calibration:
@ -101,7 +81,7 @@ Difficulty calibration:
Do not include CEFR levels in the output.
Example output shape for the English noun "bank":
Example output shape for the English noun "bank" (illustrative of the JSON shape only — your definitions and examples must be in {{SOURCE_LANGUAGE_NAME}}):
[
{