implementing phase 3 pipeline: gemini structured output, validation, sqlite staging
Prompt is now a template (fixes the hardcoded en/es leftovers in rules 2, 3, 15, 16, 26, 31). pipeline.ts replaces the pseudocode: wordlist normalization, skip-already-staged idempotency, batches of 20 against gemini-3.6-flash with responseSchema, raw responses persisted per batch, per-entry validation with rejection log, one transaction per word into db/staging.db. Flags: --langs --pos --max-batches --delay-ms --dry-run. Smoke run: 40/40 words staged (de+es, one batch each), 0 rejections. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
parent
303bb9388c
commit
37c978e230
12 changed files with 1477 additions and 66 deletions
|
|
@ -1,31 +1,11 @@
|
|||
You are a multilingual lexicographer generating vocabulary data for a language-learning app.
|
||||
|
||||
Source language: spanish
|
||||
Part of speech: noun
|
||||
Target languages: en, it, de, fr
|
||||
Source language: {{SOURCE_LANGUAGE_NAME}} ("{{SOURCE_LANGUAGE_CODE}}")
|
||||
Part of speech: {{POS}}
|
||||
Target languages: {{TARGET_LANGUAGE_CODES}}
|
||||
|
||||
Input words:
|
||||
suelo
|
||||
pared
|
||||
techo
|
||||
portón
|
||||
valla
|
||||
esquina
|
||||
centro
|
||||
borde
|
||||
superficie
|
||||
centro
|
||||
suburbio
|
||||
medioambiente
|
||||
oportunidad
|
||||
ventaja
|
||||
decisión
|
||||
paciencia
|
||||
comportamiento
|
||||
industria
|
||||
conocimiento
|
||||
solución
|
||||
|
||||
{{INPUT_WORDS}}
|
||||
|
||||
Return ONLY valid JSON.
|
||||
Do not include markdown fences.
|
||||
|
|
@ -39,8 +19,8 @@ Each word object must have this shape:
|
|||
|
||||
{
|
||||
"headword": string,
|
||||
"language": "es",
|
||||
"pos": "noun",
|
||||
"language": "{{SOURCE_LANGUAGE_CODE}}",
|
||||
"pos": "{{POS}}",
|
||||
"senses": [
|
||||
{
|
||||
"sense_index": number,
|
||||
|
|
@ -49,7 +29,7 @@ Each word object must have this shape:
|
|||
"examples": string[],
|
||||
"translations": [
|
||||
{
|
||||
"target_language": "de" | "it" | "en" | "fr",
|
||||
"target_language": {{TARGET_LANGUAGE_UNION}},
|
||||
"word": string,
|
||||
"gender": "masculine" | "feminine" | "neuter" | null,
|
||||
"difficulty": "easy" | "medium" | "hard"
|
||||
|
|
@ -62,36 +42,36 @@ Each word object must have this shape:
|
|||
Rules:
|
||||
|
||||
1. The headword must be exactly one of the input words.
|
||||
2. language must be "en".
|
||||
3. pos must be "noun".
|
||||
2. language must be "{{SOURCE_LANGUAGE_CODE}}".
|
||||
3. pos must be "{{POS}}".
|
||||
4. Include only common, learner-relevant senses.
|
||||
5. Most words should have 1 sense.
|
||||
6. Polysemous words may have 2 or 3 senses.
|
||||
7. Do not include rare, archaic, highly technical, or literary senses unless they are common.
|
||||
8. sense_index must start at 0 and increase by 1.
|
||||
9. definitions must be written in English.
|
||||
10. examples must be written in English.
|
||||
9. definitions must be written in {{SOURCE_LANGUAGE_NAME}}.
|
||||
10. examples must be written in {{SOURCE_LANGUAGE_NAME}}.
|
||||
11. Include 1 or 2 definitions per sense.
|
||||
12. Include 1 or 2 example sentences per sense.
|
||||
13. Each definition must be student-friendly and at most 15 words.
|
||||
14. Each example should naturally contain the headword or a clear form of it.
|
||||
15. Every sense must have translations for all target languages: de, it, es, fr.
|
||||
16. Do not include English as a target_language.
|
||||
15. Every sense must have translations for all target languages: {{TARGET_LANGUAGE_CODES}}.
|
||||
16. Do not include {{SOURCE_LANGUAGE_NAME}} ("{{SOURCE_LANGUAGE_CODE}}") as a target_language.
|
||||
17. You may include up to 2 translations per target language per sense if they are genuinely common synonyms or difficulty variants.
|
||||
18. Do not include more than 2 translations per target language per sense.
|
||||
19. Do not duplicate the same translation word for the same target language within one sense.
|
||||
20. Use the base dictionary form of the translated noun.
|
||||
20. Use the base dictionary form of the translated {{POS}}.
|
||||
21. Do not include articles or determiners in translations.
|
||||
22. German translation nouns must be capitalized.
|
||||
23. Spanish, French, and Italian translation nouns should be lowercase unless they are proper nouns.
|
||||
24. For target_language "de", gender must be "masculine", "feminine", or "neuter".
|
||||
25. For target_language "it", "es", or "fr", gender must be "masculine" or "feminine".
|
||||
26. For this prompt, gender must never be null.
|
||||
26. gender must be null if and only if target_language is "en".
|
||||
27. difficulty must be one of: "easy", "medium", "hard".
|
||||
28. senses.difficulty describes how common or advanced the meaning is.
|
||||
29. translations.difficulty describes how difficult the specific target-language word is for a learner.
|
||||
30. A common translation like "Bank" may be easy, while a formal synonym like "Geldinstitut" may be medium.
|
||||
31. If a word cannot be treated as a valid English noun, return it with "senses": [].
|
||||
31. If a word cannot be treated as a valid {{SOURCE_LANGUAGE_NAME}} {{POS}}, return it with "senses": [].
|
||||
|
||||
Difficulty calibration:
|
||||
|
||||
|
|
@ -101,7 +81,7 @@ Difficulty calibration:
|
|||
|
||||
Do not include CEFR levels in the output.
|
||||
|
||||
Example output shape for the English noun "bank":
|
||||
Example output shape for the English noun "bank" (illustrative of the JSON shape only — your definitions and examples must be in {{SOURCE_LANGUAGE_NAME}}):
|
||||
|
||||
[
|
||||
{
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue