Lila Data Pipeline — Technical Documentation
Multilingual dictionary enrichment pipeline for the Lila vocabulary trainer.
Last updated: 2026-07-18
Table of Contents
Quick Reference
| What |
Where |
| Entry point |
pipeline.ts |
| Interactive CLI |
utils/cli.ts |
| LLM config schema |
config/llm.ts |
| System prompt |
config/prompt.ts — buildSystemPrompt() |
| Batch config schema |
config/batch.ts |
| Shared constants |
config/constants.ts |
| Output schema |
utils/merge-enriched-data.ts |
| LLM adapters |
utils/llm-adapters/ |
| Current model |
gemma-4-E2B_q4_0-it.gguf |
| Target scale |
100,000+ words |
1. Executive Summary
The Lila Data Pipeline is a TypeScript-based batch processing system that enriches raw word lists into structured multilingual dictionary entries for the Lila vocabulary trainer. It takes a source wordlist (e.g., English nouns) and, for each word, generates:
- One or more senses (definitions)
- A natural example sentence per sense
- A CEFR-based difficulty level (
easy / medium / hard)
- Translations into all target languages except the source (as raw strings)
The pipeline is designed to scale to 100,000+ words across multiple languages and parts of speech. It supports both local inference (llama.cpp) and cloud providers via a pluggable adapter pattern.
Key Design Principles
| Principle |
Rationale |
| Quality first |
Definitions, examples, and translations must be accurate. Speed and cost are secondary. |
| Local-first, cloud-fallback |
Local LLMs are the default for cost control and data privacy. |
| Deterministic Grammar |
Grammatical gender is decoupled from the LLM and resolved via Kaikki Wiktionary dumps. |
| Resumable & idempotent |
Each word writes to its own JSON file. The pipeline skips already-processed words on restart. |
| Configurable batching |
Batch size is selected interactively at runtime. The pipeline adapts without code changes. |
| Honest metrics |
Local models report detailed prompt/completion timing. Cloud providers report total request time only. |
Resolved: Gender Accuracy
Grammatical gender is no longer generated by the LLM. Comprehensive testing across 10 models (2026-07-18) confirmed that small local models systematically hallucinate or default to neuter for Romance languages.
Decision: The LLM only outputs translation strings. A deterministic post-processing step looks up the exact grammatical gender from Kaikki Wiktionary dumps. This guarantees 100% gender accuracy and allows us to use smaller, faster, and highly nuanced local models.
Current Status (2026-07-18)
- Core pipeline: Complete (scanning, enrichment, merging, verification, writing).
- Local LLM: Gemma 4 E2B selected as production model after exhaustive 10-model evaluation.
llama.cpp server optimized with KV-cache quantization.
- Batching: 20-word batches validated for local hardware.
- Gender: Decoupled from LLM; Kaikki lookup architecture confirmed.
- Pending: Online API evaluation (Gemini free tier, DeepSeek, Groq) as speed fallback.
One-Line Architecture
source wordlists -> Interactive CLI -> LLM adapter -> merge senses -> Kaikki Gender Lookup -> verify schema -> write .json
Files at a Glance
| File |
Purpose |
pipeline.ts |
Orchestrator. Runs CLI, scans sources, loops words, coordinates all stages |
utils/cli.ts |
Interactive CLI. Provider/model/batch selection, config persistence |
config/llm.ts |
LLM config schema (provider, url, model) |
config/prompt.ts |
buildSystemPrompt() — dynamic prompt with auto-target languages |
config/batch.ts |
Batch size and max retry count schema |
config/constants.ts |
Shared LANG_MAP, POS_MAP, ALL_LANGUAGES |
utils/enrich-word.ts |
Calls LLM via adapter, parses response, validates senses, builds EnrichedSense[], retry/split logic |
utils/merge-enriched-data.ts |
Merges skeleton + enriched senses into final JSON |
utils/verify-enriched-file.ts |
Schema validation (required fields, types) |
utils/check-llm-server.ts |
Health check for local server; skipped for cloud providers |
utils/scanning-source-files.ts |
Discovers wordlists from source-data/ directory |
utils/create-base-json.ts |
Writes skeleton {word, language, pos} files |
utils/write-json-file.ts |
Atomic .tmp → rename writes |
utils/check-if-json-exists.ts |
Resumability: checks if word already has enriched senses |
utils/create-line-reader.ts |
Streaming line reader for large wordlists |
utils/create-output-dirs.ts |
Creates worddata/{language}/{pos}/ folders |
utils/delete-file.ts |
Cleanup helper for failed batches |
utils/get-word-file-path.ts |
Path construction helper |
utils/progress-tracker.ts |
[current/total] formatting for console output |
utils/pipeline-timer.ts |
Timing + token metrics; unified throughput for all providers |
utils/llm-adapters/factory.ts |
Creates the right adapter based on runtime config |
utils/llm-adapters/types.ts |
LlmAdapter interface |
utils/llm-adapters/openai-compatible.ts |
Local llama.cpp, OpenRouter, DeepSeek |
utils/llm-adapters/gemini.ts |
Google Gemini native API |
Scale Target
| Metric |
Target |
| Words |
100,000+ |
| Languages |
English (source), German, Italian, Spanish, French (targets) |
| Parts of speech |
Nouns, verbs, adjectives, adverbs |
| Output |
One .json file per word, ~2-5KB each |
2. Problem & Context
Why Build This?
Existing multilingual dictionaries and translation APIs provide raw word-to-word mappings. They do not provide the structured, pedagogical data needed for a vocabulary trainer:
| What Exists |
What is Missing |
Why It Matters |
| Word + translation |
Student-friendly definition |
Learners need explanations, not just equivalents |
| Static difficulty labels |
CEFR-calibrated difficulty |
"Easy" vs "hard" must map to learner proficiency levels |
| Isolated translations |
Natural example sentences |
Context is how vocabulary is actually acquired |
| Raw gender data |
Gender integrated with translation |
Flashcards must show gender immediately (resolved via Kaikki) |
| Bulk word lists |
Per-word structured JSON |
The trainer consumes one file per word for fast random access |
The Target User
A language learner using the Lila vocabulary trainer. They see a word, its definition, an example sentence, and translations with gender — all calibrated to their CEFR level (A1-C2).
Why Not Use Existing Dictionaries?
- Wiktionary: Rich data, but unstructured, inconsistent formatting, no CEFR levels, no student-friendly definitions
- Kaikki (Wiktionary dump): Structured JSON, excellent for gender/translation lookup, but definitions are often technical, no difficulty classification, no example curation
- Google Translate / DeepL: No definitions, no examples, no difficulty, no structured output
- Existing language learning apps: Closed data, no export, no control over content
The LLM fills the gap: it generates pedagogical content (student-friendly definitions, natural examples, difficulty classification) that no existing database provides at scale.
Language Direction
The pipeline is direction-agnostic. A wordlist is defined by:
- Source language: the language of the input words
- Target languages: all other languages in the system (auto-derived from
ALL_LANGUAGES minus source)
Current focus: English -> German/Italian/Spanish/French. Planned directions include German -> French, Italian -> Spanish, etc.
Why 100,000+ Words?
- Coverage: A learner needs ~10,000 words for B2 proficiency. The pipeline targets 100k to cover multiple languages, POS categories, and difficulty levels.
- Languages: English (source) -> German, Italian, Spanish, French (targets).
- Parts of speech: Nouns (current), verbs, adjectives, adverbs.
The Quality Challenge
| Risk |
Mitigation |
| Hallucinated definitions |
Low temperature (0.1), strict system prompt, schema validation |
| Incorrect grammatical gender |
Resolved: Decoupled from LLM; Kaikki Wiktionary lookup |
| POS bleed (verb defs for nouns) |
Explicit negative constraint in system prompt |
| Inconsistent difficulty levels |
Explicit CEFR mapping in prompt, spot-checking |
| JSON parse failures |
Retry + split logic, schema validation, cleanup on failure |
| Model drift (online APIs) |
Version pinning, local fallback |
| Provider downtime |
Adapter pattern allows hot-swapping providers |
Why TypeScript + Node?
- Familiarity: Existing project uses TypeScript (frontend in TanStack Router + React)
- Ecosystem:
readline for streaming files, fs for JSON I/O, native fetch for HTTP
- Portability: Runs on the same Debian laptop as the llama.cpp server
- No build complexity:
tsx for direct execution, no bundler needed
Why llama.cpp?
- GGUF format: Single-file models, easy to swap, quantize, and version
- OpenAI-compatible API:
/v1/chat/completions means the same adapter code works for local and online models
- No dependencies: Self-contained binary, runs on old hardware (tested on GTX 950M)
- Privacy: Local inference means no data leaves the machine
3. Architecture Overview
Pipeline Flow
Run CLI -> Scan sources -> Check LLM -> Loop wordlists -> Stream words -> Skip processed
-> Create skeletons (batch) -> Call LLM -> Parse JSON -> Validate senses
-> Retry/split on failure -> Merge -> Kaikki Gender Lookup -> Write atomically -> Verify schema -> Log metrics
Resumability
- Skip existing:
check-if-json-exists.ts checks if {word}.json exists with non-empty senses
- Atomic writes:
.tmp -> rename in write-json-file.ts, no partial files on crash
- Cleanup on failure: Deletes partially-written files for failed batches, continues to next batch
Directory Structure
data-pipeline/
|-- pipeline.ts # Entry point / orchestrator
|-- utils/
| |-- cli.ts # Interactive CLI module
| |-- enrich-word.ts # LLM call, parse, retry/split
| |-- merge-enriched-data.ts # Schema types + merge logic
| |-- verify-enriched-file.ts # Schema validation
| |-- check-llm-server.ts # Health check (local only)
| |-- scanning-source-files.ts # Source discovery
| |-- create-base-json.ts # Skeleton writer
| |-- write-json-file.ts # Atomic JSON writer
| |-- check-if-json-exists.ts # Resumability check
| |-- create-line-reader.ts # Streaming file reader
| |-- create-output-dirs.ts # Directory creation
| |-- delete-file.ts # Cleanup helper
| |-- get-word-file-path.ts # Path helper
| |-- progress-tracker.ts # Console progress formatting
| |-- pipeline-timer.ts # Timing + token metrics
| |-- llm-adapters/
| |-- factory.ts # Adapter selection
| |-- types.ts # LlmAdapter interface
| |-- openai-compatible.ts # Local, OpenRouter, DeepSeek
| |-- gemini.ts # Google Gemini
|-- config/
| |-- llm.ts # LLM config schema
| |-- prompt.ts # buildSystemPrompt()
| |-- batch.ts # Batch config schema
| |-- constants.ts # LANG_MAP, POS_MAP, ALL_LANGUAGES
|-- source-data/
| |-- {language}/
| |-- {pos} # One word per line, no extension
|-- worddata/
| |-- {language}/
| |-- {pos}/
| |-- {word}.json # One self-contained file per word
|-- kaikki-source-files/ # Wiktionary dumps for gender lookup
|-- .pipeline-config.json # Saved CLI configuration
Output Schema
Each .json file contains: word, language, pos, senses[] (each with id, sense, example, difficulty_level, translations per target language), enrichedAt, model.
Note: The translations object contains arrays of strings (e.g., {"de": ["Haus", "Gebäude"]}). Grammatical gender is appended later via the Kaikki integration step.
Error Handling
| Failure |
Behavior |
| LLM server offline |
Hard fail at startup (check-llm-server.ts, local only) |
| LLM returns bad JSON |
Retry up to 3 times, then split batch. Log and continue |
| LLM returns malformed senses |
validateSense() catches it before file write |
| Schema validation fails |
Log warnings, keep file |
| Individual batch fails |
Does not stop pipeline; cleans up partial files |
| Individual word fails (size 1) |
Log and continue to next word |
Metrics
Per-run: words processed/skipped/failed, duration, throughput, LLM token counts and speeds. See utils/pipeline-timer.ts.
- Unified throughput (all providers): total tokens / total request time
- Detailed breakdown (local only): prompt speed vs completion speed
4. Current Implementation
Tech Stack
| Layer |
Choice |
Why |
| Runtime |
Node.js + tsx |
Direct TypeScript execution, no build step |
| HTTP client |
Native fetch |
Works for local llama.cpp and online APIs |
| File I/O |
fs + readline |
Streaming line reader for large wordlists |
| JSON |
Native JSON.parse/stringify |
Simple, no schema library needed |
| CLI |
Native readline |
No external dependencies |
Current Model
| Property |
Value |
| Model |
gemma-4-E2B_q4_0-it.gguf |
| Size |
~3.2GB (file) / ~2.06GB (VRAM weights) |
| Quantization |
Q4_0 |
| Server |
llama.cpp (llama-server) |
| API |
OpenAI-compatible /v1/chat/completions |
| VRAM Usage |
~2.65GB total (weights + KV cache + compute buffers) |
| Generation Speed |
~10.9 tok/s (20-word batch) |
| Est. 100k Time |
~7 days (20-word batches, 24/7) |
llama-server Flags: History & Rationale
The server flags evolved through rigorous empirical testing on the target hardware (Intel i7-6500U, GTX 950M 4GB, 8GB RAM) across 10 different models on 2026-07-18.
Flag Evolution
| Flag |
Value Tried |
Result |
Why |
-m |
qwen3.5-4b-q4_k_m.gguf |
Works, ~6.3 tok/s |
Quality baseline. Correct translations. 2.6GB, tight on VRAM. |
-m |
qwen2.5-1.5b-instruct-q4_k_m.gguf |
Works, ~18.5 tok/s |
Fast but poor instruction following, gender wrong. |
-m |
Qwen3.5-2B-Q4_K_M.gguf |
Works, ~13.2 tok/s |
Good translations, but failed polysemy (2 identical senses). |
-m |
Llama-3.2-3B-Instruct-Q4_K_M.gguf |
Works, ~9.1 tok/s |
Dangerous false friend trap (cognates). |
-m |
Llama-3.2-3B-Instruct-UD-Q6_K_XL.gguf |
Works, ~6.7 tok/s, 3.44GB VRAM |
Higher quant did NOT fix Llama's translation issues. Slower. |
-m |
Ministral-3-3B-Instruct-2512-Q4_K_M.gguf |
Works, ~8.9 tok/s |
Fixed tokenizer bug (older version was broken). But messy translations, markdown violations. |
-m |
gemma-4-E2B_q4_0-it.gguf |
Works, ~13.1 tok/s, 2.06GB |
Winner. Perfect polysemy, false friends, nuance. Half VRAM of Qwen 4B. |
-m |
gemma-4-E4B-it-qat-UD-Q4_K_XL.gguf |
Works, ~7.7 tok/s, 3.23GB |
QAT compression is incredible. Same quality as E2B but slower. |
-m |
qwen3.5-9b-q3_k_s.gguf |
Works, ~3.3 tok/s (split CPU/GPU) |
Brilliant quality but bottlenecked by CPU/Swap. -ngl 28 max. |
-m |
qwen3.5-9b-q4_k_m.gguf |
Works, ~3.3 tok/s (split CPU/GPU) |
Same quality as 9B Q3. Not worth 2x the file size. |
-ngl |
999 |
Keeps |
Offload all layers to GPU. Required for any speed. |
-ngl |
28 |
9B models only |
Max GPU layers for 9B models before OOM. Found via binary search. |
-ngl |
30 |
OOM crash (9B) |
Pushed 2 layers too far into compute buffers. |
-c |
4096 |
Wasteful |
4K context for 300-token dictionary entries wastes VRAM. |
-c |
2048 |
Good for small batches |
Sufficient for 4-word batches. |
-c |
8192 |
Current |
Required for 20-word batches. Combined with KV cache quantization. |
-b / -ub |
512 |
Current |
Sweet spot for Maxwell memory bandwidth. |
-b / -ub |
1024 |
Tested |
Slightly faster prompt processing, but no generation speedup. |
-b / -ub |
2048 |
Slower on 950M |
Memory pressure on bandwidth-starved GPU. |
-t |
4 |
Slower |
Hyperthreading cores hurt llama.cpp performance. |
-t |
2 |
Current |
Matches 2 physical cores. |
--threads-batch |
2 |
Current |
Explicit match to -t. |
--flash-attn |
(omitted) |
Correct |
On Maxwell (compute 5.0), Flash Attention adds overhead. |
--mlock |
Tested |
Omitted for large models |
Pins model in RAM. Causes OOM on models >2.5GB with large KV cache. |
--prio |
2 |
Current |
Raises process priority. Marginal, harmless. |
--reasoning |
off |
Critical |
Mandatory for Qwen 3.5 and Gemma 4. Without this, models "think" silently, consume all max_tokens, and crash with finish_reason: length. |
--cache-type-k |
q4_0 |
Current |
Compresses KV cache keys to 4-bit. Cuts KV VRAM by ~75%. Enables 8192 context on 4GB GPU. |
--cache-type-v |
q4_0 |
Current |
Compresses KV cache values to 4-bit. Paradoxically improves translation variety (reduces "lazy duplication" bug). |
Current Production Command
cd ~/Downloads/llama.cpp
./build/bin/llama-server \
-m models/gemma-4-E2B_q4_0-it.gguf \
-ngl 999 \
-c 8192 \
-b 512 \
-ub 512 \
-t 2 \
--threads-batch 2 \
--host 127.0.0.1 \
--port 8080 \
--prio 2 \
--reasoning off \
--cache-type-k q4_0 \
--cache-type-v q4_0
VRAM Budget (Production Config)
| Component |
VRAM Usage |
| Model Weights (Gemma 4 E2B Q4_0) |
~1.50 GB |
| KV Cache (8192 ctx, q4_0 compressed) |
~0.35 GB |
| Compute Buffers (batch 512) |
~0.20 GB |
| Total |
~2.05 GB (51% of 4GB) |
Why Not Use the Remaining 2GB VRAM?
Generation speed is bottlenecked by memory bandwidth (~32 GB/s on GTX 950M), not VRAM capacity. To generate one token, the GPU must read the entire ~1.5GB model from VRAM. The theoretical maximum is ~21 tok/s. At 10.9 tok/s, the GPU is already operating at ~50% of its physical limit. Empty VRAM cannot be converted into faster generation.
Testing uncompressed f16 KV cache (2.65GB total VRAM) yielded 12.6 tok/s but caused a severe "lazy duplication" regression (model copy-pasted the same translation twice instead of providing distinct synonyms). The q4_0 compressed KV cache is the correct choice for translation quality.
Performance Baseline (20-Word Batch)
| Metric |
Gemma 4 E2B |
Qwen 3.5 4B |
| Time/batch (20 words) |
~5 min |
~9 min |
| Completion tok/s |
~10.9 |
~6.1 |
| Prompt tok/s |
~132 |
~78 |
| VRAM Usage |
2.06 GB |
3.95 GB |
| Lazy Duplication Bug |
No |
Yes (severe) |
5. The LLM Layer
5.1 Local Model Evaluation (Complete — 2026-07-18)
All 10 downloaded models were evaluated on the target hardware. Testing progressed from 4-word smoke tests to a 20-word "nightmare" torture suite covering extreme polysemy, false friends, abstract concepts, and legal/financial terminology.
Final Leaderboard
| Rank |
Model |
Size |
VRAM |
Speed (tok/s) |
Polysemy |
False Friends |
Verdict |
| 🥇 |
Gemma 4 E2B Q4_0 |
3.2G |
2.06 GB |
~13.1 |
Perfect |
Perfect |
Production model. Best quality/speed/VRAM ratio. |
| 🥈 |
Qwen 3.5 4B Q4_K_M |
2.6G |
~3.6 GB |
~6.3 |
Perfect |
Perfect |
Quality King, but 2x slower and maxes VRAM. |
| 🥉 |
Gemma 4 E4B Q4_K_XL |
4.0G |
3.23 GB |
~7.7 |
Perfect |
Perfect |
Incredible QAT compression. Same quality as E2B but slower. |
| 4 |
Qwen 3.5 2B Q4_K_M |
1.2G |
~1.6G |
~13.2 |
Failed (2 identical senses) |
Passed |
Good translations, lacks conceptual branching. |
| 5 |
Qwen 2.5 1.5B Q4_K_M |
1.1G |
~1.5G |
~18.5 |
Ignored instruction |
Failed |
Fast but easily confused. |
| 6 |
Ministral 3B 2512 Q4_K_M |
2.0G |
~2.7G |
~8.9 |
Messy translations |
Failed |
Fixed tokenizer, but outclassed. Markdown violations. |
| 7 |
Llama 3.2 3B Q4_K_M |
1.9G |
~2.6G |
~9.1 |
Good structure, bad IT/ES |
Failed (cognate trap) |
Dangerous for language learners. |
| 8 |
Llama 3.2 3B Q6_K_XL |
2.8G |
3.44G |
~6.7 |
Good structure, bad IT/ES |
Failed (cognate trap) |
Higher quant did NOT fix translation issues. |
| 9 |
Qwen 3.5 9B Q3_K_S |
4.1G |
Split CPU/GPU |
~3.3 |
Perfect |
Perfect |
Brilliant but 35-40 days for 100k words. |
| 10 |
Qwen 3.5 9B Q4_K_M |
5.3G |
Split CPU/GPU |
~3.3 |
Perfect |
Perfect |
Same quality as 9B Q3. Not worth the size. |
Key Findings
- The "Thinking" Trap: Both Qwen 3.5 and Gemma 4 have built-in Chain-of-Thought reasoning. Without
--reasoning off, they silently "think" in a hidden JSON field, consume all max_tokens, and crash with finish_reason: length. This flag is mandatory.
- The 2B vs 4B Quality Cliff: 2B models struggle with polysemy (e.g., cannot distinguish "bank" = financial vs river). 4B+ models act like professional lexicographers.
- Llama 3.2 is Unsafe for Language Learners: Consistently fell for false friend traps (e.g., translating "actual" = real to aktuell/attuale/actual/actuel = current).
- KV Cache Quantization Improves Translation Variety: Compressing the KV cache to
q4_0 introduces microscopic noise that prevents the "lazy duplication" bug (model copy-pasting the same synonym twice).
- POS Bleed is Universal: All models occasionally generate verb definitions for nouns (e.g., "run" = to move fast instead of a jogging session). Fix: explicit negative constraint in the system prompt.
5.2 Online API Options
Evaluated as fallbacks if local models fail quality or speed targets.
| Provider |
Model |
Input $/1M |
Output $/1M |
Free Tier |
Rate Limit |
Est. Cost (100k words) |
Est. Time |
| DeepSeek |
V4 Flash |
$0.14 |
$0.28 |
5M tokens |
None |
$0-15 |
~1-2 days |
| Gemini |
2.5 Flash-Lite |
$0.10 |
$0.40 |
1,500 req/day |
1M TPM |
$0 (free tier) |
~1.5 days (batched) |
| Qwen/Alibaba |
Qwen-Turbo |
$0.05 |
$0.20 |
Unknown |
600 RPM |
~$11 |
~1-2 days |
| Groq |
Llama 3.1 8B Instant |
$0.05 |
$0.08 |
Yes |
High |
~$7 |
~3-4 hours |
| OpenRouter free |
Various |
$0 |
$0 |
200 req/day |
20 RPM |
$0 |
~10 days |
5.3 Model Selection Criteria
Decision flow for 100,000 words:
Start
|
v
Gemma 4 E2B (local) — SELECTED
|
|-- Speed acceptable? (~7 days) -----> Use Gemma 4 E2B locally, $0
|
|-- Need faster? -------> Test Gemini 2.5 Flash-Lite (free)
|
|-- Quality good? --> Batch 50, free tier
| ~1.5 days, $0
|
|-- Quality meh? ---> Test Groq or DeepSeek paid
Quality gates:
- 100% JSON parse rate
- No hallucinated definitions on polysemous words
- Natural, contextually appropriate example sentences
- Sensible difficulty classification (CEFR mapping)
- Gender accuracy is no longer an LLM criterion (handled by Kaikki)
6. The Gender Problem & Kaikki Integration
The Problem (Resolved)
Grammatical gender was originally embedded in the LLM's translations object. Testing across all 10 models confirmed that small local models systematically hallucinate gender, defaulting to neuter for Romance languages (Italian, Spanish, French) which do not have a neuter grammatical gender.
The Solution: Decoupled Architecture
Decision (2026-07-18): Grammatical gender is no longer generated by the LLM. The pipeline now uses a two-stage approach:
| Stage |
Component |
Responsibility |
| 1 |
LLM (Gemma 4 E2B) |
Generates translation strings only (e.g., ["Haus", "Gebäude"]) |
| 2 |
Kaikki Lookup |
Deterministically resolves grammatical gender from Wiktionary dumps |
Benefits
- 100% deterministic gender accuracy — no hallucination possible
- Faster LLM generation — ~15-20% fewer output tokens per word
- Simpler JSON schema — translations are string arrays, not object arrays
- Model-agnostic — works with any LLM regardless of multilingual training quality
Kaikki Data
| Language |
Kaikki File |
Coverage |
| German |
kaikki.org-dictionary-German.jsonl |
High |
| Italian |
kaikki.org-dictionary-Italian.jsonl |
High |
| Spanish |
kaikki.org-dictionary-Spanish.jsonl |
High |
| French |
kaikki.org-dictionary-French.jsonl |
High |
Lookup Logic
Match on word (the translated string) -> extract gender field -> map to "masculine" | "feminine" | "neuter" | null.
7. Batching Strategy
Why Batching is Necessary
At 1 word per call, 100,000 words = 100,000 LLM requests. Each call re-processes the ~300-token system prompt. Batching amortizes this cost.
Optimal Batch Size (Local)
20 words is the validated sweet spot for the GTX 950M with Gemma 4 E2B.
| Batch Size |
VRAM |
Speed |
Quality |
Verdict |
| 1-4 |
~2.1 GB |
~13 tok/s |
Perfect |
Safe but slow (amortization waste) |
| 20 |
~2.6 GB |
~10.9 tok/s |
Perfect |
Sweet spot |
| 30-40 |
~3.0 GB (est.) |
~10 tok/s (est.) |
Likely good |
Worth testing |
| 50+ |
~3.5 GB+ (est.) |
Unknown |
Risk of JSON degradation |
Not recommended for 2B models |
| 100 |
OOM risk |
N/A |
Attention degradation |
Small models lose JSON structure past ~6000 tokens |
Retry & Split Strategy
If a batch fails (bad JSON, missing key, etc.):
Batch of 20 fails (3 retries exhausted)
|
v
Split into 2 batches of 10
|
v
If a batch of 10 fails (3 retries), split into 2 batches of 5
|
v
If a batch of 5 fails, split into batches of 1
|
v
If a single word fails (3 retries), log and skip
8. Hardware Constraints
Current Machine
| Component |
Spec |
| OS |
Debian GNU/Linux 13 (trixie) x86_64 |
| CPU |
Intel Core i7-6500U (2 physical cores, 4 threads via HT) @ 3.10 GHz |
| GPU |
NVIDIA GeForce GTX 950M (Maxwell, 2015) |
| GPU VRAM |
4GB (4037 MiB reported by CUDA) |
| GPU Bandwidth |
~32 GB/s (DDR3) |
| RAM |
8GB (~3.67GB usable at idle) |
| Swap |
5.62 GB |
| Disk |
102GB ext4 (~74GB used) |
What Fits in 4GB VRAM (Empirically Verified)
| Model |
File Size |
Total VRAM |
Fits? |
Notes |
| Qwen 2.5 1.5B Q4_K_M |
1.1G |
~1.5 GB |
✅ Easy |
|
| Qwen 3.5 2B Q4_K_M |
1.2G |
~1.6 GB |
✅ Easy |
|
| Ministral 3B 2512 Q4_K_M |
2.0G |
~2.7 GB |
✅ Yes |
|
| Llama 3.2 3B Q4_K_M |
1.9G |
~2.6 GB |
✅ Yes |
|
| Llama 3.2 3B Q6_K_XL |
2.8G |
3.44 GB |
✅ Tight |
|
| Gemma 4 E2B Q4_0 |
3.2G |
2.06 GB |
✅ Yes |
QAT compression. Production model. |
| Gemma 4 E4B Q4_K_XL |
4.0G |
3.23 GB |
✅ Yes |
QAT compression is incredible. |
| Qwen 3.5 4B Q4_K_M |
2.6G |
3.95 GB |
⚠️ Barely |
50MB headroom with 8192 ctx. |
| Qwen 3.5 9B Q3_K_S |
4.1G |
Split |
⚠️ Partial |
-ngl 28 max. Rest on CPU/Swap. |
| Qwen 3.5 9B Q4_K_M |
5.3G |
Split |
⚠️ Partial |
-ngl 20 max. Heavy swap usage. |
GPU Rental Alternatives
If local hardware becomes the bottleneck:
| Provider |
GPU |
VRAM |
Price/Hour |
Time for 100k Words |
Total Cost |
| Vast.ai |
RTX 4090 |
24GB |
~$0.30-0.60 |
~6-8 hours |
~$2-5 |
| RunPod |
RTX 4090 |
24GB |
~$0.50-0.80 |
~6-8 hours |
~$4-6 |
9. Testing & Quality Assurance
20-Word Torture Suite (Completed 2026-07-18)
Tested on Gemma 4 E2B and Qwen 3.5 4B with 20 challenging nouns:
| Category |
Words |
The Trap |
| Extreme Polysemy |
match, date, right, set, well |
Does it translate "match" as fire, sports, or dating? |
| False Friends |
sense, fabric, sympathy, eventuality, billion |
"Fabric" = material (tissu) not factory (fabrique) |
| Abstract/Cultural |
serendipity, accountability |
Concepts lacking 1:1 dictionary equivalents |
| Action-Nouns |
run, drive, play |
"Run" as jog vs tear vs campaign |
| Legal/Financial |
mortgage, lease, court, board, draft |
Requires specific legal vocabulary (Hypothek/mutuo/hipoteca) |
Results Summary
| Criterion |
Gemma 4 E2B |
Qwen 3.5 4B |
| JSON Reliability |
10/10 (raw JSON) |
10/10 (raw JSON, but added sense_index hallucination) |
| Polysemy (bank, match) |
10/10 (Ufer/riva/orilla/rive) |
10/10 |
| False Friends (fabric) |
10/10 (Stoff/tessuto/tela/tissu) |
10/10 |
| Legal Nuance (mortgage) |
10/10 (Hypothek/mutuo/hipoteca/hypothèque) |
10/10 |
| Lazy Duplication Bug |
None |
Severe (copy-pasted same word 20+ times) |
| POS Bleed |
Minor (verb defs for run/match) |
Minor (verb defs for run/match) |
| Attention Degradation (word 20 vs 1) |
None |
None |
Verification
verify-enriched-file.ts checks:
- Required top-level fields:
word, language, pos, senses
- Each sense:
sense (string), example (string), difficulty_level in {easy, medium, hard}
- Each translation: array of strings (gender appended later by Kaikki)
10. Interactive CLI
Batch Size Recommendations (Updated)
| Provider |
Recommended |
Rationale |
| local (GTX 950M) |
20 |
Validated sweet spot. 8192 ctx + q4_0 KV cache. |
| local (RTX 4090) |
50 |
Fast, more VRAM |
| gemini |
50 |
Free tier: 1,500 req/day |
| deepseek |
20 |
5M free tokens |
| groq |
50 |
Very fast |
11. Future Extensions & Roadmap
Near-Term (Next 2-4 Weeks)
| Item |
Status |
Notes |
| 10-model evaluation |
Complete |
Gemma 4 E2B selected |
| KV cache quantization |
Complete |
--cache-type-k/v q4_0 enables 8192 ctx on 4GB GPU |
| Gender decoupling |
Complete |
Kaikki lookup replaces LLM gender |
| 20-word torture suite |
Complete |
Validated on Gemma E2B and Qwen 4B |
| POS bleed fix |
Pending |
Add negative constraint to system prompt |
| Kaikki gender lookup implementation |
Pending |
Post-processing step after LLM enrichment |
| Online API testing |
Pending |
Gemini free tier, DeepSeek, Groq |
Medium-Term (1-3 Months)
| Item |
Notes |
| Multi-POS support |
Verbs, adjectives, adverbs need prompt variants |
| Multi-language source |
German -> French, Italian -> Spanish, etc. |
| Parallel wordlist processing |
Run english/nouns and english/verbs simultaneously |
| Incremental enrichment |
Only process new/changed words in a wordlist |
Long-Term (3-6 Months)
| Item |
Notes |
| GPU rental integration |
Script to spin up Vast.ai/RunPod, run pipeline, download results |
| Quality regression tests |
Run torture suite on every model change |
| Community contributions |
Open-source the pipeline for other language learners |
12. Decisions Log
| Date |
Decision |
Context |
Rationale |
| 2026-01-04 |
TanStack Router for frontend |
Previous project used React Router |
Simpler, type-safe routing |
| 2026-06-16 |
llama.cpp for local LLM |
Need local inference on old laptop |
GGUF format, OpenAI-compatible API |
| 2026-06-16 |
Q4_K_M quantization |
Balance size vs quality |
Community standard for 4-bit |
| 2026-06-16 |
-t 2 physical cores |
Default was 4 (HT threads) |
Hyperthreading hurts llama.cpp |
| 2026-06-17 |
Qwen2.5-1.5B as initial model |
Qwen3.5-4B too slow (47s/word) |
6x speedup, quality under evaluation |
| 2026-07-06 |
Adapter pattern for LLM providers |
Need to evaluate local vs cloud |
utils/llm-adapters/ with factory |
| 2026-07-06 |
Retry + split batching |
LLM JSON parse failures |
3 retries, then halve batch |
| 2026-07-06 |
Interactive CLI |
Editing config files is error-prone |
utils/cli.ts with native readline |
| 2026-07-18 |
Gemma 4 E2B as production model |
10-model evaluation completed |
2x faster than Qwen 4B, half VRAM, perfect translation quality |
| 2026-07-18 |
Gender decoupled from LLM |
All 10 models failed gender for Romance languages |
Kaikki Wiktionary lookup is deterministic and 100% accurate |
| 2026-07-18 |
--reasoning off is mandatory |
Qwen 3.5 and Gemma 4 "think" silently, consuming all tokens |
Without this flag, output crashes with finish_reason: length |
| 2026-07-18 |
KV cache quantization (q4_0) |
8192 context needed for 20-word batches |
Cuts KV VRAM by 75%, enables large batches on 4GB GPU, improves translation variety |
| 2026-07-18 |
20-word batch size for local |
Tested 4, 20 words |
Sweet spot: no attention degradation, no JSON breakage, 10.9 tok/s |
| 2026-07-18 |
Llama 3.2 discarded |
Failed false friend tests |
Translates "actual" (real) to cognates (aktuell/attuale) = "current". Dangerous for learners. |
| 2026-07-18 |
-c 8192 replaces -c 2048 |
20-word batches need more context |
Combined with q4_0 KV cache, fits in 2.65GB VRAM |
13. Known Issues & Dev Notes
Data Pipeline
| Issue |
Details |
Severity |
| POS bleed (all models) |
Models occasionally generate verb definitions for nouns (e.g., "run" = to move fast) |
Medium — fix with negative constraint in prompt |
| Lazy duplication (Qwen 4B) |
Qwen 3.5 4B copy-pastes the same translation twice to fill arrays |
Medium — use Gemma 4 E2B instead |
| Schema hallucination (Qwen 4B) |
Adds sense_index field not in schema |
Low — ignorable |
| Pre-scanning wordlists |
Entire file read into memory before processing |
Medium — streaming refactor planned |
| Single POS tested |
Only nouns validated. Verbs/adjectives need prompt changes |
Known limitation |
Hardware
| Issue |
Details |
| GTX 950M VRAM ceiling |
4GB hard limit. Models >3.5GB need KV cache quantization or CPU offloading. |
| Maxwell GPU aging |
No Flash Attention, bandwidth-starved (~32 GB/s). Theoretical max ~21 tok/s. |
| Laptop thermals |
Cannot run 24/7 for weeks unattended. Monitor temps. |
14. How to Run
Prerequisites
- Node.js + npm
tsx installed globally: npm install -g tsx
- llama.cpp built from source (for local mode)
- GGUF model downloaded to
~/Downloads/llama.cpp/models/
- API keys set as environment variables (for cloud mode)
Start the LLM Server (Local Mode)
cd ~/Downloads/llama.cpp
./build/bin/llama-server \
-m models/gemma-4-E2B_q4_0-it.gguf \
-ngl 999 \
-c 8192 \
-b 512 \
-ub 512 \
-t 2 \
--threads-batch 2 \
--host 127.0.0.1 \
--port 8080 \
--prio 2 \
--reasoning off \
--cache-type-k q4_0 \
--cache-type-v q4_0
Run the Pipeline
cd /path/to/data-pipeline
npx tsx pipeline.ts
Follow the interactive prompts to select provider, model, and batch size.
15. Roadmap
Phase 1: Batching (Complete)
| Task |
Status |
| Implement configurable batch size |
Complete |
| Implement retry + split logic |
Complete |
| Honest timing metrics |
Complete |
| Validate LLM responses |
Complete |
| Verify batching quality (20-word suite) |
Complete |
Phase 2: Interactive CLI (Complete)
| Task |
Status |
| Design prompt flow |
Complete |
| Implement CLI module |
Complete |
| Save/load config |
Complete |
| Wire into pipeline.ts |
Complete |
Phase 3: Model Selection (Complete)
| Task |
Status |
| 10-model evaluation |
Complete |
| 20-word torture suite |
Complete |
| Select production model |
Complete (Gemma 4 E2B) |
| Test online APIs |
Pending |
Phase 4: Scale
| Task |
Status |
Notes |
| Implement Kaikki gender lookup |
Pending |
Post-processing step |
| Fix POS bleed in prompt |
Pending |
Add negative constraint |
| Run 100k word pipeline |
Pending |
~7 days local (Gemma E2B, 20-word batches) |
| Spot-check output quality |
Pending |
Random sample of 100 entries |
Phase 5: Extend
| Task |
Status |
| Multi-POS support |
Pending |
| Multi-language source |
Pending |
| Parallel wordlist processing |
Pending |
| GPU rental integration |
Pending |
| Quality regression tests |
Pending |