From e534b98bc5d77a32a32df37f4eee5592b698cfc0 Mon Sep 17 00:00:00 2001 From: lila Date: Sun, 2 Aug 2026 01:01:33 +0200 Subject: [PATCH] updating documentation, prettier format --- CLAUDE.md | 60 ++ README.md | 56 +- apps/api/src/app.ts | 2 +- apps/api/src/errors/AppError.ts | 2 +- data-pipeline/pipeline.ts | 45 + docker-compose.yml | 6 +- documentation/ARCHITECTURE.md | 34 +- documentation/DATA_PIPELINE.md | 834 ++---------------- .../ai-context/00-project-overview.md | 16 +- documentation/ai-context/01-architecture.md | 2 +- documentation/ai-context/02-data-model.md | 266 +++--- documentation/ai-context/03-api-contract.md | 2 +- documentation/ai-context/05-data-pipeline.md | 226 +++-- .../archive/data-pipeline-local-llm.md | 790 +++++++++++++++++ .../llm-setup-local.md} | 0 .../model-strategy-cefr-voters.md} | 0 16 files changed, 1271 insertions(+), 1070 deletions(-) create mode 100644 CLAUDE.md create mode 100644 documentation/archive/data-pipeline-local-llm.md rename documentation/{LLM_SETUP.md => archive/llm-setup-local.md} (100%) rename documentation/{MODEL_STRATEGY.md => archive/model-strategy-cefr-voters.md} (100%) diff --git a/CLAUDE.md b/CLAUDE.md new file mode 100644 index 0000000..accb592 --- /dev/null +++ b/CLAUDE.md @@ -0,0 +1,60 @@ +# CLAUDE.md + +This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository. + +## Commands + +```bash +pnpm install # pnpm workspaces; pnpm only +docker compose up -d # postgres (5432), pipeline postgres (5433), valkey (6379) + +pnpm dev # api (:3000) + web (:5173) concurrently +pnpm build # shared → db → api, in dependency order +pnpm typecheck # tsc --build --noEmit across all projects +pnpm lint # eslint . +pnpm format # prettier --write . + +pnpm test # vitest watch, all workspace projects +pnpm test:run # single run (also what the pre-commit hook runs) +pnpm vitest run apps/api/src/services/gameService.test.ts # single file +pnpm vitest run -t "returns 400" # single test by name + +pnpm --filter @lila/db generate # drizzle-kit generate (after editing schema.ts) +pnpm --filter @lila/db migrate # apply migrations (uses DATABASE_URL_LOCAL) +pnpm --filter @lila/pipeline pipeline:run +``` + +`packages/shared` and `packages/db` are consumed as built `dist/` output — rebuild them (`pnpm --filter @lila/shared build`) after changing them, or downstream packages will typecheck against stale types. `pnpm --filter @lila/api dev` does this automatically. + +Husky pre-commit runs `lint-staged` (prettier + eslint --fix) then the full test suite. + +## Architecture + +Monorepo: `apps/api` (Express + `ws`), `apps/web` (React 19, Vite, TanStack Router, Tailwind 4), `packages/shared` (Zod contract), `packages/db` (Drizzle), `data-pipeline` (vocabulary ETL). Docs live in `documentation/` — `ARCHITECTURE.md`, `DECISIONS.md`, `DATA_PIPELINE.md`, `STATUS.md`, `BACKLOG.md`. + +**Strict layering in `apps/api`:** router → controller → service → model (`packages/db`) → PostgreSQL. Each layer talks only to the one below it. Controllers do HTTP only (Zod `safeParse`, then `next(error)`); services hold business logic and never read `req`; models hold queries and no domain semantics. `apps/api` must never import `drizzle-orm` — all queries live in `packages/db/src/models/`. + +**Errors:** `AppError` subclasses carry their own `statusCode`; a single `errorHandler` middleware maps them. `ValidationError` is thrown in controllers, `NotFoundError` in services. + +**`packages/shared` is the single source of truth** for every shape crossing the API boundary (`schemas/game.ts`, `schemas/lobby.ts`, `schemas/auth.ts`, `constants.ts`). Changing a schema breaks compilation in both api and web simultaneously — that's intended. Supported languages (`en/it/de/es/fr`) and POS values are constants here and are CHECK-constrained in the DB schema. + +**WebSocket:** upgrades on `/ws` on the same HTTP server. Better Auth cookie is validated at upgrade (`ws/auth.ts`), then `ws/router.ts` dispatches on the `type` field of a Zod discriminated union to `ws/handlers/`. Lobby membership is persisted in PostgreSQL; live game/room state lives in `InMemoryGameSessionStore` / `InMemoryLobbyGameStore` behind interfaces, so it is lost on restart (Valkey swap is planned — keep the interface boundary intact). + +**Answer evaluation is server-side.** The correct answer is never included in what is sent to the client. + +**Data model:** two vocabulary schemas coexist. The app reads `vocabulary_entries` + `entry_translations` (one row per word sense); the new pipeline targets `words` → `senses` → `translations`, which is migrated but empty. `packages/db/src/models/termModel.ts` is still on the live pair. Adding a language is rows, not schema changes. Auth tables (`user`, `session`, `account`, `verification`) are owned by Better Auth. + +## Data pipeline + +`data-pipeline/` is mid-rewrite on this branch (`refactor/gemini-only-pipeline`): the old local-LLM/adapter architecture is gone, and `pipeline.ts` currently holds the staged design as pseudocode comments — there is no executable pipeline yet. Current pieces: `source-data/{lang}/{pos}` wordlists, `prompt` (the Gemini system prompt, a plain UTF-8 file), `db/schema.sql` (SQLite staging: `words` → `senses` → `translations`), and a dedicated pipeline PostgreSQL on port 5433. + +Read `documentation/DATA_PIPELINE.md` for orientation and current phase status, `documentation/pipeline/design-doc.md` for the schema/difficulty model/Gemini JSON contract, and `documentation/pipeline/roadmap.md` for the phase plan. Everything in `documentation/archive/` is superseded — `data-pipeline-local-llm.md`, `llm-setup-local.md`, and `model-strategy-cefr-voters.md` describe the removed local-LLM / CEFR-voter pipeline and are historical only. + +The sense-based `words`/`senses`/`translations` schema is what the pipeline targets; the app is not migrated to it yet and still queries `vocabulary_entries`/`entry_translations`. + +## Conventions + +- TypeScript is maximally strict (`tsconfig.base.json`): `exactOptionalPropertyTypes`, `noUncheckedIndexedAccess`, `noPropertyAccessFromIndexSignature`, `verbatimModuleSyntax`, `erasableSyntaxOnly`. Index access yields `T | undefined` — handle it rather than loosening the config. Env vars are read as `process.env["KEY"]`. +- The base tsconfig deliberately omits `lib`/`module`/`moduleResolution`; each package sets its own (api `NodeNext`, web `ESNext`/`bundler`). +- Tests are co-located (`gameService.test.ts` beside `gameService.ts`), use vitest globals, and mock `@lila/db` with `vi.mock` — no test database. Endpoint tests use supertest against the `createApp()` factory without starting a server. +- All env config lives in the single root `.env` (see `.env.example`); `packages/db` and the pipeline both resolve it from the repo root. diff --git a/README.md b/README.md index 043d717..1bb06d6 100644 --- a/README.md +++ b/README.md @@ -26,9 +26,8 @@ docker compose up -d pnpm --filter @lila/shared build pnpm --filter @lila/db build -# 5. Run migrations and seed data +# 5. Run migrations pnpm --filter @lila/db migrate -pnpm --filter @lila/db seed # 6. Start dev servers pnpm dev @@ -36,39 +35,40 @@ pnpm dev API: `http://localhost:3000` · Web: `http://localhost:5173` -See [DEPLOYMENT.md](DEPLOYMENT.md) for production infrastructure details. +See [DEPLOYMENT.md](documentation/DEPLOYMENT.md) for production infrastructure details. --- ## Documentation Index -| Document | What you'll find there | -| -------------------------------------------- | ----------------------------------------------------------------------- | -| [STATUS.md](STATUS.md) | Current state — what's working, what's blocked, what we're building now | -| [BACKLOG.md](BACKLOG.md) | Prioritized task list: now / next / later / changelog | -| [ARCHITECTURE.md](ARCHITECTURE.md) | Monorepo structure, layered architecture, data flow | -| [DECISIONS.md](DECISIONS.md) | Why we chose X over Y — tool choices, schema design, trade-offs | -| [DATA_PIPELINE.md](DATA_PIPELINE.md) | Kaikki → CEFR enrichment → production PostgreSQL | -| [MODEL_STRATEGY.md](MODEL_STRATEGY.md) | LLM voter architecture for sense-disambiguated CEFR assignment | -| [LLM_SETUP.md](LLM_SETUP.md) | Local and cloud LLM provider configuration | -| [DEPLOYMENT.md](DEPLOYMENT.md) | Hetzner VPS, Caddy, Docker Compose, CI/CD, backups | -| [design/GAME_MODES.md](design/GAME_MODES.md) | Planned multiplayer and singleplayer game modes | +| Document | What you'll find there | +| -------------------------------------------------------------- | ----------------------------------------------------------------------- | +| [STATUS.md](documentation/STATUS.md) | Current state — what's working, what's blocked, what we're building now | +| [BACKLOG.md](documentation/BACKLOG.md) | Prioritized task list: now / next / later / changelog | +| [ARCHITECTURE.md](documentation/ARCHITECTURE.md) | Monorepo structure, layered architecture, data flow | +| [DECISIONS.md](documentation/DECISIONS.md) | Why we chose X over Y — tool choices, schema design, trade-offs | +| [DATA_PIPELINE.md](documentation/DATA_PIPELINE.md) | Gemini → SQLite staging → PostgreSQL: flow, current state, phase status | +| [pipeline/design-doc.md](documentation/pipeline/design-doc.md) | Schema design, difficulty model, query patterns, Gemini JSON contract | +| [pipeline/roadmap.md](documentation/pipeline/roadmap.md) | Phase plan for the pipeline and schema migration | +| [DEPLOYMENT.md](documentation/DEPLOYMENT.md) | Hetzner VPS, Caddy, Docker Compose, CI/CD, backups | +| [design/GAME_MODES.md](documentation/design/GAME_MODES.md) | Planned multiplayer and singleplayer game modes | +| [archive/](documentation/archive/) | Superseded docs — the removed local-LLM pipeline, CEFR voter strategy | --- ## Stack -| Layer | Technology | -| ---------- | ------------------------------------------------------------- | -| Monorepo | pnpm workspaces | -| Frontend | React 18, Vite, TanStack Router, TanStack Query, Tailwind CSS | -| Backend | Node.js, Express, TypeScript, WebSockets (`ws`) | -| Database | PostgreSQL + Drizzle ORM | -| Auth | Better Auth (Google + GitHub) | -| Validation | Zod (shared between frontend and backend) | -| Testing | Vitest, supertest | -| Deployment | Docker Compose, Caddy, Hetzner VPS | -| CI/CD | Forgejo Actions | +| Layer | Technology | +| ---------- | ----------------------------------------------- | +| Monorepo | pnpm workspaces | +| Frontend | React 19, Vite, TanStack Router, Tailwind CSS | +| Backend | Node.js, Express, TypeScript, WebSockets (`ws`) | +| Database | PostgreSQL + Drizzle ORM | +| Auth | Better Auth (Google + GitHub) | +| Validation | Zod (shared between frontend and backend) | +| Testing | Vitest, supertest | +| Deployment | Docker Compose, Caddy, Hetzner VPS | +| CI/CD | Forgejo Actions | --- @@ -78,10 +78,10 @@ See [DEPLOYMENT.md](DEPLOYMENT.md) for production infrastructure details. - ✅ Multiplayer lobby + real-time game (2–4 players, simultaneous answers, 15s timer) - ✅ Auth (Google + GitHub) - ✅ Live deployment with CI/CD -- 🔄 Migrating vocabulary data from OpenWordNet to **Kaikki** (sense-disambiguated translations) +- 🔄 Rewriting the vocabulary data pipeline around **Gemini** + a sense-based schema - 🔄 Phase 7 hardening (rate limiting, error boundaries, monitoring) -See [STATUS.md](STATUS.md) for the full picture. +See [STATUS.md](documentation/STATUS.md) for the full picture. --- @@ -95,7 +95,7 @@ lila/ ├── packages/ │ ├── shared/ — Zod schemas + constants (API/web contract) │ └── db/ — Drizzle schema, migrations, models, seeding -├── data-pipeline/ — Kaikki extraction → enrichment → PostgreSQL sync +├── data-pipeline/ — Gemini generation → SQLite staging → PostgreSQL ├── documentation/ — Project docs (this directory) └── Caddyfile, docker-compose.yml, etc. ``` diff --git a/apps/api/src/app.ts b/apps/api/src/app.ts index 47c51e6..ad9f509 100644 --- a/apps/api/src/app.ts +++ b/apps/api/src/app.ts @@ -27,7 +27,7 @@ export function createApp() { const store = new InMemoryGameSessionStore(); app.use("/api/v1", createApiRouter(store)); - + app.use(errorHandler); return app; diff --git a/apps/api/src/errors/AppError.ts b/apps/api/src/errors/AppError.ts index 7805f3e..bc5a92a 100644 --- a/apps/api/src/errors/AppError.ts +++ b/apps/api/src/errors/AppError.ts @@ -30,4 +30,4 @@ export class UnprocessableEntityError extends AppError { constructor(message: string) { super(message, 422); } -} \ No newline at end of file +} diff --git a/data-pipeline/pipeline.ts b/data-pipeline/pipeline.ts index e69de29..4379edc 100644 --- a/data-pipeline/pipeline.ts +++ b/data-pipeline/pipeline.ts @@ -0,0 +1,45 @@ +// pipeline.ts pseudo code + +/* +step 1: discover source lists + +this will give us an array of objects with this schema +sourceLanguage: 'en' | 'de' | 'es' | 'fr' | 'it' +pos: 'noun' | 'verb' | 'adjective' | 'adverb' +words: string[]; +filePath: string; + +the terminal output should be something like: + +found 5 source lists: + +de: noun +en: noun + +and so on + +later on, it will also contain it: noun, verb, adjective etc + */ + +/* + +step 2: validating source lists + +a small script that trims whitespaces, removes duplicated words etc + +terminal output: summary of how many words per pos per language were found + +*/ + +/* + +step 3: writing to database? + +my thought: ill restart the pipeline several times during testing, and when adding more wordlists with other pos or extending the exisiting noun lists +eventually the lists will contain tens or hundreds of thousands of words +how do we prevent reading and processing the same words multiple times? +if we read and validate+normalize the wordlists and write them to the database, we could then read from the database fill the missing translations etc +and not read the same words from the same text files multiple times? + +if we do this, we have to adjust the database schema because there are several notNull() rows inside +*/ diff --git a/docker-compose.yml b/docker-compose.yml index 3dccd4f..f8de7a7 100644 --- a/docker-compose.yml +++ b/docker-compose.yml @@ -33,7 +33,11 @@ services: - pipeline-db:/var/lib/postgresql/data restart: unless-stopped healthcheck: - test: ["CMD-SHELL", "pg_isready -U ${PIPELINE_POSTGRES_USER} -d ${PIPELINE_POSTGRES_DB}"] + test: + [ + "CMD-SHELL", + "pg_isready -U ${PIPELINE_POSTGRES_USER} -d ${PIPELINE_POSTGRES_DB}", + ] interval: 5s timeout: 5s retries: 5 diff --git a/documentation/ARCHITECTURE.md b/documentation/ARCHITECTURE.md index 74be351..fd28af3 100644 --- a/documentation/ARCHITECTURE.md +++ b/documentation/ARCHITECTURE.md @@ -14,7 +14,7 @@ lila/ ├── packages/ │ ├── shared/ — Zod schemas + constants (API/web contract) │ └── db/ — Drizzle schema, migrations, models, seeding -├── data-pipeline/ — Kaikki extraction → enrichment → PostgreSQL sync +├── data-pipeline/ — Gemini generation → SQLite staging → PostgreSQL ├── documentation/ — Project docs ├── Caddyfile — Reverse proxy routing ├── docker-compose.yml — Local dev stack @@ -98,17 +98,21 @@ In-memory stores (lobby game state, game session state) ## Database Schema (Core) -**Concept:** Words are language-neutral concepts (`terms`) with per-language `translations`. Adding a new language requires no schema changes — only new rows. +**Concept:** Vocabulary is stored per word sense, with translations attached to a sense rather than to a bare headword. Adding a new language requires no schema changes — only new rows. ### Core Tables -| Table | Purpose | -| -------------- | -------------------------------------------------------------------------------- | -| `terms` | Language-neutral concept: `id`, `pos` (noun/verb/adj/adv), `source`, `source_id` | -| `translations` | Per-language word: `term_id` (FK), `language_code`, `text`, `cefr_level` (A1–C2) | -| `term_glosses` | Per-language definition: `term_id` (FK), `language_code`, `text` | -| `decks` | Curated wordlists: `source_language`, `validated_languages`, frequency tier | -| `deck_terms` | Junction: which terms belong to which deck | +The database currently holds **two** vocabulary schemas — the app reads the first, the new pipeline writes the second. + +| Table | Purpose | +| -------------------- | -------------------------------------------------------------------------------------------- | +| `vocabulary_entries` | **Live.** One row per word sense: `headword`, `language_code`, `pos`, `sense_index`, `gloss` | +| `entry_translations` | **Live.** Per-entry translation: `entry_id` (FK), `target_language_code`, `translation` | +| `words` | **Target.** `headword`, `language_code`, `pos` | +| `senses` | **Target.** `word_id` (FK), `sense_index`, `difficulty`, `definitions[]`, `examples[]` | +| `translations` | **Target.** `sense_id` (FK), `target_language_code`, `translation`, `gender`, `difficulty` | + +The target tables are migrated but empty. `packages/db/src/models/termModel.ts` still queries the live pair; roadmap Phase 5 rewrites it. Full column-level detail: `ai-context/02-data-model.md`. ### Auth Tables (managed by Better Auth) @@ -123,8 +127,9 @@ In-memory stores (lobby game state, game session state) - `language_code` is CHECK-constrained against `SUPPORTED_LANGUAGE_CODES` (`en`, `it`, `de`, `es`, `fr`) - `pos` is CHECK-constrained against `SUPPORTED_POS` (`noun`, `verb`, `adjective`, `adverb`) -- `cefr_level` is nullable `varchar(2)` with CHECK `A1`–`C2` -- `translations` has UNIQUE `(term_id, language_code, text)` — allows synonyms, prevents exact duplicates +- `difficulty` is CHECK-constrained against `DIFFICULTY_LEVELS` (`easy`, `medium`, `hard`) +- `translations.gender` is nullable with CHECK against `NOUN_GENDERS` +- `translations` has UNIQUE `(sense_id, target_language_code, translation)` — allows synonyms, prevents exact duplicates --- @@ -216,14 +221,15 @@ The same pattern applies to `LobbyGameStore` (lobby state). | Why `ws` over Socket.io | `DECISIONS.md` → WebSocket | | Why server-side answer evaluation | `DECISIONS.md` → Architecture | | Why Better Auth over Keycloak | `DECISIONS.md` → Auth | -| Why terms/translations schema | `DECISIONS.md` → Data Model | +| Why the sense-based schema | `pipeline/design-doc.md` → §3 | | Why Caddy over Nginx/Traefik | `DECISIONS.md` → Deployment | --- ## Further Reading -- [DATA_PIPELINE.md](DATA_PIPELINE.md) — How vocabulary data gets from Kaikki into PostgreSQL +- [DATA_PIPELINE.md](DATA_PIPELINE.md) — How vocabulary data is generated and gets into PostgreSQL +- [pipeline/design-doc.md](pipeline/design-doc.md) — Schema design, difficulty model, Gemini output contract +- [pipeline/roadmap.md](pipeline/roadmap.md) — Phase plan for the pipeline and schema migration - [DEPLOYMENT.md](DEPLOYMENT.md) — Production infrastructure and ops -- [MODEL_STRATEGY.md](MODEL_STRATEGY.md) — LLM voter architecture for CEFR assignment - [design/GAME_MODES.md](design/GAME_MODES.md) — Planned multiplayer modes diff --git a/documentation/DATA_PIPELINE.md b/documentation/DATA_PIPELINE.md index b7a1bf6..80fde25 100644 --- a/documentation/DATA_PIPELINE.md +++ b/documentation/DATA_PIPELINE.md @@ -1,790 +1,144 @@ -# Lila Data Pipeline — Technical Documentation +# Lila Data Pipeline -Multilingual dictionary enrichment pipeline for the Lila vocabulary trainer. -Last updated: 2026-07-18 +> How vocabulary data is generated and gets into PostgreSQL. +> Last updated: 2026-08-01 · Branch: `refactor/gemini-only-pipeline` -## Table of Contents +**Authoritative detail lives in two companion docs:** -- [1. Executive Summary](#1-executive-summary) -- [2. Problem & Context](#2-problem--context) -- [3. Architecture Overview](#3-architecture-overview) -- [4. Current Implementation](#4-current-implementation) -- [5. The LLM Layer](#5-the-llm-layer) - - 5.1 [Local Model Evaluation](#51-local-model-evaluation) - - 5.2 [Online API Options](#52-online-api-options) - - 5.3 [Model Selection Criteria](#53-model-selection-criteria) -- [6. The Gender Problem & Kaikki Integration](#6-the-gender-problem--kaikki-integration) -- [7. Batching Strategy](#7-batching-strategy) -- [8. Hardware Constraints](#8-hardware-constraints) -- [9. Testing & Quality Assurance](#9-testing--quality-assurance) -- [10. Interactive CLI](#10-interactive-cli) -- [11. Future Extensions & Roadmap](#11-future-extensions--roadmap) -- [12. Decisions Log](#12-decisions-log) -- [13. Known Issues & Dev Notes](#13-known-issues--dev-notes) -- [14. How to Run](#14-how-to-run) -- [15. Roadmap](#15-roadmap) +| Doc | What's in it | +| ------------------------------------------------ | ------------------------------------------------------------------------------ | +| [pipeline/design-doc.md](pipeline/design-doc.md) | Schema design, difficulty model, query patterns, Gemini JSON contract, indexes | +| [pipeline/roadmap.md](pipeline/roadmap.md) | Phase-by-phase plan with task checklists and acceptance criteria | -## Quick Reference +This file is the orientation layer: what the pipeline is, what exists on disk today, and what is not built yet. -| What | Where | -| ------------------- | ------------------------------------------ | -| Entry point | `pipeline.ts` | -| Interactive CLI | `utils/cli.ts` | -| LLM config schema | `config/llm.ts` | -| System prompt | `config/prompt.ts` — `buildSystemPrompt()` | -| Batch config schema | `config/batch.ts` | -| Shared constants | `config/constants.ts` | -| Output schema | `utils/merge-enriched-data.ts` | -| LLM adapters | `utils/llm-adapters/` | -| **Current model** | **`gemma-4-E2B_q4_0-it.gguf`** | -| Target scale | 100,000+ words | +The previous local-LLM pipeline (llama.cpp, adapter pattern, 10-model evaluation, CEFR voter ensemble, Kaikki gender lookup) has been removed from the codebase. Its documentation is preserved under `archive/` and describes `utils/` and `config/` modules that **no longer exist**: + +- [archive/data-pipeline-local-llm.md](archive/data-pipeline-local-llm.md) — the old pipeline stages and file layout +- [archive/llm-setup-local.md](archive/llm-setup-local.md) — llama.cpp / cloud provider configuration +- [archive/model-strategy-cefr-voters.md](archive/model-strategy-cefr-voters.md) — the multi-model voter architecture for sense-disambiguated CEFR assignment --- -## 1. Executive Summary +## What changed, and why -The Lila Data Pipeline is a TypeScript-based batch processing system that enriches raw word lists into structured multilingual dictionary entries for the Lila vocabulary trainer. It takes a source wordlist (e.g., English nouns) and, for each word, generates: +The old pipeline ran small local models and needed a deterministic Kaikki Wiktionary lookup to patch grammatical gender, because local models hallucinated it. The rewrite drops local inference entirely in favour of the Gemini API: one provider, no adapter layer, gender produced directly by the model and enforced by validation instead of by a second data source. -- One or more senses (definitions) -- A natural example sentence per sense -- A CEFR-based difficulty level (`easy` / `medium` / `hard`) -- Translations into all target languages except the source (as raw strings) +The data model changed with it. The live `vocabulary_entries` / `entry_translations` tables (one row per word sense, populated from Kaikki) are replaced by `words` → `senses` → `translations`, where translations hang off a **sense**, not off a flat entry. That is the whole point of the rewrite: a quiz question can now be tied to one specific meaning of a word. -The pipeline is designed to scale to 100,000+ words across multiple languages and parts of speech. It supports both local inference (`llama.cpp`) and cloud providers via a pluggable adapter pattern. +--- -### Key Design Principles +## Flow -| Principle | Rationale | -| ------------------------------- | ------------------------------------------------------------------------------------------------------ | -| **Quality first** | Definitions, examples, and translations must be accurate. Speed and cost are secondary. | -| **Local-first, cloud-fallback** | Local LLMs are the default for cost control and data privacy. | -| **Deterministic Grammar** | Grammatical gender is decoupled from the LLM and resolved via Kaikki Wiktionary dumps. | -| **Resumable & idempotent** | Each word writes to its own JSON file. The pipeline skips already-processed words on restart. | -| **Configurable batching** | Batch size is selected interactively at runtime. The pipeline adapts without code changes. | -| **Honest metrics** | Local models report detailed prompt/completion timing. Cloud providers report total request time only. | - -### Resolved: Gender Accuracy - -Grammatical gender is **no longer generated by the LLM**. Comprehensive testing across 10 models (2026-07-18) confirmed that small local models systematically hallucinate or default to `neuter` for Romance languages. -**Decision:** The LLM only outputs translation strings. A deterministic post-processing step looks up the exact grammatical gender from Kaikki Wiktionary dumps. This guarantees 100% gender accuracy and allows us to use smaller, faster, and highly nuanced local models. - -### Current Status (2026-07-18) - -- **Core pipeline:** Complete (scanning, enrichment, merging, verification, writing). -- **Local LLM:** Gemma 4 E2B selected as production model after exhaustive 10-model evaluation. `llama.cpp` server optimized with KV-cache quantization. -- **Batching:** 20-word batches validated for local hardware. -- **Gender:** Decoupled from LLM; Kaikki lookup architecture confirmed. -- **Pending:** Online API evaluation (Gemini free tier, DeepSeek, Groq) as speed fallback. - -### One-Line Architecture - -```text -source wordlists -> Interactive CLI -> LLM adapter -> merge senses -> Kaikki Gender Lookup -> verify schema -> write .json +``` +source-data/{lang}/{pos} frequency wordlists, one word per line, UTF-8 + │ + ▼ +Gemini API batches of 20 words, one language at a time + │ + ▼ +validation per-entry; invalid entries → rejection log, not the DB + │ + ▼ +db/staging.db SQLite staging (words, senses, translations) + │ + ▼ +import script SQLite → PostgreSQL via Drizzle, transaction per batch + │ + ▼ +PostgreSQL (dev :5432, then prod) ``` -### Files at a Glance +Each language is processed independently so definitions and examples are written **in that language** — a German word gets a German definition, not a translation of an English one. Only the translations cross language boundaries. -| File | Purpose | -| ----------------------------------------- | ----------------------------------------------------------------------------------------------------- | -| `pipeline.ts` | Orchestrator. Runs CLI, scans sources, loops words, coordinates all stages | -| `utils/cli.ts` | Interactive CLI. Provider/model/batch selection, config persistence | -| `config/llm.ts` | LLM config schema (provider, url, model) | -| `config/prompt.ts` | `buildSystemPrompt()` — dynamic prompt with auto-target languages | -| `config/batch.ts` | Batch size and max retry count schema | -| `config/constants.ts` | Shared `LANG_MAP`, `POS_MAP`, `ALL_LANGUAGES` | -| `utils/enrich-word.ts` | Calls LLM via adapter, parses response, validates senses, builds `EnrichedSense[]`, retry/split logic | -| `utils/merge-enriched-data.ts` | Merges skeleton + enriched senses into final JSON | -| `utils/verify-enriched-file.ts` | Schema validation (required fields, types) | -| `utils/check-llm-server.ts` | Health check for local server; skipped for cloud providers | -| `utils/scanning-source-files.ts` | Discovers wordlists from `source-data/` directory | -| `utils/create-base-json.ts` | Writes skeleton `{word, language, pos}` files | -| `utils/write-json-file.ts` | Atomic `.tmp` → rename writes | -| `utils/check-if-json-exists.ts` | Resumability: checks if word already has enriched senses | -| `utils/create-line-reader.ts` | Streaming line reader for large wordlists | -| `utils/create-output-dirs.ts` | Creates `worddata/{language}/{pos}/` folders | -| `utils/delete-file.ts` | Cleanup helper for failed batches | -| `utils/get-word-file-path.ts` | Path construction helper | -| `utils/progress-tracker.ts` | `[current/total]` formatting for console output | -| `utils/pipeline-timer.ts` | Timing + token metrics; unified throughput for all providers | -| `utils/llm-adapters/factory.ts` | Creates the right adapter based on runtime config | -| `utils/llm-adapters/types.ts` | `LlmAdapter` interface | -| `utils/llm-adapters/openai-compatible.ts` | Local llama.cpp, OpenRouter, DeepSeek | -| `utils/llm-adapters/gemini.ts` | Google Gemini native API | - -### Scale Target - -| Metric | Target | -| --------------- | ------------------------------------------------------------ | -| Words | 100,000+ | -| Languages | English (source), German, Italian, Spanish, French (targets) | -| Parts of speech | Nouns, verbs, adjectives, adverbs | -| Output | One `.json` file per word, ~2-5KB each | +The app always reads from PostgreSQL. SQLite exists purely as a staging file so re-runs, prompt tweaks, and spot-checks never touch a real database. --- -## 2. Problem & Context +## What exists on disk today -### Why Build This? +| Path | State | +| ---------------------------------------- | ------------------------------------------------------------------------- | +| `data-pipeline/source-data/{lang}/{pos}` | ✅ Noun lists for `de`, `en`, `es`, `fr`, `it` | +| `data-pipeline/prompt` | ✅ The Gemini system prompt (plain UTF-8 text, not a module) | +| `data-pipeline/db/schema.sql` | ✅ SQLite staging schema | +| `data-pipeline/db/staging.db` | ✅ Created, tables present, **0 rows** — gitignored | +| `data-pipeline/pipeline.ts` | 🚧 Design pseudocode in comments. No executable pipeline code yet. | +| Validation module | ❌ Not written (rules specced in design-doc §6.4) | +| SQLite → PostgreSQL import script | ❌ Not written | +| `data-pipeline/kaikki-source-files/` | ⚠️ Leftover JSONL dumps from the old pipeline; nothing reads them anymore | +| `data-pipeline/worddata/english/nouns/` | ⚠️ Empty leftover output directory from the old per-word-JSON design | -Existing multilingual dictionaries and translation APIs provide raw word-to-word mappings. They do not provide the structured, pedagogical data needed for a vocabulary trainer: +Directory naming follows the language/POS codes used in `packages/shared/src/constants.ts` (`de/noun`, not `german/nouns`) so no name mapping is needed anywhere in the pipeline. -| What Exists | What is Missing | Why It Matters | -| ------------------------ | ---------------------------------- | ------------------------------------------------------------- | -| Word + translation | Student-friendly definition | Learners need explanations, not just equivalents | -| Static difficulty labels | CEFR-calibrated difficulty | "Easy" vs "hard" must map to learner proficiency levels | -| Isolated translations | Natural example sentences | Context is how vocabulary is actually acquired | -| Raw gender data | Gender integrated with translation | Flashcards must show gender immediately (resolved via Kaikki) | -| Bulk word lists | Per-word structured JSON | The trainer consumes one file per word for fast random access | - -### The Target User - -A language learner using the Lila vocabulary trainer. They see a word, its definition, an example sentence, and translations with gender — all calibrated to their CEFR level (A1-C2). - -### Why Not Use Existing Dictionaries? - -- **Wiktionary:** Rich data, but unstructured, inconsistent formatting, no CEFR levels, no student-friendly definitions -- **Kaikki (Wiktionary dump):** Structured JSON, excellent for gender/translation lookup, but definitions are often technical, no difficulty classification, no example curation -- **Google Translate / DeepL:** No definitions, no examples, no difficulty, no structured output -- **Existing language learning apps:** Closed data, no export, no control over content - -The LLM fills the gap: it generates pedagogical content (student-friendly definitions, natural examples, difficulty classification) that no existing database provides at scale. - -### Language Direction - -The pipeline is direction-agnostic. A wordlist is defined by: - -- **Source language:** the language of the input words -- **Target languages:** all other languages in the system (auto-derived from `ALL_LANGUAGES` minus source) - -Current focus: English -> German/Italian/Spanish/French. Planned directions include German -> French, Italian -> Spanish, etc. - -### Why 100,000+ Words? - -- **Coverage:** A learner needs ~10,000 words for B2 proficiency. The pipeline targets 100k to cover multiple languages, POS categories, and difficulty levels. -- **Languages:** English (source) -> German, Italian, Spanish, French (targets). -- **Parts of speech:** Nouns (current), verbs, adjectives, adverbs. - -### The Quality Challenge - -| Risk | Mitigation | -| ------------------------------- | -------------------------------------------------------------- | -| Hallucinated definitions | Low temperature (0.1), strict system prompt, schema validation | -| Incorrect grammatical gender | **Resolved:** Decoupled from LLM; Kaikki Wiktionary lookup | -| POS bleed (verb defs for nouns) | Explicit negative constraint in system prompt | -| Inconsistent difficulty levels | Explicit CEFR mapping in prompt, spot-checking | -| JSON parse failures | Retry + split logic, schema validation, cleanup on failure | -| Model drift (online APIs) | Version pinning, local fallback | -| Provider downtime | Adapter pattern allows hot-swapping providers | - -### Why TypeScript + Node? - -- **Familiarity:** Existing project uses TypeScript (frontend in TanStack Router + React) -- **Ecosystem:** `readline` for streaming files, `fs` for JSON I/O, native `fetch` for HTTP -- **Portability:** Runs on the same Debian laptop as the llama.cpp server -- **No build complexity:** `tsx` for direct execution, no bundler needed - -### Why llama.cpp? - -- **GGUF format:** Single-file models, easy to swap, quantize, and version -- **OpenAI-compatible API:** `/v1/chat/completions` means the same adapter code works for local and online models -- **No dependencies:** Self-contained binary, runs on old hardware (tested on GTX 950M) -- **Privacy:** Local inference means no data leaves the machine +Note that `data-pipeline/vitest.config.ts` looks for tests in `tests/**/*.test.ts` — that directory does not exist yet. --- -## 3. Architecture Overview +## Staging schema -### Pipeline Flow +`data-pipeline/db/schema.sql` mirrors the PostgreSQL schema with two SQLite concessions: IDs are `TEXT` (`crypto.randomUUID()`), and `definitions` / `examples` are JSON-encoded strings because SQLite has no array type. The import script parses them back into PostgreSQL `TEXT[]`. -```text -Run CLI -> Scan sources -> Check LLM -> Loop wordlists -> Stream words -> Skip processed - -> Create skeletons (batch) -> Call LLM -> Parse JSON -> Validate senses - -> Retry/split on failure -> Merge -> Kaikki Gender Lookup -> Write atomically -> Verify schema -> Log metrics +``` +words id, headword, language_code, pos UNIQUE(headword, language_code, pos) +senses id, word_id→words, sense_index, UNIQUE(word_id, sense_index) + difficulty, definitions, examples +translations id, sense_id→senses, target_language_code, UNIQUE(sense_id, target_language_code, translation) + translation, gender, difficulty ``` -### Resumability - -- **Skip existing:** `check-if-json-exists.ts` checks if `{word}.json` exists with non-empty `senses` -- **Atomic writes:** `.tmp` -> rename in `write-json-file.ts`, no partial files on crash -- **Cleanup on failure:** Deletes partially-written files for failed batches, continues to next batch - -### Directory Structure - -```text -data-pipeline/ - |-- pipeline.ts # Entry point / orchestrator - |-- utils/ - | |-- cli.ts # Interactive CLI module - | |-- enrich-word.ts # LLM call, parse, retry/split - | |-- merge-enriched-data.ts # Schema types + merge logic - | |-- verify-enriched-file.ts # Schema validation - | |-- check-llm-server.ts # Health check (local only) - | |-- scanning-source-files.ts # Source discovery - | |-- create-base-json.ts # Skeleton writer - | |-- write-json-file.ts # Atomic JSON writer - | |-- check-if-json-exists.ts # Resumability check - | |-- create-line-reader.ts # Streaming file reader - | |-- create-output-dirs.ts # Directory creation - | |-- delete-file.ts # Cleanup helper - | |-- get-word-file-path.ts # Path helper - | |-- progress-tracker.ts # Console progress formatting - | |-- pipeline-timer.ts # Timing + token metrics - | |-- llm-adapters/ - | |-- factory.ts # Adapter selection - | |-- types.ts # LlmAdapter interface - | |-- openai-compatible.ts # Local, OpenRouter, DeepSeek - | |-- gemini.ts # Google Gemini - |-- config/ - | |-- llm.ts # LLM config schema - | |-- prompt.ts # buildSystemPrompt() - | |-- batch.ts # Batch config schema - | |-- constants.ts # LANG_MAP, POS_MAP, ALL_LANGUAGES - |-- source-data/ - | |-- {language}/ - | |-- {pos} # One word per line, no extension - |-- worddata/ - | |-- {language}/ - | |-- {pos}/ - | |-- {word}.json # One self-contained file per word - |-- kaikki-source-files/ # Wiktionary dumps for gender lookup - |-- .pipeline-config.json # Saved CLI configuration -``` - -### Output Schema - -Each `.json` file contains: `word`, `language`, `pos`, `senses[]` (each with `id`, `sense`, `example`, `difficulty_level`, `translations` per target language), `enrichedAt`, `model`. -_Note: The `translations` object contains arrays of strings (e.g., `{"de": ["Haus", "Gebäude"]}`). Grammatical gender is appended later via the Kaikki integration step._ - -### Error Handling - -| Failure | Behavior | -| ------------------------------ | -------------------------------------------------------- | -| LLM server offline | Hard fail at startup (`check-llm-server.ts`, local only) | -| LLM returns bad JSON | Retry up to 3 times, then split batch. Log and continue | -| LLM returns malformed senses | `validateSense()` catches it before file write | -| Schema validation fails | Log warnings, keep file | -| Individual batch fails | Does not stop pipeline; cleans up partial files | -| Individual word fails (size 1) | Log and continue to next word | - -### Metrics - -Per-run: words processed/skipped/failed, duration, throughput, LLM token counts and speeds. See `utils/pipeline-timer.ts`. - -- **Unified throughput (all providers):** total tokens / total request time -- **Detailed breakdown (local only):** prompt speed vs completion speed +`difficulty` is `easy | medium | hard` on both `senses` and `translations`, and they mean different things — sense difficulty is "is this _meaning_ appropriate for the level", translation difficulty is "is this _word_ an acceptable answer". design-doc §4 explains how queries use sense difficulty as a ceiling and translation difficulty as the target. --- -## 4. Current Implementation +## The prompt -### Tech Stack +`data-pipeline/prompt` is the current working prompt, checked in as a plain text file and edited by hand. It is currently pinned to a concrete sample run (Spanish nouns, 20 words inlined) rather than templated — source language, POS, target languages, and the word batch will need to become substitutions when `pipeline.ts` is implemented. -| Layer | Choice | Why | -| ----------- | ----------------------------- | ------------------------------------------ | -| Runtime | Node.js + `tsx` | Direct TypeScript execution, no build step | -| HTTP client | Native `fetch` | Works for local llama.cpp and online APIs | -| File I/O | `fs` + `readline` | Streaming line reader for large wordlists | -| JSON | Native `JSON.parse/stringify` | Simple, no schema library needed | -| CLI | Native `readline` | No external dependencies | +What it enforces, beyond the JSON shape in design-doc §6.3: -### Current Model +- Raw JSON only — no markdown fences, comments, or trailing commas; one object per input word, in input order. +- Definitions and examples in the **source** language. +- Gender required for `de` (m/f/n) and `it`/`es`/`fr` (m/f); always `null` for `en`. +- German translation nouns capitalized; Romance-language nouns lowercase unless proper nouns. +- Base dictionary form, no articles or determiners. +- 1–3 senses per word, most words 1; skip rare, archaic, and technical senses. +- Up to 2 translations per target language per sense, only genuine synonyms or difficulty variants. +- A translation's difficulty may never be lower than its sense's difficulty. +- A word that isn't a valid noun in that language comes back with `"senses": []`. -| Property | Value | -| ---------------- | ---------------------------------------------------- | -| **Model** | **`gemma-4-E2B_q4_0-it.gguf`** | -| Size | ~3.2GB (file) / ~2.06GB (VRAM weights) | -| Quantization | Q4_0 | -| Server | `llama.cpp` (`llama-server`) | -| API | OpenAI-compatible `/v1/chat/completions` | -| VRAM Usage | ~2.65GB total (weights + KV cache + compute buffers) | -| Generation Speed | ~10.9 tok/s (20-word batch) | -| Est. 100k Time | ~7 days (20-word batches, 24/7) | +⚠️ **Known inconsistencies in the current prompt file** — it was adapted from the English version and some hardcoded values were not updated: rules 2 and 3 still say `language` must be `"en"` and there is a stray "valid English noun" in rule 31, while the header correctly says Spanish. Rule 15 lists target languages `de, it, es, fr` while the header says `en, it, de, fr`. Fix these when templating the prompt. -### `llama-server` Flags: History & Rationale +Validation is the safety net, not the prompt — every entry is checked before it reaches SQLite, and rejects go to a log for review rather than silently disappearing. Target reject rate is under 10%. -The server flags evolved through rigorous empirical testing on the target hardware (Intel i7-6500U, GTX 950M 4GB, 8GB RAM) across 10 different models on 2026-07-18. +--- -#### Flag Evolution - -| Flag | Value Tried | Result | Why | -| ----------------- | ------------------------------------------ | --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ | -| `-m` | `qwen3.5-4b-q4_k_m.gguf` | Works, ~6.3 tok/s | Quality baseline. Correct translations. 2.6GB, tight on VRAM. | -| `-m` | `qwen2.5-1.5b-instruct-q4_k_m.gguf` | Works, ~18.5 tok/s | Fast but poor instruction following, gender wrong. | -| `-m` | `Qwen3.5-2B-Q4_K_M.gguf` | Works, ~13.2 tok/s | Good translations, but failed polysemy (2 identical senses). | -| `-m` | `Llama-3.2-3B-Instruct-Q4_K_M.gguf` | Works, ~9.1 tok/s | Dangerous false friend trap (cognates). | -| `-m` | `Llama-3.2-3B-Instruct-UD-Q6_K_XL.gguf` | Works, ~6.7 tok/s, 3.44GB VRAM | Higher quant did NOT fix Llama's translation issues. Slower. | -| `-m` | `Ministral-3-3B-Instruct-2512-Q4_K_M.gguf` | Works, ~8.9 tok/s | Fixed tokenizer bug (older version was broken). But messy translations, markdown violations. | -| `-m` | `gemma-4-E2B_q4_0-it.gguf` | **Works, ~13.1 tok/s, 2.06GB** | **Winner.** Perfect polysemy, false friends, nuance. Half VRAM of Qwen 4B. | -| `-m` | `gemma-4-E4B-it-qat-UD-Q4_K_XL.gguf` | Works, ~7.7 tok/s, 3.23GB | QAT compression is incredible. Same quality as E2B but slower. | -| `-m` | `qwen3.5-9b-q3_k_s.gguf` | Works, ~3.3 tok/s (split CPU/GPU) | Brilliant quality but bottlenecked by CPU/Swap. `-ngl 28` max. | -| `-m` | `qwen3.5-9b-q4_k_m.gguf` | Works, ~3.3 tok/s (split CPU/GPU) | Same quality as 9B Q3. Not worth 2x the file size. | -| `-ngl` | 999 | Keeps | Offload all layers to GPU. Required for any speed. | -| `-ngl` | 28 | 9B models only | Max GPU layers for 9B models before OOM. Found via binary search. | -| `-ngl` | 30 | OOM crash (9B) | Pushed 2 layers too far into compute buffers. | -| `-c` | 4096 | Wasteful | 4K context for 300-token dictionary entries wastes VRAM. | -| `-c` | 2048 | Good for small batches | Sufficient for 4-word batches. | -| `-c` | **8192** | **Current** | Required for 20-word batches. Combined with KV cache quantization. | -| `-b` / `-ub` | 512 | **Current** | Sweet spot for Maxwell memory bandwidth. | -| `-b` / `-ub` | 1024 | Tested | Slightly faster prompt processing, but no generation speedup. | -| `-b` / `-ub` | 2048 | Slower on 950M | Memory pressure on bandwidth-starved GPU. | -| `-t` | 4 | Slower | Hyperthreading cores hurt llama.cpp performance. | -| `-t` | **2** | **Current** | Matches 2 physical cores. | -| `--threads-batch` | **2** | **Current** | Explicit match to `-t`. | -| `--flash-attn` | (omitted) | Correct | On Maxwell (compute 5.0), Flash Attention adds overhead. | -| `--mlock` | Tested | Omitted for large models | Pins model in RAM. Causes OOM on models >2.5GB with large KV cache. | -| `--prio` | **2** | **Current** | Raises process priority. Marginal, harmless. | -| `--reasoning` | **off** | **Critical** | **Mandatory for Qwen 3.5 and Gemma 4.** Without this, models "think" silently, consume all `max_tokens`, and crash with `finish_reason: length`. | -| `--cache-type-k` | **q4_0** | **Current** | Compresses KV cache keys to 4-bit. Cuts KV VRAM by ~75%. Enables 8192 context on 4GB GPU. | -| `--cache-type-v` | **q4_0** | **Current** | Compresses KV cache values to 4-bit. Paradoxically improves translation variety (reduces "lazy duplication" bug). | - -#### Current Production Command +## Running it ```bash -cd ~/Downloads/llama.cpp -./build/bin/llama-server \ - -m models/gemma-4-E2B_q4_0-it.gguf \ - -ngl 999 \ - -c 8192 \ - -b 512 \ - -ub 512 \ - -t 2 \ - --threads-batch 2 \ - --host 127.0.0.1 \ - --port 8080 \ - --prio 2 \ - --reasoning off \ - --cache-type-k q4_0 \ - --cache-type-v q4_0 +docker compose up -d pipeline-database # dedicated PostgreSQL on :5433 +pnpm --filter @lila/pipeline pipeline:run # tsx --env-file .env pipeline.ts (currently a no-op) +pnpm --filter @lila/pipeline test ``` -#### VRAM Budget (Production Config) - -| Component | VRAM Usage | -| ------------------------------------ | ------------------------- | -| Model Weights (Gemma 4 E2B Q4_0) | ~1.50 GB | -| KV Cache (8192 ctx, q4_0 compressed) | ~0.35 GB | -| Compute Buffers (batch 512) | ~0.20 GB | -| **Total** | **~2.05 GB (51% of 4GB)** | - -#### Why Not Use the Remaining 2GB VRAM? - -Generation speed is bottlenecked by **memory bandwidth** (~32 GB/s on GTX 950M), not VRAM capacity. To generate one token, the GPU must read the entire ~1.5GB model from VRAM. The theoretical maximum is ~21 tok/s. At 10.9 tok/s, the GPU is already operating at ~50% of its physical limit. Empty VRAM cannot be converted into faster generation. - -Testing uncompressed `f16` KV cache (2.65GB total VRAM) yielded 12.6 tok/s but caused a severe "lazy duplication" regression (model copy-pasted the same translation twice instead of providing distinct synonyms). The `q4_0` compressed KV cache is the correct choice for translation quality. - -### Performance Baseline (20-Word Batch) - -| Metric | Gemma 4 E2B | Qwen 3.5 4B | -| --------------------- | ----------- | ------------ | -| Time/batch (20 words) | ~5 min | ~9 min | -| Completion tok/s | ~10.9 | ~6.1 | -| Prompt tok/s | ~132 | ~78 | -| VRAM Usage | 2.06 GB | 3.95 GB | -| Lazy Duplication Bug | No | Yes (severe) | +The pipeline reads `.env` from the repo root: `GEMINI_API_KEY`, plus `PIPELINE_POSTGRES_USER` / `PIPELINE_POSTGRES_PASSWORD` / `PIPELINE_POSTGRES_DB` / `PIPELINE_DATABASE_URL`. The pipeline database is deliberately separate from the app database (`:5432`) so pipeline work can never damage dev data. --- -## 5. The LLM Layer +## Phase status -### 5.1 Local Model Evaluation (Complete — 2026-07-18) +Full breakdown in [pipeline/roadmap.md](pipeline/roadmap.md). -All 10 downloaded models were evaluated on the target hardware. Testing progressed from 4-word smoke tests to a 20-word "nightmare" torture suite covering extreme polysemy, false friends, abstract concepts, and legal/financial terminology. +| Phase | State | +| ------------------------------------------------- | -------------------------------------------------------------- | +| 1 — Drizzle schema (words/senses/translations) | ✅ Complete | +| 2 — Preparation (wordlists, DBs, prompt) | ✅ Complete | +| 3 — Build the pipeline → SQLite | 🔄 **Current.** Validation + `pipeline.ts` + first run | +| 4 — Migration & SQLite → PostgreSQL import | ⬜ Not started | +| 5 — App integration (`getGameTerms`, distractors) | ⬜ Not started | +| 6 — Production deploy | ⬜ Not started | +| 7 — Extend to verbs, adjectives, adverbs | ⬜ Not started — new wordlists + prompt only, no schema change | -#### Final Leaderboard +Target for Phase 3: ~1000 nouns × 5 languages in staging, reject rate under 10%, 50 entries spot-checked by hand. -| Rank | Model | Size | VRAM | Speed (tok/s) | Polysemy | False Friends | Verdict | -| ------ | ------------------------ | ---- | ------------- | ------------- | --------------------------- | ------------------------- | ----------------------------------------------------------- | -| **🥇** | **Gemma 4 E2B Q4_0** | 3.2G | **2.06 GB** | **~13.1** | **Perfect** | **Perfect** | **Production model.** Best quality/speed/VRAM ratio. | -| 🥈 | Qwen 3.5 4B Q4_K_M | 2.6G | ~3.6 GB | ~6.3 | Perfect | Perfect | Quality King, but 2x slower and maxes VRAM. | -| 🥉 | Gemma 4 E4B Q4_K_XL | 4.0G | 3.23 GB | ~7.7 | Perfect | Perfect | Incredible QAT compression. Same quality as E2B but slower. | -| 4 | Qwen 3.5 2B Q4_K_M | 1.2G | ~1.6G | ~13.2 | Failed (2 identical senses) | Passed | Good translations, lacks conceptual branching. | -| 5 | Qwen 2.5 1.5B Q4_K_M | 1.1G | ~1.5G | ~18.5 | Ignored instruction | Failed | Fast but easily confused. | -| 6 | Ministral 3B 2512 Q4_K_M | 2.0G | ~2.7G | ~8.9 | Messy translations | Failed | Fixed tokenizer, but outclassed. Markdown violations. | -| 7 | Llama 3.2 3B Q4_K_M | 1.9G | ~2.6G | ~9.1 | Good structure, bad IT/ES | **Failed (cognate trap)** | Dangerous for language learners. | -| 8 | Llama 3.2 3B Q6_K_XL | 2.8G | 3.44G | ~6.7 | Good structure, bad IT/ES | **Failed (cognate trap)** | Higher quant did NOT fix translation issues. | -| 9 | Qwen 3.5 9B Q3_K_S | 4.1G | Split CPU/GPU | ~3.3 | Perfect | Perfect | Brilliant but 35-40 days for 100k words. | -| 10 | Qwen 3.5 9B Q4_K_M | 5.3G | Split CPU/GPU | ~3.3 | Perfect | Perfect | Same quality as 9B Q3. Not worth the size. | - -#### Key Findings - -1. **The "Thinking" Trap:** Both Qwen 3.5 and Gemma 4 have built-in Chain-of-Thought reasoning. Without `--reasoning off`, they silently "think" in a hidden JSON field, consume all `max_tokens`, and crash with `finish_reason: length`. This flag is **mandatory**. -2. **The 2B vs 4B Quality Cliff:** 2B models struggle with polysemy (e.g., cannot distinguish "bank" = financial vs river). 4B+ models act like professional lexicographers. -3. **Llama 3.2 is Unsafe for Language Learners:** Consistently fell for false friend traps (e.g., translating "actual" = _real_ to _aktuell/attuale/actual/actuel_ = _current_). -4. **KV Cache Quantization Improves Translation Variety:** Compressing the KV cache to `q4_0` introduces microscopic noise that prevents the "lazy duplication" bug (model copy-pasting the same synonym twice). -5. **POS Bleed is Universal:** All models occasionally generate verb definitions for nouns (e.g., "run" = _to move fast_ instead of _a jogging session_). Fix: explicit negative constraint in the system prompt. - -### 5.2 Online API Options - -Evaluated as fallbacks if local models fail quality or speed targets. - -| Provider | Model | Input $/1M | Output $/1M | Free Tier | Rate Limit | Est. Cost (100k words) | Est. Time | -| --------------- | -------------------- | ---------- | ----------- | ------------- | ---------- | ---------------------- | ------------------- | -| DeepSeek | V4 Flash | $0.14 | $0.28 | 5M tokens | None | $0-15 | ~1-2 days | -| Gemini | 2.5 Flash-Lite | $0.10 | $0.40 | 1,500 req/day | 1M TPM | $0 (free tier) | ~1.5 days (batched) | -| Qwen/Alibaba | Qwen-Turbo | $0.05 | $0.20 | Unknown | 600 RPM | ~$11 | ~1-2 days | -| Groq | Llama 3.1 8B Instant | $0.05 | $0.08 | Yes | High | ~$7 | ~3-4 hours | -| OpenRouter free | Various | $0 | $0 | 200 req/day | 20 RPM | $0 | ~10 days | - -### 5.3 Model Selection Criteria - -Decision flow for 100,000 words: - -```text -Start - | - v - Gemma 4 E2B (local) — SELECTED - | - |-- Speed acceptable? (~7 days) -----> Use Gemma 4 E2B locally, $0 - | - |-- Need faster? -------> Test Gemini 2.5 Flash-Lite (free) - | - |-- Quality good? --> Batch 50, free tier - | ~1.5 days, $0 - | - |-- Quality meh? ---> Test Groq or DeepSeek paid -``` - -Quality gates: - -- 100% JSON parse rate -- No hallucinated definitions on polysemous words -- Natural, contextually appropriate example sentences -- Sensible difficulty classification (CEFR mapping) -- Gender accuracy is no longer an LLM criterion (handled by Kaikki) - ---- - -## 6. The Gender Problem & Kaikki Integration - -### The Problem (Resolved) - -Grammatical gender was originally embedded in the LLM's `translations` object. Testing across all 10 models confirmed that small local models systematically hallucinate gender, defaulting to `neuter` for Romance languages (Italian, Spanish, French) which do not have a neuter grammatical gender. - -### The Solution: Decoupled Architecture - -**Decision (2026-07-18):** Grammatical gender is no longer generated by the LLM. The pipeline now uses a two-stage approach: - -| Stage | Component | Responsibility | -| ----- | ----------------- | -------------------------------------------------------------------- | -| 1 | LLM (Gemma 4 E2B) | Generates translation **strings only** (e.g., `["Haus", "Gebäude"]`) | -| 2 | Kaikki Lookup | Deterministically resolves grammatical gender from Wiktionary dumps | - -### Benefits - -- **100% deterministic gender accuracy** — no hallucination possible -- **Faster LLM generation** — ~15-20% fewer output tokens per word -- **Simpler JSON schema** — translations are string arrays, not object arrays -- **Model-agnostic** — works with any LLM regardless of multilingual training quality - -### Kaikki Data - -| Language | Kaikki File | Coverage | -| -------- | ------------------------------------- | -------- | -| German | `kaikki.org-dictionary-German.jsonl` | High | -| Italian | `kaikki.org-dictionary-Italian.jsonl` | High | -| Spanish | `kaikki.org-dictionary-Spanish.jsonl` | High | -| French | `kaikki.org-dictionary-French.jsonl` | High | - -### Lookup Logic - -Match on `word` (the translated string) -> extract `gender` field -> map to `"masculine" | "feminine" | "neuter" | null`. - ---- - -## 7. Batching Strategy - -### Why Batching is Necessary - -At 1 word per call, 100,000 words = 100,000 LLM requests. Each call re-processes the ~300-token system prompt. Batching amortizes this cost. - -### Optimal Batch Size (Local) - -**20 words** is the validated sweet spot for the GTX 950M with Gemma 4 E2B. - -| Batch Size | VRAM | Speed | Quality | Verdict | -| ---------- | --------------- | ---------------- | ------------------------ | -------------------------------------------------- | -| 1-4 | ~2.1 GB | ~13 tok/s | Perfect | Safe but slow (amortization waste) | -| **20** | **~2.6 GB** | **~10.9 tok/s** | **Perfect** | **Sweet spot** | -| 30-40 | ~3.0 GB (est.) | ~10 tok/s (est.) | Likely good | Worth testing | -| 50+ | ~3.5 GB+ (est.) | Unknown | Risk of JSON degradation | **Not recommended for 2B models** | -| 100 | OOM risk | N/A | Attention degradation | Small models lose JSON structure past ~6000 tokens | - -### Retry & Split Strategy - -If a batch fails (bad JSON, missing key, etc.): - -```text -Batch of 20 fails (3 retries exhausted) - | - v -Split into 2 batches of 10 - | - v -If a batch of 10 fails (3 retries), split into 2 batches of 5 - | - v -If a batch of 5 fails, split into batches of 1 - | - v -If a single word fails (3 retries), log and skip -``` - ---- - -## 8. Hardware Constraints - -### Current Machine - -| Component | Spec | -| ------------- | ------------------------------------------------------------------- | -| OS | Debian GNU/Linux 13 (trixie) x86_64 | -| CPU | Intel Core i7-6500U (2 physical cores, 4 threads via HT) @ 3.10 GHz | -| GPU | NVIDIA GeForce GTX 950M (Maxwell, 2015) | -| GPU VRAM | 4GB (4037 MiB reported by CUDA) | -| GPU Bandwidth | ~32 GB/s (DDR3) | -| RAM | 8GB (~3.67GB usable at idle) | -| Swap | 5.62 GB | -| Disk | 102GB ext4 (~74GB used) | - -### What Fits in 4GB VRAM (Empirically Verified) - -| Model | File Size | Total VRAM | Fits? | Notes | -| ------------------------ | --------- | ----------- | ---------- | -------------------------------------- | -| Qwen 2.5 1.5B Q4_K_M | 1.1G | ~1.5 GB | ✅ Easy | | -| Qwen 3.5 2B Q4_K_M | 1.2G | ~1.6 GB | ✅ Easy | | -| Ministral 3B 2512 Q4_K_M | 2.0G | ~2.7 GB | ✅ Yes | | -| Llama 3.2 3B Q4_K_M | 1.9G | ~2.6 GB | ✅ Yes | | -| Llama 3.2 3B Q6_K_XL | 2.8G | 3.44 GB | ✅ Tight | | -| **Gemma 4 E2B Q4_0** | **3.2G** | **2.06 GB** | **✅ Yes** | **QAT compression. Production model.** | -| Gemma 4 E4B Q4_K_XL | 4.0G | 3.23 GB | ✅ Yes | QAT compression is incredible. | -| Qwen 3.5 4B Q4_K_M | 2.6G | 3.95 GB | ⚠️ Barely | 50MB headroom with 8192 ctx. | -| Qwen 3.5 9B Q3_K_S | 4.1G | Split | ⚠️ Partial | `-ngl 28` max. Rest on CPU/Swap. | -| Qwen 3.5 9B Q4_K_M | 5.3G | Split | ⚠️ Partial | `-ngl 20` max. Heavy swap usage. | - -### GPU Rental Alternatives - -If local hardware becomes the bottleneck: -| Provider | GPU | VRAM | Price/Hour | Time for 100k Words | Total Cost | -| --- | --- | --- | --- | --- | --- | -| Vast.ai | RTX 4090 | 24GB | ~$0.30-0.60 | ~6-8 hours | ~$2-5 | -| RunPod | RTX 4090 | 24GB | ~$0.50-0.80 | ~6-8 hours | ~$4-6 | - ---- - -## 9. Testing & Quality Assurance - -### 20-Word Torture Suite (Completed 2026-07-18) - -Tested on Gemma 4 E2B and Qwen 3.5 4B with 20 challenging nouns: - -| Category | Words | The Trap | -| ----------------- | ------------------------------------------------------- | -------------------------------------------------------------- | -| Extreme Polysemy | `match`, `date`, `right`, `set`, `well` | Does it translate "match" as fire, sports, or dating? | -| False Friends | `sense`, `fabric`, `sympathy`, `eventuality`, `billion` | "Fabric" = material (_tissu_) not factory (_fabrique_) | -| Abstract/Cultural | `serendipity`, `accountability` | Concepts lacking 1:1 dictionary equivalents | -| Action-Nouns | `run`, `drive`, `play` | "Run" as jog vs tear vs campaign | -| Legal/Financial | `mortgage`, `lease`, `court`, `board`, `draft` | Requires specific legal vocabulary (_Hypothek/mutuo/hipoteca_) | - -### Results Summary - -| Criterion | Gemma 4 E2B | Qwen 3.5 4B | -| ------------------------------------ | ------------------------------------------ | ------------------------------------------------------- | -| JSON Reliability | 10/10 (raw JSON) | 10/10 (raw JSON, but added `sense_index` hallucination) | -| Polysemy (bank, match) | 10/10 (Ufer/riva/orilla/rive) | 10/10 | -| False Friends (fabric) | 10/10 (Stoff/tessuto/tela/tissu) | 10/10 | -| Legal Nuance (mortgage) | 10/10 (Hypothek/mutuo/hipoteca/hypothèque) | 10/10 | -| Lazy Duplication Bug | None | **Severe** (copy-pasted same word 20+ times) | -| POS Bleed | Minor (verb defs for run/match) | Minor (verb defs for run/match) | -| Attention Degradation (word 20 vs 1) | None | None | - -### Verification - -`verify-enriched-file.ts` checks: - -- Required top-level fields: `word`, `language`, `pos`, `senses` -- Each sense: `sense` (string), `example` (string), `difficulty_level` in {easy, medium, hard} -- Each translation: array of strings (gender appended later by Kaikki) - ---- - -## 10. Interactive CLI - -### Batch Size Recommendations (Updated) - -| Provider | Recommended | Rationale | -| -------------------- | ----------- | ----------------------------------------------- | -| **local (GTX 950M)** | **20** | Validated sweet spot. 8192 ctx + q4_0 KV cache. | -| local (RTX 4090) | 50 | Fast, more VRAM | -| gemini | 50 | Free tier: 1,500 req/day | -| deepseek | 20 | 5M free tokens | -| groq | 50 | Very fast | - ---- - -## 11. Future Extensions & Roadmap - -### Near-Term (Next 2-4 Weeks) - -| Item | Status | Notes | -| ----------------------------------- | ------------ | --------------------------------------------------- | -| 10-model evaluation | **Complete** | Gemma 4 E2B selected | -| KV cache quantization | **Complete** | `--cache-type-k/v q4_0` enables 8192 ctx on 4GB GPU | -| Gender decoupling | **Complete** | Kaikki lookup replaces LLM gender | -| 20-word torture suite | **Complete** | Validated on Gemma E2B and Qwen 4B | -| POS bleed fix | **Pending** | Add negative constraint to system prompt | -| Kaikki gender lookup implementation | **Pending** | Post-processing step after LLM enrichment | -| Online API testing | **Pending** | Gemini free tier, DeepSeek, Groq | - -### Medium-Term (1-3 Months) - -| Item | Notes | -| ---------------------------- | ------------------------------------------------------ | -| Multi-POS support | Verbs, adjectives, adverbs need prompt variants | -| Multi-language source | German -> French, Italian -> Spanish, etc. | -| Parallel wordlist processing | Run `english/nouns` and `english/verbs` simultaneously | -| Incremental enrichment | Only process new/changed words in a wordlist | - -### Long-Term (3-6 Months) - -| Item | Notes | -| ------------------------ | ---------------------------------------------------------------- | -| GPU rental integration | Script to spin up Vast.ai/RunPod, run pipeline, download results | -| Quality regression tests | Run torture suite on every model change | -| Community contributions | Open-source the pipeline for other language learners | - ---- - -## 12. Decisions Log - -| Date | Decision | Context | Rationale | -| -------------- | ----------------------------------- | --------------------------------------------------------------- | ------------------------------------------------------------------------------------------------- | -| 2026-01-04 | TanStack Router for frontend | Previous project used React Router | Simpler, type-safe routing | -| 2026-06-16 | llama.cpp for local LLM | Need local inference on old laptop | GGUF format, OpenAI-compatible API | -| 2026-06-16 | Q4_K_M quantization | Balance size vs quality | Community standard for 4-bit | -| 2026-06-16 | `-t 2` physical cores | Default was 4 (HT threads) | Hyperthreading hurts llama.cpp | -| 2026-06-17 | Qwen2.5-1.5B as initial model | Qwen3.5-4B too slow (47s/word) | 6x speedup, quality under evaluation | -| 2026-07-06 | Adapter pattern for LLM providers | Need to evaluate local vs cloud | `utils/llm-adapters/` with factory | -| 2026-07-06 | Retry + split batching | LLM JSON parse failures | 3 retries, then halve batch | -| 2026-07-06 | Interactive CLI | Editing config files is error-prone | `utils/cli.ts` with native readline | -| **2026-07-18** | **Gemma 4 E2B as production model** | **10-model evaluation completed** | **2x faster than Qwen 4B, half VRAM, perfect translation quality** | -| **2026-07-18** | **Gender decoupled from LLM** | **All 10 models failed gender for Romance languages** | **Kaikki Wiktionary lookup is deterministic and 100% accurate** | -| **2026-07-18** | **`--reasoning off` is mandatory** | **Qwen 3.5 and Gemma 4 "think" silently, consuming all tokens** | **Without this flag, output crashes with `finish_reason: length`** | -| **2026-07-18** | **KV cache quantization (`q4_0`)** | **8192 context needed for 20-word batches** | **Cuts KV VRAM by 75%, enables large batches on 4GB GPU, improves translation variety** | -| **2026-07-18** | **20-word batch size for local** | **Tested 4, 20 words** | **Sweet spot: no attention degradation, no JSON breakage, 10.9 tok/s** | -| **2026-07-18** | **Llama 3.2 discarded** | **Failed false friend tests** | **Translates "actual" (real) to cognates (aktuell/attuale) = "current". Dangerous for learners.** | -| **2026-07-18** | **`-c 8192` replaces `-c 2048`** | **20-word batches need more context** | **Combined with q4_0 KV cache, fits in 2.65GB VRAM** | - ---- - -## 13. Known Issues & Dev Notes - -### Data Pipeline - -| Issue | Details | Severity | -| ------------------------------ | -------------------------------------------------------------------------------------- | ----------------------------------------------- | -| POS bleed (all models) | Models occasionally generate verb definitions for nouns (e.g., "run" = _to move fast_) | Medium — fix with negative constraint in prompt | -| Lazy duplication (Qwen 4B) | Qwen 3.5 4B copy-pastes the same translation twice to fill arrays | Medium — use Gemma 4 E2B instead | -| Schema hallucination (Qwen 4B) | Adds `sense_index` field not in schema | Low — ignorable | -| Pre-scanning wordlists | Entire file read into memory before processing | Medium — streaming refactor planned | -| Single POS tested | Only nouns validated. Verbs/adjectives need prompt changes | Known limitation | - -### Hardware - -| Issue | Details | -| --------------------- | ---------------------------------------------------------------------------- | -| GTX 950M VRAM ceiling | 4GB hard limit. Models >3.5GB need KV cache quantization or CPU offloading. | -| Maxwell GPU aging | No Flash Attention, bandwidth-starved (~32 GB/s). Theoretical max ~21 tok/s. | -| Laptop thermals | Cannot run 24/7 for weeks unattended. Monitor temps. | - ---- - -## 14. How to Run - -### Prerequisites - -- Node.js + npm -- `tsx` installed globally: `npm install -g tsx` -- llama.cpp built from source (for local mode) -- GGUF model downloaded to `~/Downloads/llama.cpp/models/` -- API keys set as environment variables (for cloud mode) - -### Start the LLM Server (Local Mode) - -```bash -cd ~/Downloads/llama.cpp -./build/bin/llama-server \ - -m models/gemma-4-E2B_q4_0-it.gguf \ - -ngl 999 \ - -c 8192 \ - -b 512 \ - -ub 512 \ - -t 2 \ - --threads-batch 2 \ - --host 127.0.0.1 \ - --port 8080 \ - --prio 2 \ - --reasoning off \ - --cache-type-k q4_0 \ - --cache-type-v q4_0 -``` - -### Run the Pipeline - -```bash -cd /path/to/data-pipeline -npx tsx pipeline.ts -``` - -Follow the interactive prompts to select provider, model, and batch size. - ---- - -## 15. Roadmap - -### Phase 1: Batching (Complete) - -| Task | Status | -| --------------------------------------- | ------------ | -| Implement configurable batch size | Complete | -| Implement retry + split logic | Complete | -| Honest timing metrics | Complete | -| Validate LLM responses | Complete | -| Verify batching quality (20-word suite) | **Complete** | - -### Phase 2: Interactive CLI (Complete) - -| Task | Status | -| --------------------- | -------- | -| Design prompt flow | Complete | -| Implement CLI module | Complete | -| Save/load config | Complete | -| Wire into pipeline.ts | Complete | - -### Phase 3: Model Selection (Complete) - -| Task | Status | -| ----------------------- | -------------------------- | -| 10-model evaluation | **Complete** | -| 20-word torture suite | **Complete** | -| Select production model | **Complete (Gemma 4 E2B)** | -| Test online APIs | Pending | - -### Phase 4: Scale - -| Task | Status | Notes | -| ------------------------------ | ------- | ------------------------------------------ | -| Implement Kaikki gender lookup | Pending | Post-processing step | -| Fix POS bleed in prompt | Pending | Add negative constraint | -| Run 100k word pipeline | Pending | ~7 days local (Gemma E2B, 20-word batches) | -| Spot-check output quality | Pending | Random sample of 100 entries | - -### Phase 5: Extend - -| Task | Status | -| ---------------------------- | ------- | -| Multi-POS support | Pending | -| Multi-language source | Pending | -| Parallel wordlist processing | Pending | -| GPU rental integration | Pending | -| Quality regression tests | Pending | +Phase 5 is where this becomes visible in the app: `packages/db/src/models/termModel.ts` still queries `vocabulary_entries`/`entry_translations` and must be rewritten against the sense-based schema. Until then, production runs on the old data. diff --git a/documentation/ai-context/00-project-overview.md b/documentation/ai-context/00-project-overview.md index 6db8d8b..d506799 100644 --- a/documentation/ai-context/00-project-overview.md +++ b/documentation/ai-context/00-project-overview.md @@ -31,7 +31,7 @@ Live at [lilastudy.com](https://lilastudy.com). ### What's In Progress / Blocked -- **Kaikki data pipeline migration** — Replacing OpenWordNet/OMW with sense-disambiguated Kaikki data. Stage 1 (extract) and Stage 2 (reverse link) complete on sample data. Stage 3 (enrich) being rewritten for sub-stage architecture. +- **Data pipeline rewrite** — The Kaikki/local-LLM pipeline was replaced by a Gemini-only pipeline writing a new sense-based schema (`words` → `senses` → `translations`). Schema and prompt are done; the pipeline script itself is not written yet. See 05-data-pipeline.md. - **Guest play** — No try-before-signup flow yet. Auth required for all game routes. - **Game session store** — Still in-memory. Valkey container exists locally but not wired up. - **Media ingestion** — Not started. No pipeline for subtitles/lyrics → vocab extraction yet. @@ -40,7 +40,7 @@ Live at [lilastudy.com](https://lilastudy.com). The app is currently a **generic vocabulary quiz**. The media-based practice feature (the differentiator) does not exist yet. It depends on: -1. Kaikki pipeline reaching production (fixes translation quality) +1. The new data pipeline reaching production (fixes translation quality) 2. A media ingestion prototype (subtitles/lyrics → text → vocab extraction → quiz) --- @@ -50,7 +50,7 @@ The app is currently a **generic vocabulary quiz**. The media-based practice fea | Layer | Technology | | ------------- | -------------------------------------------------------------- | | Monorepo | pnpm workspaces | -| Frontend | React 18, Vite, TanStack Router, TanStack Query, Tailwind CSS | +| Frontend | React 19, Vite, TanStack Router, Tailwind CSS | | Backend | Node.js, Express, TypeScript, WebSockets (`ws` library) | | Database | PostgreSQL + Drizzle ORM | | Auth | Better Auth (Google + GitHub) | @@ -58,7 +58,7 @@ The app is currently a **generic vocabulary quiz**. The media-based practice fea | Testing | Vitest, supertest | | Deployment | Docker Compose, Caddy, Hetzner VPS | | CI/CD | Forgejo Actions | -| Data Pipeline | Kaikki (Wiktionary) → SQLite (`pipeline.db`) → PostgreSQL | +| Data Pipeline | Gemini API → SQLite staging (`db/staging.db`) → PostgreSQL | --- @@ -72,7 +72,7 @@ lila/ ├── packages/ │ ├── shared/ — Zod schemas + constants (API/web contract) │ └── db/ — Drizzle schema, migrations, models, seeding -├── data-pipeline/ — Kaikki extraction → enrichment → PostgreSQL sync +├── data-pipeline/ — Gemini generation → SQLite staging → PostgreSQL └── documentation/ — Project docs (human + AI-context branches) ``` @@ -86,7 +86,7 @@ lila/ 2. **Server-side answer evaluation** — The correct answer is never sent to the frontend. All evaluation happens server-side. 3. **Zod discriminated unions for WebSockets** — All WS messages are typed via Zod schemas in `packages/shared`. The router switches on the `type` field. 4. **GameSessionStore abstraction** — Session state is stored through an interface (`InMemoryGameSessionStore` now, `ValkeyGameSessionStore` planned). -5. **Language-neutral data model** — `terms` are concepts; `translations` are per-language words. Adding a language requires no schema changes. +5. **Sense-based data model** — `words` have `senses`, and translations hang off a sense. Adding a language requires no schema changes. Note the app still queries the older `vocabulary_entries` tables — see 02-data-model.md. --- @@ -98,7 +98,7 @@ lila/ | WebSocket | `ws` library, not Socket.io | 2–4 players, explicit Zod protocol sufficient | | Auth | Better Auth, not Keycloak | Embedded middleware, no separate service | | Answer eval | Server-side only | Correct answer never sent to frontend | -| Data source | Kaikki, not OMW | Sense-disambiguated translations | +| Data source | Gemini-generated, not OMW | Sense-disambiguated, language-native glosses | --- @@ -110,7 +110,7 @@ lila/ | [02-data-model.md](02-data-model.md) | Database schema, tables, relationships, constraints | | [03-api-contract.md](03-api-contract.md) | REST endpoints, request/response schemas, Zod types | | [04-websocket-protocol.md](04-websocket-protocol.md) | WS message types, game flow, auth, state management | -| [05-data-pipeline.md](05-data-pipeline.md) | Kaikki pipeline stages, enrich sub-stages, sync | +| [05-data-pipeline.md](05-data-pipeline.md) | Gemini pipeline flow, output contract, validation, blockers | | [06-deployment.md](06-deployment.md) | Docker, Caddy, CI/CD, backups | | [prompts/meta.md](prompts/meta.md) | How to work with LLMs on this codebase | | [99-current-task.md](99-current-task.md) | Template: fill this out before giving a task to an LLM | diff --git a/documentation/ai-context/01-architecture.md b/documentation/ai-context/01-architecture.md index 03596d6..5974be0 100644 --- a/documentation/ai-context/01-architecture.md +++ b/documentation/ai-context/01-architecture.md @@ -16,7 +16,7 @@ lila/ ├── packages/ │ ├── shared/ — Zod schemas, constants, derived types. THE CONTRACT. │ └── db/ — Drizzle schema, migrations, models (termModel, lobbyModel), seeding -├── data-pipeline/ — Kaikki extraction → enrichment → sync to PostgreSQL +├── data-pipeline/ — Gemini generation → SQLite staging → sync to PostgreSQL └── documentation/ — Human docs + ai-context/ ``` diff --git a/documentation/ai-context/02-data-model.md b/documentation/ai-context/02-data-model.md index c9e1ad4..5fa250e 100644 --- a/documentation/ai-context/02-data-model.md +++ b/documentation/ai-context/02-data-model.md @@ -1,136 +1,123 @@ # 02 — Data Model > **Purpose:** Database schema reference for LLMs working on features that query or modify data. Concatenate with 00-project-overview.md and 99-current-task.md. -> **Last updated:** 2026-05-15 +> **Last updated:** 2026-08-01 > **Depends on:** 00-project-overview.md +> **Source of truth:** `packages/db/src/db/schema.ts`. If this file and the schema disagree, the schema wins. --- -## Core Tables +## Two vocabulary schemas exist right now -### `terms` — Language-neutral concepts +The database is mid-migration and contains **both** vocabulary schemas. This is the most important thing to know before writing a query. -| Column | Type | Constraints | Notes | -| ------------ | --------- | -------------------------------------------- | ------------------------------------------------------ | -| `id` | uuid | PK | | -| `pos` | varchar | CHECK: `noun`, `verb`, `adjective`, `adverb` | Part of speech | -| `source` | varchar | | Pipeline that created this term (e.g. `kaikki`, `omw`) | -| `source_id` | varchar | UNIQUE(`source`, `source_id`) | Idempotency key for imports | -| `synset_id` | varchar | nullable | WordNet synset ID. Nullable for non-WordNet terms. | -| `created_at` | timestamp | default now() | | +| Schema | Status | +| ------------------------------------------- | --------------------------------------------------------------------------------------- | +| `vocabulary_entries` + `entry_translations` | **Live.** What the app queries today (`packages/db/src/models/termModel.ts`). | +| `words` → `senses` → `translations` | **Target.** Migrated and empty. The new pipeline writes here; no app code reads it yet. | -**Rule:** One row per concept. The word "cat" (animal) and "cat" (nautical) are separate rows because they have different `source_id` values. +The `terms` / `term_glosses` / `decks` / `deck_terms` tables described in earlier versions of this doc **no longer exist**. + +Migration path: the pipeline fills `words`/`senses`/`translations`, then `termModel.ts` is rewritten against it (roadmap Phase 5), then the `vocabulary_entries` tables are dropped. --- -### `translations` — Per-language words +## Live schema (what the app queries) -| Column | Type | Constraints | Notes | -| --------------- | ---------- | ----------------------------------- | ---------------------------------------- | -| `id` | uuid | PK | | -| `term_id` | uuid | FK → terms.id | | -| `language_code` | varchar(2) | CHECK: `en`, `it`, `de`, `es`, `fr` | | -| `text` | varchar | | The actual word | -| `cefr_level` | varchar(2) | nullable, CHECK: `A1`–`C2` | Difficulty of THIS word in THIS language | -| `created_at` | timestamp | default now() | | +### `vocabulary_entries` — one row per word sense -**Unique constraint:** (`term_id`, `language_code`, `text`) — allows synonyms (e.g. "dog" and "hound" for same term), prevents exact duplicates. +| Column | Type | Constraints | Notes | +| --------------- | ----------- | ----------------------------------- | ----------------------------------------- | +| `id` | uuid | PK, default random | | +| `headword` | text | NOT NULL | The word itself | +| `language_code` | varchar(10) | CHECK `SUPPORTED_LANGUAGE_CODES` | | +| `pos` | varchar(20) | CHECK `SUPPORTED_POS` | | +| `sense_index` | smallint | NOT NULL, default 0 | Distinguishes senses of the same headword | +| `gloss` | text | nullable | Definition | +| `examples` | text[] | NOT NULL, default `[]` | | +| `cefr_level` | varchar(2) | nullable, CHECK `A1`–`C2` | | +| `difficulty` | varchar(20) | nullable, CHECK `DIFFICULTY_LEVELS` | | +| `source` | varchar(50) | NOT NULL, default `"kaikki"` | | +| `created_at` | timestamptz | NOT NULL, default now() | | -**Key design:** `cefr_level` is on `translations`, not `terms`. "House" in English is A1; "domicile" is also English but B2 — same concept, different words, different difficulty. +UNIQUE (`headword`, `language_code`, `pos`, `sense_index`) · INDEX (`language_code`, `pos`, `difficulty`) + +### `entry_translations` + +| Column | Type | Constraints | +| ---------------------- | ----------- | ----------------------------------------------- | +| `id` | uuid | PK | +| `entry_id` | uuid | FK → `vocabulary_entries.id`, ON DELETE CASCADE | +| `target_language_code` | varchar(10) | CHECK `SUPPORTED_LANGUAGE_CODES` | +| `translation` | text | NOT NULL | +| `sense_hint` | text | nullable | +| `cefr_level` | varchar(2) | nullable, CHECK `A1`–`C2` | +| `difficulty` | varchar(20) | nullable, CHECK `DIFFICULTY_LEVELS` | +| `source` | varchar(50) | NOT NULL, default `"kaikki"` | +| `created_at` | timestamptz | NOT NULL, default now() | + +UNIQUE (`entry_id`, `target_language_code`, `translation`) · INDEX (`target_language_code`, `difficulty`, `entry_id`) --- -### `term_glosses` — Definitions per language +## Target schema (what the new pipeline writes) -| Column | Type | Constraints | Notes | -| --------------- | ---------- | ----------------------------------- | ---------------------- | -| `id` | uuid | PK | | -| `term_id` | uuid | FK → terms.id | | -| `language_code` | varchar(2) | CHECK: `en`, `it`, `de`, `es`, `fr` | | -| `text` | text | | Definition/explanation | -| `created_at` | timestamp | default now() | | +Three levels: a **word** has **senses**, and translations hang off a **sense**, not off the word. That is the point of the redesign — a quiz question is tied to one specific meaning. -**Unique constraint:** (`term_id`, `language_code`) — one gloss per term per language. Prevents left joins from multiplying question rows. +### `words` -**Note:** Italian gloss coverage is sparse (~2% of terms have Italian glosses). UI falls back to English gloss when no gloss exists for the user's language. +| Column | Type | Constraints | +| --------------- | ----------- | -------------------------------- | +| `id` | uuid | PK | +| `headword` | text | NOT NULL | +| `language_code` | varchar(10) | CHECK `SUPPORTED_LANGUAGE_CODES` | +| `pos` | varchar(20) | CHECK `SUPPORTED_POS` | +| `created_at` | timestamptz | NOT NULL, default now() | ---- +UNIQUE `unique_word_per_language_and_pos` (`headword`, `language_code`, `pos`) · INDEX `idx_language_code_pos` -### `decks` — Curated wordlists +### `senses` -| Column | Type | Constraints | Notes | -| --------------------- | ------------ | ------------------------------------------------- | ------------------------------------------------------- | -| `id` | uuid | PK | | -| `name` | varchar | | e.g. `en-core-1000` | -| `source_language` | varchar(2) | CHECK | Language the wordlist was built from | -| `validated_languages` | varchar(2)[] | CHECK: source_language NOT IN validated_languages | Languages with complete translations for all deck terms | -| `description` | text | nullable | | -| `created_at` | timestamp | default now() | | +| Column | Type | Constraints | +| ------------- | ----------- | ----------------------------------- | +| `id` | uuid | PK | +| `word_id` | uuid | FK → `words.id`, ON DELETE CASCADE | +| `sense_index` | smallint | NOT NULL, default 0 | +| `difficulty` | varchar(20) | NOT NULL, CHECK `DIFFICULTY_LEVELS` | +| `definitions` | text[] | NOT NULL, default `[]` | +| `examples` | text[] | NOT NULL, default `[]` | +| `created_at` | timestamptz | NOT NULL, default now() | -**Design:** One deck per frequency tier per source language. POS, difficulty, and category are query filters, not separate decks. Decks must not overlap — each term appears in exactly one tier. +UNIQUE `unique_sense_per_word` (`word_id`, `sense_index`) · INDEX `idx_word_sense_difficulty` -**Source:** SUBTLEX frequency lists (per-language editions, same methodology). +### `translations` ---- +| Column | Type | Constraints | +| ---------------------- | ----------- | -------------------------------------- | +| `id` | uuid | PK | +| `sense_id` | uuid | FK → `senses.id`, ON DELETE CASCADE | +| `target_language_code` | varchar(10) | CHECK `SUPPORTED_LANGUAGE_CODES` | +| `translation` | text | NOT NULL | +| `gender` | varchar(20) | nullable, CHECK NULL or `NOUN_GENDERS` | +| `difficulty` | varchar(20) | NOT NULL, CHECK `DIFFICULTY_LEVELS` | +| `created_at` | timestamptz | NOT NULL, default now() | -### `deck_terms` — Junction table +UNIQUE `unique_translation_per_sense` (`sense_id`, `target_language_code`, `translation`) · INDEX `idx_translations_sense_language_difficulty` -| Column | Type | Constraints | Notes | -| ------------ | --------- | ------------- | ----- | -| `deck_id` | uuid | FK → decks.id | | -| `term_id` | uuid | FK → terms.id | | -| `created_at` | timestamp | default now() | | +**Two difficulty columns, two meanings.** `senses.difficulty` = is this _meaning_ appropriate for the level. `translations.difficulty` = is this _word_ an acceptable answer. Queries use sense difficulty as a ceiling and translation difficulty as the target; a translation's difficulty is never lower than its sense's. See `documentation/pipeline/design-doc.md` §4. -**PK:** (`deck_id`, `term_id`) +**Note:** `DIFFICULTY_LEVELS` is `easy | medium | hard`. `"intermediate"` was renamed to `"medium"` and no longer exists anywhere. --- ## Auth Tables (managed by Better Auth) -Better Auth creates and owns these tables. Do not modify directly. +Better Auth creates and owns `user`, `session`, `account`, and `verification`. Do not modify them directly — changes come from Better Auth config. `user.id` is `text`, not uuid, so foreign keys to it must also be `text`. -### `user` - -| Column | Type | Notes | -| ---------------- | --------- | -------------------- | -| `id` | varchar | PK | -| `name` | varchar | Display name | -| `email` | varchar | | -| `email_verified` | boolean | | -| `image` | varchar | nullable, avatar URL | -| `created_at` | timestamp | | -| `updated_at` | timestamp | | - -### `session` - -| Column | Type | Notes | -| ------------ | --------- | ------------- | -| `id` | varchar | PK | -| `user_id` | varchar | FK → user.id | -| `token` | varchar | Session token | -| `expires_at` | timestamp | | -| `ip_address` | varchar | nullable | -| `user_agent` | text | nullable | -| `created_at` | timestamp | | - -### `account` — Social provider links - -| Column | Type | Notes | -| --------------- | --------- | -------------------- | -| `id` | varchar | PK | -| `user_id` | varchar | FK → user.id | -| `account_id` | varchar | Provider's user ID | -| `provider_id` | varchar | `google` or `github` | -| `access_token` | text | nullable | -| `refresh_token` | text | nullable | -| `id_token` | text | nullable | -| `expires_at` | timestamp | nullable | - -**Note:** One user can have multiple accounts (Google + GitHub linked to same user). - -### `verification` - -Email verification tokens. Unused for social-only auth but managed by Better Auth. +- `user` — `id`, `name`, `email` (unique), `email_verified`, `image`, timestamps +- `session` — `id`, `user_id`, `token`, `expires_at`, `ip_address`, `user_agent` +- `account` — social provider links; one user can have both Google and GitHub +- `verification` — email verification tokens; managed but unused for social-only auth --- @@ -138,84 +125,61 @@ Email verification tokens. Unused for social-only auth but managed by Better Aut ### `lobbies` -| Column | Type | Constraints | Notes | -| ------------- | --------- | ------------------------------------------- | -------------------------------------------- | -| `id` | uuid | PK | | -| `code` | varchar | UNIQUE | Human-readable room code (e.g. `WOLF-42`) | -| `host_id` | varchar | FK → user.id | | -| `status` | varchar | CHECK: `waiting`, `in_progress`, `finished` | | -| `max_players` | integer | default 4 | | -| `settings` | jsonb | nullable | Game mode, round count, timer duration, etc. | -| `created_at` | timestamp | default now() | | -| `updated_at` | timestamp | default now() | Used for stale recovery | +| Column | Type | Constraints | +| -------------- | ----------- | --------------------------------------------------- | +| `id` | uuid | PK | +| `code` | varchar(10) | NOT NULL, UNIQUE — room code | +| `host_user_id` | text | FK → `user.id`, ON DELETE CASCADE | +| `status` | varchar(20) | NOT NULL, default `waiting`, CHECK `LOBBY_STATUSES` | +| `created_at` | timestamptz | NOT NULL, default now() | ### `lobby_players` -| Column | Type | Constraints | Notes | -| -------------- | --------- | --------------- | ---------------------------- | -| `id` | uuid | PK | | -| `lobby_id` | uuid | FK → lobbies.id | | -| `user_id` | varchar | FK → user.id | | -| `display_name` | varchar | | Player's shown name in lobby | -| `is_host` | boolean | default false | | -| `joined_at` | timestamp | default now() | | +| Column | Type | Constraints | +| ----------- | ----------- | ------------------------------------ | +| `lobby_id` | uuid | FK → `lobbies.id`, ON DELETE CASCADE | +| `user_id` | text | FK → `user.id`, ON DELETE CASCADE | +| `score` | integer | NOT NULL, default 0 | +| `joined_at` | timestamptz | NOT NULL, default now() | -**Unique constraint:** (`lobby_id`, `user_id`) — one entry per player per lobby. +**Composite PK:** (`lobby_id`, `user_id`) — no surrogate `id` column, one row per player per lobby. + +Only lobby _membership_ is persisted. Live game state (questions, timers, per-round answers) lives in the in-memory stores in `apps/api`, not in these tables. Max players is the `MAX_LOBBY_PLAYERS` constant in `packages/shared`, not a column. --- ## Key Relationships ``` -terms (1) ←──→ (N) translations -terms (1) ←──→ (N) term_glosses -terms (N) ←──→ (N) decks via deck_terms -user (1) ←──→ (N) sessions -user (1) ←──→ (N) accounts +vocabulary_entries (1) ←──→ (N) entry_translations ← live +words (1) ←──→ (N) senses (1) ←──→ (N) translations ← target +user (1) ←──→ (N) session +user (1) ←──→ (N) account user (1) ←──→ (N) lobbies (as host) -user (1) ←──→ (N) lobby_players -lobbies (1) ←──→ (N) lobby_players +lobbies (1) ←──→ (N) lobby_players (N) ←──→ (1) user ``` --- ## Query Patterns -### Get quiz terms (singleplayer) +All queries live in `packages/db/src/models/` — `apps/api` never imports `drizzle-orm`. -```sql -SELECT t.id, t.pos, src.text AS source_text, tgt.text AS target_text, g.text AS gloss -FROM terms t -JOIN translations src ON src.term_id = t.id AND src.language_code = ? -JOIN translations tgt ON tgt.term_id = t.id AND tgt.language_code = ? -LEFT JOIN term_glosses g ON g.term_id = t.id AND g.language_code = ? -WHERE t.pos = ? AND tgt.cefr_level IN (?) -LIMIT ? -``` +`termModel.ts` queries the live schema: it self-joins `vocabulary_entries` (aliased source and target), joins `entry_translations` for the answer, and fetches distractors with a separate query per question that excludes both the current entry id and the correct answer text — different entries can share a translation string. -### Get distractors +The distractor query is N+1, one round trip per question. Batching is a known BACKLOG item. -```sql -SELECT text FROM translations -WHERE language_code = ? AND pos = ? AND cefr_level IN (?) -AND term_id != ? AND text != ? -ORDER BY RANDOM() -LIMIT 3 -``` - -**Note:** This is the N+1 query mentioned in BACKLOG.md. Each question fetches 3 distractors separately. Batching is planned. +Phase 5 rewrites these against `words`/`senses`/`translations`, which changes the shape: filter on `senses.difficulty` as a ceiling, then select `translations` at the requested difficulty. Target queries are sketched in `documentation/pipeline/design-doc.md` §5. --- ## Deferred Schema Extensions (Not Yet Implemented) -These tables are planned but do not exist yet. All are additive — they reference existing `terms` rows via FK. +Planned, additive, and keyed off the **target** schema: -| Table | Purpose | Trigger | -| --------------------- | ----------------------------------------------- | ----------------------- | -| `noun_forms` | Gender, singular, plural, articles per language | Grammar quiz mode | -| `verb_forms` | Conjugation tables per language | Grammar quiz mode | -| `term_pronunciations` | IPA + audio URLs per language | Pronunciation quiz mode | -| `user_decks` | Which decks a user studies | User customization | -| `user_term_progress` | Spaced repetition state per user/term/language | SRS review queue | -| `quiz_answers` | Answer history for stats/analytics | User stats dashboard | +| Table | Purpose | Trigger | +| -------------------- | ---------------------------------------------- | ----------------------- | +| `inflection_forms` | Gender, plural, conjugation/declension tables | Grammar quiz mode | +| `pronunciations` | IPA + audio URLs per language | Pronunciation quiz mode | +| `user_word_progress` | Spaced repetition state per user/word/language | SRS review queue | +| `quiz_answers` | Answer history for stats/analytics | User stats dashboard | diff --git a/documentation/ai-context/03-api-contract.md b/documentation/ai-context/03-api-contract.md index ab384b3..074ea25 100644 --- a/documentation/ai-context/03-api-contract.md +++ b/documentation/ai-context/03-api-contract.md @@ -33,7 +33,7 @@ POST /api/v1/game/start source_language: SupportedLanguageCode, // "en" | "it" | "de" | "es" | "fr" target_language: SupportedLanguageCode, pos: SupportedPos, // "noun" | "verb" | "adjective" | "adverb" - difficulty: DifficultyLevel, // "easy" | "intermediate" | "hard" + difficulty: DifficultyLevel, // "easy" | "medium" | "hard" rounds: GameRounds // "3" | "10" (string enum, converted to number in service) } ``` diff --git a/documentation/ai-context/05-data-pipeline.md b/documentation/ai-context/05-data-pipeline.md index e0ae12c..50ab136 100644 --- a/documentation/ai-context/05-data-pipeline.md +++ b/documentation/ai-context/05-data-pipeline.md @@ -1,173 +1,151 @@ # 05 — Data Pipeline -> **Purpose:** Condensed reference for LLMs working on the Kaikki data pipeline. Covers stages, data flow, and current blockers. For full operational details (llama.cpp setup, provider configs, hardware specs), see the human-readable DATA_PIPELINE.md. -> **Last updated:** 2026-05-15 -> **Depends on:** 00-project-overview.md +> **Purpose:** Condensed reference for LLMs working on the vocabulary data pipeline. Covers the flow, what exists, and what is still unwritten. Full detail: `documentation/DATA_PIPELINE.md`, `documentation/pipeline/design-doc.md`, `documentation/pipeline/roadmap.md`. +> **Last updated:** 2026-08-01 +> **Depends on:** 00-project-overview.md, 02-data-model.md + +--- + +## Read this first + +The pipeline was **completely rewritten**. The old Kaikki/local-LLM architecture — six stages, `pipeline.db`, `stage-1-extract/`, `stage-3-enrich/`, multi-model CEFR voters, llama.cpp — is **gone from the codebase**. Any reference you see to those stages, directories, or the voter strategy is historical (`documentation/archive/`), not something you can call or modify. + +The current pipeline is Gemini-only, and most of it **is not written yet**. Do not assume a module exists because a doc names it. --- ## Pipeline Overview ``` -Kaikki JSONL (Wiktionary extracts) +source-data/{lang}/{pos} frequency wordlists, one word per line, UTF-8 ↓ -Stage 1: Extract → Parse into pipeline.db (SQLite) +Gemini API batches of 20 words, one language at a time ↓ -Stage 2: Reverse Link → Insert missing reverse translations +validation per entry; rejects go to a log, never to the DB ↓ -Stage 3: Enrich → LLMs review glosses, examples, translations, assign CEFR +db/staging.db SQLite staging (words, senses, translations) ↓ -Stage 4: Merge → Resolve LLM votes into final values +import script SQLite → PostgreSQL via Drizzle, transaction per batch ↓ -Stage 4b: Tiebreak → Run unused models on flagged entries - ↓ -Stage 5: Compare / QA → Generate COVERAGE.md quality report - ↓ -Stage 6: Sync → Upsert resolved records into production PostgreSQL +PostgreSQL dev (:5432), then production ``` -**Current state:** Stage 1 and 2 complete on sample data. Stage 3 enrich script being rewritten for sub-stage architecture. Stages 4–6 not started. +Each language is processed independently so that definitions and examples are written **in that language**. Only translations cross language boundaries. + +The app always reads PostgreSQL. SQLite is a staging file only, so re-runs and prompt tweaks never touch a real database. + +**Current state:** Phases 1–2 complete (schema, wordlists, databases, prompt). Phase 3 in progress — validation module, pipeline script, and first real run are all still to be written. Phases 4–7 not started. --- -## Stage 1: Extract +## What exists in `data-pipeline/` -**Input:** `data-pipeline/stage-1-extract/sources/*.jsonl` (Kaikki files, not in git) -**Output:** `pipeline.db` — `vocabulary_entries` and `entry_translations` tables +| Path | State | +| --------------------------- | ------------------------------------------------------------------ | +| `source-data/{lang}/{pos}/` | ✅ Noun lists for `de`, `en`, `es`, `fr`, `it` | +| `prompt` | ✅ Gemini system prompt — a plain UTF-8 text file, not a TS module | +| `db/schema.sql` | ✅ SQLite staging schema | +| `db/staging.db` | ✅ Tables created, **0 rows**, gitignored | +| `pipeline.ts` | 🚧 Design pseudocode in comments only — no executable code | +| validation module | ❌ Not written (rules in design-doc §6.4) | +| SQLite → PostgreSQL import | ❌ Not written | +| `kaikki-source-files/` | ⚠️ Leftover JSONL dumps; nothing reads them | +| `worddata/english/nouns/` | ⚠️ Empty leftover output dir from the old per-word-JSON design | -**What it does:** +Directory names use the codes from `packages/shared/src/constants.ts` (`de/noun`, not `german/nouns`) so no mapping layer is needed. -- Parses Kaikki JSONL for all 5 languages (en, de, es, fr, it) -- Filters to 4 POS: noun, verb, adjective, adverb -- Each Kaikki sense becomes one `vocabulary_entries` row -- Translations stored in `entry_translations` with sense hints - -**Key design:** Kaikki is structured per word sense. Each headword has multiple senses, and translations are linked to a specific sense. This prevents the sense-disambiguation problems of OpenWordNet/OMW. +`data-pipeline/vitest.config.ts` looks for `tests/**/*.test.ts`; that directory does not exist yet. --- -## Stage 2: Reverse Link Sync +## Gemini output contract -**Pure script, no LLMs.** +The model returns a JSON array, one object per input word, in input order — no markdown fences, comments, or trailing commas. -For each translation pair (e.g., English "thrill" → German "begeistern"), checks if the reverse exists (German "begeistern" → English "thrill"). If the German entry exists but lacks the English back-link, inserts it automatically. + +```json +[ + { + "headword": "Haus", + "language": "de", + "pos": "noun", + "senses": [ + { + "sense_index": 0, + "difficulty": "easy", + "definitions": ["Ein Gebäude zum Wohnen."], + "examples": ["Sie kauften ein Haus in der Stadt."], + "translations": [ + { "target_language": "en", "word": "house", "gender": null, "difficulty": "easy" }, + { "target_language": "es", "word": "casa", "gender": "feminine", "difficulty": "easy" } + ] + } + ] + } +] +``` -**Why:** Ensures LLMs in Stage 3 only generate translations that are genuinely missing — not translations findable by simple reverse lookup. +Prompt rules that the output depends on: + +- Definitions and examples in the **source** language, not English. +- Gender required for `de` (masculine/feminine/neuter) and `it`/`es`/`fr` (masculine/feminine); always `null` for `en`. +- German nouns capitalized; Romance-language nouns lowercase unless proper nouns. +- Base dictionary form, no articles or determiners. +- 1–3 senses per word, most words 1; no rare, archaic, or technical senses. +- Max 2 translations per target language per sense, only genuine synonyms or difficulty variants. +- A translation's difficulty is never lower than its sense's difficulty. +- A word that is not a valid noun in that language returns `"senses": []`. + +⚠️ The checked-in `prompt` file still has hardcoded English leftovers (rules 2, 3, and 31 say `"en"` / "English noun" while the header says Spanish) and its target-language list disagrees with its own header. It is also pinned to one sample batch rather than templated. Fix when implementing `pipeline.ts`. --- -## Stage 3: Enrich (In Progress — Being Rewritten) +## Validation rules (design-doc §6.4) -**Current blocker:** The original single-prompt design had problems (skipped invalid translations, triggered reasoning mode, 20% manual review). Being rewritten as four ordered sub-stages. +Run per entry before anything is written to SQLite. Invalid entries go to a rejection log for review, not to the database. Target reject rate: under 10%. -### Sub-Stage Architecture - -Each model processes every entry through four sub-stages in order: - -1. **`round1_gloss`** — Review existing gloss. Confirm if clear, generate better one if not. -2. **`round1_example`** — Review examples. Confirm if natural, generate one better sentence. -3. **`round1_translations`** — Validate translations with verified gloss as context. Confirm valid, reject invalid, generate missing. -4. **`round1_cefr`** — Assign CEFR level (A1–C2) to headword and each confirmed translation. - -**Why this order:** CEFR sub-stage only sees clean, verified data. Bad translations are rejected before reaching CEFR assignment. - -**Voter strategy:** Multiple models vote independently. Each model = one vote per sub-stage. Current plan: - -- Primary: Local Qwen3.5-9B (overnight runs, unlimited) -- Secondary: Groq Llama 3.3 70B (cloud, batched) -- Tertiary: Gemini AI Studio (cloud, batched) - -**Context enrichment:** Before calling models for gloss/example, pipeline queries Wiktionary API for the headword. Full entry (all senses, usage notes) added to prompt. Fixes category header glosses and short ambiguous glosses. +- `headword` non-empty; `language` in the 5 supported codes; `pos` in the 4 supported values +- at least one sense; each sense has ≥1 definition, ≥1 example, ≥1 translation +- `sense_index` a non-negative integer, starting at 0 and increasing by 1 +- `difficulty` in `easy | medium | hard` on both senses and translations +- `target_language` supported and never equal to the word's own language +- `gender` valid for the target language (see above); `null` for English +- no duplicate (headword, language, pos, sense_index) --- -## Stage 4: Merge +## Constants -Resolves LLM votes into final values per entry. +| Constant | Values | Source | +| ---------- | ------------------------------------- | -------------------------- | +| Languages | `en`, `it`, `de`, `es`, `fr` | `SUPPORTED_LANGUAGE_CODES` | +| POS | `noun`, `verb`, `adjective`, `adverb` | `SUPPORTED_POS` | +| Difficulty | `easy`, `medium`, `hard` | `DIFFICULTY_LEVELS` | +| Gender | `masculine`, `feminine`, `neuter` | `NOUN_GENDERS` | -**Rules:** +All live in `packages/shared/src/constants.ts` and are CHECK-constrained in the PostgreSQL schema. Adding a value means updating the constant **and** a Drizzle migration before re-running the pipeline. -- Kaikki source data wins automatically (never overridden) -- CEFR: level with most votes wins -- Text fields (gloss, example, translation): candidate with most votes wins -- No majority → flag for tiebreaker - -**Difficulty mapping:** -| CEFR | Difficulty | -|------|-----------| -| A1, A2 | easy | -| B1, B2 | intermediate | -| C1, C2 | hard | +CEFR levels still exist as a constant and as columns on the old `vocabulary_entries` tables, but the new pipeline does not produce them — it produces the three-level difficulty directly. The prompt calibrates difficulty against CEFR bands internally (easy ≈ A1/A2, medium ≈ B1/B2, hard ≈ C1/C2) but is explicitly told not to emit CEFR levels. --- -## Stage 4b: Tiebreak +## Running it -Runs automatically after merge if flagged entries remain. Queries unused models (not yet voted) and re-runs merge. Repeats until resolved or no unused models remain. +```bash +docker compose up -d pipeline-database # dedicated PostgreSQL on :5433 +pnpm --filter @lila/pipeline pipeline:run # tsx --env-file .env pipeline.ts (currently a no-op) +pnpm --filter @lila/pipeline test +``` -**If still unresolved:** Sync is blocked. Add more models to config and re-run. +Env vars come from the repo-root `.env`: `GEMINI_API_KEY`, `PIPELINE_POSTGRES_USER`, `PIPELINE_POSTGRES_PASSWORD`, `PIPELINE_POSTGRES_DB`, `PIPELINE_DATABASE_URL`. The pipeline database is deliberately separate from the app database (`:5432`). ---- - -## Stage 5: Compare / QA - -Read-only. Generates `COVERAGE.md` with per-language breakdown: - -- Total entries, POS distribution -- Translation coverage per language pair -- CEFR coverage and difficulty breakdown -- Gloss/example coverage by source (Kaikki vs LLM) -- Per-model contribution stats - -Run this before syncing to production. - ---- - -## Stage 6: Sync - -Upserts all `status = "final"` entries from `pipeline.db` to production PostgreSQL. - -**Behavior:** - -- Missing → insert -- Present but changed → update -- Present and unchanged → skip - -**Idempotent.** Safe to re-run. - ---- - -## Key Constraints - -| Constant | Values | -| ---------- | ------------------------------------- | -| Languages | `en`, `it`, `de`, `es`, `fr` | -| POS | `noun`, `verb`, `adjective`, `adverb` | -| CEFR | `A1`, `A2`, `B1`, `B2`, `C1`, `C2` | -| Difficulty | `easy`, `intermediate`, `hard` | - -Adding a new value requires updating `packages/shared/src/constants.ts` AND a database migration before re-running the pipeline. +Implementation notes from the roadmap: `better-sqlite3` for staging (synchronous), `crypto.randomUUID()` for ids, `JSON.stringify` for the definitions/examples arrays (SQLite has no array type — the import script parses them back into PostgreSQL `text[]`), batches of 20 with a 1s sleep between calls. --- ## Current Blockers -1. **Enrich sub-stage rewrite** — Stage 3 script needs redesign and testing -2. **Cloud provider integration** — Groq and Gemini not yet wired into pipeline -3. **Batching prompt design** — 5–10 entries per API call for efficiency; not yet designed -4. **Full dataset scale unknown** — Currently running on 500-entry samples. Full Kaikki English file has ~1.3M entries. Exact filtered count and runtime estimate not yet known. - ---- - -## Key Files - -| File | Purpose | -| ------------------------------------------------------------ | --------------------------------------------------------- | -| `data-pipeline/pipeline.ts` | Orchestrator — runs stages in order, handles resumability | -| `data-pipeline/stage-1-extract/scripts/extract.ts` | Parse Kaikki JSONL | -| `data-pipeline/stage-2-reverse-link/scripts/reverse-link.ts` | Insert reverse translations | -| `data-pipeline/stage-3-enrich/scripts/enrich.ts` | LLM enrichment (being rewritten) | -| `data-pipeline/stage-3-enrich/config.ts` | Provider configs (local, OpenRouter, etc.) | -| `data-pipeline/db/schema.sql` | pipeline.db schema | -| `data-pipeline/db/import.ts` | Import stage 1 output into pipeline.db | -| `packages/shared/src/constants.ts` | Language codes, POS, CEFR, difficulty constants | +1. **Validation module and `pipeline.ts` are unwritten** — this is Phase 3, the active work. +2. **Prompt is not templated** — source language, POS, target languages, and the word batch are hardcoded for one sample run. +3. **No import script** — nothing moves staging rows into PostgreSQL yet (Phase 4). +4. **App still reads the old schema** — `termModel.ts` queries `vocabulary_entries`/`entry_translations`. Until Phase 5 rewrites it, pipeline output is invisible to the app. diff --git a/documentation/archive/data-pipeline-local-llm.md b/documentation/archive/data-pipeline-local-llm.md new file mode 100644 index 0000000..b7a1bf6 --- /dev/null +++ b/documentation/archive/data-pipeline-local-llm.md @@ -0,0 +1,790 @@ +# Lila Data Pipeline — Technical Documentation + +Multilingual dictionary enrichment pipeline for the Lila vocabulary trainer. +Last updated: 2026-07-18 + +## Table of Contents + +- [1. Executive Summary](#1-executive-summary) +- [2. Problem & Context](#2-problem--context) +- [3. Architecture Overview](#3-architecture-overview) +- [4. Current Implementation](#4-current-implementation) +- [5. The LLM Layer](#5-the-llm-layer) + - 5.1 [Local Model Evaluation](#51-local-model-evaluation) + - 5.2 [Online API Options](#52-online-api-options) + - 5.3 [Model Selection Criteria](#53-model-selection-criteria) +- [6. The Gender Problem & Kaikki Integration](#6-the-gender-problem--kaikki-integration) +- [7. Batching Strategy](#7-batching-strategy) +- [8. Hardware Constraints](#8-hardware-constraints) +- [9. Testing & Quality Assurance](#9-testing--quality-assurance) +- [10. Interactive CLI](#10-interactive-cli) +- [11. Future Extensions & Roadmap](#11-future-extensions--roadmap) +- [12. Decisions Log](#12-decisions-log) +- [13. Known Issues & Dev Notes](#13-known-issues--dev-notes) +- [14. How to Run](#14-how-to-run) +- [15. Roadmap](#15-roadmap) + +## Quick Reference + +| What | Where | +| ------------------- | ------------------------------------------ | +| Entry point | `pipeline.ts` | +| Interactive CLI | `utils/cli.ts` | +| LLM config schema | `config/llm.ts` | +| System prompt | `config/prompt.ts` — `buildSystemPrompt()` | +| Batch config schema | `config/batch.ts` | +| Shared constants | `config/constants.ts` | +| Output schema | `utils/merge-enriched-data.ts` | +| LLM adapters | `utils/llm-adapters/` | +| **Current model** | **`gemma-4-E2B_q4_0-it.gguf`** | +| Target scale | 100,000+ words | + +--- + +## 1. Executive Summary + +The Lila Data Pipeline is a TypeScript-based batch processing system that enriches raw word lists into structured multilingual dictionary entries for the Lila vocabulary trainer. It takes a source wordlist (e.g., English nouns) and, for each word, generates: + +- One or more senses (definitions) +- A natural example sentence per sense +- A CEFR-based difficulty level (`easy` / `medium` / `hard`) +- Translations into all target languages except the source (as raw strings) + +The pipeline is designed to scale to 100,000+ words across multiple languages and parts of speech. It supports both local inference (`llama.cpp`) and cloud providers via a pluggable adapter pattern. + +### Key Design Principles + +| Principle | Rationale | +| ------------------------------- | ------------------------------------------------------------------------------------------------------ | +| **Quality first** | Definitions, examples, and translations must be accurate. Speed and cost are secondary. | +| **Local-first, cloud-fallback** | Local LLMs are the default for cost control and data privacy. | +| **Deterministic Grammar** | Grammatical gender is decoupled from the LLM and resolved via Kaikki Wiktionary dumps. | +| **Resumable & idempotent** | Each word writes to its own JSON file. The pipeline skips already-processed words on restart. | +| **Configurable batching** | Batch size is selected interactively at runtime. The pipeline adapts without code changes. | +| **Honest metrics** | Local models report detailed prompt/completion timing. Cloud providers report total request time only. | + +### Resolved: Gender Accuracy + +Grammatical gender is **no longer generated by the LLM**. Comprehensive testing across 10 models (2026-07-18) confirmed that small local models systematically hallucinate or default to `neuter` for Romance languages. +**Decision:** The LLM only outputs translation strings. A deterministic post-processing step looks up the exact grammatical gender from Kaikki Wiktionary dumps. This guarantees 100% gender accuracy and allows us to use smaller, faster, and highly nuanced local models. + +### Current Status (2026-07-18) + +- **Core pipeline:** Complete (scanning, enrichment, merging, verification, writing). +- **Local LLM:** Gemma 4 E2B selected as production model after exhaustive 10-model evaluation. `llama.cpp` server optimized with KV-cache quantization. +- **Batching:** 20-word batches validated for local hardware. +- **Gender:** Decoupled from LLM; Kaikki lookup architecture confirmed. +- **Pending:** Online API evaluation (Gemini free tier, DeepSeek, Groq) as speed fallback. + +### One-Line Architecture + +```text +source wordlists -> Interactive CLI -> LLM adapter -> merge senses -> Kaikki Gender Lookup -> verify schema -> write .json +``` + +### Files at a Glance + +| File | Purpose | +| ----------------------------------------- | ----------------------------------------------------------------------------------------------------- | +| `pipeline.ts` | Orchestrator. Runs CLI, scans sources, loops words, coordinates all stages | +| `utils/cli.ts` | Interactive CLI. Provider/model/batch selection, config persistence | +| `config/llm.ts` | LLM config schema (provider, url, model) | +| `config/prompt.ts` | `buildSystemPrompt()` — dynamic prompt with auto-target languages | +| `config/batch.ts` | Batch size and max retry count schema | +| `config/constants.ts` | Shared `LANG_MAP`, `POS_MAP`, `ALL_LANGUAGES` | +| `utils/enrich-word.ts` | Calls LLM via adapter, parses response, validates senses, builds `EnrichedSense[]`, retry/split logic | +| `utils/merge-enriched-data.ts` | Merges skeleton + enriched senses into final JSON | +| `utils/verify-enriched-file.ts` | Schema validation (required fields, types) | +| `utils/check-llm-server.ts` | Health check for local server; skipped for cloud providers | +| `utils/scanning-source-files.ts` | Discovers wordlists from `source-data/` directory | +| `utils/create-base-json.ts` | Writes skeleton `{word, language, pos}` files | +| `utils/write-json-file.ts` | Atomic `.tmp` → rename writes | +| `utils/check-if-json-exists.ts` | Resumability: checks if word already has enriched senses | +| `utils/create-line-reader.ts` | Streaming line reader for large wordlists | +| `utils/create-output-dirs.ts` | Creates `worddata/{language}/{pos}/` folders | +| `utils/delete-file.ts` | Cleanup helper for failed batches | +| `utils/get-word-file-path.ts` | Path construction helper | +| `utils/progress-tracker.ts` | `[current/total]` formatting for console output | +| `utils/pipeline-timer.ts` | Timing + token metrics; unified throughput for all providers | +| `utils/llm-adapters/factory.ts` | Creates the right adapter based on runtime config | +| `utils/llm-adapters/types.ts` | `LlmAdapter` interface | +| `utils/llm-adapters/openai-compatible.ts` | Local llama.cpp, OpenRouter, DeepSeek | +| `utils/llm-adapters/gemini.ts` | Google Gemini native API | + +### Scale Target + +| Metric | Target | +| --------------- | ------------------------------------------------------------ | +| Words | 100,000+ | +| Languages | English (source), German, Italian, Spanish, French (targets) | +| Parts of speech | Nouns, verbs, adjectives, adverbs | +| Output | One `.json` file per word, ~2-5KB each | + +--- + +## 2. Problem & Context + +### Why Build This? + +Existing multilingual dictionaries and translation APIs provide raw word-to-word mappings. They do not provide the structured, pedagogical data needed for a vocabulary trainer: + +| What Exists | What is Missing | Why It Matters | +| ------------------------ | ---------------------------------- | ------------------------------------------------------------- | +| Word + translation | Student-friendly definition | Learners need explanations, not just equivalents | +| Static difficulty labels | CEFR-calibrated difficulty | "Easy" vs "hard" must map to learner proficiency levels | +| Isolated translations | Natural example sentences | Context is how vocabulary is actually acquired | +| Raw gender data | Gender integrated with translation | Flashcards must show gender immediately (resolved via Kaikki) | +| Bulk word lists | Per-word structured JSON | The trainer consumes one file per word for fast random access | + +### The Target User + +A language learner using the Lila vocabulary trainer. They see a word, its definition, an example sentence, and translations with gender — all calibrated to their CEFR level (A1-C2). + +### Why Not Use Existing Dictionaries? + +- **Wiktionary:** Rich data, but unstructured, inconsistent formatting, no CEFR levels, no student-friendly definitions +- **Kaikki (Wiktionary dump):** Structured JSON, excellent for gender/translation lookup, but definitions are often technical, no difficulty classification, no example curation +- **Google Translate / DeepL:** No definitions, no examples, no difficulty, no structured output +- **Existing language learning apps:** Closed data, no export, no control over content + +The LLM fills the gap: it generates pedagogical content (student-friendly definitions, natural examples, difficulty classification) that no existing database provides at scale. + +### Language Direction + +The pipeline is direction-agnostic. A wordlist is defined by: + +- **Source language:** the language of the input words +- **Target languages:** all other languages in the system (auto-derived from `ALL_LANGUAGES` minus source) + +Current focus: English -> German/Italian/Spanish/French. Planned directions include German -> French, Italian -> Spanish, etc. + +### Why 100,000+ Words? + +- **Coverage:** A learner needs ~10,000 words for B2 proficiency. The pipeline targets 100k to cover multiple languages, POS categories, and difficulty levels. +- **Languages:** English (source) -> German, Italian, Spanish, French (targets). +- **Parts of speech:** Nouns (current), verbs, adjectives, adverbs. + +### The Quality Challenge + +| Risk | Mitigation | +| ------------------------------- | -------------------------------------------------------------- | +| Hallucinated definitions | Low temperature (0.1), strict system prompt, schema validation | +| Incorrect grammatical gender | **Resolved:** Decoupled from LLM; Kaikki Wiktionary lookup | +| POS bleed (verb defs for nouns) | Explicit negative constraint in system prompt | +| Inconsistent difficulty levels | Explicit CEFR mapping in prompt, spot-checking | +| JSON parse failures | Retry + split logic, schema validation, cleanup on failure | +| Model drift (online APIs) | Version pinning, local fallback | +| Provider downtime | Adapter pattern allows hot-swapping providers | + +### Why TypeScript + Node? + +- **Familiarity:** Existing project uses TypeScript (frontend in TanStack Router + React) +- **Ecosystem:** `readline` for streaming files, `fs` for JSON I/O, native `fetch` for HTTP +- **Portability:** Runs on the same Debian laptop as the llama.cpp server +- **No build complexity:** `tsx` for direct execution, no bundler needed + +### Why llama.cpp? + +- **GGUF format:** Single-file models, easy to swap, quantize, and version +- **OpenAI-compatible API:** `/v1/chat/completions` means the same adapter code works for local and online models +- **No dependencies:** Self-contained binary, runs on old hardware (tested on GTX 950M) +- **Privacy:** Local inference means no data leaves the machine + +--- + +## 3. Architecture Overview + +### Pipeline Flow + +```text +Run CLI -> Scan sources -> Check LLM -> Loop wordlists -> Stream words -> Skip processed + -> Create skeletons (batch) -> Call LLM -> Parse JSON -> Validate senses + -> Retry/split on failure -> Merge -> Kaikki Gender Lookup -> Write atomically -> Verify schema -> Log metrics +``` + +### Resumability + +- **Skip existing:** `check-if-json-exists.ts` checks if `{word}.json` exists with non-empty `senses` +- **Atomic writes:** `.tmp` -> rename in `write-json-file.ts`, no partial files on crash +- **Cleanup on failure:** Deletes partially-written files for failed batches, continues to next batch + +### Directory Structure + +```text +data-pipeline/ + |-- pipeline.ts # Entry point / orchestrator + |-- utils/ + | |-- cli.ts # Interactive CLI module + | |-- enrich-word.ts # LLM call, parse, retry/split + | |-- merge-enriched-data.ts # Schema types + merge logic + | |-- verify-enriched-file.ts # Schema validation + | |-- check-llm-server.ts # Health check (local only) + | |-- scanning-source-files.ts # Source discovery + | |-- create-base-json.ts # Skeleton writer + | |-- write-json-file.ts # Atomic JSON writer + | |-- check-if-json-exists.ts # Resumability check + | |-- create-line-reader.ts # Streaming file reader + | |-- create-output-dirs.ts # Directory creation + | |-- delete-file.ts # Cleanup helper + | |-- get-word-file-path.ts # Path helper + | |-- progress-tracker.ts # Console progress formatting + | |-- pipeline-timer.ts # Timing + token metrics + | |-- llm-adapters/ + | |-- factory.ts # Adapter selection + | |-- types.ts # LlmAdapter interface + | |-- openai-compatible.ts # Local, OpenRouter, DeepSeek + | |-- gemini.ts # Google Gemini + |-- config/ + | |-- llm.ts # LLM config schema + | |-- prompt.ts # buildSystemPrompt() + | |-- batch.ts # Batch config schema + | |-- constants.ts # LANG_MAP, POS_MAP, ALL_LANGUAGES + |-- source-data/ + | |-- {language}/ + | |-- {pos} # One word per line, no extension + |-- worddata/ + | |-- {language}/ + | |-- {pos}/ + | |-- {word}.json # One self-contained file per word + |-- kaikki-source-files/ # Wiktionary dumps for gender lookup + |-- .pipeline-config.json # Saved CLI configuration +``` + +### Output Schema + +Each `.json` file contains: `word`, `language`, `pos`, `senses[]` (each with `id`, `sense`, `example`, `difficulty_level`, `translations` per target language), `enrichedAt`, `model`. +_Note: The `translations` object contains arrays of strings (e.g., `{"de": ["Haus", "Gebäude"]}`). Grammatical gender is appended later via the Kaikki integration step._ + +### Error Handling + +| Failure | Behavior | +| ------------------------------ | -------------------------------------------------------- | +| LLM server offline | Hard fail at startup (`check-llm-server.ts`, local only) | +| LLM returns bad JSON | Retry up to 3 times, then split batch. Log and continue | +| LLM returns malformed senses | `validateSense()` catches it before file write | +| Schema validation fails | Log warnings, keep file | +| Individual batch fails | Does not stop pipeline; cleans up partial files | +| Individual word fails (size 1) | Log and continue to next word | + +### Metrics + +Per-run: words processed/skipped/failed, duration, throughput, LLM token counts and speeds. See `utils/pipeline-timer.ts`. + +- **Unified throughput (all providers):** total tokens / total request time +- **Detailed breakdown (local only):** prompt speed vs completion speed + +--- + +## 4. Current Implementation + +### Tech Stack + +| Layer | Choice | Why | +| ----------- | ----------------------------- | ------------------------------------------ | +| Runtime | Node.js + `tsx` | Direct TypeScript execution, no build step | +| HTTP client | Native `fetch` | Works for local llama.cpp and online APIs | +| File I/O | `fs` + `readline` | Streaming line reader for large wordlists | +| JSON | Native `JSON.parse/stringify` | Simple, no schema library needed | +| CLI | Native `readline` | No external dependencies | + +### Current Model + +| Property | Value | +| ---------------- | ---------------------------------------------------- | +| **Model** | **`gemma-4-E2B_q4_0-it.gguf`** | +| Size | ~3.2GB (file) / ~2.06GB (VRAM weights) | +| Quantization | Q4_0 | +| Server | `llama.cpp` (`llama-server`) | +| API | OpenAI-compatible `/v1/chat/completions` | +| VRAM Usage | ~2.65GB total (weights + KV cache + compute buffers) | +| Generation Speed | ~10.9 tok/s (20-word batch) | +| Est. 100k Time | ~7 days (20-word batches, 24/7) | + +### `llama-server` Flags: History & Rationale + +The server flags evolved through rigorous empirical testing on the target hardware (Intel i7-6500U, GTX 950M 4GB, 8GB RAM) across 10 different models on 2026-07-18. + +#### Flag Evolution + +| Flag | Value Tried | Result | Why | +| ----------------- | ------------------------------------------ | --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ | +| `-m` | `qwen3.5-4b-q4_k_m.gguf` | Works, ~6.3 tok/s | Quality baseline. Correct translations. 2.6GB, tight on VRAM. | +| `-m` | `qwen2.5-1.5b-instruct-q4_k_m.gguf` | Works, ~18.5 tok/s | Fast but poor instruction following, gender wrong. | +| `-m` | `Qwen3.5-2B-Q4_K_M.gguf` | Works, ~13.2 tok/s | Good translations, but failed polysemy (2 identical senses). | +| `-m` | `Llama-3.2-3B-Instruct-Q4_K_M.gguf` | Works, ~9.1 tok/s | Dangerous false friend trap (cognates). | +| `-m` | `Llama-3.2-3B-Instruct-UD-Q6_K_XL.gguf` | Works, ~6.7 tok/s, 3.44GB VRAM | Higher quant did NOT fix Llama's translation issues. Slower. | +| `-m` | `Ministral-3-3B-Instruct-2512-Q4_K_M.gguf` | Works, ~8.9 tok/s | Fixed tokenizer bug (older version was broken). But messy translations, markdown violations. | +| `-m` | `gemma-4-E2B_q4_0-it.gguf` | **Works, ~13.1 tok/s, 2.06GB** | **Winner.** Perfect polysemy, false friends, nuance. Half VRAM of Qwen 4B. | +| `-m` | `gemma-4-E4B-it-qat-UD-Q4_K_XL.gguf` | Works, ~7.7 tok/s, 3.23GB | QAT compression is incredible. Same quality as E2B but slower. | +| `-m` | `qwen3.5-9b-q3_k_s.gguf` | Works, ~3.3 tok/s (split CPU/GPU) | Brilliant quality but bottlenecked by CPU/Swap. `-ngl 28` max. | +| `-m` | `qwen3.5-9b-q4_k_m.gguf` | Works, ~3.3 tok/s (split CPU/GPU) | Same quality as 9B Q3. Not worth 2x the file size. | +| `-ngl` | 999 | Keeps | Offload all layers to GPU. Required for any speed. | +| `-ngl` | 28 | 9B models only | Max GPU layers for 9B models before OOM. Found via binary search. | +| `-ngl` | 30 | OOM crash (9B) | Pushed 2 layers too far into compute buffers. | +| `-c` | 4096 | Wasteful | 4K context for 300-token dictionary entries wastes VRAM. | +| `-c` | 2048 | Good for small batches | Sufficient for 4-word batches. | +| `-c` | **8192** | **Current** | Required for 20-word batches. Combined with KV cache quantization. | +| `-b` / `-ub` | 512 | **Current** | Sweet spot for Maxwell memory bandwidth. | +| `-b` / `-ub` | 1024 | Tested | Slightly faster prompt processing, but no generation speedup. | +| `-b` / `-ub` | 2048 | Slower on 950M | Memory pressure on bandwidth-starved GPU. | +| `-t` | 4 | Slower | Hyperthreading cores hurt llama.cpp performance. | +| `-t` | **2** | **Current** | Matches 2 physical cores. | +| `--threads-batch` | **2** | **Current** | Explicit match to `-t`. | +| `--flash-attn` | (omitted) | Correct | On Maxwell (compute 5.0), Flash Attention adds overhead. | +| `--mlock` | Tested | Omitted for large models | Pins model in RAM. Causes OOM on models >2.5GB with large KV cache. | +| `--prio` | **2** | **Current** | Raises process priority. Marginal, harmless. | +| `--reasoning` | **off** | **Critical** | **Mandatory for Qwen 3.5 and Gemma 4.** Without this, models "think" silently, consume all `max_tokens`, and crash with `finish_reason: length`. | +| `--cache-type-k` | **q4_0** | **Current** | Compresses KV cache keys to 4-bit. Cuts KV VRAM by ~75%. Enables 8192 context on 4GB GPU. | +| `--cache-type-v` | **q4_0** | **Current** | Compresses KV cache values to 4-bit. Paradoxically improves translation variety (reduces "lazy duplication" bug). | + +#### Current Production Command + +```bash +cd ~/Downloads/llama.cpp +./build/bin/llama-server \ + -m models/gemma-4-E2B_q4_0-it.gguf \ + -ngl 999 \ + -c 8192 \ + -b 512 \ + -ub 512 \ + -t 2 \ + --threads-batch 2 \ + --host 127.0.0.1 \ + --port 8080 \ + --prio 2 \ + --reasoning off \ + --cache-type-k q4_0 \ + --cache-type-v q4_0 +``` + +#### VRAM Budget (Production Config) + +| Component | VRAM Usage | +| ------------------------------------ | ------------------------- | +| Model Weights (Gemma 4 E2B Q4_0) | ~1.50 GB | +| KV Cache (8192 ctx, q4_0 compressed) | ~0.35 GB | +| Compute Buffers (batch 512) | ~0.20 GB | +| **Total** | **~2.05 GB (51% of 4GB)** | + +#### Why Not Use the Remaining 2GB VRAM? + +Generation speed is bottlenecked by **memory bandwidth** (~32 GB/s on GTX 950M), not VRAM capacity. To generate one token, the GPU must read the entire ~1.5GB model from VRAM. The theoretical maximum is ~21 tok/s. At 10.9 tok/s, the GPU is already operating at ~50% of its physical limit. Empty VRAM cannot be converted into faster generation. + +Testing uncompressed `f16` KV cache (2.65GB total VRAM) yielded 12.6 tok/s but caused a severe "lazy duplication" regression (model copy-pasted the same translation twice instead of providing distinct synonyms). The `q4_0` compressed KV cache is the correct choice for translation quality. + +### Performance Baseline (20-Word Batch) + +| Metric | Gemma 4 E2B | Qwen 3.5 4B | +| --------------------- | ----------- | ------------ | +| Time/batch (20 words) | ~5 min | ~9 min | +| Completion tok/s | ~10.9 | ~6.1 | +| Prompt tok/s | ~132 | ~78 | +| VRAM Usage | 2.06 GB | 3.95 GB | +| Lazy Duplication Bug | No | Yes (severe) | + +--- + +## 5. The LLM Layer + +### 5.1 Local Model Evaluation (Complete — 2026-07-18) + +All 10 downloaded models were evaluated on the target hardware. Testing progressed from 4-word smoke tests to a 20-word "nightmare" torture suite covering extreme polysemy, false friends, abstract concepts, and legal/financial terminology. + +#### Final Leaderboard + +| Rank | Model | Size | VRAM | Speed (tok/s) | Polysemy | False Friends | Verdict | +| ------ | ------------------------ | ---- | ------------- | ------------- | --------------------------- | ------------------------- | ----------------------------------------------------------- | +| **🥇** | **Gemma 4 E2B Q4_0** | 3.2G | **2.06 GB** | **~13.1** | **Perfect** | **Perfect** | **Production model.** Best quality/speed/VRAM ratio. | +| 🥈 | Qwen 3.5 4B Q4_K_M | 2.6G | ~3.6 GB | ~6.3 | Perfect | Perfect | Quality King, but 2x slower and maxes VRAM. | +| 🥉 | Gemma 4 E4B Q4_K_XL | 4.0G | 3.23 GB | ~7.7 | Perfect | Perfect | Incredible QAT compression. Same quality as E2B but slower. | +| 4 | Qwen 3.5 2B Q4_K_M | 1.2G | ~1.6G | ~13.2 | Failed (2 identical senses) | Passed | Good translations, lacks conceptual branching. | +| 5 | Qwen 2.5 1.5B Q4_K_M | 1.1G | ~1.5G | ~18.5 | Ignored instruction | Failed | Fast but easily confused. | +| 6 | Ministral 3B 2512 Q4_K_M | 2.0G | ~2.7G | ~8.9 | Messy translations | Failed | Fixed tokenizer, but outclassed. Markdown violations. | +| 7 | Llama 3.2 3B Q4_K_M | 1.9G | ~2.6G | ~9.1 | Good structure, bad IT/ES | **Failed (cognate trap)** | Dangerous for language learners. | +| 8 | Llama 3.2 3B Q6_K_XL | 2.8G | 3.44G | ~6.7 | Good structure, bad IT/ES | **Failed (cognate trap)** | Higher quant did NOT fix translation issues. | +| 9 | Qwen 3.5 9B Q3_K_S | 4.1G | Split CPU/GPU | ~3.3 | Perfect | Perfect | Brilliant but 35-40 days for 100k words. | +| 10 | Qwen 3.5 9B Q4_K_M | 5.3G | Split CPU/GPU | ~3.3 | Perfect | Perfect | Same quality as 9B Q3. Not worth the size. | + +#### Key Findings + +1. **The "Thinking" Trap:** Both Qwen 3.5 and Gemma 4 have built-in Chain-of-Thought reasoning. Without `--reasoning off`, they silently "think" in a hidden JSON field, consume all `max_tokens`, and crash with `finish_reason: length`. This flag is **mandatory**. +2. **The 2B vs 4B Quality Cliff:** 2B models struggle with polysemy (e.g., cannot distinguish "bank" = financial vs river). 4B+ models act like professional lexicographers. +3. **Llama 3.2 is Unsafe for Language Learners:** Consistently fell for false friend traps (e.g., translating "actual" = _real_ to _aktuell/attuale/actual/actuel_ = _current_). +4. **KV Cache Quantization Improves Translation Variety:** Compressing the KV cache to `q4_0` introduces microscopic noise that prevents the "lazy duplication" bug (model copy-pasting the same synonym twice). +5. **POS Bleed is Universal:** All models occasionally generate verb definitions for nouns (e.g., "run" = _to move fast_ instead of _a jogging session_). Fix: explicit negative constraint in the system prompt. + +### 5.2 Online API Options + +Evaluated as fallbacks if local models fail quality or speed targets. + +| Provider | Model | Input $/1M | Output $/1M | Free Tier | Rate Limit | Est. Cost (100k words) | Est. Time | +| --------------- | -------------------- | ---------- | ----------- | ------------- | ---------- | ---------------------- | ------------------- | +| DeepSeek | V4 Flash | $0.14 | $0.28 | 5M tokens | None | $0-15 | ~1-2 days | +| Gemini | 2.5 Flash-Lite | $0.10 | $0.40 | 1,500 req/day | 1M TPM | $0 (free tier) | ~1.5 days (batched) | +| Qwen/Alibaba | Qwen-Turbo | $0.05 | $0.20 | Unknown | 600 RPM | ~$11 | ~1-2 days | +| Groq | Llama 3.1 8B Instant | $0.05 | $0.08 | Yes | High | ~$7 | ~3-4 hours | +| OpenRouter free | Various | $0 | $0 | 200 req/day | 20 RPM | $0 | ~10 days | + +### 5.3 Model Selection Criteria + +Decision flow for 100,000 words: + +```text +Start + | + v + Gemma 4 E2B (local) — SELECTED + | + |-- Speed acceptable? (~7 days) -----> Use Gemma 4 E2B locally, $0 + | + |-- Need faster? -------> Test Gemini 2.5 Flash-Lite (free) + | + |-- Quality good? --> Batch 50, free tier + | ~1.5 days, $0 + | + |-- Quality meh? ---> Test Groq or DeepSeek paid +``` + +Quality gates: + +- 100% JSON parse rate +- No hallucinated definitions on polysemous words +- Natural, contextually appropriate example sentences +- Sensible difficulty classification (CEFR mapping) +- Gender accuracy is no longer an LLM criterion (handled by Kaikki) + +--- + +## 6. The Gender Problem & Kaikki Integration + +### The Problem (Resolved) + +Grammatical gender was originally embedded in the LLM's `translations` object. Testing across all 10 models confirmed that small local models systematically hallucinate gender, defaulting to `neuter` for Romance languages (Italian, Spanish, French) which do not have a neuter grammatical gender. + +### The Solution: Decoupled Architecture + +**Decision (2026-07-18):** Grammatical gender is no longer generated by the LLM. The pipeline now uses a two-stage approach: + +| Stage | Component | Responsibility | +| ----- | ----------------- | -------------------------------------------------------------------- | +| 1 | LLM (Gemma 4 E2B) | Generates translation **strings only** (e.g., `["Haus", "Gebäude"]`) | +| 2 | Kaikki Lookup | Deterministically resolves grammatical gender from Wiktionary dumps | + +### Benefits + +- **100% deterministic gender accuracy** — no hallucination possible +- **Faster LLM generation** — ~15-20% fewer output tokens per word +- **Simpler JSON schema** — translations are string arrays, not object arrays +- **Model-agnostic** — works with any LLM regardless of multilingual training quality + +### Kaikki Data + +| Language | Kaikki File | Coverage | +| -------- | ------------------------------------- | -------- | +| German | `kaikki.org-dictionary-German.jsonl` | High | +| Italian | `kaikki.org-dictionary-Italian.jsonl` | High | +| Spanish | `kaikki.org-dictionary-Spanish.jsonl` | High | +| French | `kaikki.org-dictionary-French.jsonl` | High | + +### Lookup Logic + +Match on `word` (the translated string) -> extract `gender` field -> map to `"masculine" | "feminine" | "neuter" | null`. + +--- + +## 7. Batching Strategy + +### Why Batching is Necessary + +At 1 word per call, 100,000 words = 100,000 LLM requests. Each call re-processes the ~300-token system prompt. Batching amortizes this cost. + +### Optimal Batch Size (Local) + +**20 words** is the validated sweet spot for the GTX 950M with Gemma 4 E2B. + +| Batch Size | VRAM | Speed | Quality | Verdict | +| ---------- | --------------- | ---------------- | ------------------------ | -------------------------------------------------- | +| 1-4 | ~2.1 GB | ~13 tok/s | Perfect | Safe but slow (amortization waste) | +| **20** | **~2.6 GB** | **~10.9 tok/s** | **Perfect** | **Sweet spot** | +| 30-40 | ~3.0 GB (est.) | ~10 tok/s (est.) | Likely good | Worth testing | +| 50+ | ~3.5 GB+ (est.) | Unknown | Risk of JSON degradation | **Not recommended for 2B models** | +| 100 | OOM risk | N/A | Attention degradation | Small models lose JSON structure past ~6000 tokens | + +### Retry & Split Strategy + +If a batch fails (bad JSON, missing key, etc.): + +```text +Batch of 20 fails (3 retries exhausted) + | + v +Split into 2 batches of 10 + | + v +If a batch of 10 fails (3 retries), split into 2 batches of 5 + | + v +If a batch of 5 fails, split into batches of 1 + | + v +If a single word fails (3 retries), log and skip +``` + +--- + +## 8. Hardware Constraints + +### Current Machine + +| Component | Spec | +| ------------- | ------------------------------------------------------------------- | +| OS | Debian GNU/Linux 13 (trixie) x86_64 | +| CPU | Intel Core i7-6500U (2 physical cores, 4 threads via HT) @ 3.10 GHz | +| GPU | NVIDIA GeForce GTX 950M (Maxwell, 2015) | +| GPU VRAM | 4GB (4037 MiB reported by CUDA) | +| GPU Bandwidth | ~32 GB/s (DDR3) | +| RAM | 8GB (~3.67GB usable at idle) | +| Swap | 5.62 GB | +| Disk | 102GB ext4 (~74GB used) | + +### What Fits in 4GB VRAM (Empirically Verified) + +| Model | File Size | Total VRAM | Fits? | Notes | +| ------------------------ | --------- | ----------- | ---------- | -------------------------------------- | +| Qwen 2.5 1.5B Q4_K_M | 1.1G | ~1.5 GB | ✅ Easy | | +| Qwen 3.5 2B Q4_K_M | 1.2G | ~1.6 GB | ✅ Easy | | +| Ministral 3B 2512 Q4_K_M | 2.0G | ~2.7 GB | ✅ Yes | | +| Llama 3.2 3B Q4_K_M | 1.9G | ~2.6 GB | ✅ Yes | | +| Llama 3.2 3B Q6_K_XL | 2.8G | 3.44 GB | ✅ Tight | | +| **Gemma 4 E2B Q4_0** | **3.2G** | **2.06 GB** | **✅ Yes** | **QAT compression. Production model.** | +| Gemma 4 E4B Q4_K_XL | 4.0G | 3.23 GB | ✅ Yes | QAT compression is incredible. | +| Qwen 3.5 4B Q4_K_M | 2.6G | 3.95 GB | ⚠️ Barely | 50MB headroom with 8192 ctx. | +| Qwen 3.5 9B Q3_K_S | 4.1G | Split | ⚠️ Partial | `-ngl 28` max. Rest on CPU/Swap. | +| Qwen 3.5 9B Q4_K_M | 5.3G | Split | ⚠️ Partial | `-ngl 20` max. Heavy swap usage. | + +### GPU Rental Alternatives + +If local hardware becomes the bottleneck: +| Provider | GPU | VRAM | Price/Hour | Time for 100k Words | Total Cost | +| --- | --- | --- | --- | --- | --- | +| Vast.ai | RTX 4090 | 24GB | ~$0.30-0.60 | ~6-8 hours | ~$2-5 | +| RunPod | RTX 4090 | 24GB | ~$0.50-0.80 | ~6-8 hours | ~$4-6 | + +--- + +## 9. Testing & Quality Assurance + +### 20-Word Torture Suite (Completed 2026-07-18) + +Tested on Gemma 4 E2B and Qwen 3.5 4B with 20 challenging nouns: + +| Category | Words | The Trap | +| ----------------- | ------------------------------------------------------- | -------------------------------------------------------------- | +| Extreme Polysemy | `match`, `date`, `right`, `set`, `well` | Does it translate "match" as fire, sports, or dating? | +| False Friends | `sense`, `fabric`, `sympathy`, `eventuality`, `billion` | "Fabric" = material (_tissu_) not factory (_fabrique_) | +| Abstract/Cultural | `serendipity`, `accountability` | Concepts lacking 1:1 dictionary equivalents | +| Action-Nouns | `run`, `drive`, `play` | "Run" as jog vs tear vs campaign | +| Legal/Financial | `mortgage`, `lease`, `court`, `board`, `draft` | Requires specific legal vocabulary (_Hypothek/mutuo/hipoteca_) | + +### Results Summary + +| Criterion | Gemma 4 E2B | Qwen 3.5 4B | +| ------------------------------------ | ------------------------------------------ | ------------------------------------------------------- | +| JSON Reliability | 10/10 (raw JSON) | 10/10 (raw JSON, but added `sense_index` hallucination) | +| Polysemy (bank, match) | 10/10 (Ufer/riva/orilla/rive) | 10/10 | +| False Friends (fabric) | 10/10 (Stoff/tessuto/tela/tissu) | 10/10 | +| Legal Nuance (mortgage) | 10/10 (Hypothek/mutuo/hipoteca/hypothèque) | 10/10 | +| Lazy Duplication Bug | None | **Severe** (copy-pasted same word 20+ times) | +| POS Bleed | Minor (verb defs for run/match) | Minor (verb defs for run/match) | +| Attention Degradation (word 20 vs 1) | None | None | + +### Verification + +`verify-enriched-file.ts` checks: + +- Required top-level fields: `word`, `language`, `pos`, `senses` +- Each sense: `sense` (string), `example` (string), `difficulty_level` in {easy, medium, hard} +- Each translation: array of strings (gender appended later by Kaikki) + +--- + +## 10. Interactive CLI + +### Batch Size Recommendations (Updated) + +| Provider | Recommended | Rationale | +| -------------------- | ----------- | ----------------------------------------------- | +| **local (GTX 950M)** | **20** | Validated sweet spot. 8192 ctx + q4_0 KV cache. | +| local (RTX 4090) | 50 | Fast, more VRAM | +| gemini | 50 | Free tier: 1,500 req/day | +| deepseek | 20 | 5M free tokens | +| groq | 50 | Very fast | + +--- + +## 11. Future Extensions & Roadmap + +### Near-Term (Next 2-4 Weeks) + +| Item | Status | Notes | +| ----------------------------------- | ------------ | --------------------------------------------------- | +| 10-model evaluation | **Complete** | Gemma 4 E2B selected | +| KV cache quantization | **Complete** | `--cache-type-k/v q4_0` enables 8192 ctx on 4GB GPU | +| Gender decoupling | **Complete** | Kaikki lookup replaces LLM gender | +| 20-word torture suite | **Complete** | Validated on Gemma E2B and Qwen 4B | +| POS bleed fix | **Pending** | Add negative constraint to system prompt | +| Kaikki gender lookup implementation | **Pending** | Post-processing step after LLM enrichment | +| Online API testing | **Pending** | Gemini free tier, DeepSeek, Groq | + +### Medium-Term (1-3 Months) + +| Item | Notes | +| ---------------------------- | ------------------------------------------------------ | +| Multi-POS support | Verbs, adjectives, adverbs need prompt variants | +| Multi-language source | German -> French, Italian -> Spanish, etc. | +| Parallel wordlist processing | Run `english/nouns` and `english/verbs` simultaneously | +| Incremental enrichment | Only process new/changed words in a wordlist | + +### Long-Term (3-6 Months) + +| Item | Notes | +| ------------------------ | ---------------------------------------------------------------- | +| GPU rental integration | Script to spin up Vast.ai/RunPod, run pipeline, download results | +| Quality regression tests | Run torture suite on every model change | +| Community contributions | Open-source the pipeline for other language learners | + +--- + +## 12. Decisions Log + +| Date | Decision | Context | Rationale | +| -------------- | ----------------------------------- | --------------------------------------------------------------- | ------------------------------------------------------------------------------------------------- | +| 2026-01-04 | TanStack Router for frontend | Previous project used React Router | Simpler, type-safe routing | +| 2026-06-16 | llama.cpp for local LLM | Need local inference on old laptop | GGUF format, OpenAI-compatible API | +| 2026-06-16 | Q4_K_M quantization | Balance size vs quality | Community standard for 4-bit | +| 2026-06-16 | `-t 2` physical cores | Default was 4 (HT threads) | Hyperthreading hurts llama.cpp | +| 2026-06-17 | Qwen2.5-1.5B as initial model | Qwen3.5-4B too slow (47s/word) | 6x speedup, quality under evaluation | +| 2026-07-06 | Adapter pattern for LLM providers | Need to evaluate local vs cloud | `utils/llm-adapters/` with factory | +| 2026-07-06 | Retry + split batching | LLM JSON parse failures | 3 retries, then halve batch | +| 2026-07-06 | Interactive CLI | Editing config files is error-prone | `utils/cli.ts` with native readline | +| **2026-07-18** | **Gemma 4 E2B as production model** | **10-model evaluation completed** | **2x faster than Qwen 4B, half VRAM, perfect translation quality** | +| **2026-07-18** | **Gender decoupled from LLM** | **All 10 models failed gender for Romance languages** | **Kaikki Wiktionary lookup is deterministic and 100% accurate** | +| **2026-07-18** | **`--reasoning off` is mandatory** | **Qwen 3.5 and Gemma 4 "think" silently, consuming all tokens** | **Without this flag, output crashes with `finish_reason: length`** | +| **2026-07-18** | **KV cache quantization (`q4_0`)** | **8192 context needed for 20-word batches** | **Cuts KV VRAM by 75%, enables large batches on 4GB GPU, improves translation variety** | +| **2026-07-18** | **20-word batch size for local** | **Tested 4, 20 words** | **Sweet spot: no attention degradation, no JSON breakage, 10.9 tok/s** | +| **2026-07-18** | **Llama 3.2 discarded** | **Failed false friend tests** | **Translates "actual" (real) to cognates (aktuell/attuale) = "current". Dangerous for learners.** | +| **2026-07-18** | **`-c 8192` replaces `-c 2048`** | **20-word batches need more context** | **Combined with q4_0 KV cache, fits in 2.65GB VRAM** | + +--- + +## 13. Known Issues & Dev Notes + +### Data Pipeline + +| Issue | Details | Severity | +| ------------------------------ | -------------------------------------------------------------------------------------- | ----------------------------------------------- | +| POS bleed (all models) | Models occasionally generate verb definitions for nouns (e.g., "run" = _to move fast_) | Medium — fix with negative constraint in prompt | +| Lazy duplication (Qwen 4B) | Qwen 3.5 4B copy-pastes the same translation twice to fill arrays | Medium — use Gemma 4 E2B instead | +| Schema hallucination (Qwen 4B) | Adds `sense_index` field not in schema | Low — ignorable | +| Pre-scanning wordlists | Entire file read into memory before processing | Medium — streaming refactor planned | +| Single POS tested | Only nouns validated. Verbs/adjectives need prompt changes | Known limitation | + +### Hardware + +| Issue | Details | +| --------------------- | ---------------------------------------------------------------------------- | +| GTX 950M VRAM ceiling | 4GB hard limit. Models >3.5GB need KV cache quantization or CPU offloading. | +| Maxwell GPU aging | No Flash Attention, bandwidth-starved (~32 GB/s). Theoretical max ~21 tok/s. | +| Laptop thermals | Cannot run 24/7 for weeks unattended. Monitor temps. | + +--- + +## 14. How to Run + +### Prerequisites + +- Node.js + npm +- `tsx` installed globally: `npm install -g tsx` +- llama.cpp built from source (for local mode) +- GGUF model downloaded to `~/Downloads/llama.cpp/models/` +- API keys set as environment variables (for cloud mode) + +### Start the LLM Server (Local Mode) + +```bash +cd ~/Downloads/llama.cpp +./build/bin/llama-server \ + -m models/gemma-4-E2B_q4_0-it.gguf \ + -ngl 999 \ + -c 8192 \ + -b 512 \ + -ub 512 \ + -t 2 \ + --threads-batch 2 \ + --host 127.0.0.1 \ + --port 8080 \ + --prio 2 \ + --reasoning off \ + --cache-type-k q4_0 \ + --cache-type-v q4_0 +``` + +### Run the Pipeline + +```bash +cd /path/to/data-pipeline +npx tsx pipeline.ts +``` + +Follow the interactive prompts to select provider, model, and batch size. + +--- + +## 15. Roadmap + +### Phase 1: Batching (Complete) + +| Task | Status | +| --------------------------------------- | ------------ | +| Implement configurable batch size | Complete | +| Implement retry + split logic | Complete | +| Honest timing metrics | Complete | +| Validate LLM responses | Complete | +| Verify batching quality (20-word suite) | **Complete** | + +### Phase 2: Interactive CLI (Complete) + +| Task | Status | +| --------------------- | -------- | +| Design prompt flow | Complete | +| Implement CLI module | Complete | +| Save/load config | Complete | +| Wire into pipeline.ts | Complete | + +### Phase 3: Model Selection (Complete) + +| Task | Status | +| ----------------------- | -------------------------- | +| 10-model evaluation | **Complete** | +| 20-word torture suite | **Complete** | +| Select production model | **Complete (Gemma 4 E2B)** | +| Test online APIs | Pending | + +### Phase 4: Scale + +| Task | Status | Notes | +| ------------------------------ | ------- | ------------------------------------------ | +| Implement Kaikki gender lookup | Pending | Post-processing step | +| Fix POS bleed in prompt | Pending | Add negative constraint | +| Run 100k word pipeline | Pending | ~7 days local (Gemma E2B, 20-word batches) | +| Spot-check output quality | Pending | Random sample of 100 entries | + +### Phase 5: Extend + +| Task | Status | +| ---------------------------- | ------- | +| Multi-POS support | Pending | +| Multi-language source | Pending | +| Parallel wordlist processing | Pending | +| GPU rental integration | Pending | +| Quality regression tests | Pending | diff --git a/documentation/LLM_SETUP.md b/documentation/archive/llm-setup-local.md similarity index 100% rename from documentation/LLM_SETUP.md rename to documentation/archive/llm-setup-local.md diff --git a/documentation/MODEL_STRATEGY.md b/documentation/archive/model-strategy-cefr-voters.md similarity index 100% rename from documentation/MODEL_STRATEGY.md rename to documentation/archive/model-strategy-cefr-voters.md