617 lines
24 KiB
Markdown
617 lines
24 KiB
Markdown
# Vocabulary Trainer — Data Schema Design Document
|
||
|
||
> **Project:** PERN-stack vocabulary trainer with Gemini-powered data pipeline
|
||
> **Author:** [Your Name]
|
||
> **Date:** July 2026
|
||
> **Status:** Approved — ready for implementation
|
||
|
||
---
|
||
|
||
## 1. Overview
|
||
|
||
This document describes the database schema, data pipeline, and query
|
||
patterns for the vocabulary trainer application. The app supports 5
|
||
languages (English, German, Italian, French, Spanish) and tests learners
|
||
by showing a word with a definition and example sentence, then asking
|
||
them to pick the correct translation from 4 options (1 correct, 3
|
||
distractors).
|
||
|
||
The previous data source (OpenWordNet / kaikki.org) produced
|
||
low-quality and inaccurate entries. The new approach uses a batch
|
||
pipeline: word frequency lists are fed to the Gemini API in groups of
|
||
20, which generates structured metadata (definitions, examples,
|
||
translations, difficulty levels). The output is validated, stored in a
|
||
local SQLite staging database, and mirrored into the production
|
||
Postgres database.
|
||
|
||
---
|
||
|
||
## 2. Database Choice
|
||
|
||
**Postgres** (production and development) with **SQLite** (pipeline
|
||
staging).
|
||
|
||
### Why Postgres
|
||
|
||
- The data is inherently relational: words → senses → translations.
|
||
Foreign keys enforce referential integrity at the database level.
|
||
- The query pattern (filter → join → random → limit) is exactly what
|
||
SQL is designed for.
|
||
- Postgres provides JSONB and native arrays for semi-structured fields
|
||
(definitions, examples, inflection tags) without sacrificing
|
||
relational structure.
|
||
- ACID transactions ensure batch imports are atomic.
|
||
- Already part of the PERN stack. No additional infrastructure.
|
||
|
||
### Why not MongoDB
|
||
|
||
- The data has real, meaningful relationships (not arbitrary nested
|
||
documents). The distractor query ("exclude translations from the
|
||
same sense") is a single `WHERE sense_id != X` in SQL but requires
|
||
a complex aggregation pipeline in MongoDB.
|
||
- No foreign key enforcement. Data integrity would depend entirely on
|
||
application code — risky with LLM-generated data.
|
||
- Future inflection tables (one word → 30–50 forms) are a natural
|
||
relational fit, not a document-store fit.
|
||
|
||
### Why not DynamoDB
|
||
|
||
- Designed for simple key-value lookups at massive scale (billions of
|
||
rows). Cannot do ad-hoc filtering, joins, or `ORDER BY RANDOM()`.
|
||
- The query pattern (filter by language + pos + difficulty, then
|
||
randomize) would require pre-building indexes for every combination.
|
||
- Massive overkill for 500k words.
|
||
|
||
### Why SQLite for staging
|
||
|
||
- Zero-config, file-based. Ideal for the single-writer batch pipeline.
|
||
- The pipeline writes to SQLite, a separate import script mirrors the
|
||
data into Postgres. The application (dev and prod) always reads
|
||
from Postgres to avoid SQLite/Postgres dialect differences.
|
||
|
||
---
|
||
|
||
## 3. Schema
|
||
|
||
### 3.1 Entity Relationship
|
||
|
||
```
|
||
┌─────────────┐ ┌─────────────┐ ┌──────────────────┐
|
||
│ words │ │ senses │ │ translations │
|
||
├─────────────┤ ├─────────────┤ ├──────────────────┤
|
||
│ id (PK) │──┐ │ id (PK) │──┐ │ id (PK) │
|
||
│ headword │ └───>│ word_id(FK) │ └───>│ sense_id (FK) │
|
||
│ language_code│ │ sense_index │ │ target_lang_code │
|
||
│ pos │ │ difficulty │ │ translation │
|
||
│ │ │ cefr_level │ │ gender │
|
||
│ │ │ definitions │ │ difficulty │
|
||
│ │ │ examples │ │ │
|
||
└─────────────┘ └─────────────┘ └──────────────────┘
|
||
|
||
Future (not yet implemented):
|
||
┌──────────────────┐
|
||
│ inflection_forms │
|
||
├──────────────────┤
|
||
│ id (PK) │
|
||
│ word_id (FK) ────────> words.id
|
||
│ form │
|
||
│ tags (JSONB) │
|
||
└──────────────────┘
|
||
```
|
||
|
||
### 3.2 Table: `words`
|
||
|
||
One row per unique word in a specific language.
|
||
|
||
| Column | Type | Constraints | Notes |
|
||
| ------------- | ----------- | --------------------------------- | ------------------------------------- |
|
||
| id | UUID | PK, default random | |
|
||
| headword | TEXT | NOT NULL | "Haus", "casa", "house" |
|
||
| language_code | VARCHAR(10) | NOT NULL, CHECK in supported list | "de", "es", "en", "fr", "it" |
|
||
| pos | VARCHAR(20) | NOT NULL, CHECK in supported list | "noun", "verb", "adjective", "adverb" |
|
||
| created_at | TIMESTAMPTZ | NOT NULL, default now() | |
|
||
|
||
**Constraints:**
|
||
|
||
- `UNIQUE (headword, language_code, pos)` — prevents duplicate entries.
|
||
- `CHECK (language_code IN ('en','de','it','fr','es'))`
|
||
- `CHECK (pos IN ('noun','verb','adjective','adverb'))`
|
||
|
||
**Index:**
|
||
|
||
- `idx_words_lang_pos ON (language_code, pos)` — accelerates the
|
||
primary game query filter.
|
||
|
||
**Design note:** Each language gets its own headword entries. "Haus"
|
||
is a German word row. "casa" is a Spanish word row. They are separate
|
||
entries, linked through the translations table. This is what enables
|
||
any language pair as source/target.
|
||
|
||
### 3.3 Table: `senses`
|
||
|
||
One row per distinct meaning of a word. This is where polysemy is
|
||
handled: "bank" (financial institution) and "bank" (river edge) are
|
||
two senses of one word.
|
||
|
||
| Column | Type | Constraints | Notes |
|
||
| ----------- | ----------- | -------------------------------- | -------------------------------------------- |
|
||
| id | UUID | PK, default random | |
|
||
| word_id | UUID | FK → words.id, ON DELETE CASCADE | |
|
||
| sense_index | SMALLINT | NOT NULL, default 0 | 0 = primary meaning, 1 = secondary, etc. |
|
||
| difficulty | VARCHAR(20) | NOT NULL, CHECK in allowed list | "easy", "medium", "hard" |
|
||
| cefr_level | VARCHAR(2) | nullable, CHECK in allowed list | "A1","A2","B1","B2","C1","C2" |
|
||
| definitions | TEXT[] | NOT NULL, default '{}' | 1–3 definitions in the word's language |
|
||
| examples | TEXT[] | NOT NULL, default '{}' | 1–3 example sentences in the word's language |
|
||
| created_at | TIMESTAMPTZ | NOT NULL, default now() | |
|
||
|
||
**Constraints:**
|
||
|
||
- `UNIQUE (word_id, sense_index)` — one sense per index per word.
|
||
- `CHECK (difficulty IN ('easy','medium','hard'))`
|
||
- `CHECK (cefr_level IS NULL OR cefr_level IN ('A1','A2','B1','B2','C1','C2'))`
|
||
|
||
**CEFR → difficulty mapping:**
|
||
|
||
- A1, A2 → easy
|
||
- B1, B2 → medium
|
||
- C1, C2 → hard
|
||
|
||
**Index:**
|
||
|
||
- `idx_senses_word_diff ON (word_id, difficulty)` — accelerates the
|
||
join from words and the difficulty filter.
|
||
|
||
**Design note — definitions and examples as arrays:**
|
||
Definitions and examples are stored as `TEXT[]` arrays on the sense
|
||
row rather than in separate tables. Rationale:
|
||
|
||
- Each sense has at most 2–3 definitions/examples (1-to-few).
|
||
- They are always fetched together with the sense (no independent
|
||
querying needed).
|
||
- Separate tables would add 2 JOINs to the hottest query for no
|
||
practical benefit.
|
||
- The exercise generator picks one definition and one example
|
||
randomly in application code:
|
||
`arr[Math.floor(Math.random() * arr.length)]`.
|
||
|
||
### 3.4 Table: `translations`
|
||
|
||
One row per translation of a sense into another language. A single
|
||
sense can have multiple translations into the same language at
|
||
different difficulty levels (e.g., "Bank" easy, "Geldinstitut" medium).
|
||
|
||
| Column | Type | Constraints | Notes |
|
||
| -------------------- | ----------- | --------------------------------- | ------------------------------------- |
|
||
| id | UUID | PK, default random | |
|
||
| sense_id | UUID | FK → senses.id, ON DELETE CASCADE | |
|
||
| target_language_code | VARCHAR(10) | NOT NULL, CHECK in supported list | Language of the translation |
|
||
| translation | TEXT | NOT NULL | "casa", "Haus", "maison" |
|
||
| gender | VARCHAR(20) | nullable | "masculine","feminine","neuter", NULL |
|
||
| difficulty | VARCHAR(20) | NOT NULL, CHECK in allowed list | Can differ from sense difficulty |
|
||
| created_at | TIMESTAMPTZ | NOT NULL, default now() | |
|
||
|
||
**Constraints:**
|
||
|
||
- `UNIQUE (sense_id, target_language_code, translation)` — allows
|
||
multiple translations per language (synonyms) but no exact
|
||
duplicates.
|
||
- `CHECK (target_language_code IN ('en','de','it','fr','es'))`
|
||
- `CHECK (difficulty IN ('easy','medium','hard'))`
|
||
- `CHECK (gender IS NULL OR gender IN ('masculine','feminine','neuter'))`
|
||
|
||
**Index:**
|
||
|
||
- `idx_translations_sense_lang_diff ON (sense_id, target_language_code, difficulty)`
|
||
— accelerates the join from senses and the language/difficulty filter.
|
||
|
||
**Design note — gender as a real column:**
|
||
Gender is stored as a column, not embedded in a JSON blob, because
|
||
future gender exercises will need to filter and group by gender.
|
||
English nouns have `gender = NULL`.
|
||
|
||
### 3.5 Future Table: `inflection_forms` (not yet implemented)
|
||
|
||
Will be added when verb conjugation and adjective declension exercises
|
||
are built.
|
||
|
||
| Column | Type | Constraints | Notes |
|
||
| ------- | ----- | -------------------------------- | ------------------------------------- |
|
||
| id | UUID | PK, default random | |
|
||
| word_id | UUID | FK → words.id, ON DELETE CASCADE | |
|
||
| form | TEXT | NOT NULL | "Häuser", "ginge", "grüner" |
|
||
| tags | JSONB | NOT NULL, default '{}' | {"case":"genitive","number":"plural"} |
|
||
|
||
This table is the primary reason the schema is normalized rather than
|
||
using JSONB documents. A German verb has 30–50 inflected forms. An
|
||
adjective has 12+. These are one-to-many relationships that require
|
||
their own table with a foreign key.
|
||
|
||
---
|
||
|
||
## 4. Difficulty Model
|
||
|
||
There are two difficulty columns. They answer different questions.
|
||
|
||
### `senses.difficulty` — "Is this meaning appropriate for the level?"
|
||
|
||
Controls which _meanings_ of a word are shown. A beginner should not
|
||
be tested on "Haus = noble dynasty" because the concept itself is
|
||
advanced, regardless of how hard the translation word is.
|
||
|
||
### `translations.difficulty` — "Is this word an appropriate answer?"
|
||
|
||
Controls which _translation word_ is the correct answer. The concept
|
||
of "bank" is easy, but "Geldinstitut" is a harder word than "Bank"
|
||
for that same concept.
|
||
|
||
### Query filter logic: sense as ceiling, translation as target
|
||
|
||
```sql
|
||
WHERE s.difficulty IN ('easy', 'medium') -- sense at or below level
|
||
AND t.difficulty = 'medium' -- translation exactly at level
|
||
```
|
||
|
||
- The sense difficulty acts as a **ceiling**: don't show meanings
|
||
harder than the selected level.
|
||
- The translation difficulty acts as the **target**: test the learner
|
||
on a word at exactly this level.
|
||
|
||
This ensures that easy concepts with medium-level synonyms (e.g.,
|
||
"bank" → "Geldinstitut") are reachable at the medium level, while
|
||
advanced concepts (e.g., "Haus" → "dynasty") remain gated.
|
||
|
||
### Example data
|
||
|
||
| Word | Sense | Sense Diff. | Translation | Trans. Diff. |
|
||
| ---- | --------------- | ----------- | ------------------- | ------------ |
|
||
| Haus | building | easy | casa (es) | easy |
|
||
| Haus | noble dynasty | hard | dinastía (es) | hard |
|
||
| bank | financial inst. | easy | Bank (de) | easy |
|
||
| bank | financial inst. | easy | Geldinstitut (de) | medium |
|
||
| bank | financial inst. | easy | Kreditinstitut (de) | hard |
|
||
| bank | river edge | medium | Ufer (de) | medium |
|
||
|
||
At **medium** level, the query returns: Geldinstitut, Ufer.
|
||
At **hard** level: Kreditinstitut, dinastía.
|
||
"Bank" (easy translation) never appears at medium/hard as a correct
|
||
answer. "dinastía" (hard sense) never appears at easy/medium.
|
||
|
||
---
|
||
|
||
## 5. Query Patterns
|
||
|
||
### 5.1 Game query — get N random words
|
||
|
||
Scenario: user picks German → Spanish, nouns, medium, 20 rounds.
|
||
|
||
```sql
|
||
SELECT
|
||
w.id AS word_id,
|
||
w.headword,
|
||
s.id AS sense_id,
|
||
s.definitions,
|
||
s.examples,
|
||
t.translation,
|
||
t.gender
|
||
FROM words w
|
||
INNER JOIN senses s
|
||
ON s.word_id = w.id
|
||
INNER JOIN translations t
|
||
ON t.sense_id = s.id
|
||
WHERE w.language_code = 'de'
|
||
AND w.pos = 'noun'
|
||
AND s.difficulty IN ('easy', 'medium')
|
||
AND t.target_language_code = 'es'
|
||
AND t.difficulty = 'medium'
|
||
ORDER BY RANDOM()
|
||
LIMIT 20;
|
||
```
|
||
|
||
The application then picks one random definition and one random
|
||
example from the arrays for each word.
|
||
|
||
### 5.2 Distractor query — get 3 wrong answers
|
||
|
||
For a given correct answer, fetch 3 distractors from the same
|
||
language, pos, and difficulty pool.
|
||
|
||
```sql
|
||
SELECT t.translation, t.gender
|
||
FROM translations t
|
||
INNER JOIN senses s
|
||
ON t.sense_id = s.id
|
||
INNER JOIN words w
|
||
ON s.word_id = w.id
|
||
WHERE w.language_code = 'de'
|
||
AND w.pos = 'noun'
|
||
AND s.difficulty IN ('easy', 'medium')
|
||
AND t.target_language_code = 'es'
|
||
AND t.difficulty = 'medium'
|
||
AND t.sense_id != :current_sense_id
|
||
AND t.translation != :correct_answer
|
||
ORDER BY RANDOM()
|
||
LIMIT 3;
|
||
```
|
||
|
||
### 5.3 Distractor exclusion rule
|
||
|
||
**Distractors must come from a different sense than the correct
|
||
answer.** Not just a different word — a different sense.
|
||
|
||
Rationale: multiple translations of the same sense are all valid
|
||
answers. "Bank" and "Geldinstitut" are both correct translations of
|
||
the financial-institution sense. Showing one as a distractor for the
|
||
other would confuse the learner and break trust.
|
||
|
||
The `sense_id != :current_sense_id` filter excludes all synonyms of
|
||
the same sense in one condition. No synonym table needed.
|
||
|
||
Translations from a _different_ sense of the same word are valid
|
||
distractors (e.g., "Ufer" from the river-bank sense is a fine
|
||
distractor for the financial-institution sense — the definition makes
|
||
it clearly wrong).
|
||
|
||
### 5.4 Edge case: identical translation text across senses
|
||
|
||
Two different senses of different words may share the same translation
|
||
text (e.g., "Schloss" = castle and "Schloss" = lock). The
|
||
`t.translation != :correct_answer` filter handles this by excluding
|
||
the exact text regardless of sense.
|
||
|
||
---
|
||
|
||
## 6. Data Pipeline
|
||
|
||
### 6.1 Flow
|
||
|
||
```
|
||
Word frequency lists (per language, per POS)
|
||
│
|
||
▼
|
||
Gemini API (batches of 20 words)
|
||
│
|
||
▼
|
||
Validation script (reject/flag bad entries)
|
||
│
|
||
▼
|
||
SQLite staging database (local file)
|
||
│
|
||
▼
|
||
Import script (SQLite → Postgres, batch inserts)
|
||
│
|
||
▼
|
||
Postgres (dev) → test full game flow
|
||
│
|
||
▼
|
||
Postgres (prod) via Drizzle migration
|
||
```
|
||
|
||
### 6.2 Wordlist source
|
||
|
||
Frequency-based word lists, one per language. Example: "1000 most
|
||
common German nouns." Sources: Leipzig Corpora, Wiktionary frequency
|
||
lists, or similar open-source frequency data.
|
||
|
||
Each language is processed independently. This ensures language-native
|
||
definitions and examples (a German word gets a German definition, not
|
||
a translated English one).
|
||
|
||
### 6.3 Gemini output JSON contract
|
||
|
||
The API is prompted to return an array of objects. Expected shape:
|
||
|
||
```json
|
||
[
|
||
{
|
||
"headword": "Haus",
|
||
"language": "de",
|
||
"pos": "noun",
|
||
"senses": [
|
||
{
|
||
"sense_index": 0,
|
||
"cefr_level": "A1",
|
||
"difficulty": "easy",
|
||
"definitions": ["Ein Gebäude zum Wohnen."],
|
||
"examples": ["Sie kauften ein Haus in der Stadt."],
|
||
"translations": [
|
||
{
|
||
"target_language": "en",
|
||
"word": "house",
|
||
"gender": null,
|
||
"difficulty": "easy"
|
||
},
|
||
{
|
||
"target_language": "es",
|
||
"word": "casa",
|
||
"gender": "feminine",
|
||
"difficulty": "easy"
|
||
},
|
||
{
|
||
"target_language": "fr",
|
||
"word": "maison",
|
||
"gender": "feminine",
|
||
"difficulty": "easy"
|
||
},
|
||
{
|
||
"target_language": "it",
|
||
"word": "casa",
|
||
"gender": "feminine",
|
||
"difficulty": "easy"
|
||
}
|
||
]
|
||
},
|
||
{
|
||
"sense_index": 1,
|
||
"cefr_level": "C1",
|
||
"difficulty": "hard",
|
||
"definitions": ["Ein Adelsgeschlecht, eine Dynastie."],
|
||
"examples": ["Das Haus der Merowinger herrschte über Franken."],
|
||
"translations": [
|
||
{
|
||
"target_language": "en",
|
||
"word": "house",
|
||
"gender": null,
|
||
"difficulty": "hard"
|
||
},
|
||
{
|
||
"target_language": "es",
|
||
"word": "dinastía",
|
||
"gender": "feminine",
|
||
"difficulty": "hard"
|
||
}
|
||
]
|
||
}
|
||
]
|
||
}
|
||
]
|
||
```
|
||
|
||
### 6.4 Validation rules
|
||
|
||
Before writing to SQLite, every entry is checked:
|
||
|
||
- `headword`, `language`, `pos` are present and valid.
|
||
- At least one sense per word.
|
||
- Each sense has at least one definition and one example.
|
||
- `difficulty` is one of: `easy`, `medium`, `hard`.
|
||
- `cefr_level` is one of: `A1`, `A2`, `B1`, `B2`, `C1`, `C2`.
|
||
- CEFR → difficulty mapping is consistent.
|
||
- `gender` is valid for the target language:
|
||
- German: masculine, feminine, neuter
|
||
- French, Spanish, Italian: masculine, feminine
|
||
- English: null
|
||
- Translations exist for at least the 4 other supported languages.
|
||
- No duplicate entries (headword + language + pos + sense_index).
|
||
|
||
Invalid entries are logged and excluded. They can be reviewed and
|
||
re-processed manually.
|
||
|
||
### 6.5 Import: SQLite → Postgres
|
||
|
||
A Node.js script reads from SQLite (via `better-sqlite3`) and
|
||
batch-inserts into Postgres (via Drizzle). Inserts are wrapped in
|
||
transactions per batch (20 words) for atomicity.
|
||
|
||
The application (dev and prod) always reads from Postgres. SQLite is
|
||
used only as a pipeline staging file.
|
||
|
||
---
|
||
|
||
## 7. Indexes
|
||
|
||
Three indexes cover the game and distractor queries:
|
||
|
||
```sql
|
||
CREATE INDEX idx_words_lang_pos
|
||
ON words (language_code, pos);
|
||
|
||
CREATE INDEX idx_senses_word_diff
|
||
ON senses (word_id, difficulty);
|
||
|
||
CREATE INDEX idx_translations_sense_lang_diff
|
||
ON translations (sense_id, target_language_code, difficulty);
|
||
```
|
||
|
||
---
|
||
|
||
## 8. Performance
|
||
|
||
### Estimated data volume at target scale
|
||
|
||
| Table | Rows | Derivation |
|
||
| ------------ | ----- | -------------------------------- |
|
||
| words | 500k | ~100k per language × 5 languages |
|
||
| senses | ~750k | ~1.5 senses per word average |
|
||
| translations | ~3M | ~4 translations per sense |
|
||
|
||
### Query performance
|
||
|
||
| Scale | Game query (LIMIT 20) | Distractor query (LIMIT 3) |
|
||
| ---------- | --------------------- | -------------------------- |
|
||
| 10k words | < 10 ms | < 10 ms |
|
||
| 500k words | 20–80 ms | 15–60 ms |
|
||
| 5M words | 100–300 ms | 80–200 ms |
|
||
|
||
The bottleneck at scale is `ORDER BY RANDOM()`, which sorts the
|
||
entire filtered result set before applying LIMIT. At 500k words, the
|
||
filtered set per query is ~5k–15k rows — well within comfortable
|
||
range.
|
||
|
||
**Future optimization** (if filtered sets exceed ~100k rows):
|
||
|
||
```sql
|
||
WHERE ... AND random() < 0.05 -- pre-filter to ~5% of rows
|
||
ORDER BY RANDOM()
|
||
LIMIT 20;
|
||
```
|
||
|
||
Or use `TABLESAMPLE`. Not needed at current scale.
|
||
|
||
### Import performance
|
||
|
||
| Method | Time for 3.5M rows |
|
||
| ---------------------------- | ------------------ |
|
||
| Individual INSERT | ~30–60 min |
|
||
| Batch INSERT (1000 per stmt) | ~2–5 min |
|
||
| Postgres COPY (CSV) | ~10–30 sec |
|
||
|
||
The pipeline uses batch inserts via Drizzle. The full import is a
|
||
one-time or occasional operation.
|
||
|
||
---
|
||
|
||
## 9. Implementation Plan
|
||
|
||
```
|
||
1. ✅ Design doc (this document)
|
||
2. Acquire word frequency lists (5 languages, nouns first)
|
||
3. Build + test Gemini prompt (5 words → 20 words → full batch)
|
||
4. Validation script (Gemini output → clean JSON)
|
||
5. SQLite staging schema + pipeline write
|
||
6. Import script (SQLite → local Postgres)
|
||
7. Drizzle schema: words, senses, translations + indexes
|
||
8. Drizzle migration on local Postgres
|
||
9. Dev branch: new queries (game + distractor), full game flow test
|
||
10. Drizzle migration on prod Postgres + data import + verify
|
||
11. Extend pipeline: verbs, adjectives, adverbs (schema unchanged)
|
||
12. Later: inflection_forms table + conjugation/declension exercises
|
||
```
|
||
|
||
---
|
||
|
||
## 10. Future Extensions
|
||
|
||
### Verb conjugation / adjective declension exercises
|
||
|
||
The `inflection_forms` table (section 3.5) will store inflected forms
|
||
with grammatical tags as JSONB. The schema is normalized specifically
|
||
to support this: one word → many forms, each independently queryable.
|
||
|
||
### Gender exercises
|
||
|
||
The `gender` column on `translations` enables filtering and grouping
|
||
by grammatical gender for dedicated gender practice rounds.
|
||
|
||
### Additional POS
|
||
|
||
The `pos` column already supports noun, verb, adjective, adverb.
|
||
Adding a new POS requires no schema change — only a new wordlist and
|
||
an adjusted Gemini prompt.
|
||
|
||
---
|
||
|
||
## 11. Key Design Decisions — Summary
|
||
|
||
| Decision | Choice | Rationale |
|
||
| ----------------------- | ------------------------------------------------------------ | ------------------------------------------------------------------------------------ |
|
||
| Database | Postgres + SQLite staging | Relational data, FK integrity, SQL query pattern, already in stack |
|
||
| Schema structure | 3 normalized tables | Matches query pattern, supports any language pair, extensible for inflections |
|
||
| Definitions / examples | `TEXT[]` arrays on `senses` | 1-to-few relationship, always fetched with sense, avoids 2 extra JOINs |
|
||
| Gender | Column on `translations` | Needed for filtering in future gender exercises |
|
||
| Difficulty | Two columns: `senses.difficulty` + `translations.difficulty` | Sense = concept gate, translation = word-level target |
|
||
| Difficulty filter logic | Sense as ceiling, translation as exact match | Ensures easy concepts with hard synonyms are reachable; advanced concepts stay gated |
|
||
| Distractor exclusion | `sense_id != current` | Prevents valid synonyms from appearing as wrong answers |
|
||
| Language direction | Any of 5 languages as source or target | Each language has its own headword entries; translations link them |
|
||
| Pipeline | Gemini → validate → SQLite → Postgres | Batch-generated data, LLM output needs validation, SQLite for staging simplicity |
|
||
| Table-per-language/pos | Rejected | Anti-pattern: 40+ tables, exponential maintenance |
|
||
| Single JSONB blob | Rejected | Cannot support inflection tables, cannot index gender, no FK integrity |
|