lila/documentation/ai-context/02-data-model.md

11 KiB
Raw Blame History

02 — Data Model

Purpose: Database schema reference for LLMs working on features that query or modify data. Concatenate with 00-project-overview.md and 99-current-task.md. Last updated: 2026-08-01 Depends on: 00-project-overview.md Source of truth: packages/db/src/db/schema.ts. If this file and the schema disagree, the schema wins.


Two vocabulary schemas exist right now

The database is mid-migration and contains both vocabulary schemas. This is the most important thing to know before writing a query.

Schema Status
vocabulary_entries + entry_translations Live. What the app queries today (packages/db/src/models/termModel.ts).
wordssensestranslations Target. Migrated and empty. The new pipeline writes here; no app code reads it yet.

The terms / term_glosses / decks / deck_terms tables described in earlier versions of this doc no longer exist.

Migration path: the pipeline fills words/senses/translations, then termModel.ts is rewritten against it (roadmap Phase 5), then the vocabulary_entries tables are dropped.


Live schema (what the app queries)

vocabulary_entries — one row per word sense

Column Type Constraints Notes
id uuid PK, default random
headword text NOT NULL The word itself
language_code varchar(10) CHECK SUPPORTED_LANGUAGE_CODES
pos varchar(20) CHECK SUPPORTED_POS
sense_index smallint NOT NULL, default 0 Distinguishes senses of the same headword
gloss text nullable Definition
examples text[] NOT NULL, default []
cefr_level varchar(2) nullable, CHECK A1C2
difficulty varchar(20) nullable, CHECK DIFFICULTY_LEVELS
source varchar(50) NOT NULL, default "kaikki"
created_at timestamptz NOT NULL, default now()

UNIQUE (headword, language_code, pos, sense_index) · INDEX (language_code, pos, difficulty)

entry_translations

Column Type Constraints
id uuid PK
entry_id uuid FK → vocabulary_entries.id, ON DELETE CASCADE
target_language_code varchar(10) CHECK SUPPORTED_LANGUAGE_CODES
translation text NOT NULL
sense_hint text nullable
cefr_level varchar(2) nullable, CHECK A1C2
difficulty varchar(20) nullable, CHECK DIFFICULTY_LEVELS
source varchar(50) NOT NULL, default "kaikki"
created_at timestamptz NOT NULL, default now()

UNIQUE (entry_id, target_language_code, translation) · INDEX (target_language_code, difficulty, entry_id)


Target schema (what the new pipeline writes)

Three levels: a word has senses, and translations hang off a sense, not off the word. That is the point of the redesign — a quiz question is tied to one specific meaning.

words

Column Type Constraints
id uuid PK
headword text NOT NULL
language_code varchar(10) CHECK SUPPORTED_LANGUAGE_CODES
pos varchar(20) CHECK SUPPORTED_POS
created_at timestamptz NOT NULL, default now()

UNIQUE unique_word_per_language_and_pos (headword, language_code, pos) · INDEX idx_language_code_pos

senses

Column Type Constraints
id uuid PK
word_id uuid FK → words.id, ON DELETE CASCADE
sense_index smallint NOT NULL, default 0
difficulty varchar(20) NOT NULL, CHECK DIFFICULTY_LEVELS
definitions text[] NOT NULL, default []
examples text[] NOT NULL, default []
created_at timestamptz NOT NULL, default now()

UNIQUE unique_sense_per_word (word_id, sense_index) · INDEX idx_word_sense_difficulty

translations

Column Type Constraints
id uuid PK
sense_id uuid FK → senses.id, ON DELETE CASCADE
target_language_code varchar(10) CHECK SUPPORTED_LANGUAGE_CODES
translation text NOT NULL
gender varchar(20) nullable, CHECK NULL or NOUN_GENDERS
difficulty varchar(20) NOT NULL, CHECK DIFFICULTY_LEVELS
created_at timestamptz NOT NULL, default now()

UNIQUE unique_translation_per_sense (sense_id, target_language_code, translation) · INDEX idx_translations_sense_language_difficulty

Two difficulty columns, two meanings. senses.difficulty = is this meaning appropriate for the level. translations.difficulty = is this word an acceptable answer. Queries use sense difficulty as a ceiling and translation difficulty as the target; a translation's difficulty is never lower than its sense's. See documentation/pipeline/design-doc.md §4.

Note: DIFFICULTY_LEVELS is easy | medium | hard. "intermediate" was renamed to "medium" and no longer exists anywhere.


Auth Tables (managed by Better Auth)

Better Auth creates and owns user, session, account, and verification. Do not modify them directly — changes come from Better Auth config. user.id is text, not uuid, so foreign keys to it must also be text.

  • userid, name, email (unique), email_verified, image, timestamps
  • sessionid, user_id, token, expires_at, ip_address, user_agent
  • account — social provider links; one user can have both Google and GitHub
  • verification — email verification tokens; managed but unused for social-only auth

Lobby Tables (Multiplayer)

lobbies

Column Type Constraints
id uuid PK
code varchar(10) NOT NULL, UNIQUE — room code
host_user_id text FK → user.id, ON DELETE CASCADE
status varchar(20) NOT NULL, default waiting, CHECK LOBBY_STATUSES
created_at timestamptz NOT NULL, default now()

lobby_players

Column Type Constraints
lobby_id uuid FK → lobbies.id, ON DELETE CASCADE
user_id text FK → user.id, ON DELETE CASCADE
score integer NOT NULL, default 0
joined_at timestamptz NOT NULL, default now()

Composite PK: (lobby_id, user_id) — no surrogate id column, one row per player per lobby.

Only lobby membership is persisted. Live game state (questions, timers, per-round answers) lives in the in-memory stores in apps/api, not in these tables. Max players is the MAX_LOBBY_PLAYERS constant in packages/shared, not a column.


Key Relationships

vocabulary_entries (1) ←──→ (N) entry_translations        ← live
words (1) ←──→ (N) senses (1) ←──→ (N) translations       ← target
user (1) ←──→ (N) session
user (1) ←──→ (N) account
user (1) ←──→ (N) lobbies (as host)
lobbies (1) ←──→ (N) lobby_players (N) ←──→ (1) user

Query Patterns

All queries live in packages/db/src/models/apps/api never imports drizzle-orm.

termModel.ts queries the live schema: it self-joins vocabulary_entries (aliased source and target), joins entry_translations for the answer, and fetches distractors with a separate query per question that excludes both the current entry id and the correct answer text — different entries can share a translation string.

The distractor query is N+1, one round trip per question. Batching is a known BACKLOG item.

Phase 5 rewrites these against words/senses/translations, which changes the shape: filter on senses.difficulty as a ceiling, then select translations at the requested difficulty. Target queries are sketched in documentation/pipeline/design-doc.md §5.


Deferred Schema Extensions (Not Yet Implemented)

Planned, additive, and keyed off the target schema:

Table Purpose Trigger
inflection_forms Gender, plural, conjugation/declension tables Grammar quiz mode
pronunciations IPA + audio URLs per language Pronunciation quiz mode
user_word_progress Spaced repetition state per user/word/language SRS review queue
quiz_answers Answer history for stats/analytics User stats dashboard