41 lines
1.5 KiB
Markdown
41 lines
1.5 KiB
Markdown
# trial run — implementation roadmap (pure vertical)
|
||
|
||
goal: english + italian nouns through the full pipeline on local models, emit
|
||
coverage report. build a thin end-to-end slice first, then widen each pass.
|
||
|
||
## phase 0 — foundation
|
||
|
||
- confirm @lila/shared types (lang codes, POS) exist and match
|
||
- package.json deps for what phase 1 needs
|
||
- one hardcoded test word to carry through the slice (e.g. "house")
|
||
|
||
## phase 1 — thin vertical slice (ONE word, 1a→3, crudest possible)
|
||
|
||
goal: prove a single card can travel the whole pipeline and come out the end.
|
||
allowed to be ugly — hardcode, skip voting, one model, fake CEFRLex.
|
||
|
||
- 1a: one lemma record, by hand or trivial read
|
||
- 2a: look it up in Kaikki, confirm it exists
|
||
- 2b: extract its senses → card(s) with gloss
|
||
- 2c: fill one missing translation with ONE local model, no voting
|
||
- 3: assign a difficulty crudely (even hardcoded "easy")
|
||
- OUT: one finished card. the pipeline has a shape.
|
||
|
||
## phase 2 — widen: real deterministic front (1a–2b, all words)
|
||
|
||
- 1a: real frequency list, verified
|
||
- 2a: real existence gate + miss-list triage
|
||
- 2b: real sense extraction, gender from tags, gloss-required drop
|
||
- still JSON output. inspect cards by hand.
|
||
|
||
## phase 3 — widen: real LLM back (2c + 3)
|
||
|
||
- storage seam: JSON → SQLite, schema, resumability
|
||
- 2c: real generate→vote, three families, all sub-passes
|
||
- 3: real CEFRLex single-sense + 3-model vote multi-sense + exclude-on-split
|
||
|
||
## phase 4 — coverage report + runs
|
||
|
||
- emit report
|
||
- run english, then italian
|
||
- decide: local good enough? rent vs API?
|