updating docs to match the implemented phase 3 pipeline
The pipeline docs still described pipeline.ts as pseudocode and the validation module as unwritten. Both have been implemented and run. - CLAUDE.md: replace the "no executable pipeline yet" description with the actual module flow, plus the two invariants worth preserving (resumability via headword diffing, raw responses saved before parsing) - DATA_PIPELINE.md: mark the seven implemented modules, add a module responsibility map and the CLI flag table, drop the resolved warning about hardcoded prompt values - roadmap.md: check off phase 3 tasks 3.2/3.3/3.4, record where the build diverged from the plan, note that validate.ts is stricter than its own spec - STATUS.md: phase 3 is data work now, not code work Also records two open issues: the systemic difficulty-ordering rejection cause, and the hard-tier shortfall pending a full run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
37c978e230
commit
da9cdbfa1b
4 changed files with 219 additions and 49 deletions
|
|
@ -46,7 +46,9 @@ Monorepo: `apps/api` (Express + `ws`), `apps/web` (React 19, Vite, TanStack Rout
|
|||
|
||||
## Data pipeline
|
||||
|
||||
`data-pipeline/` is mid-rewrite on this branch (`refactor/gemini-only-pipeline`): the old local-LLM/adapter architecture is gone, and `pipeline.ts` currently holds the staged design as pseudocode comments — there is no executable pipeline yet. Current pieces: `source-data/{lang}/{pos}` wordlists, `prompt` (the Gemini system prompt, a plain UTF-8 file), `db/schema.sql` (SQLite staging: `words` → `senses` → `translations`), and a dedicated pipeline PostgreSQL on port 5433.
|
||||
`data-pipeline/` was rewritten on this branch (`refactor/gemini-only-pipeline`): the old local-LLM/adapter architecture is gone, replaced by a Gemini-only pipeline that is implemented and running (Phase 3). Flow: `sourceLists.ts` (discover + dedup `source-data/{lang}/{pos}`) → `promptTemplate.ts` (render `prompt` placeholders) → `gemini.ts` (structured output with a per-batch `responseSchema`, retry/backoff) → `validate.ts` (per-entry, three-way `valid`/`empty`/`invalid`) → `staging.ts` (SQLite, one transaction per word), orchestrated by `pipeline.ts`. A dedicated pipeline PostgreSQL runs on port 5433; the SQLite → PostgreSQL import script is Phase 4 and does not exist yet.
|
||||
|
||||
Two invariants to preserve when touching it: runs are **resumable** because already-staged headwords are diffed out before batching, and **raw responses are written to `responses/` before parsing**, so validation changes can be replayed without re-spending API quota. Rejected entries go to `rejections/{lang}-{pos}.jsonl`; `"senses": []` means "not a valid word of this POS" and is skipped, not rejected.
|
||||
|
||||
Read `documentation/DATA_PIPELINE.md` for orientation and current phase status, `documentation/pipeline/design-doc.md` for the schema/difficulty model/Gemini JSON contract, and `documentation/pipeline/roadmap.md` for the phase plan. Everything in `documentation/archive/` is superseded — `data-pipeline-local-llm.md`, `llm-setup-local.md`, and `model-strategy-cefr-voters.md` describe the removed local-LLM / CEFR-voter pipeline and are historical only.
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue