updating documentation, prettier format
This commit is contained in:
parent
039ed50567
commit
e534b98bc5
16 changed files with 1271 additions and 1070 deletions
|
|
@ -0,0 +1,45 @@
|
|||
// pipeline.ts pseudo code
|
||||
|
||||
/*
|
||||
step 1: discover source lists
|
||||
|
||||
this will give us an array of objects with this schema
|
||||
sourceLanguage: 'en' | 'de' | 'es' | 'fr' | 'it'
|
||||
pos: 'noun' | 'verb' | 'adjective' | 'adverb'
|
||||
words: string[];
|
||||
filePath: string;
|
||||
|
||||
the terminal output should be something like:
|
||||
|
||||
found 5 source lists:
|
||||
|
||||
de: noun
|
||||
en: noun
|
||||
|
||||
and so on
|
||||
|
||||
later on, it will also contain it: noun, verb, adjective etc
|
||||
*/
|
||||
|
||||
/*
|
||||
|
||||
step 2: validating source lists
|
||||
|
||||
a small script that trims whitespaces, removes duplicated words etc
|
||||
|
||||
terminal output: summary of how many words per pos per language were found
|
||||
|
||||
*/
|
||||
|
||||
/*
|
||||
|
||||
step 3: writing to database?
|
||||
|
||||
my thought: ill restart the pipeline several times during testing, and when adding more wordlists with other pos or extending the exisiting noun lists
|
||||
eventually the lists will contain tens or hundreds of thousands of words
|
||||
how do we prevent reading and processing the same words multiple times?
|
||||
if we read and validate+normalize the wordlists and write them to the database, we could then read from the database fill the missing translations etc
|
||||
and not read the same words from the same text files multiple times?
|
||||
|
||||
if we do this, we have to adjust the database schema because there are several notNull() rows inside
|
||||
*/
|
||||
Loading…
Add table
Add a link
Reference in a new issue