Emmio is an experimental project focused on language learning. It provides learning and testing algorithms:
The project manages four kinds of artifacts:
Emmio is primarily a terminal application, but it also includes an experimental web interface and an iOS application built on top of a JSON API (see Other Interfaces).
Requires Python 3.12 or later and a PostgreSQL server: learning and
lexicon records are stored in a PostgreSQL database (see scripts/schema.sql
for the table definitions). The connection string is taken from the
EMMIO_DATABASE_URL or DATABASE_URL environment variable and defaults to
postgresql:///emmio.
pip install git+https://github.com/enzet/emmioAudio playback additionally requires the mpv player to be installed on the system.
To run Emmio, simply execute
emmioYou can specify a data directory with the --data option and a username with
the --user option. If not specified, the data directory defaults to
~/.emmio and the username defaults to your current system username.
> learn
The learning process is based on spaced repetition with increasing intervals. Emmio picks new words from the word lists configured for the learning process (optionally checking them against your lexicons and dictionaries first), shows translations of the word with the word itself hidden, and asks you to type the word in the learning language.
- If the answer is correct, the word is scheduled for repetition: after the first correct answer it comes back in about 5 minutes, after the second in a day, and each subsequent correct answer doubles the interval.
- If the answer is wrong, type
n— the word and its translations are shown, and repetition starts over with short intervals. - Press Enter on empty input to see an example sentence with the word (press again for more sentences).
- Type
pto postpone the question,/skipto exclude the word from the learning process entirely, or/stopto finish the session.
Learning processes are configured in the user configuration file
(users/<user name>/config.json).
> lexicon
The algorithm randomly presents words from the target language (weighted by frequency). For each word, you need to indicate
- whether you know at least one meaning of this word (press y or Enter),
- whether you don’t know any meaning of the word (press n),
- whether the word is commonly used as a proper name or isn’t valid (press -).
Press q to finish.
After completion, the algorithm provides a non-negative number called rate, that indicates your vocabulary level. A score of 0 means you don’t know any words in the language, while infinity means you know every word in the frequency list. This rate is most useful for tracking your learning progress and comparing vocabulary levels between different users using the same frequency list.
There is no upper limit for the rate, if you know meanings of all words in a language, the rate is infinity. It’s finite though, if you don’t know meaning of at least one word.
Column “Don’t understand” means how much you don’t understand of an arbitrary text. Column “Understand” means how much you understand of an arbitrary text.
| Rate | Don’t understand | Understand | Level * |
|---|---|---|---|
| 10 | 0.10 % | 99.90 % | |
| 9 | 0.20 % | 99.80 % | |
| 8 | 0.39 % | 99.61 % | C2 |
| 7 | 0.78 % | 99.22 % | C1 |
| 6 | 1.56 % | 98.44 % | B2 |
| 5 | 3.12 % | 96.88 % | B1 |
| 4 | 6.25 % | 93.75 % | A2 |
| 3 | 12.50 % | 87.50 % | A1 |
| 2 | 25.00 % | 75.00 % | |
| 1 | 50.00 % | 50.00 % | |
| 0 | 100.00 % | 0.00 % |
* Level is a very rough approximation of a level of the Common European Framework of Reference for Languages. Please, don't take it as a strict definition.
Lexicon checking is run with the lexicon command or lexicon <language>,
where <language> is a 2-letter ISO 639-1 language code. Important: a
lexicon needs a full (not stripped)
frequency list.
Two command-line tools compute how the rate estimate r̂ evolves as answers accumulate and write the trajectories to CSV files; a third draws those CSV files in a shared visual style.
To replay the recorded checking process of a real lexicon, run
python -m emmio.lexicon.replay --user $USER_ID --lexicon $LEXICON_IDIt reads the lexicon's answer log from the database, re-runs the production
rate estimator after every answer, and writes the estimate trajectory to a
CSV file (by default ${USER_ID}_${LEXICON_ID}.csv, override with --csv).
The frequency list configured for the lexicon is loaded from the Emmio data
directory (EMMIO_DATA_DIR, default ~/.emmio).
To compare how fast the checking methods converge on a synthetic language, run
python -m emmio.lexicon.simulation --rate 3 --trials 5It builds a fake Zipfian language, draws random knowledge profiles with a
chosen true rate, drives every checking method through the production
selection and estimation code, reports how fast each converges, and writes
the per-trial trajectories to a CSV file (by default simulation.csv,
override with --csv).
To visualize a CSV file written by either tool, run
python -m emmio.lexicon.visualize replay ${USER_ID}_${LEXICON_ID}.csv
python -m emmio.lexicon.visualize simulation simulation.csvThe replay plot shows the estimate trajectory with its ±1 standard deviation
band and ticks along the bottom marking the "don't know" answers; the
simulation plot shows per-trial and mean trajectories with the RMSE against
the realized rates. The output image file defaults to the CSV file with a
.png suffix, override with --plot.
With --format tikz the same picture (without the "don't know" ticks) is
written as TikZ (pgfplots) code in a .tex file instead, ready to be embedded
into a LaTeX document with \input. The document preamble needs
\usepackage{pgfplots} and \pgfplotsset{compat=1.18}.
Emmio data directory is located by default in ~/.emmio and contains all
downloaded artifacts and their configuration files and collected user data.
dictionaries— single word translations.sentences— whole sentence translations.lists— frequency and word lists.audio— audio files with word pronunciations.users— user data.<user name>config.json— user configuration file.learn— user learning process data.lexicon— user lexicon checking data.
Dictionaries are entities that provide definitions and translations for single
words. Artifacts are controlled by configuration file
dictionaries/config.json.
Emmio supports:
- dictionaries stored in JSON files,
- English Wiktionary (through Kaikki.org website, containing dictionaries extracted from English Wiktionary using wiktextract).
Sentences are example parts of texts in one language, translated into another
language. Artifacts are controlled by configuration file
sentences/config.json.
Emmio supports:
- sentences and translations stored in simple text files (odd lines are sentences, even lines are translations),
- sentences from Tatoeba project.
Frequency list is a relation between unique words and the number of their
occurrences in some text or a corpus of texts. Some frequency lists are
stripped (e.g. 6,500-lemma list based on the New Corpus for Ireland). Lists
are controlled by configuration file lists/config.json.
Emmio supports:
- word lists stored in text files,
- FrequencyWords (Hermit Dave’s project, which contains full and stripped frequency lists extracted from subtitles in Opensubtitles project).
Wiktionary project contains frequency lists for different languages.
Audio artifacts are word pronunciations, played during learning. Artifacts are
controlled by configuration file audio/config.json.
Emmio supports:
- audio files stored in local directories,
- pronunciations from Wikimedia Commons.
Besides the terminal application, the repository contains experimental clients built on top of a FastAPI backend:
- API (emmio/api/) — a JSON API with token-based
authentication. Install its dependencies with
pip install -e .[api]and run it withuvicorn emmio.api.app:app --reload. - Web (web/) — a minimal JavaScript client served by the API at
http://localhost:8000/app/; see web/README.md. - iOS (ios/) — a SwiftUI application; see
ios/run.shfor building and running it in the simulator.
Before contributing, please follow these steps:
- Install development dependencies with
pip install -e .[dev]. - Enable Git hooks with
git config core.hooksPath .githooks.
The pre-commit hook checks formatting and linting with Ruff, types with ty, and runs the tests with pytest. To run the same checks manually:
ruff format emmio/ tests/
ruff check emmio/ tests/
ty check emmio/ tests/
pytestFor commit messages we use 50/72 rule and each commit should start with a verb in the present tense. We also use Markdown for commit message formatting.