Local OpenAI-compatible speech-to-text server that wraps
mlx-audio on Apple Silicon.
Exposes the subset of OpenAI's /v1/audio/transcriptions API that most
dictation clients, transcription tools, and openai SDK wrappers call,
so you can point any OpenAI-compatible client at http://127.0.0.1:18765
and run inference locally on an M-series Mac with no cloud round-trip.
Default model is Cohere Transcribe
(CohereLabs/cohere-transcribe-03-2026),
a 2B-parameter multilingual ASR model, but any mlx-audio STT model
works via --model — Whisper, Parakeet, Qwen3-ASR, Voxtral, Canary,
Moonshine, etc.
- macOS on Apple Silicon (M1 or newer)
- Python 3.10+
- ~5 GB free disk for the default Cohere model (downloaded on first run)
- A Hugging Face account with access to the model you want to use
# 1. Clone
git clone https://github.com/piercecohen1/mlx-stt-server.git
cd mlx-stt-server
# 2. Create a virtualenv (recommended — the start script auto-detects
# .venv/bin/python and venv/bin/python)
python3 -m venv .venv
source .venv/bin/activate
# 3. Install dependencies
pip install -r requirements.txtThe default model (CohereLabs/cohere-transcribe-03-2026) is a gated
model on Hugging Face — you need an HF account, you need to accept the
model's terms on its page, and you need a token available to
huggingface_hub. Whisper and some other mlx-audio models are public and
don't require this step.
# Option A: interactive login (browser + paste token)
huggingface-cli login
# Option B: set a token env var (useful for scripts and the Mac Mini)
export HF_TOKEN="hf_xxx..."
# Option C: long-term, add to your shell rc file
echo 'export HF_TOKEN="hf_xxx..."' >> ~/.zshrcIf you skip this, the first request to load Cohere Transcribe will fail
with a 401 from Hugging Face. You can swap to a non-gated model with
--model, e.g. --model mlx-community/whisper-large-v3-turbo.
python server.py # preloads default model on :18765
python server.py --port 8123
python server.py --no-preload # load on first request instead
python server.py --model mlx-community/whisper-large-v3-turboThe first run downloads the model from Hugging Face (~4 GB for Cohere Transcribe). Subsequent runs load from the local HF cache in ~1 second.
A wrapper script at scripts/stt-server backgrounds the server via
nohup, tracks the PID in ~/.cache/mlx-stt-server/server.pid, and
appends logs to ~/.cache/mlx-stt-server/server.log. Before reporting
or stopping that PID, it verifies the process command and working
directory and automatically discards stale PID files. One-time setup:
./scripts/install-aliases.sh
source ~/.zshrc # or ~/.bashrc; or just open a new terminal tabThat adds a single stt-server alias pointing at the wrapper script.
Then from any directory:
stt-server --start # nohup the server, return immediately
stt-server --status # shows PID + health check against /healthz
stt-server --stop # SIGTERM the recorded PID (SIGKILL after 10s)
stt-server --restart # stop + start
stt-server --logs # tail -f the log fileThe wrapper sources ~/.config/mlx-stt-server/env (if present) at the
top of the script, so you can set STT_SERVER_PORT, STT_SERVER_MODEL,
STT_SERVER_PYTHON, or any MLX_STT_* server-side variable in one
place.
./scripts/install-launch-agent.sh # install / update
./scripts/install-launch-agent.sh --uninstall # remove autostartThis writes a launchd LaunchAgent
(~/Library/LaunchAgents/com.mlx-stt-server.plist) that runs
stt-server --start when you log in, and loads it immediately, so the
server also starts right away if it isn't running. The agent's own
output goes to ~/.cache/mlx-stt-server/launchd.log.
launchd jobs run without your shell environment, so the installer
resolves the Python interpreter and the directory containing ffmpeg
at install time and pins both into the plist. Re-run the installer
after moving the repo, switching Python versions, or reinstalling
ffmpeg.
The agent only starts the server; it does not supervise it. A crashed
server stays down until the next login or a manual stt-server --start, and stt-server --stop keeps it stopped the same way.
Any tool that speaks the OpenAI audio API works. Base URL is
http://127.0.0.1:18765 — most clients append /v1/audio/transcriptions
themselves, so paste just the host. A few examples:
from openai import OpenAI
client = OpenAI(
base_url="http://127.0.0.1:18765/v1",
api_key="local", # anything works — not validated in local mode
)
with open("audio.wav", "rb") as f:
result = client.audio.transcriptions.create(
model="CohereLabs/cohere-transcribe-03-2026",
file=f,
)
print(result.text)curl -X POST http://127.0.0.1:18765/v1/audio/transcriptions \
-F "file=@samples/test.m4a" \
-F "model=CohereLabs/cohere-transcribe-03-2026" \
-F "language=en" \
-F "response_format=json"| Field | Value |
|---|---|
| Base URL / Endpoint | http://127.0.0.1:18765 |
| API Key | any non-empty string (e.g. local) |
| Model | CohereLabs/cohere-transcribe-03-2026 |
Some clients require the full /v1 suffix (e.g. http://127.0.0.1:18765/v1),
some append it automatically — if one form 404s, try the other.
POST /v1/audio/transcriptions— multipart form-datafile— audio file; any format mlx-audio / librosa can read (wav, mp3, m4a, flac, ogg, ...)model— any mlx-audio STT model IDlanguage— ISO-639-1 code (defaulten)response_format—json(default),text,verbose_json,srt,vttprompt,temperature— accepted for client compatibility but ignored by the default Cohere model (it has no prompt slot). Whisper does consumepromptif you use--model ...whisper...; wiring that through is a ~10-line change toserver.py.
POST /v1/audio/translations— forwards to transcription withlanguage=en. Cohere Transcribe is ASR-only, not a translator; this endpoint exists so clients that default to/translationsstill get something usable.GET /v1/models— lists loaded + default models in OpenAI format.GET /healthz— plaintextok(used by the wrapper's health probe).
The repo ships with a tiny synthesized test clip at samples/test.m4a
(generated via macOS say, ~28 KB) so the curl example above works
out of the box without you supplying your own audio.
curl http://127.0.0.1:18765/v1/models
curl -X POST http://127.0.0.1:18765/v1/audio/transcriptions \
-F "file=@samples/test.m4a" \
-F "model=CohereLabs/cohere-transcribe-03-2026" \
-F "language=en"For local-only use, skip this section — the server listens on
127.0.0.1 with no auth by default, which is the right setup for
dictation on your own machine.
To expose the server beyond loopback (LAN, Tailscale, Cloudflare Tunnel, etc.), enable bearer-token auth by creating a key file:
umask 077
mkdir -p ~/.config/mlx-stt-server
openssl rand -hex 32 > ~/.config/mlx-stt-server/api-key
chmod 600 ~/.config/mlx-stt-server/api-keyThe server auto-detects the key file at startup. To refuse startup when the file is missing or empty (recommended on any tunneled host):
python server.py --require-key
# or, equivalently
MLX_STT_REQUIRE_KEY=1 python server.pyClients send the token in the Authorization header:
curl -X POST https://your-host/v1/audio/transcriptions \
-H "Authorization: Bearer $(cat ~/.config/mlx-stt-server/api-key)" \
-F "file=@samples/test.m4a" \
-F "model=CohereLabs/cohere-transcribe-03-2026"/healthz is always open (so tunnel health probes don't need
credentials). Every other endpoint returns 401 without a valid
bearer token when a key is loaded.
Environment knobs:
| Variable | Default | Purpose |
|---|---|---|
MLX_STT_API_KEY_FILE |
~/.config/mlx-stt-server/api-key |
Override the key-file path. |
MLX_STT_REQUIRE_KEY |
0 |
Set to 1 to refuse startup without a valid key file. |
MLX_STT_MAX_UPLOAD_MB |
500 |
Body-size cap enforced before form parsing (returns 413 over cap). |
MLX_STT_LOG_TRANSCRIPTS |
0 |
Set to 1 to log each request's transcribed text (also --log-transcripts). Off by default so routine dictation isn't written to the log. |
MLX_STT_LOG_VERBOSE |
0 |
Set to 1 for a colorized, multi-line per-request log block — request lifecycle + inference metrics, no transcript (also --log-verbose). Pairs well with stt-server --logs for a live demo. Honors NO_COLOR. |
Rotate the token by atomically overwriting the key file (mktemp in
the same dir → chmod 600 → mv -f); the server re-reads it at
startup only, so you need to restart after rotation.
Thin FastAPI wrapper around mlx_audio.stt.load(...).generate(path, language=...):
- On startup the model is loaded eagerly via FastAPI
lifespanand cached in a module-level dict keyed by model ID. - On each request the uploaded audio is written to a
tempfilein$TMPDIR, passed as a path tostt_model.generate(path, language=lang), and deleted in afinallyblock. - The returned
STTOutputis shaped into whicheverresponse_formatthe client asked for.json→{"text": ...},verbose_jsonmirrors OpenAI's segment shape,srt/vttare generated fromresult.segments.
The model only loads once per process, so first-request latency is just inference (~0.7 s for a 35 s clip on an M-series Mac with the default Cohere model).
Expect the server process to sit at roughly 4 GB with the default Cohere model loaded: about 3.9 GB of model weights resident in unified memory, plus a couple hundred MB of Python runtime. The footprint stays flat across requests because the server releases MLX's Metal buffer cache after every transcription. That cache is bounded only by a default limit near total system RAM, and variable-length audio means cached buffers are rarely reused, so an uncleared cache grows by gigabytes over a day of dictation and eventually gets swapped out, which shows up as multi-second latency on the first request after an idle period.
To check the live footprint, use Activity Monitor or:
top -l 1 -pid $(cat ~/.cache/mlx-stt-server/server.pid) -stats memmacOS ps under-reports it because the weights live in Metal regions.
A process climbing well past ~5 GB means the per-request
mx.clear_cache() in server.py has stopped doing its job.
- Cohere Transcribe is ASR only —
/v1/audio/translationsis a pass-through that forceslanguage=en. promptandtemperatureare accepted but ignored for the default Cohere model. Other mlx-audio models may consume them (Whisper'sinitial_prompt, Qwen3-ASR's context field); wire-through requires a small patch toserver.py.- No streaming; Cohere's
generate()explicitly raises onstream=True. - 14 languages for the default Cohere model: en, fr, de, it, es, pt, el, nl, pl, zh, ja, ko, vi, ar. Switch models for other languages.
scripts/install-aliases.shwrites to~/.zshrcor~/.bashrc. If you use fish, nushell, or any other shell, you'll need to install the alias by hand — the script handles bash and zsh only.
MIT — see LICENSE. mlx-audio, which this server wraps, is also MIT-licensed.