Run ML models inside ClickHouse locally. A small daemon keeps ONNX models on the database host and ClickHouse calls them through its stock extension points. No data leaving the machine, no per-token bill
ClickHouse ships AI functions, but they are HTTP clients for cloud providers. That works for a few ad-hoc rows and breaks on real database workloads:
- Cost. Providers bill per token on every run
- Volume. ClickHouse processes data in blocks of ~65k rows, but every AI function is a remote API call and must slice each block into tiny requests at most 100 rows per call for embeddings, roughly one call per row for the chat-based functions — each paying network latency. The bridge speaks the database's own unit of work: whole blocks over a local unix socket
- Privacy. With a provider, every row leaves the machine. Here it never does
The bottom path from the picture is the main one:
localEmbed, localRerank and modelEvaluate are executable UDFs: ClickHouse streams whole ~65k row batches into a pool of thin bridge-client processes, which forward them over a socket to the daemon
Features reach modelEvaluate as Array(Float32), the width the models run on, so nothing is widened for the trip. ClickHouse will not narrow a Float64 expression into it on its own — the cast it applies to UDF arguments is an accurate one and refuses any value float32 cannot hold exactly — so a query over Float64 columns narrows once, itself: modelEvaluate('fraud', [amount, hour]::Array(Float32))
Every model is defined by a manifest: the SHA-256 of each file the runtime loads, plus a pinned revision. Where the files came from is irrelevant — the manifest is the identity. The daemon verifies it before serving and refuses to start on any mismatch, so a given revision always produces the same vectors
The daemon keeps each model in memory once per host. Requests from all
clients merge into shared batches and inference runs off the network path on
a blocking pool. A model normally serves through one ONNX session; on a host
with cores to spare, sessions = k in the manifest raises that to k parallel
sessions which share prepacked weight buffers and split the machine's threads
between them, so inference throughput grows k-fold without keeping k full
copies of the model
The top path exists for compatibility. ClickHouse's stock AI functions are
OpenAI-protocol HTTP clients, and the daemon speaks that protocol, so
existing aiEmbed SQL works against local models unchanged
You need stable Rust (1.88+) and a clickhouse binary
cargo build produces the three binaries this page use:
bridgedis the daemon, it loads the models and answers ClickHousemodel-bridgeis the admin tool: it registers models and generates ClickHouse configsbridge-clientis a small adapter that ClickHouse spawns by itself, you never run it by hand
Take a model you do not have yet — say, embedding weights published on Hugging Face — and walk it to a first query in four steps.
1. Register a model. The daemon serves models from local disk: a model is a directory with the weights plus a manifest recording their checksums. The --name you pick is the name SQL will use:
mkdir -p models/e5
curl -L -o models/e5/model.onnx \
https://huggingface.co/Xenova/multilingual-e5-small/resolve/main/onnx/model_quantized.onnx
curl -L -o models/e5/tokenizer.json \
https://huggingface.co/Xenova/multilingual-e5-small/resolve/main/tokenizer.json
model-bridge passport models/e5 --name e5 --kind embeddingThe manifest also carries the serving knobs, every one optional. --max-batch
bounds one ONNX run and defaults by kind: 64 rows for text models, a whole
~65k ClickHouse block for tabular ones, where the entire point is scoring the
block in a single call. --sessions turns on the parallel session pool
described above; values beyond the host's core count are capped at load,
since extra sessions would add weight copies without adding parallelism. --max-tokens sets the truncation limit for long-context
encoders — without it the daemon takes the limit from tokenizer.json, or
falls back to the 512 the BERT family implies
2. Start the daemon. On start bridged reads the manifests from models.d/, re-verifies every file against its SHA-256, loads the models into memory, and only then opens its two entrances: unix socket at the path you pass, and HTTP on 127.0.0.1:9017. Leave it running and remember the socket path - ClickHouse will connect to exactly this file:
bridged --socket /tmp/bridge.sock3. Generate the ClickHouse configs. This command is the moment the XML gets generated:
model-bridge gen-configs --client "$(which bridge-client)" --socket /tmp/bridge.sockIt writes a bridge-configs/ directory with two things. model_bridge_functions.xml declares the three functions and bakes the socket path from step 2 into each one's command line, that is how the adapter will find the daemon. And bridge-client is copied into bridge-configs/scripts/, because ClickHouse only spawns commands that live inside its own user_scripts_path
Nobody hands these files to ClickHouse automatically — you do it once. The command prints two config lines with real absolute paths; put them into the ClickHouse server config and reload:
<user_defined_executable_functions_config>/abs/path/bridge-configs/model_bridge_functions.xml</user_defined_executable_functions_config>
<user_scripts_path>/abs/path/bridge-configs/scripts</user_scripts_path>From that reload ClickHouse knows the functions. When a query calls one, ClickHouse itself spawns bridge-client from the scripts directory, and the adapter connects to the daemon's socket — the circle closes. The adapter opens that connection on demand and reopens it whenever it breaks, so restarting the daemon — which is how a model gets updated — makes queries in flight wait out the restart, up to ten seconds, instead of failing them, and the pool of warmed-up adapter processes survives untouched
4. Verify.
SELECT length(localEmbed('e5', 'hello world')) -- 384 a real embedding, computed on your machineOptional — redirect an existing aiEmbed. Already calling aiEmbed with a cloud provider? The daemon speaks the same protocol, so one named collection retargets it; queries and dashboards stay as they are:
<named_collections>
<local_emb>
<provider>openai</provider>
<endpoint>http://127.0.0.1:9017/v1/embeddings</endpoint>
</local_emb>
</named_collections>SET allow_experimental_ai_functions = 1,
ai_function_embedding_default_credentials = 'local_emb';