xdna.cpp is an experimental llama.cpp backend for AMD XDNA1 NPUs. It offloads Q4_0 decode GEMV to the 16 AIE2 tiles available on Phoenix and Hawk Point APUs, while keeping model loading, attention, normalization, sampling, and unsupported operations on the normal llama.cpp path.
The backend is currently focused on Ryzen 7040 and 8040 series APUs running Linux.
The Linux x86_64 release includes llama-cli, llama-bench, llama-server, and libggml-xdna.so.
wget https://github.com/random-unknown-username/xdna.cpp/releases/download/v1.0.0/xdna-llama-linux-x86_64.tar.gz
tar -xzf xdna-llama-linux-x86_64.tar.gz
cd xdna-llama-linux-x86_64Download a Q4_0 model:
wget https://github.com/random-unknown-username/xdna.cpp/releases/download/v1.0.0/qwen2.5-0.5b-instruct-q4_0.ggufRun it on the XDNA backend:
./run-chat.sh qwen2.5-0.5b-instruct-q4_0.ggufPhoenix and Hawk Point expose 16 AIE2 compute tiles arranged as 4 columns by 4 rows.
xdna.cpp registers as a GGML backend and handles supported GGML_OP_MUL_MAT operations using Q4_0 weights.
For each supported operation:
- Q4_0 weights are read from the normal mmap backed GGUF.
- Weights are copied through a bounded host staging buffer.
- XRT DMA transfers them to the NPU.
- The AIE2 kernel unpacks the 4 bit values, applies the Q4_0 scale in BF16, and executes the vector MAC.
- Output activations return to host memory and llama.cpp continues execution.
GGUF Q4_0 weights
|
v
file backed mmap
|
v
2 x 64 MiB staging BOs
|
v
XRT DMA
|
v
16 AIE2 tiles
|
v
Q4_0 unpack
BF16 scaling
vector MAC
|
v
llama.cpp CPU path
The larger streamed matrices currently sustain roughly 32 to 38 GB/s through the NPU path.
Not every supported matrix is worth sending to the NPU.
Small projections can complete on the CPU faster than the XRT dispatch and synchronization overhead required to execute them on XDNA. The backend therefore uses a crossover policy and leaves smaller operations on the CPU.
This is visible on Qwen2.5 0.5B:
llama.cpp CPU, 8 threads
106.5 ± 4.9 tok/s
XDNA hybrid backend
72.6 ± 1.9 tok/s
At 3B the larger projections give the NPU enough work for the offload to pay off:
Qwen2.5 3B Q4_0
llama.cpp CPU, 8 threads
9.8 ± 0.4 tok/s
XDNA hybrid backend
16.5 ± 1.3 tok/s
That is about a 68.8% improvement over the measured CPU baseline.
The largest model tested so far is Qwen3.8 27B Q4_0:
Parameters 27.3B
Streamed weights ~12.6 GiB
Decode speed 1.72 ± 0.05 tok/s
The model weights are not permanently allocated as XRT buffer objects.
An earlier implementation created persistent xrt::bo allocations for model tensors. On the 27B model this pinned roughly 12.6 GB as unevictable kernel memory, heavily reducing available page cache and causing swap traffic.
Decode performance dropped to roughly:
0.15 tok/s
The current runtime instead keeps weights in normal file backed memory and allocates two reusable 64 MiB host staging BOs:
staging BO 0 64 MiB
staging BO 1 64 MiB
total pinned 128 MiB
Weights are streamed through the two buffers during layer execution, keeping the pinned allocation fixed regardless of total model size.
This is what allows the 27B model to stream roughly 12.6 GiB of weights while keeping only 128 MiB permanently pinned for staging.
XDNA output has been compared against separate pure CPU reference runs using logit similarity and token agreement.
Evaluated tokens 32
Top 1 match 32 / 32
Logit cosine 0.999511
Relative L2 delta 3.12%
Evaluated tokens 16
Top 1 match 16 / 16
Logit cosine 0.997227
Relative L2 delta 7.08%
Evaluated tokens 16
Top 1 match 16 / 16
Logit cosine 0.999906
Relative L2 delta 1.41%
The XDNA path is not expected to be bit identical to the CPU implementation because the AIE2 kernel uses BF16 internally. All evaluated samples still matched the CPU reference on the top 1 token.
git clone https://github.com/random-unknown-username/xdna.cpp.git
cd xdna.cpp
cmake -B build -S . -DCMAKE_BUILD_TYPE=Release
cmake --build build -j$(nproc)Probe the XDNA device:
./build/xdna-cli probeRun the test suite:
ctest --test-dir build --output-on-failuregit clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build -S . \
-DGGML_XDNA=ON \
-DCMAKE_BUILD_TYPE=Release
cmake --build build -j$(nproc)Run a model using the XDNA device:
./build/bin/llama-cli \
-m /path/to/model.gguf \
--device XDNA0 \
-t 8The current backend requires:
- AMD Ryzen 7040 Phoenix or Ryzen 8040 Hawk Point APU
- Linux with the
amdxdnadriver /dev/accel/accel0- AMD XRT 2.18 or newer
- access to the
rendergroup - CMake and a C++ compiler
Add the current user to the render group if required:
sudo usermod -a -G render $USERLog out and back in after changing group membership.
The optimized path currently targets:
Operation GGML_OP_MUL_MAT
Weight format Q4_0
Primary use decode GEMV
Hardware XDNA1 / AIE2
APUs Phoenix / Hawk Point
Unsupported GGML operations remain on the normal host backend.
The current work is mostly around decode performance. More quant formats, prefill acceleration, lower dispatch overhead, and better DMA overlap still need work.
MIT.
See THIRD_PARTY_NOTICES.md for llama.cpp and AMD acknowledgements.