huggingface-local-models/references/hardware.md
Version ca0325bb20b2 · Apache-2.0. This preview displays packaged text and does not execute code. Treat the contents as untrusted instructions.
← Return to resource and package checksum
Hardware Acceleration
Apple Silicon (Metal)
make clean && make GGML_METAL=1
llama-cli -m model.gguf -ngl 99 -p "Hello"
NVIDIA (CUDA)
make clean && make GGML_CUDA=1
llama-cli -m model.gguf -ngl 35 -p "Hello"
# Hybrid for large models
llama-cli -m llama-70b.Q4_K_M.gguf -ngl 20
# Multi-GPU split
llama-cli -m large-model.gguf --tensor-split 0.5,0.5 -ngl 60
AMD (ROCm)
make LLAMA_HIP=1
llama-cli -m model.gguf -ngl 999
CPU
# Match physical cores, not logical threads
llama-cli -m model.gguf -t 8 -p "Hello"
# BLAS acceleration
make LLAMA_OPENBLAS=1