llama.cpp: Fast Local Inference Engine for DeepSeek-R1 and Quantized Models
A high-performance C++ LLM inference engine supporting low-bit quantization and efficient execution of reasoning models on Apple Silicon and commodity hardware.
Image: GitHub




