Hugging Face just showed how to turn Pi into a fully private, offline development environment using llama.cpp and the new /llama command. The tutorial covers installing the engine, using Hugging Face's hardware compatibility tools to find the perfect GGUF quantization level for your specific machine, and running Qwen3 8B locally.
No data, prompt, or code ever leaves your hardware, and there are zero token costs. Perfect for private code assistants.
Have you tried local workflows in Pi yet?