30-SECOND SUMMARY
What to take away
- Choose releases, Docker, or source build—not all three.
- Check the model card and license.
- Make the CLI work before starting a server.
The shortest GGUF path
Make CLI work before the server.
- 01INSTALL
Official package
- 02GGUF
Model and license
- 03CLI
Single prompt
- 04SERVER
Local API
Choose an install path
The official repository provides prebuilt releases, Docker instructions, and source builds. Match CUDA, Metal, Vulkan, or CPU support to the machine.
The working rule for “Choose an install path” is: Choose releases, Docker, or source build—not all three. Do not treat an open window as proof of a complete installation; also record versions, process state, ports, and persistent data paths.
For verification, save the model and runtime versions, source input, relevant settings, and observed output together. Repeat the step while changing only one factor, and record unexpected results and untested limits as carefully as successes before applying the guidance to private or production data.
Select a GGUF
GGUF stores tensors and metadata. Start near Q4 and confirm the chat template, publisher, and license.
The working rule for “Select a GGUF” is: Check the model card and license. Do not treat an open window as proof of a complete installation; also record versions, process state, ports, and persistent data paths.
Preserve the before-and-after state and the time of the check so that another run can reproduce the result. Include at least one failure condition—such as empty input, constrained resources, or a restart—to reveal the boundary of the step rather than documenting only the happy path.
Run the CLI
The current quick start can download a Hugging Face model directly.
The working rule for “Run the CLI” is: Make the CLI work before starting a server. Do not treat an open window as proof of a complete installation; also record versions, process state, ports, and persistent data paths.
One successful run is not enough: repeat it after a restart and send one invalid input to confirm a controlled failure. Before connecting production data, test timeouts and cleanup so that an interrupted command does not leave stale processes, files, or application state.
llama-cli -hf ggml-org/Qwen3.5-0.8B-GGUFStart locally
Use the compatible server on localhost first. External access needs a separate authentication layer.
The working rule for “Start locally” is: Choose releases, Docker, or source build—not all three. Do not treat an open window as proof of a complete installation; also record versions, process state, ports, and persistent data paths.
After running the command or code, inspect the exit status, logs, and the file, process, or response it was meant to create. If it fails, change one input, version, permission, or resource condition at a time and repeat the same check so that the cause remains attributable.
llama-server -hf ggml-org/Qwen3.5-0.8B-GGUFFrequently asked questions
Is GGUF a model name?
No. It is a file format for tensors and metadata. For a practical check, follow the “Choose an install path” section, change one condition at a time, and record the result.
Must I compile from source?
No. Official binaries and Docker are alternatives. For a practical check, follow the “Select a GGUF” section, change one condition at a time, and record the result.
Can it run without a GPU?
Yes, but speed depends heavily on model and CPU. Make the CLI work before starting a server. For a practical check, follow the “Run the CLI” section, change one condition at a time, and record the result.
Primary sources
Check the original documentation for version-specific details.
llama.cpp repository Hugging Face GGUF