30-SECOND SUMMARY
What to take away
- Start with VRAM capacity and runtime support.
- Include PSU, case, cooling, motherboard, RAM, and storage in total cost.
- Run a sustained load and the real model before purchase when possible.
Verify used hardware
A real load matters more than a listing.
- 01IDENTIFY
Model and VRAM
- 02FIT
Power, space, cooling
- 03STRESS
Heat and errors
- 04MODEL
Run the intended GGUF
Work backward from the model
Define model size, context, concurrency, and memory headroom. Training or image generation may require different software support.
The working rule for “Work backward from the model” is: Start with VRAM capacity and runtime support. Keep the model, quantization, context length, and concurrency fixed so that a device or model comparison has a clear cause.
For verification, save the model and runtime versions, source input, relevant settings, and observed output together. Repeat the step while changing only one factor, and record unexpected results and untested limits as carefully as successes before applying the guidance to private or production data.
Cross-check identity
Request the exact model, VRAM, serial information, ports, and power connectors. Compare compute capability and runtime support with official sources.
The working rule for “Cross-check identity” is: Include PSU, case, cooling, motherboard, RAM, and storage in total cost. Keep the model, quantization, context length, and concurrency fixed so that a device or model comparison has a clear cause.
Preserve the before-and-after state and the time of the check so that another run can reproduce the result. Include at least one failure condition—such as empty input, constrained resources, or a restart—to reveal the boundary of the step rather than documenting only the happy path.
Calculate system cost
Check physical clearance, PSU rating and connectors, slots, CPU, RAM, storage, cooling, and noise. Avoid unsafe adapter chains.
The working rule for “Calculate system cost” is: Run a sustained load and the real model before purchase when possible. Keep the model, quantization, context length, and concurrency fixed so that a device or model comparison has a clear cause.
Define completion with an observable result instead of a general impression. Repeat the same input, and if the output changes, isolate whether the model, runtime settings, or source data changed before moving to the next stage.
Test under load
Inspect artifacts, fan noise, temperature, errors, and throttling, then run the intended GGUF model for 20–30 minutes. Document warranty and return terms.
The working rule for “Test under load” is: Start with VRAM capacity and runtime support. Keep the model, quantization, context length, and concurrency fixed so that a device or model comparison has a clear cause.
Use the checklist with a date and observed result beside every item. When one check fails, record its scope before continuing, then repeat the same input after the change; that turns a list of advice into evidence that the step actually works in this environment.
- Serial and seals
- VRAM identity and errors
- Fans, temperature, noise
- No shutdown under load
- Real model loading and speed
- Warranty and return terms
Frequently asked questions
Should every mining GPU be avoided?
No. Verify temperature, errors, fans, and sustained stability instead of relying on history alone. For a practical check, follow the “Work backward from the model” section, change one condition at a time, and record the result.
Is VRAM all that matters?
No. Software support, bandwidth, power, cooling, and system RAM matter too. For a practical check, follow the “Cross-check identity” section, change one condition at a time, and record the result.
What should I request in person?
Exact identification, a load test, the intended model run, and purchase evidence. Run a sustained load and the real model before purchase when possible. For a practical check, follow the “Calculate system cost” section, change one condition at a time, and record the result.
Primary sources
Check the original documentation for version-specific details.
Ollama hardware support NVIDIA CUDA GPUs llama.cpp repository