Identify which part is slow
“One reply took a minute” combines possible delays. Was the model downloading, loading from disk, processing a long input or generating a long answer? Distinguish initial startup from repeat runs with the same model already loaded. Keep the input and output limit fixed during diagnosis.
A language model's pause before the first token differs from slow output after it starts. Prefill handles supplied context; decode generates successive tokens. NVIDIA explains those phases. A large document can make a short visible question expensive to process.
Check actual processor placement
In Ollama, run ollama ps while the model is loaded. Its PROCESSOR column identifies CPU, GPU or split placement; see the FAQ. Compare placement with the runtime's current hardware and driver support. Do not infer acceleration simply because the PC has a graphics card.
CPU use can be a valid fallback. A split model may fit but involve transfers or slower computation. For image pipelines, Hugging Face documents offloading trade-offs. Enabling every memory-saving option can slow an otherwise fitting pipeline. Diagnose the shortage before adding an optimisation.
Keep the workload modest and valid
Check the actual context setting. A model's maximum advertised window is not a setting every task needs. Ollama documents configured context and memory. Start with enough context for the job and increase only when required. Long chats can supply much more input than a fresh conversation.
For images and video, record dimensions, frame count, batch size and sampling settings. Stay within supported combinations. Reducing work may help, but arbitrary changes to required frames or resolution can cause errors. Follow the exact example before adapting it.
Measure one change at a time
Record model and precision, runtime version, hardware, input size, output limit, placement, elapsed time and whether the result is useful. Repeat the same input after warm-up, then change one setting. Comparing a two-line answer with a long essay does not isolate model speed.
For readers already using Ollama's API, its generation response documentation includes loading, prompt-processing and generation timing fields. You need not build an API tool to begin: a stopwatch and consistent task are enough for a first comparison. Record exactly which duration you measured.
Choose the next step from evidence
If loading dominates, inspect storage and whether the runtime unloads between requests. If input processing dominates, remove irrelevant history or supply the needed passage. For slow generation, inspect placement and try a supported smaller model or precision variant. For one failing workflow, inspect dependencies before replacing the installation.
Keep the baseline's settings and output so you can reverse an unhelpful change. This is diagnosis guidance, not a speed benchmark or hardware-upgrade recommendation. We have not executed these models for this page.
Sources
Primary references checked 11 October 2026. This explanation is not a benchmark of a particular model or computer.