Fifty engineers open an internal assistant at once. The first ten get fast answers, and the rest wait even though the GPU dashboard shows 60% utilization.
You have two plausible fixes: buy a larger GPU or change the serving runtime, the software that schedules model work. The dashboard doesn't tell you which one addresses the wait. Requests can pile up when work is admitted in the wrong order, memory is reserved poorly, or prompts vary in length.
Start with the deployment surface
First separate a packaged local product from a library embedded in your application and a shared GPU server. Then check whether each candidate supports your exact model. Those filters remove incompatible choices before any speed comparison.
Selection rule: Deployment surface narrows the field. Exact checkpoint support decides what can run. Your latency service-level objective (SLO), throughput, and operating cost decide what should run.
Ollama packages model download, lifecycle, a local daemon, command-line interface (CLI), and application programming interfaces (APIs) into one local product. One supported execution backend is llama.cpp. Ollama also provides an MLX (Apple's machine-learning framework) path on Apple Silicon. The March 2026 announcement introduced it as a preview; check the installed release and model rather than assuming every Ollama model uses the same backend.[1][2]
llama.cpp is the lower-level runtime and library. It reads GGUF model files (a format for model tensors and metadata), builds across CPU and GPU backends, and includes llama-server, an OpenAI-compatible HTTP server.[3] Ollama packages the local workflow. llama.cpp gives an application more control over embedding the runtime, builds, and runtime flags. Here, embedding means incorporating a library, not generating an embedding vector. Local LLM Deployment covers hardware and memory sizing; Run Gemma 4 Locally with Ollama works through local model tags.
The server-oriented engines expose more cluster scheduling, parallelism, and model-serving controls. That extra control only helps when the deployment needs it, so choose this branch because you have shared traffic to serve, not because the product name sounds more serious.
Their release state changes quickly. Pin a version before comparing behavior.
| Project | Stable release inspected on September 23, 2026 | Code license | Main deployment surfaces | Hardware scope to verify |
|---|---|---|---|---|
| vLLM | v0.30.0[4] | Apache-2.0[5] | Python offline API, OpenAI-compatible server, distributed serving | Native NVIDIA, AMD, Intel GPU, and CPU paths plus hardware plugins such as TPU and Gaudi[6][7][8] |
| SGLang | v0.5.20[9] | Apache-2.0[10] | SGLang language frontend, Python engine, OpenAI-compatible server, distributed serving | Official docs list NVIDIA, AMD, Intel Xeon, Google TPU, Ascend NPU, and Moore Threads MUSA; feature parity varies[11] |
| TensorRT-LLM | v1.2.1[12] | Apache-2.0 codebase; bundled components can differ[13] | Python LLM, trtllm-serve, benchmarking and distributed NVIDIA serving | NVIDIA GPUs; check model-by-feature and GPU-generation support matrix |
| Ollama | v0.34.3[14] | MIT[1] | Local daemon, CLI, REST API, OpenAI-compatible endpoints[15] | macOS, Windows, Linux, NVIDIA/AMD GPUs, and Apple paths; exact backend depends on platform |
| llama.cpp | v0.4.1[16] | MIT[3] | C/C++ library, CLI, local OpenAI-compatible server | CPU plus Metal, CUDA, HIP, Vulkan, SYCL, and other compiled backends |
These are code licenses, not model licenses. A runtime's license doesn't grant permission to use a checkpoint, tokenizer, dataset, or remote model code.
Treat these stable tags as inspection snapshots, not upgrade recommendations. Prereleases, nightlies, and unversioned documentation can describe different behavior. Record the deployed tag and container digest, not just the project name.
Release numbers still don't prove compatibility. Validate the exact checkpoint, quantization, attention backend, tool parser, structured-output path, and parallel layout. "Supports the architecture" can still mean your chosen feature combination is untested.
What an inference engine controls
Follow one request through the runtime. First, prefill processes all prompt tokens and writes the request's key-value (KV) state, the attention data reused at each output step. Then decode reads that state while producing one new token at each request step.
That split gives us two useful clocks. Client-observed time to first token (TTFT) includes network transit, queueing, scheduling, prefill, and delivery of the first output token. Time per output token (TPOT) averages the remaining decode duration over the remaining output tokens. For 101 output tokens whose first and last arrivals are 5 seconds apart, TPOT is 5/100 = 0.05 seconds, or 50 ms. Its reciprocal is 20 tokens/s for that request after the first token. Aggregate output tokens per second also depends on concurrent requests and batching.
An average can hide pauses. Inter-token latency measures the gaps between successive streamed tokens; if the client receives chunks rather than token-level events, record that limitation. A smooth-looking average TPOT doesn't prove that every token arrived smoothly.
The Inference: TTFT, TPS & KV Cache lesson derives those metrics and the cache-size formula. For this comparison, remember that a runtime controls four coupled resources:
- Admission and scheduling: which prefills and decodes enter each step.
- KV memory: which request owns each block and when blocks can be reused.
- Kernels and parallelism: how model work maps to target hardware.
- API and operations: model loading, streaming, metrics, failure recovery, and rollout artifacts.

Now return to the 60% dashboard. Low utilization can come from a host-side scheduler gap, a long prefill blocking decode, insufficient queued work, collective communication, or memory admission. GPU utilization alone can't distinguish them.
Use a request trace to choose an engine. The name on the server matters less than which part of this path is holding the queue.
Count work that meets the latency target
Suppose the assistant's target is TTFT at most 500 ms and average TPOT at most 50 ms. These are illustrative requirements, not universal limits. Compare two fictional 10-second runs on the same hardware and request mix, each with ten offered requests. Every completed response contains 100 output tokens; failed requests emit none in this simplified example.
| Result over 10 seconds | Configuration A | Configuration B |
|---|---|---|
| Completed requests | 8 | 10 |
| Completed within both latency targets | 6 | 4 |
| Failed requests | 2 | 0 |
| Output throughput | 80 tokens/s | 100 tokens/s |
| SLO-compliant output throughput | 60 tokens/s | 40 tokens/s |
| SLO-compliant request throughput | 0.6 requests/s | 0.4 requests/s |
B produces more tokens, but A completes more requests within the target. Goodput counts work that satisfies a declared success contract. Here the contract includes completion and both latency limits; an actual deployment must also check answer quality and endpoint correctness. Neither result automatically justifies release: A's 20% failure rate may violate a separate reliability gate.

The short calculation below reconstructs the table from per-request records. Each tuple contains completion status, output-token count, TTFT in milliseconds, and average TPOT in milliseconds. Missing timing fields on failures never qualify as success. All ten offered requests remain in the accounting.
1from math import isfinite
2
3def summarize(rows, seconds, ttft_limit=500, tpot_limit=50):
4 if not isfinite(seconds) or seconds <= 0:
5 raise ValueError("Measurement duration must be positive and finite")
6 completed = good = tokens = good_tokens = 0
7 for ok, output_tokens, ttft_ms, tpot_ms in rows:
8 if output_tokens < 0:
9 raise ValueError("Output-token count must be nonnegative")
10 tokens += output_tokens # Includes partial output from failed requests.
11 completed += int(ok)
12 timing_valid = all(
13 value is not None and isfinite(value) and value >= 0
14 for value in (ttft_ms, tpot_ms)
15 )
16 if ok and timing_valid and ttft_ms <= ttft_limit and tpot_ms <= tpot_limit:
17 good += 1
18 good_tokens += output_tokens
19 return {
20 "offered": len(rows), "completed": completed,
21 "failed": len(rows) - completed, "good": good,
22 "output_tps": tokens / seconds,
23 "good_output_tps": good_tokens / seconds,
24 "good_rps": good / seconds,
25 }
26
27runs = {
28 "A": [(True, 100, 400, 40)] * 6
29 + [(True, 100, 700, 40)] * 2
30 + [(False, 0, None, None)] * 2,
31 "B": [(True, 100, 400, 40)] * 4
32 + [(True, 100, 400, 70)] * 6,
33}
34for name, rows in runs.items():
35 result = summarize(rows, seconds=10)
36 print(f"{name}: completed={result['completed']}/10, good={result['good']}/10, "
37 f"output={result['output_tps']:.0f} tok/s, "
38 f"goodput={result['good_output_tps']:.0f} tok/s, "
39 f"goodput={result['good_rps']:.1f} req/s")1A: completed=8/10, good=6/10, output=80 tok/s, goodput=60 tok/s, goodput=0.6 req/s
2B: completed=10/10, good=4/10, output=100 tok/s, goodput=40 tok/s, goodput=0.4 req/sFor a real run, choose a measurement-window policy before collecting data. If requests cross window boundaries, don't divide whole-request token counts by an interval that contains only part of their execution. Use token timestamps within a steady-state window, or a complete cohort with its full elapsed duration. Report offered, completed, failed, and cancelled requests alongside the throughput numbers.
Benchmark boundary: Same model name isn't enough. Match model revision, precision, tokenizer, prompt template, output limits, hardware, workload, and correctness checks before attributing a difference to the runtime.
Those fifty engineers may also share a long runbook prefix. Unique prompts, repeated prefixes, and memory pressure exercise different parts of the serving path. Before comparing cache hit rates, separate two mechanisms: packing request memory and reusing work across requests.
Why block allocation changes capacity
Return to the shared runbook. If many requests repeat the same long prefix, the runtime needs enough KV memory to keep them active at once. vLLM's original PagedAttention work made that allocation problem a first-class serving concern. Instead of reserving one maximum-length contiguous region per request, the runtime assigns fixed-size physical blocks as tokens arrive and keeps a logical block table.[17]
Use the same 18-block pool on both sides. A fixed policy reserves six blocks for each request. Requests A, B, and C currently need three, four, and two blocks: nine used, nine reserved but empty, and no unreserved capacity. On-demand allocation gives them only their nine live blocks. It can then admit D and E using three and two blocks, leaving four blocks free.
The diagram compares a hypothetical fixed-reservation policy with paged allocation, not two named modern engines. No shared prefix is needed for this benefit. Prefix caching is a separate optimization that can reuse already computed KV blocks across matching requests. Real allocators also face partially filled tails, metadata, eviction, and future growth. Four free blocks now don't guarantee all five requests can finish without preemption. KV Cache & PagedAttention derives those capacity constraints.

With the memory problem visible, the engine comparison has a sharper question: which runtime gives this workload the scheduling, caching, and hardware behavior it needs?
What each runtime optimizes
vLLM
vLLM is a strong first server candidate when you need broad model coverage, familiar APIs, several accelerator paths, and a large operational ecosystem.[18]
Its design combines paged KV memory, continuous batching, chunked prefill, parallel execution, and automatic prefix caching (APC). APC reuses matching prompt KV state to skip prefill work; it doesn't eliminate decoding the new answer.[19] Shared-prefix reuse isn't exclusive to SGLang.
A fair comparison therefore measures both engines with prefix caching enabled, compatible cache isolation, and cache-aware routing. Prefix Caching and Prompt Caching covers the hit-rate mechanics. Use Deep Dive - vLLM when scheduler, block-pool, or preemption behavior is part of the decision.
SGLang
SGLang combines a high-performance server runtime with a programming frontend for multi-call language-model programs. Its founding work introduced RadixAttention, a radix-tree design that reuses exact token prefixes across requests.[20]
That mechanism can help chat, retrieval-augmented generation (RAG), few-shot prompts, and agent loops that repeat large prefixes. It doesn't guarantee a win. Template changes, random identifiers near the prompt front, replica routing, or early branch divergence can erase reuse.
SGLang also offers broad hardware paths, structured generation, speculative decoding, and distributed serving. "Prefix specialist" is now too narrow a label.[11] Deep Dive - SGLang traces those mechanisms through its runtime.
TensorRT-LLM
TensorRT-LLM is an NVIDIA-focused serving stack with optimized attention, quantization, CUDA graphs (reusable GPU execution plans), parallelism, KV-cache reuse, speculative decoding, and disaggregated-serving paths that place prefill and decode on separate workers.[13][21]
Its decision boundary is support-matrix fit, not a universal compile-versus-flexibility story. NVIDIA's current migration guide says the TensorRT engine backend has been removed: the PyTorch backend loads supported Hugging Face checkpoints without the old trtllm-build step.[22] That rolling guide is not the same artifact as the stable 1.2.1 tag in the table. Don't apply a legacy compiled-engine recipe to a newer release without checking its migration notes.
Pin one line, check its exact model and feature matrix, then benchmark it against vLLM and SGLang on the same NVIDIA system. A strong result for one dense model doesn't transfer automatically to another architecture, quantization, or GPU generation.
Ollama and llama.cpp
Ollama is the shortest path to a managed local daemon, model tags, REST APIs, and OpenAI-compatible local endpoints.[15] It can import GGUF artifacts, but model conversion and quantization quality remain separate decisions.[23]
Ollama also offers cloud models. A cloud model selected through the local daemon still executes remotely, so check the chosen tag before treating a local endpoint as a local-compute or data-residency boundary.[24]
Choose llama.cpp directly when you need an embedded C/C++ library, a custom build, explicit backend control, or its lightweight HTTP server. Choose Ollama when model lifecycle and local integration matter more than exposing every runtime knob.
Neither is automatically disqualified from a small service. Measure its concurrency, memory behavior, authentication boundary, cancellation handling, and endpoint compatibility under that service's load.
Exact checkpoints still need a recipe
"Supports the architecture" is not the same as a vendor launch command for the checkpoint you actually downloaded.
GLM-5.2 documents a 1M-token context and lists official SGLang (v0.5.13.post1+) and vLLM (v0.23.0+) recipes on its model card.[25] Use those recipes as the first compatibility tests.
DeepSeek-V4-Flash-0731 ships a DSpark speculative module in the same checkpoint. Its published recipes name vLLM method: dspark and SGLang --speculative-algorithm DSPARK explicitly.[26] The V4 technical report frames that family around million-token context.[27]
Generic Transformers support, or "we served an older DeepSeek," isn't proof that sparse attention, speculative heads, quantization, and the tokenizer path work together. Start from the vendor recipe, pin the runtime and checkpoint, then verify load, tool calls, long context, and throughput on the GPUs you'll actually use. Ollama can import community GGUF files, but these checkpoints don't belong in the laptop row above.
Recipe gate: Architecture support is not a serving recipe. Pin the vendor command, then prove load and correctness on your GPUs.
Keep adjacent layers separate
Text Generation Inference (TGI) is in maintenance mode. Hugging Face directs new work toward engines including vLLM, SGLang, llama.cpp, and MLX.[28] That makes TGI a migration consideration, not another row in this engine bake-off.
NVIDIA Dynamo sits above an engine. It can coordinate TensorRT-LLM, vLLM, or SGLang across workers and supports patterns such as disaggregated prefill and decode.[29] Comparing Dynamo with one single-node engine mixes orchestration and execution layers, so put it in a separate architecture decision.
Run a bake-off that can reverse the decision
Once the candidate set is small, write down the workload before running commands. This contract keeps a benchmark from quietly changing the question.
First record:
- exact model revision, tokenizer, chat template, quantization, and allowed remote code;
- target GPU or accelerator, driver, container, engine version, and parallel topology;
- prompt and output length distributions, not one average;
- concurrency, arrival bursts, shared-prefix fraction, tool or schema usage, and cancellation rate;
- latency SLOs for p50 and tail TTFT, TPOT or inter-token latency, and end-to-end completion;
- quality, correctness, memory headroom, deployment time, rollback time, and staff effort.
Run at least three traffic shapes: mostly unique prompts, repeated prefixes, and the real production mix. Unique prompts isolate general serving efficiency. Shared prefixes test cache reuse. The production mix tells you whether either result matters to users.
Distinguish a fixed-concurrency test, where a new request follows a completion, from an arrival-rate test, where requests arrive on a schedule. Fixed concurrency reduces offered traffic when a server slows; it can hide queue growth under a real burst. Sweep offered load, show tail latency and failures at each point, and report repeated runs rather than selecting the best run.
Warm each engine the same way and hold model semantics constant. Test cold starts and cold caches separately from warmed operation. Report output tokens per second separately from total tokens per second, and pair aggregate throughput with SLO-compliant throughput. Keep failed and cancelled requests in the offered count, and report preemptions as a separate operational signal. Compare cost per successful request only after the candidates pass the same quality and reliability gates.

Reversal test: Write down what result would make you keep the incumbent before running the benchmark. A migration without a reversal threshold is a preference disguised as measurement.
Which engine to start with
| Need | Start here | What must be proved before committing |
|---|---|---|
| Packaged local model workflow | Ollama | Artifact quality, memory fit, endpoint behavior, latency, and rollback |
| Embeddable or highly portable local runtime | llama.cpp | Backend build, GGUF quality, context capacity, and target-device speed |
| Shared GPU service with broad model churn | vLLM and SGLang bake-off | Exact model-feature compatibility, SLO-compliant throughput, and operational fit |
| Prefix-heavy chat, RAG, or agent traffic | SGLang and vLLM with caching enabled | Cache hit rate, routing locality, TTFT, memory, and correctness under real templates |
| Stable NVIDIA deployment | Add TensorRT-LLM | Version-line support matrix, same-workload performance, failure recovery, and upgrade path |
| No meaningful measured gap | Keep incumbent | Confirm capacity headroom and document next review trigger |
The table turns the earlier trade-offs into a starting rule. It doesn't replace the bake-off: every row still asks you to prove compatibility, SLO behavior, and operational fit.
The engine is only one control surface. Scaling LLM Inference connects runtime choice to batching, KV-cache sizing, quantization, disaggregation, and the throughput-latency-cost operating point.