Local LLM Deployment: Ollama vs llama.cpp vs vLLM
Choose the operating model first: Ollama favors a convenient local API, llama.cpp emphasizes portable inference and quantized GGUF models, and vLLM targets higher-throughput serving. Model quality and hardware fit still require your own evaluation.
“Run it locally” can mean a developer laptop, an offline workstation, an on-premises service, or a multi-user GPU server. Those environments have different security and performance requirements. This guide compares deployment roles rather than declaring one runtime universally best.
Decision summary
| Need | Reasonable starting point | Validate before adopting |
|---|---|---|
| Simple developer setup and local API | Ollama | Model availability, memory use, API behavior, update process |
| Portable CPU/GPU inference and GGUF control | llama.cpp | Build/backend support, quantization quality, server hardening |
| Shared GPU service with concurrency | vLLM | Supported hardware/model, throughput, scheduling, container operations |
| Strict offline environment | Any after dependency and model mirroring | No hidden network calls, signed artifacts, patch process |
This table is a shortlist, not a benchmark. Run the exact model, context length, quantization, hardware, and workload you plan to deploy.
Ollama: convenient local model API
Ollama exposes a local API after installation and provides official Python and JavaScript libraries. It is a practical entry point for personal tools and development environments where operators want model management and an API without assembling a serving stack.
Convenience does not remove production responsibilities. Bind interfaces deliberately, add authentication or a protected gateway when access extends beyond localhost, control which models may be pulled, and plan updates. Confirm whether an application's client expects Ollama's native API or an OpenAI-compatible interface.
llama.cpp: portable and configurable inference
llama.cpp provides local inference across CPU and GPU backends and commonly uses quantized GGUF model files. Its server supports web and API interfaces, including OpenAI-compatible routes. It is useful when hardware portability, quantization choices, or a compact native runtime matter.
More control means more testing. Quantization can change quality, context settings affect memory, and backend/build options affect performance. The project's security guidance recommends isolating untrusted models. Do not expose experimental or unauthenticated endpoints to an open network.
vLLM: service-oriented GPU inference
vLLM provides an OpenAI-compatible HTTP server and official deployment documentation, including container images. It is generally evaluated for shared serving where batching, concurrency, and GPU utilization matter more than a minimal laptop setup.
API compatibility is not identity. Supported endpoints and parameters differ, chat templates must match the model, and model repositories can provide generation configuration that changes defaults. Pin the runtime, model revision, tokenizer, chat template, and launch arguments.
Hardware sizing begins with memory
Model weights are only part of the footprint. Budget memory for the runtime, KV cache, context length, concurrent sequences, activations or temporary buffers, and safety margin. Quantization reduces weight memory but may affect quality and speed differently across hardware.
Measure:
- cold-start and model-load time;
- time to first token and tokens per second;
- maximum stable concurrency;
- memory at representative and worst-case context lengths;
- quality on your evaluation set at the selected quantization;
- power draw and idle utilization if cost matters.
Privacy is a deployment property
A locally executed model does not automatically make the whole application private. Prompts may still leave the machine through telemetry, embedding APIs, web-search tools, error reporting, package downloads, or remote model fallbacks. Document the full data flow and test egress controls.
- Bind services to localhost by default.
- Require authentication, TLS, rate limits, and network policy for shared access.
- Minimize prompt and response logging; define retention.
- Scan model and container artifacts and pin trustworthy sources.
- Keep secrets outside prompts and model files.
- Sandbox code execution and untrusted tool output separately from inference.
Model and license due diligence
The runtime's open-source license does not determine the model's license. Review the model card, allowed uses, redistribution terms, geographic restrictions, and any acceptable-use policy. Record the exact model revision and hash. A model name alone is insufficient for reproducibility.
Production architecture
Place a controlled application layer between users and the inference server. That layer should authenticate callers, enforce request and output limits, select approved models, validate structured outputs, apply content and tool policies, and emit privacy-safe metrics. Do not allow users to select arbitrary local file paths or model repositories.
A reproducible three-runtime lab
Protocol, not claimed benchmark results: the following lab is designed so a reader can produce evidence on their own hardware. It deliberately avoids publishing invented tokens-per-second numbers. Use one legally obtained model family at comparable precision, pin every artifact, and record any runtime-specific conversion.
1. Freeze the test manifest
{
"test_date": "YYYY-MM-DD",
"host": "CPU, RAM, GPU, VRAM, driver, operating system",
"runtime": "name and exact version or commit",
"model": "repository and immutable revision",
"artifact_sha256": "hash of the served weights",
"quantization": "exact format, or none",
"context_limit": 8192,
"generation": {"temperature": 0, "max_output_tokens": 512},
"concurrency_levels": [1, 4, 8]
}
Do not compare a four-bit GGUF build in one runtime with full-precision weights in another and label the result a runtime benchmark. If identical artifacts are impossible, report the artifact difference as a limitation and evaluate task quality separately.
2. Start each server with explicit boundaries
These minimal commands illustrate the variables that must be recorded. Replace model identifiers and paths with approved artifacts, keep the services on loopback during the lab, and use the current project documentation for installation.
# llama.cpp: local GGUF, explicit context and loopback binding
llama-server -m /models/model.gguf -c 8192 --host 127.0.0.1 --port 8080
# vLLM: immutable repository revision should be recorded separately
vllm serve organization/model --api-key local-test-token --generation-config vllm
# Ollama exposes its local API on the configured local service;
# record the exact model manifest returned by your installation.
The vLLM documentation notes that a model repository's generation_config.json can override sampling defaults; --generation-config vllm avoids silently inheriting those values during a controlled comparison. API compatibility also varies by supported endpoint and parameter, so run contract tests instead of assuming every OpenAI client feature is identical.
3. Use a balanced prompt set
| Bucket | Suggested cases | What it exposes |
|---|---|---|
| Short interactive | 20 prompts, 200–800 input tokens | Cold start and time to first token |
| Document reasoning | 15 prompts, 4K–8K tokens | KV-cache pressure and long-input latency |
| Structured output | 10 schema-constrained tasks | JSON validity and client compatibility |
| Adversarial operations | 5 oversized or malformed requests | Limits, error handling, and recovery |
Run each case at least three times after a warm-up. Randomize order, preserve failures, and separate cold-start results from steady state. Quality scoring should use a fixed rubric or deterministic validator where possible.
4. Publish a result table readers can interpret
| Runtime / artifact | Concurrency | TTFT p50 / p95 | Output tok/s | Peak memory | Accepted tasks | Errors |
|---|---|---|---|---|---|---|
| Runtime A / revision | 1 | record result | record result | record result | x / 50 | count + cause |
| Runtime A / revision | 4 | record result | record result | record result | x / 50 | count + cause |
| Runtime B / revision | 1 | record result | record result | record result | x / 50 | count + cause |
Report both successful throughput and rejection behavior. A server that accepts unlimited work and crashes is not more capable than one that applies a predictable queue or returns a clear overload response.
Estimate the memory envelope before downloading
A rough weight-only estimate is parameter count × bits per parameter ÷ 8, but it is only a lower bound. Add runtime overhead, KV cache, concurrent sequences, temporary buffers, multimodal encoders, and headroom. Measure the chosen build rather than treating the formula as capacity proof.
| Memory component | Driven by | How to bound it |
|---|---|---|
| Weights | Parameters, precision, quantization metadata | Artifact size plus load-time expansion |
| KV cache | Model architecture, context, concurrency, cache precision | Test target and maximum context at each concurrency |
| Runtime buffers | Backend, batching, kernels, graph capture | Observe steady state and peak during warm-up |
| Safety margin | Fragmentation and co-located processes | Keep measured headroom; do not allocate to the limit |
Failure-injection checklist
- Send a request above the configured context limit and verify a bounded, documented error.
- Fill the concurrency queue and confirm overload does not crash the server.
- Terminate a client mid-stream and check that server resources are released.
- Restart during model load and verify the service returns to a known configuration.
- Present malformed JSON, an unknown model name, and unsupported parameters.
- Block outbound network access and confirm inference still works with mirrored dependencies.
- Rotate the API credential and confirm the previous credential stops working.
- Attempt access from outside the intended network boundary.
Evaluation plan
- Define privacy, latency, throughput, quality, and availability requirements.
- Select one or two models whose licenses and capabilities fit.
- Test each runtime on the target hardware with fixed model artifacts.
- Use representative prompt lengths and concurrent traffic, not a one-line demo.
- Evaluate task success and safety at the selected quantization.
- Simulate restarts, out-of-memory errors, malformed requests, and overload.
- Estimate total cost: hardware, electricity, idle capacity, storage, and operator time.
When a hosted API may be better
Self-hosting is not automatically cheaper. A hosted API may be preferable when traffic is low or bursty, the team lacks GPU operations experience, or managed capability and reliability outweigh data-residency needs. A hybrid design can keep sensitive workloads local and route approved tasks externally, but the route must be explicit and auditable.
Primary references
- Ollama API introduction
- llama.cpp server documentation
- llama.cpp security guidance
- vLLM OpenAI-compatible server
- vLLM container deployment
Bottom line
Start with the runtime that matches your deployment shape, then verify the exact model, quantization, hardware, concurrency, security, and license. Local inference is an operational system—not just a model download.
The API cost you are comparing against
Self-hosting is usually justified by avoiding API spend, so the comparison needs the API number to be real. Below: the API equivalent of one local query at 20K input, 12K cached, 1.5K output, at twenty thousand queries a month. Of the 20,000 input tokens, 12,000 are billed at the cache-read rate and 8,000 at full input rate.
| Model | Cost per query | Monthly at 20,000 querys |
|---|---|---|
| DeepSeek V4.1 Flash | $0.0021 | $42.72 |
| GPT-5.6 Luna | $0.0036 | $72.80 |
| Gemini 3.8 Flash | $0.013 | $250 |
| Claude Haiku 4.5 | $0.017 | $334 |
On the cheap models the API equivalent of a fairly busy month is in the tens of dollars, which is worth knowing before buying a GPU to avoid it. Self-hosting wins on privacy, latency and fixed-cost predictability; on pure token economics at this volume it usually does not.
Rates verified against provider documentation on September 18, 2026. Promotional rates expire, so re-check before budgeting: LLM API cost planning · September 2026 pricing update. Run your own numbers in the cost calculator.