Deployment guide · Reviewed September 18, 2026

Local LLM Deployment: Ollama vs llama.cpp vs vLLM

By AI Agent Hub Editorial Desk · Review method · Corrections

Choose the operating model first: Ollama favors a convenient local API, llama.cpp emphasizes portable inference and quantized GGUF models, and vLLM targets higher-throughput serving. Model quality and hardware fit still require your own evaluation.

“Run it locally” can mean a developer laptop, an offline workstation, an on-premises service, or a multi-user GPU server. Those environments have different security and performance requirements. This guide compares deployment roles rather than declaring one runtime universally best.

Decision summary

NeedReasonable starting pointValidate before adopting
Simple developer setup and local APIOllamaModel availability, memory use, API behavior, update process
Portable CPU/GPU inference and GGUF controlllama.cppBuild/backend support, quantization quality, server hardening
Shared GPU service with concurrencyvLLMSupported hardware/model, throughput, scheduling, container operations
Strict offline environmentAny after dependency and model mirroringNo hidden network calls, signed artifacts, patch process

This table is a shortlist, not a benchmark. Run the exact model, context length, quantization, hardware, and workload you plan to deploy.

Ollama: convenient local model API

Ollama exposes a local API after installation and provides official Python and JavaScript libraries. It is a practical entry point for personal tools and development environments where operators want model management and an API without assembling a serving stack.

Convenience does not remove production responsibilities. Bind interfaces deliberately, add authentication or a protected gateway when access extends beyond localhost, control which models may be pulled, and plan updates. Confirm whether an application's client expects Ollama's native API or an OpenAI-compatible interface.

llama.cpp: portable and configurable inference

llama.cpp provides local inference across CPU and GPU backends and commonly uses quantized GGUF model files. Its server supports web and API interfaces, including OpenAI-compatible routes. It is useful when hardware portability, quantization choices, or a compact native runtime matter.

More control means more testing. Quantization can change quality, context settings affect memory, and backend/build options affect performance. The project's security guidance recommends isolating untrusted models. Do not expose experimental or unauthenticated endpoints to an open network.

vLLM: service-oriented GPU inference

vLLM provides an OpenAI-compatible HTTP server and official deployment documentation, including container images. It is generally evaluated for shared serving where batching, concurrency, and GPU utilization matter more than a minimal laptop setup.

API compatibility is not identity. Supported endpoints and parameters differ, chat templates must match the model, and model repositories can provide generation configuration that changes defaults. Pin the runtime, model revision, tokenizer, chat template, and launch arguments.

Hardware sizing begins with memory

Model weights are only part of the footprint. Budget memory for the runtime, KV cache, context length, concurrent sequences, activations or temporary buffers, and safety margin. Quantization reduces weight memory but may affect quality and speed differently across hardware.

Measure:

Privacy is a deployment property

A locally executed model does not automatically make the whole application private. Prompts may still leave the machine through telemetry, embedding APIs, web-search tools, error reporting, package downloads, or remote model fallbacks. Document the full data flow and test egress controls.

Model and license due diligence

The runtime's open-source license does not determine the model's license. Review the model card, allowed uses, redistribution terms, geographic restrictions, and any acceptable-use policy. Record the exact model revision and hash. A model name alone is insufficient for reproducibility.

Production architecture

Place a controlled application layer between users and the inference server. That layer should authenticate callers, enforce request and output limits, select approved models, validate structured outputs, apply content and tool policies, and emit privacy-safe metrics. Do not allow users to select arbitrary local file paths or model repositories.

A reproducible three-runtime lab

Protocol, not claimed benchmark results: the following lab is designed so a reader can produce evidence on their own hardware. It deliberately avoids publishing invented tokens-per-second numbers. Use one legally obtained model family at comparable precision, pin every artifact, and record any runtime-specific conversion.

1. Freeze the test manifest

{
  "test_date": "YYYY-MM-DD",
  "host": "CPU, RAM, GPU, VRAM, driver, operating system",
  "runtime": "name and exact version or commit",
  "model": "repository and immutable revision",
  "artifact_sha256": "hash of the served weights",
  "quantization": "exact format, or none",
  "context_limit": 8192,
  "generation": {"temperature": 0, "max_output_tokens": 512},
  "concurrency_levels": [1, 4, 8]
}

Do not compare a four-bit GGUF build in one runtime with full-precision weights in another and label the result a runtime benchmark. If identical artifacts are impossible, report the artifact difference as a limitation and evaluate task quality separately.

2. Start each server with explicit boundaries

These minimal commands illustrate the variables that must be recorded. Replace model identifiers and paths with approved artifacts, keep the services on loopback during the lab, and use the current project documentation for installation.

# llama.cpp: local GGUF, explicit context and loopback binding
llama-server -m /models/model.gguf -c 8192 --host 127.0.0.1 --port 8080

# vLLM: immutable repository revision should be recorded separately
vllm serve organization/model --api-key local-test-token --generation-config vllm

# Ollama exposes its local API on the configured local service;
# record the exact model manifest returned by your installation.

The vLLM documentation notes that a model repository's generation_config.json can override sampling defaults; --generation-config vllm avoids silently inheriting those values during a controlled comparison. API compatibility also varies by supported endpoint and parameter, so run contract tests instead of assuming every OpenAI client feature is identical.

3. Use a balanced prompt set

BucketSuggested casesWhat it exposes
Short interactive20 prompts, 200–800 input tokensCold start and time to first token
Document reasoning15 prompts, 4K–8K tokensKV-cache pressure and long-input latency
Structured output10 schema-constrained tasksJSON validity and client compatibility
Adversarial operations5 oversized or malformed requestsLimits, error handling, and recovery

Run each case at least three times after a warm-up. Randomize order, preserve failures, and separate cold-start results from steady state. Quality scoring should use a fixed rubric or deterministic validator where possible.

4. Publish a result table readers can interpret

Runtime / artifactConcurrencyTTFT p50 / p95Output tok/sPeak memoryAccepted tasksErrors
Runtime A / revision1record resultrecord resultrecord resultx / 50count + cause
Runtime A / revision4record resultrecord resultrecord resultx / 50count + cause
Runtime B / revision1record resultrecord resultrecord resultx / 50count + cause

Report both successful throughput and rejection behavior. A server that accepts unlimited work and crashes is not more capable than one that applies a predictable queue or returns a clear overload response.

Estimate the memory envelope before downloading

A rough weight-only estimate is parameter count × bits per parameter ÷ 8, but it is only a lower bound. Add runtime overhead, KV cache, concurrent sequences, temporary buffers, multimodal encoders, and headroom. Measure the chosen build rather than treating the formula as capacity proof.

Memory componentDriven byHow to bound it
WeightsParameters, precision, quantization metadataArtifact size plus load-time expansion
KV cacheModel architecture, context, concurrency, cache precisionTest target and maximum context at each concurrency
Runtime buffersBackend, batching, kernels, graph captureObserve steady state and peak during warm-up
Safety marginFragmentation and co-located processesKeep measured headroom; do not allocate to the limit

Failure-injection checklist

Evaluation plan

  1. Define privacy, latency, throughput, quality, and availability requirements.
  2. Select one or two models whose licenses and capabilities fit.
  3. Test each runtime on the target hardware with fixed model artifacts.
  4. Use representative prompt lengths and concurrent traffic, not a one-line demo.
  5. Evaluate task success and safety at the selected quantization.
  6. Simulate restarts, out-of-memory errors, malformed requests, and overload.
  7. Estimate total cost: hardware, electricity, idle capacity, storage, and operator time.

When a hosted API may be better

Self-hosting is not automatically cheaper. A hosted API may be preferable when traffic is low or bursty, the team lacks GPU operations experience, or managed capability and reliability outweigh data-residency needs. A hybrid design can keep sensitive workloads local and route approved tasks externally, but the route must be explicit and auditable.

Primary references

Bottom line

Start with the runtime that matches your deployment shape, then verify the exact model, quantization, hardware, concurrency, security, and license. Local inference is an operational system—not just a model download.

The API cost you are comparing against

Self-hosting is usually justified by avoiding API spend, so the comparison needs the API number to be real. Below: the API equivalent of one local query at 20K input, 12K cached, 1.5K output, at twenty thousand queries a month. Of the 20,000 input tokens, 12,000 are billed at the cache-read rate and 8,000 at full input rate.

Model Cost per query Monthly at 20,000 querys
DeepSeek V4.1 Flash$0.0021$42.72
GPT-5.6 Luna$0.0036$72.80
Gemini 3.8 Flash$0.013$250
Claude Haiku 4.5$0.017$334

On the cheap models the API equivalent of a fairly busy month is in the tens of dollars, which is worth knowing before buying a GPU to avoid it. Self-hosting wins on privacy, latency and fixed-cost predictability; on pure token economics at this volume it usually does not.

Rates verified against provider documentation on September 18, 2026. Promotional rates expire, so re-check before budgeting: LLM API cost planning · September 2026 pricing update. Run your own numbers in the cost calculator.