Local DeepSeek-R1 & V3 Deployment with Ollama

By AI Agent Hub Editorial Desk · Review method · Corrections

Tutorial · 5 min read · Reviewed September 18, 2026

Model currency: this guide targets DeepSeek-R1 and V3, which remain widely deployed open-weights models. DeepSeek’s current generation is V4.1, also released with open weights under the MIT license. Check the current model card and repository before sizing hardware or committing to a variant.

DeepSeek-R1 and V3 were widely adopted open-weights reasoning releases and remain common choices for local hosting. Running DeepSeek models locally gives developers complete privacy, zero API rate limits, and zero recurring token costs. This guide covers local deployment using Ollama across NVIDIA GPUs and Apple Silicon MacBooks.

1. Selecting the Right DeepSeek-R1 Model Distill Flag

Depending on your hardware VRAM capacity, select the appropriate parameter quantization size:

Model Variant Minimum VRAM Recommended Hardware Use Case
deepseek-r1:7b 6 GB VRAM RTX 3060 / M1 Mac (8GB) Lightweight coding & autocomplete
deepseek-r1:14b 12 GB VRAM RTX 4070 / M2 Mac (16GB) Daily code refactoring & logic reasoning
deepseek-r1:32b 24 GB VRAM RTX 4090 / M3 Max (36GB) Complex multi-file architectural reasoning
deepseek-r1:70b 48 GB VRAM 2x RTX 3090 / M2 Ultra (64GB) Strong local reasoning on complex tasks

2. Installation & One-Command CLI Run

Install Ollama and pull the target DeepSeek-R1 model quantized build:

# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh

# Run 14B DeepSeek reasoning model
ollama run deepseek-r1:14b

# Verify REST API endpoint
curl http://localhost:11434/api/generate -d '{
  "model": "deepseek-r1:14b",
  "prompt": "Write a Python script for asynchronous web scraping."
}'

3. Connecting to VS Code & Cursor IDE

Set Ollama's base URL in VS Code extensions (like Continue.dev or OpenCode) to http://localhost:11434/v1 for full local inline autocomplete and agent chat.

4. Verify the Artifact Before You Benchmark It

A library tag may point to a distilled model, a quantized conversion, or a later revision rather than the original full checkpoint described in a research paper. Record the exact tag, digest, parameter scale, base model, quantization, license, Ollama version, and model metadata. Do not describe a small distilled variant as the full DeepSeek-R1 system.

ollama pull deepseek-r1:8b
ollama show deepseek-r1:8b
ollama ps
ollama --version

Check the official Ollama library page before copying a command because available tags change. Confirm the model's context configuration and increase it only after measuring memory. KV cache grows with context and concurrency, so a model that loads successfully can still run out of memory under realistic traffic.

5. A Reproducible Local Evaluation

  1. Freeze 30–50 prompts from the actual task: code review, extraction, question answering, or drafting.
  2. Record CPU, GPU, VRAM, system RAM, operating system, driver, runtime, model digest, and context setting.
  3. Warm the model, then run several repetitions at concurrency one and at the expected concurrent load.
  4. Measure accepted-task quality, time to first token, generation rate, end-to-end latency, memory peak, and failure rate.
  5. Test long prompts, cancellation, malformed input, model reload, and memory pressure.
  6. Compare with a cloud or alternative local baseline using the same rubric; include hardware and operations cost.

Do not publish a speed result without the full configuration. CPU offload and unified memory can make a model run but may change latency dramatically. The correct model size is the smallest one that meets the quality floor under the real context and concurrency requirement.

6. Local Does Not Automatically Mean Private

Bind the server to the minimum required interface, add authentication at a reverse proxy when other machines can reach it, and restrict browser origins. Review IDE extensions and web interfaces because they may send telemetry or prompts elsewhere. Protect model files and logs, patch the runtime, and avoid loading unreviewed custom code. For business data, define retention and deletion even if inference never leaves the device.

7. Update and Rollback Discipline

A tag can move to a newer artifact. Pull updates into a staging environment, record the new digest, rerun the frozen quality and load suite, and verify the chat template and API contract. Keep the previous digest locally until the new version survives a canary period. Never update the runtime and model in the same change if you need to identify the cause of a regression.

8. Primary References

9. Related Hosting & Benchmark Guides

What the hosted equivalent costs

Running DeepSeek locally is a reasonable choice for privacy and for fixed costs, but it helps to know what the same traffic would cost on the hosted API before deciding. Below: one request at 15K input, 8K cached, 1.5K output, at ten thousand a month. Of the 15,000 input tokens, 8,000 are billed at the cache-read rate and 7,000 at full input rate.

Model Cost per request Monthly at 10,000 requests
DeepSeek V4.1 Flash$0.0020$19.74
GPT-5.6 Luna$0.0034$33.60
DeepSeek V4 Pro$0.0078$77.66
Gemini 3.8 Flash$0.011$115

Note that the two DeepSeek entries differ by roughly fourfold for the same workload: the model choice inside one provider's catalogue matters more than the choice between local and hosted at this volume.

Rates verified against provider documentation on September 18, 2026. Promotional rates expire, so re-check before budgeting: LLM API cost planning · September 2026 pricing update. Run your own numbers in the cost calculator.