Mixture of Experts (MoE) Architecture: Inside DeepSeek-V3 and Qwen3
Editorial note: Reviewed August 5, 2026. Concepts are separated from fast-changing product claims; see our methodology and report corrections through contact.
Dense and Mixture-of-Experts architectures coexist in modern language models. Public technical reports for DeepSeek-V3 and members of the Qwen3 family describe sparse MoE designs. MoE increases total parameter capacity while routing each token through only part of the feed-forward network, but it does not automatically guarantee lower latency, price, or higher quality. This article separates architecture facts from deployment claims.
Dense vs. MoE: The Fundamental Difference
In a standard Dense LLM (like GPT-3.5 or Llama 2), every token is processed by every parameter in the neural network. To make a dense model smarter, you must add more parameters, which increases the computational cost (FLOPs) and inference latency for every single token generated.
In an MoE model, the feed-forward neural network (FFN) layers are split into multiple independent "experts." A routing network (or gating function) dynamically directs each input token to only a few selected experts. Consequently, although the total model size might be hundreds of billions of parameters, only a small fraction is active at any one time.
Dense vs. MoE Architecture Trade-offs
| Dimension | Dense Architecture | MoE Architecture |
|---|---|---|
| Inference Compute Cost | High (Proportional to total parameters) | Low (Proportional to active parameters only) |
| Memory Footprint | Moderate (Fits in smaller GPU rigs) | Very High (All experts must sit in memory) |
| Training Efficiency | Standard scaling laws | Highly efficient (Scales parameter capacity quickly) |
| Serving economics | Depends on model, hardware, batching, and provider | May reduce compute per token but adds routing and communication costs |
Core Pillars of MoE Design
An MoE model relies on three key mechanisms to route tokens and manage experts efficiently:
1. Gating Network (The Router)
The router is a lightweight neural network that sits at the entrance of each MoE layer. When a token (e.g., "code") enters the layer, the router evaluates it and calculates probability scores across all experts. It then forwards the token to the Top-K experts (usually Top-1 or Top-2) with the highest compatibility scores.
2. Expert Specialization
Experts can learn different activation patterns, but assigning a human-readable domain such as “Python” or “grammar” to a specific expert requires interpretability evidence. The safe claim is that the router learns token-dependent allocation. Load-balancing objectives and capacity controls are used to prevent a small number of experts from becoming overloaded.
3. Multi-Head Latent Attention (MLA)
To prevent memory bottlenecks during inference, modern MoE models like DeepSeek-V3 implement Multi-Head Latent Attention. MLA compresses the Key-Value (KV) cache matrix into a low-dimensional space, dramatically reducing VRAM usage. This allows the model to support 128K context windows while keeping memory usage manageable.
DeepSeek-V3: Pushing MoE to the Limit
DeepSeek-V3 is a massive MoE model containing 671 billion total parameters, yet it activates only 37 billion parameters per token. To achieve this efficiency, DeepSeek introduced two innovations:
- Fine-Grained Experts: Instead of having 8 or 16 large experts, DeepSeek splits the FFN layers into 256 tiny experts. This allows the router to allocate parameters with much higher precision.
- Shared Experts: Several experts are kept permanently active for all tokens to capture redundant, baseline knowledge. This frees up the remaining experts to focus on specialized, domain-specific tasks.
Why This Matters for Developers
Sparse activation can reduce arithmetic per token relative to a dense model with the same total parameter count. API price still depends on provider strategy, hardware, utilization, memory, networking, context, caching, and competition. Developers should compare dated official price pages and measure cost per accepted task rather than inferring price from architecture alone.
Activated Parameters Are Not a Hardware Quote
Only a subset of experts may contribute to each token, but the serving system still needs access to all expert weights assigned to the deployment. Memory capacity, communication topology, expert placement, batching, and load balance determine whether sparse computation becomes practical speed. “Activated parameters” is an architecture description, not a statement that the full model fits like a dense model of that size.
Expert parallelism distributes experts across devices, reducing per-device weight storage while adding all-to-all communication. Router imbalance can create stragglers. Quantization may reduce weight memory while leaving KV cache and communication as bottlenecks.
A Reproducible MoE Serving Study
- Pin the exact model revision, tokenizer, precision, runtime, drivers, and hardware topology.
- Record total and activated parameter counts from the model's primary technical report.
- Use realistic prompt/output distributions and increase concurrency to the expected service load.
- Measure time to first token, inter-token latency, request/token throughput, memory, communication utilization, queue time, and errors.
- Inspect expert load balance and compare quality before and after any quantization.
- Report power and cost per accepted task, not only peak token throughput.
Questions to Ask When Reading an MoE Claim
- Is the number about training compute, inference FLOPs, weight memory, latency, or quality?
- How many experts exist, how many are selected, and is there a shared expert?
- Were results measured with expert parallelism, tensor parallelism, or offload?
- What hardware interconnect, batch size, context, and precision were used?
- Does the benchmark include routing, communication, and serving overhead?
- Are dense and sparse models compared at equal quality, latency, cost, or parameter count?
Primary References
- DeepSeek-V3 official repository and technical report
- Qwen3 technical report
- AI Agent Hub: Local LLM deployment guide
The Verdict
MoE separates total parameter capacity from per-token sparse computation. Its real value depends on training quality and a serving system that manages weights, routing, communication, cache, and load efficiently. Verify each product claim with its technical report and a workload-specific benchmark.
What MoE pricing actually looks like
Mixture-of-Experts architectures activate only a subset of parameters per token, and the pricing reflects it: MoE models tend to sit well below dense models of similar capability. Below: one request at 100K input, 50K cached, 10K output. Of the 100,000 input tokens, 50,000 are billed at the cache-read rate and 50,000 at full input rate.
| Model | Cost per request | Monthly at 3,000 requests |
|---|---|---|
| DeepSeek V4.1 Flash | $0.014 | $40.95 |
| DeepSeek V4 Pro | $0.054 | $162 |
| Claude Sonnet 5 | $0.210 | $630 |
| GPT-5.6 Terra | $0.230 | $690 |
DeepSeek's MoE line is roughly a fifth of the dense mid-tier at this shape, and the gap is structural rather than promotional. That is the commercial reason MoE has become the default for high-volume workloads: the architecture is the discount.
Rates verified against provider documentation on September 18, 2026. Promotional rates expire, so re-check before budgeting: LLM API cost planning · September 2026 pricing update. Run your own numbers in the cost calculator.