DSPy Programmatic Prompt Optimization & Compilation Guide
Hand-crafted prompt engineering is brittle, non-reproducible, and fragile across LLM model upgrades. Stanford's DSPy framework replaces manual prompt tweaking with declarative programming. DSPy compiles your system prompts, automatically selects optimal few-shot examples, and tunes model instructions against a metric evaluation function.
1. Why Programmatic Prompt Compilation Beats Manual Prompting
- Model portability: a program can be recompiled for another supported model, but the new candidate still requires a full holdout and safety evaluation.
- Automatic Few-Shot Bootstrapping: Teleprompters automatically discover and synthesize effective few-shot prompt demonstrations.
- Metric-Driven Optimization: Tunes prompts against measurable metrics (accuracy, JSON validity, exact match).
2. Python DSPy Program Example
import dspy
# 1. Configure Language Model Backend
lm = dspy.LM('openai/deepseek-v3', api_key='YOUR_API_KEY')
dspy.configure(lm=lm)
# 2. Define Signature Input/Output Spec
class RAGTask(dspy.Signature):
"""Answer questions using retrieved context snippets."""
context = dspy.InputField(desc="Retrieved document chunks")
question = dspy.InputField()
answer = dspy.OutputField(desc="Concise factual answer")
# 3. Define Chain-of-Thought Module
class RAGModule(dspy.Module):
def __init__(self):
super().__init__()
self.generate_answer = dspy.ChainOfThought(RAGTask)
def forward(self, context, question):
return self.generate_answer(context=context, question=question)
# 4. Compile with BootstrapFewShot Teleprompter
from dspy.teleprompt import BootstrapFewShot
teleprompter = BootstrapFewShot(metric=dspy.evaluate.answer_exact_match)
compiled_rag = teleprompter.compile(RAGModule(), trainset=train_data)
3. What DSPy Actually Optimizes
A DSPy program declares input and output fields, composes modules, and evaluates predictions with a metric. An optimizer searches for better instructions, demonstrations, or model parameters using a training set and traces. It does not remove the need to define the task. If the metric rewards the wrong behavior, compilation can produce a prompt that scores well while failing users.
Start with a narrow signature and a metric that checks the real contract. For a cited support answer, the metric might require the correct disposition, verify that every citation belongs to the supplied evidence, and award partial credit for an appropriate abstention. A generic semantic-similarity score would miss fabricated citations and policy errors.
4. Reproducible Optimization Workflow
- Version the examples: record source, label provenance, consent, and deduplication decisions.
- Split by entity or time: keep near-duplicate tickets, documents, or users out of different splits.
- Freeze a baseline: measure the uncompiled program with the same model, tools, and decoding settings.
- Compile on training data: set an explicit call and cost budget; save the resulting program artifact.
- Select with validation data: compare candidate programs without touching the final test set.
- Evaluate once on holdout: report uncertainty and inspect failures, not just the aggregate score.
- Canary in production: watch quality, latency, cost, and safety separately.
5. Minimal Metric Design
def metric(example, prediction, trace=None):
label_ok = prediction.label == example.label
grounded = all(c in example.allowed_citations for c in prediction.citations)
abstention_ok = example.answerable or prediction.label == "needs_review"
return float(label_ok and grounded and abstention_ok)
This example intentionally uses deterministic checks. For open-ended quality, add a rubric-based human or calibrated model grader, but keep safety and authorization conditions as hard failures. Store optimizer version, random seed, model snapshot, program hash, dataset hash, metric code, and total optimization calls so the result can be reproduced.
6. Failure Modes to Check
- Leakage: demonstrations contain the same entities or answers as the holdout.
- Metric gaming: a program learns formatting shortcuts that satisfy the scorer without solving the task.
- Overfitting: validation improves while a later time slice degrades.
- Cost blindness: a small quality gain requires many more calls or much larger prompts.
- Provider drift: a compiled prompt is assumed to transfer unchanged to another model.
- Unsafe examples: secrets or personal data are copied into demonstrations and traces.
7. Primary References
8. Change Control in CI
Store the compiled program as a versioned artifact and block silent recompilation in production. A pull request that changes the model, optimizer, metric, examples, or prompt budget should run the same evaluation suite and show paired results. Require review when quality drops on any safety-critical slice even if the average improves. Keep the previous artifact available for rollback and monitor whether live input distribution drifts away from the optimization data.
9. Related Builder Tools & Tutorials
- 🛠️ Interactive AI Agent Builder Lab & Code Studio
- 🔄 Agentic RAG Architecture & Self-Correction Pipelines
What compiling a program costs
Optimisation is the expensive part of DSPy, not inference: the compiler evaluates many candidate prompts over a training set before it emits anything. It is also the part people budget least carefully. Below: one compilation at 2M input with no cache benefit and 200K output.
| Model | Cost per compilation | Monthly at 20 compilations |
|---|---|---|
| GPT-5.6 Luna | $0.640 | $12.80 |
| Claude Sonnet 5 | $6.00 | $120 |
| GPT-5.6 Terra | $6.40 | $128 |
Output dominates because the compiler is generating and scoring candidates, not just reading them. Compiling against a small model and deploying the result to a large one is the usual answer, and the arithmetic above is why: the same compilation is roughly thirty times cheaper on Luna.
Rates verified against provider documentation on September 18, 2026. Promotional rates expire, so re-check before budgeting: LLM API cost planning · September 2026 pricing update. Run your own numbers in the cost calculator.