How to Build an AI Agent from Scratch: Complete Developer Guide
Building an AI agent goes far beyond sending a prompt to a language model. An autonomous agent combines a reasoning LLM brain, memory systems, tool execution routines, and a self-correcting loop. This tutorial walks you through building a production-ready AI agent from scratch.
1. What Distinguishes an AI Agent from a Standard LLM Call?
A standard LLM call operates as a single-turn function: Input Prompt → Model → Text Output. In contrast, an AI Agent operates autonomously inside an execution loop:
Perception (User Input) → Reasoning (Thought) → Tool Selection (Action) → Sandbox Execution (Observation) → Evaluation → Self-Correction
2. The 4 Pillars of Agent Architecture
- Proposal Layer: A currently supported model processes bounded observations and proposes a typed next step; application policy decides whether it may execute.
- Memory Subsystem: Short-term conversation history buffer combined with long-term vector search (ChromaDB / Pinecone) or knowledge graphs.
- Tool Execution Routines: Functions registered via JSON schema or Model Context Protocol (MCP) allowing the model to perform web searches, file edits, or terminal operations.
- Control Flow Pattern: The execution state machine—such as a ReAct (Reason + Act) loop, Plan-and-Solve framework, or LangGraph cyclic graph.
3. Complete Step-by-Step Python Implementation
Below is a clean, dependency-minimal Python script demonstrating a working ReAct agent loop:
import json
from openai import OpenAI
# Initialize client (works with DeepSeek, OpenAI, or local vLLM)
client = OpenAI(api_key="your-api-key")
# 1. Define Tool Functions
def calculate_budget(daily_cost: float, days: int) -> str:
return json.dumps({"monthly_total": daily_cost * days})
tools_schema = [
{
"type": "function",
"function": {
"name": "calculate_budget",
"description": "Calculate total monthly budget from daily spend",
"parameters": {
"type": "object",
"properties": {
"daily_cost": {"type": "number"},
"days": {"type": "integer"}
},
"required": ["daily_cost", "days"]
}
}
}
]
# 2. Agent Execution Loop
def run_agent_loop(prompt: str):
messages = [{"role": "user", "content": prompt}]
while True:
response = client.chat.completions.create(
model="deepseek-v3",
messages=messages,
tools=tools_schema
)
msg = response.choices[0].message
# Check if model wants to invoke a tool
if msg.tool_calls:
tool_call = msg.tool_calls[0]
args = json.loads(tool_call.function.arguments)
print(f"Agent Action: Calling {tool_call.function.name} with {args}")
# Execute tool locally
result = calculate_budget(**args)
# Append observation to history
messages.append(msg)
messages.append({"role": "tool", "tool_call_id": tool_call.id, "content": result})
else:
# Model finished reasoning
print(f"Final Agent Answer: {msg.content}")
break
run_agent_loop("If our API spend is $5.5 per day, what is our 30-day budget?")
4. Define the Task Contract Before Adding Autonomy
Write down the trigger, allowed inputs, evidence sources, tools, side effects, stop conditions, and human approval points. A useful first agent has one bounded job. For example: read one bug report, search an approved repository, propose a patch in an isolated branch, run tests, and stop for review. It may not merge, deploy, message a customer, or access another repository.
| Contract item | Example | Enforced by |
|---|---|---|
| Identity | Signed-in repository contributor | Application authentication |
| Read scope | One repository at a pinned revision | Workspace and API allowlist |
| Write scope | Ephemeral branch only | Scoped credential and branch policy |
| Budget | Eight turns, twelve tool calls, ninety seconds | Orchestrator counters |
| Success | Patch produced and specified tests pass | Deterministic verifier |
| Approval | Human reviews diff before publication | Separate privileged action |
5. Build the Smallest Observable Loop
- Construct context: include durable policy, the current task, and only the evidence needed for the next decision.
- Ask for a typed proposal: answer, stop, request input, or call one of a narrow set of tools.
- Validate: check schema, authorization, budgets, paths, URLs, and action class outside the model.
- Execute: run the tool in an isolated environment and return a bounded structured result.
- Record: persist the state transition, evidence IDs, decision, result, cost, and stop reason.
- Verify: use tests and business rules; require approval before consequential side effects.
Do not add multiple agents merely to imitate an organization chart. Add a second role only when it provides a measurable boundary, such as an independent verifier with no write access or a specialist that uses a distinct evidence source. A deterministic function is often a better router than another model call.
6. Evaluation and Failure Handling
Create fixtures before launch: ordinary success, incomplete input, tool timeout, stale state, conflicting evidence, permission denial, prompt injection inside a file, and budget exhaustion. Measure accepted-task success, unsafe-action rate, unnecessary tool calls, human override rate, latency, and cost. Preserve the complete trajectory so a failure can be attributed to context, model decision, policy, tool, or verifier.
Every run needs a visible terminal state. Retry only typed transient failures and use idempotency keys for operations that may already have succeeded. After repeated equivalent failures, stop and ask for help instead of rewriting the prompt indefinitely.
7. Production Readiness Checklist
- Credentials are scoped, short-lived, and unavailable to the model text.
- Untrusted documents and tool output cannot change policy or authorization.
- Filesystem and network access use explicit allowlists.
- High-impact actions show exact targets and require fresh approval.
- Logs contain identifiers and outcomes without unnecessary sensitive content.
- Model, prompt, tool schema, dataset, and policy versions are recorded.
- Rollback and incident response do not depend on the agent.
8. Primary References
- OpenAI practical guide to building agents
- Anthropic: Building effective agents
- OWASP prompt injection guidance
9. Try the Interactive Studio
Ready to customize your agent architecture and generate TypeScript or Python boilerplate? Try our interactive AI Agent Builder Lab tool.
What a working prototype costs to run
Prototypes are usually built on a flagship model and never re-priced, which is how teams end up with a bill that surprises them at launch. Below: one agent run at 40K input, 20K cached, 5K output. Of the 40,000 input tokens, 20,000 are billed at the cache-read rate and 20,000 at full input rate.
| Model | Cost per run | Monthly at 2,000 runs |
|---|---|---|
| DeepSeek V4.1 Flash | $0.0061 | $12.12 |
| Claude Sonnet 5 | $0.094 | $188 |
| GPT-5.6 Terra | $0.104 | $208 |
| Gemini 3.1 Pro | $0.104 | $208 |
The gap between the default choice and the cheap end is roughly twentyfold at this shape. Most prototypes never need the flagship for the majority of their calls, which is exactly the case the multi-model routing guide makes.
Rates verified against provider documentation on September 18, 2026. Promotional rates expire, so re-check before budgeting: LLM API cost planning · September 2026 pricing update. Run your own numbers in the cost calculator.