How to Build an AI Agent from Scratch: Complete Developer Guide

By AI Agent Hub Editorial Desk · Review method · Corrections

Tutorial · 6 min read · Reviewed September 18, 2026

Building an AI agent goes far beyond sending a prompt to a language model. An autonomous agent combines a reasoning LLM brain, memory systems, tool execution routines, and a self-correcting loop. This tutorial walks you through building a production-ready AI agent from scratch.

1. What Distinguishes an AI Agent from a Standard LLM Call?

A standard LLM call operates as a single-turn function: Input Prompt → Model → Text Output. In contrast, an AI Agent operates autonomously inside an execution loop:

Perception (User Input) → Reasoning (Thought) → Tool Selection (Action) → Sandbox Execution (Observation) → Evaluation → Self-Correction

2. The 4 Pillars of Agent Architecture

3. Complete Step-by-Step Python Implementation

Below is a clean, dependency-minimal Python script demonstrating a working ReAct agent loop:

import json
from openai import OpenAI

# Initialize client (works with DeepSeek, OpenAI, or local vLLM)
client = OpenAI(api_key="your-api-key")

# 1. Define Tool Functions
def calculate_budget(daily_cost: float, days: int) -> str:
    return json.dumps({"monthly_total": daily_cost * days})

tools_schema = [
    {
        "type": "function",
        "function": {
            "name": "calculate_budget",
            "description": "Calculate total monthly budget from daily spend",
            "parameters": {
                "type": "object",
                "properties": {
                    "daily_cost": {"type": "number"},
                    "days": {"type": "integer"}
                },
                "required": ["daily_cost", "days"]
            }
        }
    }
]

# 2. Agent Execution Loop
def run_agent_loop(prompt: str):
    messages = [{"role": "user", "content": prompt}]
    
    while True:
        response = client.chat.completions.create(
            model="deepseek-v3",
            messages=messages,
            tools=tools_schema
        )
        msg = response.choices[0].message
        
        # Check if model wants to invoke a tool
        if msg.tool_calls:
            tool_call = msg.tool_calls[0]
            args = json.loads(tool_call.function.arguments)
            print(f"Agent Action: Calling {tool_call.function.name} with {args}")
            
            # Execute tool locally
            result = calculate_budget(**args)
            
            # Append observation to history
            messages.append(msg)
            messages.append({"role": "tool", "tool_call_id": tool_call.id, "content": result})
        else:
            # Model finished reasoning
            print(f"Final Agent Answer: {msg.content}")
            break

run_agent_loop("If our API spend is $5.5 per day, what is our 30-day budget?")

4. Define the Task Contract Before Adding Autonomy

Write down the trigger, allowed inputs, evidence sources, tools, side effects, stop conditions, and human approval points. A useful first agent has one bounded job. For example: read one bug report, search an approved repository, propose a patch in an isolated branch, run tests, and stop for review. It may not merge, deploy, message a customer, or access another repository.

Contract itemExampleEnforced by
IdentitySigned-in repository contributorApplication authentication
Read scopeOne repository at a pinned revisionWorkspace and API allowlist
Write scopeEphemeral branch onlyScoped credential and branch policy
BudgetEight turns, twelve tool calls, ninety secondsOrchestrator counters
SuccessPatch produced and specified tests passDeterministic verifier
ApprovalHuman reviews diff before publicationSeparate privileged action

5. Build the Smallest Observable Loop

  1. Construct context: include durable policy, the current task, and only the evidence needed for the next decision.
  2. Ask for a typed proposal: answer, stop, request input, or call one of a narrow set of tools.
  3. Validate: check schema, authorization, budgets, paths, URLs, and action class outside the model.
  4. Execute: run the tool in an isolated environment and return a bounded structured result.
  5. Record: persist the state transition, evidence IDs, decision, result, cost, and stop reason.
  6. Verify: use tests and business rules; require approval before consequential side effects.

Do not add multiple agents merely to imitate an organization chart. Add a second role only when it provides a measurable boundary, such as an independent verifier with no write access or a specialist that uses a distinct evidence source. A deterministic function is often a better router than another model call.

6. Evaluation and Failure Handling

Create fixtures before launch: ordinary success, incomplete input, tool timeout, stale state, conflicting evidence, permission denial, prompt injection inside a file, and budget exhaustion. Measure accepted-task success, unsafe-action rate, unnecessary tool calls, human override rate, latency, and cost. Preserve the complete trajectory so a failure can be attributed to context, model decision, policy, tool, or verifier.

Every run needs a visible terminal state. Retry only typed transient failures and use idempotency keys for operations that may already have succeeded. After repeated equivalent failures, stop and ask for help instead of rewriting the prompt indefinitely.

7. Production Readiness Checklist

8. Primary References

9. Try the Interactive Studio

Ready to customize your agent architecture and generate TypeScript or Python boilerplate? Try our interactive AI Agent Builder Lab tool.

What a working prototype costs to run

Prototypes are usually built on a flagship model and never re-priced, which is how teams end up with a bill that surprises them at launch. Below: one agent run at 40K input, 20K cached, 5K output. Of the 40,000 input tokens, 20,000 are billed at the cache-read rate and 20,000 at full input rate.

Model Cost per run Monthly at 2,000 runs
DeepSeek V4.1 Flash$0.0061$12.12
Claude Sonnet 5$0.094$188
GPT-5.6 Terra$0.104$208
Gemini 3.1 Pro$0.104$208

The gap between the default choice and the cheap end is roughly twentyfold at this shape. Most prototypes never need the flagship for the majority of their calls, which is exactly the case the multi-model routing guide makes.

Rates verified against provider documentation on September 18, 2026. Promotional rates expire, so re-check before budgeting: LLM API cost planning · September 2026 pricing update. Run your own numbers in the cost calculator.