If you've searched for how to build AI agents tutorial, you've probably hit the same wall I did — most guides are either too abstract or show you toy examples that fall apart the moment you try something real. This guide fixes that. You'll build a working autonomous agent using Python and the Anthropic SDK that can reason through problems, call tools, and loop until it finishes a task — not just generate a single response.
What You'll Build
By the end of this tutorial, you'll have a production-ready AI agent that uses Claude Sonnet 4.6 to autonomously answer questions by calling real tools — a web search simulator and a calculator. The agent runs a loop, decides which tools to use, executes them, and keeps going until it has a final answer.
This is the same architecture we use at Naples AI when building custom agent systems for local businesses. It's minimal, but it's real — and you can extend it to production in an afternoon.
The complete working code is built step-by-step in the sections below. Each step adds one piece of the agent — by Step 4 you'll have the full system running. No placeholder logic, no pseudocode. Copy each block in order and you're done.
Prerequisites
- Python 3.10 or higher installed
- An Anthropic API key (get one at console.anthropic.com)
- Basic familiarity with Python classes and functions
- The
anthropicPython SDK installed (pip install anthropic) - A terminal and a code editor — nothing else needed
Step 1: Set Up Claude API and Anthropic SDK
First, install the Anthropic SDK and confirm your API key is working. I always test the connection with a bare-minimum call before building anything on top of it — saves a lot of debugging later.
Set your API key as an environment variable so it never touches your source code. On Mac/Linux run export ANTHROPIC_API_KEY=your_key_here. On Windows use set ANTHROPIC_API_KEY=your_key_here.
import anthropic
import os
# Confirm the SDK can reach the API before building anything on top of it
client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY"))
response = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=64,
messages=[{"role": "user", "content": "Say hello in one sentence."}]
)
print(response.content[0].text)
If your key is set correctly you'll see something like: "Hello! It's great to meet you." If you get an AuthenticationError, your key isn't loaded — double-check the environment variable name and restart your terminal.
Never hardcode your API key directly in source files. If you push to GitHub with a key in the code, Anthropic will auto-revoke it and you'll need a new one.
Step 2: Define Your Agent's Tools and Capabilities
Tools are how Claude takes action in the world. You define them as a JSON schema — Claude reads that schema and decides when and how to call each tool during a conversation. Think of it like a function signature with a description Claude can actually understand.
For this tutorial, I'm defining two tools: a search tool and a calculator. In a real production system these would hit actual APIs, but the agent logic is identical regardless of what the tool does under the hood.
tools.py
import anthropic
import json
import math
# Tool definitions tell Claude what capabilities it has and how to invoke them
TOOLS = [
{
"name": "web_search",
"description": (
"Search the web for current information about a topic. "
"Use this when you need facts, news, or data you don't already know."
),
"input_schema": {
"type": "object",
"properties": {
"query": {
"type": "string",
"description": "The search query string"
}
},
"required": ["query"]
}
},
{
"name": "calculator",
"description": (
"Evaluate a mathematical expression and return the result. "
"Use this for any arithmetic, algebra, or numeric computation."
),
"input_schema": {
"type": "object",
"properties": {
"expression": {
"type": "string",
"description": "A valid Python math expression, e.g. '(15 * 4) / 2 + math.sqrt(9)'"
}
},
"required": ["expression"]
}
}
]
def execute_tool(tool_name: str, tool_input: dict) -> str:
"""Run the requested tool and return the result as a string."""
if tool_name == "web_search":
query = tool_input["query"]
# Simulated search results — swap this for a real search API in production
simulated_results = {
"naples florida population": "Naples, Florida has a population of approximately 22,000 in the city proper, with the greater Naples metro area exceeding 390,000 residents as of 2025.",
"anthropic claude models": "Anthropic's current model lineup includes Claude Sonnet 4-6, Claude Opus 4, and Claude Haiku 3-5, released throughout 2025 and 2026.",
"average restaurant profit margin": "The average restaurant profit margin is 3–9% for full-service restaurants and 6–9% for fast-casual concepts.",
}
# Return closest match or a default message
for key, value in simulated_results.items():
if any(word in query.lower() for word in key.split()):
return value
return f"Search results for '{query}': No specific data found. Please refine your query."
elif tool_name == "calculator":
expression = tool_input["expression"]
try:
# Allow math module functions inside expressions
result = eval(expression, {"__builtins__": {}}, {"math": math})
return str(result)
except Exception as e:
return f"Calculation error: {e}"
return f"Unknown tool: {tool_name}"
The input_schema field is what Claude actually reads to understand what arguments to pass. If your schema is vague, Claude will guess — and it'll sometimes guess wrong. Be specific in your descriptions.
Step 3: Implement the Main Agent Class with Tool Use
Now we build the agent class itself. This is the core of the whole system — it holds the conversation history, sends messages to Claude, and handles the tool call responses that come back. The class design keeps everything in one place so the orchestration logic in Step 4 stays clean.
agent.py
import anthropic
import json
import os
from tools import TOOLS, execute_tool
class ClaudeAgent:
"""
A single-agent system that uses Claude Sonnet 4-6 with tool use.
Maintains conversation history across turns in the run loop.
"""
def __init__(self, system_prompt: str = None):
self.client = anthropic.Anthropic(
api_key=os.environ.get("ANTHROPIC_API_KEY")
)
self.model = "claude-sonnet-4-6"
self.tools = TOOLS
self.conversation_history = []
# Default system prompt instructs the agent to use tools proactively
self.system_prompt = system_prompt or (
"You are a helpful research assistant with access to web search and a calculator. "
"When answering questions, use your tools to get accurate, up-to-date information. "
"Think step-by-step and always verify numbers with the calculator tool. "
"Be concise in your final answer."
)
def _send_message(self, user_message: str) -> anthropic.types.Message:
"""Append the user message to history and call the Claude API."""
self.conversation_history.append({
"role": "user",
"content": user_message
})
response = self.client.messages.create(
model=self.model,
max_tokens=4096,
system=self.system_prompt,
tools=self.tools,
messages=self.conversation_history
)
return response
def _handle_tool_calls(self, response: anthropic.types.Message) -> None:
"""
Process all tool_use blocks in the response.
Appends the assistant's tool calls and our tool results back into history
so Claude sees the full context on the next turn.
"""
# Record what the assistant decided to do
self.conversation_history.append({
"role": "assistant",
"content": response.content
})
# Build a tool_result block for every tool_use block in the response
tool_results = []
for block in response.content:
if block.type == "tool_use":
print(f" → Calling tool: {block.name}({json.dumps(block.input)})")
result = execute_tool(block.name, block.input)
print(f" ← Result: {result}\n")
tool_results.append({
"type": "tool_result",
"tool_use_id": block.id,
"content": result
})
# Feed all tool results back as a single user turn
self.conversation_history.append({
"role": "user",
"content": tool_results
})
Notice that _handle_tool_calls appends both the assistant's tool-use blocks and the tool results back into conversation_history. This is the part most beginners miss. Claude needs to see its own tool calls and the results in the history to continue reasoning correctly.
If you forget to append the assistant's
tool_use content blocks to history before sending tool results, Claude throws a 400 Bad Request error. The history must be: user → assistant (with tool_use) → user (with tool_result). That exact order matters.
Step 4: Create the Run Loop and Orchestration Logic
This is where the agent becomes autonomous. The run loop keeps calling Claude, handling tool calls, and feeding results back in — until Claude returns a final text response instead of another tool call. That's how you know it's done.
agent.py (continued — add this method to ClaudeAgent)
def run(self, user_message: str, max_iterations: int = 10) -> str:
"""
Main agent loop. Runs until Claude produces a final text answer
or until max_iterations is reached to prevent infinite loops.
"""
print(f"\n{'='*60}")
print(f"USER: {user_message}")
print(f"{'='*60}\n")
response = self._send_message(user_message)
iterations = 0
# Keep looping as long as Claude wants to use tools
while response.stop_reason == "tool_use" and iterations < max_iterations:
iterations += 1
print(f"[Iteration {iterations}] Claude is using tools...\n")
self._handle_tool_calls(response)
# Send an empty continuation — Claude reads the tool results from history
response = self.client.messages.create(
model=self.model,
max_tokens=4096,
system=self.system_prompt,
tools=self.tools,
messages=self.conversation_history
)
# Extract the final plain-text answer from the response
final_answer = ""
for block in response.content:
if hasattr(block, "text"):
final_answer += block.text
# Add the final assistant response to history for multi-turn conversations
self.conversation_history.append({
"role": "assistant",
"content": response.content
})
print(f"\nAGENT: {final_answer}")
return final_answer
The max_iterations guard is not optional — it's essential. Without it, a misconfigured tool can cause the agent to loop forever and burn through your API credits. Ten iterations is generous for most tasks; you can lower it for simpler use cases.
Step 5: Test with Real-World Scenarios
Now let's wire everything together and run the agent against a few real questions. This file is your entry point — run it and you'll see the full reasoning loop in your terminal.
main.py
import anthropic
import json
import os
import math
from tools import TOOLS, execute_tool
from agent import ClaudeAgent
def main():
agent = ClaudeAgent()
# Test 1: Research question that needs web search
agent.run(
"What is the population of Naples, Florida, "
"and how many restaurants would serve 10% of that population daily "
"if each restaurant seats 80 people and turns tables 3 times a day?"
)
# Test 2: Pure calculation to verify the calculator tool works
agent.run("What is the square root of 1764 multiplied by 15?")
# Test 3: Multi-tool question requiring both search and math
agent.run(
"What is the average restaurant profit margin, "
"and if a Naples restaurant does $1.2 million in annual revenue, "
"what is the expected annual profit at the midpoint margin?"
)
if __name__ == "__main__":
main()
When you run python main.py, you'll see output that looks like this:
============================================================
USER: What is the population of Naples, Florida, and how many restaurants would
serve 10% of that population daily if each restaurant seats 80 people and turns
tables 3 times a day?
============================================================
[Iteration 1] Claude is using tools...
→ Calling tool: web_search({"query": "naples florida population"})
← Result: Naples, Florida has a population of approximately 22,000 in the city
proper, with the greater Naples metro area exceeding 390,000 residents as of 2025.
[Iteration 2] Claude is using tools...
→ Calling tool: calculator({"expression": "390000 * 0.10 / (80 * 3)"})
← Result: 162.5
AGENT: Naples, Florida has a metro population of about 390,000 people.
To serve 10% of that population (39,000 people) daily, where each restaurant
seats 80 people and turns tables 3 times a day (240 covers per day),
you would need approximately 163 restaurants.
============================================================
USER: What is the square root of 1764 multiplied by 15?
============================================================
[Iteration 1] Claude is using tools...
→ Calling tool: calculator({"expression": "math.sqrt(1764) * 15"})
← Result: 630.0
AGENT: The square root of 1764 is 42, and 42 multiplied by 15 equals 630.
How It Works: Agent Decision Flow Explained
Here's what's actually happening under the hood on every loop iteration. Claude doesn't "run" your tools — it just tells you which tool to run and with what arguments. Your Python code runs the tool, hands the result back, and Claude decides what to do next.
The decision flow looks like this:
- User sends a message → appended to
conversation_history - Claude responds → either with a
tool_useblock (stop_reason: "tool_use") or a final text answer (stop_reason: "end_turn") - If tool_use: your code runs the tool, appends both the assistant's tool call and your tool result to history, then calls Claude again
- If end_turn: extract the text, print it, done
The conversation history is the agent's working memory. Every tool call, every result, every response — it's all in there. Claude reads the full history on every API call, which is how it maintains context across multiple tool uses in a single task.
Claude doesn't have persistent memory between separate
ClaudeAgent instances. If you want the agent to remember things across sessions, you need to save and reload self.conversation_history to a database or file. That's a common upgrade we add for production systems.
Common Errors and Fixes
Error 1: anthropic.BadRequestError — messages must alternate between user and assistant roles
This is the most common mistake when building the tool loop. It happens when you send tool results without first recording the assistant's tool_use response in history.
# WRONG — sending tool results without the assistant's prior response in history
conversation_history.append({"role": "user", "content": tool_results})
# RIGHT — always append the assistant response first, THEN the tool results
conversation_history.append({"role": "assistant", "content": response.content})
conversation_history.append({"role": "user", "content": tool_results})
Error 2: anthropic.AuthenticationError — invalid x-api-key
Your API key isn't loading. The environment variable name is case-sensitive and the terminal session matters — setting it in one terminal tab doesn't carry to another.
# Check if the key is actually available before running the agent
import os
api_key = os.environ.get("ANTHROPIC_API_KEY")
if not api_key:
raise ValueError(
"ANTHROPIC_API_KEY environment variable not set. "
"Run: export ANTHROPIC_API_KEY=your_key_here"
)
# Then pass it explicitly to the client
client = anthropic.Anthropic(api_key=api_key)
Error 3: Agent loops forever / hits max_iterations and stops mid-task
This usually means your tool is returning an error string instead of a useful result — so Claude keeps trying. Log every tool call result during development so you can see what's going wrong.
# Add this inside execute_tool() to debug tool failures during development
def execute_tool(tool_name: str, tool_input: dict) -> str:
print(f"[DEBUG] Tool: {tool_name} | Input: {tool_input}") # log every call
result = ""
if tool_name == "calculator":
try:
result = str(eval(tool_input["expression"], {"__builtins__": {}}, {"math": math}))
except Exception as e:
result = f"Error: {e}" # return the error as a string, not an exception
print(f"[DEBUG] Result: {result}") # log every result
return result
Next Steps: Scaling to Multi-Agent Systems
A single agent with a handful of tools is a solid starting point. Here's where to take it next once you have the basics working:
1. Add Real Tool Integrations
Swap the simulated search for a real API — Brave Search, Serper, or Tavily all have Python SDKs and free tiers. Add a tool that reads from a Google Sheet or queries a database. That's when the agent becomes genuinely useful for business workflows.
2. Build a Multi-Agent System with Orchestration
Anthropic's multi-agent architecture lets one "orchestrator" agent spin up specialized "subagent" instances — a research agent, a writing agent, a data agent. Each handles its own domain, and the orchestrator coordinates the output. We use this pattern for complex automation workflows at Naples AI.
3. Add Persistent Memory
Store conversation_history in a database like PostgreSQL or Redis between sessions. This lets your agent remember past interactions, user preferences, and task context — which is the difference between a demo and a production product.
4. Wrap It in an API with FastAPI
Put your ClaudeAgent class behind a FastAPI endpoint so any frontend or other service can talk to it over HTTP. Add a task queue like Celery if you need the agent to run long jobs in the background without blocking the request. That's the standard production setup.