Tool Calling vs. Code Execution for AI Agents: Choosing the Right Action Primitive - MachineLearningMastery.com Tool Calling vs. Code Execution for AI Agents: Choosing the Right Action Primitive - MachineLearningMastery.com In this article, you will learn what tool calling and code execution are as agent action primitives, how they differ mechanically, and when to choose one over the other. Topics we will cover include: How tool calling works under the hood, and why it remains the right choice for single, time-sensitive lookups. How code execution via Programmatic Tool Calling differs from standard tool calling, and what measurable benefits it offers for fan-out and aggregation tasks. A practical decision framework for choosing between the two primitives based on call count, data sensitivity, latency, infrastructure, and auditability needs. Picture an agent asking one simple-sounding question: which of twenty employees went over their Q3 travel budget. To answer it, the agent needs each person’s expense line items, every flight, hotel, and meal receipt, compared against a budget limit tied to their level. Built the obvious way, with the model calling a tool for each person’s expenses one at a time, that’s twenty separate tool calls, each returning fifty to a hundred line items, and every single one of those items has to pass through the model’s context just so it can be added up. That’s over 2,000 line items and more than 50KB of raw data the model never actually needed to read — it needed a sum. That’s the real cost hiding behind a design decision most agent tutorials skip past entirely: how does an agent actually take action in the world. There are two real answers, tool calling and code execution, and which one you reach for isn’t a style preference — it’s an architectural choice with measurable consequences for cost, latency, and accuracy. This article breaks down both action primitives for AI agents in detail, builds a real, runnable example of each using the same underlying tool, and closes with an honest, numbers-backed framework for choosing between them. If you haven’t built a basic tool-calling agent yet, check out this article, Easy Agentic Tool Calling with Gemma 4 — it is the natural place to start before this one. What Is an Action Primitive, and Why Does the Choice Matter? An action primitive is the fundamental mechanism by which a language model turns a decision into a real effect in the world — a database write, an API call, a file read. Every agent framework, whatever else it does, is built on top of one of these primitives at its core. Tool calling is the primitive most people learn first: the model produces one structured request at a time, a host application executes it, and the result comes back into the conversation before the model decides what to do next. Code execution is the newer alternative: instead of requesting one action and waiting, the model writes an actual program — in Python or TypeScript — that performs several actions in sequence or in parallel, and only the program’s final output returns to the model. Neither one is a wrapper around the other, and neither has quietly replaced the other. They’re genuinely different mechanisms with different failure modes, different infrastructure requirements, and different cost profiles, and the rest of this article is about understanding both well enough to pick correctly. Tool Calling It’s worth understanding what’s actually happening underneath a tool call, because the mechanics explain both its strengths and its real limitations. According to Cloudflare’s detailed breakdown of the process, a model generating a tool call doesn’t produce ordinary text. It’s been specifically trained to output a pair of special tokens — one signaling “the following is a tool call” and another marking its end — with a JSON payload describing the tool name and arguments sitting between them. The application running the model watches for those tokens, pauses generation the moment it sees the closing one, parses the JSON against a schema you defined, actually executes the call, and feeds the result back into the conversation as though it were the next thing the user said. That’s a clean, auditable, one-step-at-a-time loop, and it’s exactly why tool calling became the default. Every action is a discrete, loggable event. Every result is something the model directly sees and can reason about in natural language before deciding what happens next. Code Execution Code execution takes a different starting position entirely: instead of asking the model to describe an action in a constrained JSON format, you let it write actual code that performs the action, running in a sandboxed environment separate from the model itself. Anthropic’s original code-execution-with-MCP pattern frames this precisely as presenting your tools as a code API rather than a set of directly callable functions, so the model can write a script that imports exactly the tools it needs and calls them the way it would call any other function. The mechanism that makes this genuinely different — not just a relabeled tool call — is what Anthropic now calls Programmatic Tool Calling, released alongside two companion features in November 2025. Rather than each tool result flowing back through the model one at a time, you mark specific tools as callable from code by adding an allowed_callers field to their definition, and add a code_execution tool to the request. When the model wants to act, it writes a full script — loops, conditionals, error handling, and all — that calls those tools directly inside a sandboxed execution environment. Each individual tool call the script makes still executes exactly the way it would in ordinary tool calling; you still receive a request and return a result, but that result is intercepted and processed by the running script rather than being pushed into the model’s context. Only when the script finishes does its final output — and nothing else — return to the model. That’s the entire difference in one sentence: tool calling puts every intermediate result in front of the model; code execution lets the model decide, through the code it writes, exactly what makes it back. A side-by-side flow diagram of Tool Calling and Code Execution (click to enlarge) Tool Calling for a Single, Time-Sensitive Lookup Theory is easier to trust once it’s running against a real API, so both examples in this article use the same tool — a get_weather function backed by Open-Meteo, a free weather API that needs no API key at all, only an Anthropic API key to run the agent itself. Start with the case tool calling is obviously right for: a single question that needs one lookup and a natural-language answer — “what’s the weather like in London right now.” 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 import json import requests from anthropic import Anthropic client = Anthropic() # reads ANTHROPIC_API_KEY from the environment def get_weather(city: str) -> dict: """Look up a city's coordinates, then fetch its current temperature and this week's daily highs from Open-Meteo's free, keyless API.""" geo = requests.get( "https://geocoding-api.open-meteo.com/v1/search", params={"name": city, "count": 1}, ).json() if not geo.get("results"): return {"error": f"Could not find a location named '{city}'"} lat = geo["results"][0]["latitude"] lon = geo["results"][0]["longitude"] forecast = requests.get( "https://api.open-meteo.com/v1/forecast", params={ "latitude": lat, "longitude": lon, "current": "temperature_2m", "daily": "temperature_2m_max", "timezone": "auto", }, ).json() return { "city": city, "current_temp_c": forecast["current"]["temperature_2m"], "week_high_temps_c": forecast["daily"]["temperature_2m_max"], "unit": "celsius", } weather_tool = { "name": "get_weather", "description": ( "Get the current temperature and this week's daily high " "temperatures for a city. Returns JSON with city, " "current_temp_c, week_high_temps_c (7 daily highs), and unit." ), "input_schema": { "type": "object", "properties": { "city": {"type": "string", "description": "City name, e.g. 'Lagos'"} }, "required": ["city"], }, } messages = [{"role": "user", "content": "What's the weather like in London right now?"}] response = client.messages.create( model="claude-sonnet-5", max_tokens=1024, tools=[weather_tool], messages=messages, ) # Keep resolving tool calls until Claude produces a final text answer while response.stop_reason == "tool_use": messages.append({"role": "assistant", "content": response.content}) tool_results = [] for block in response.content: if block.type == "tool_use" and block.name == "get_weather": result = get_weather(**block.input) tool_results.append({ "type": "tool_result", "tool_use_id": block.id, "content": json.dumps(result), }) messages.append({"role": "user", "content": tool_results}) response = client.messages.create( model="claude-sonnet-5", max_tokens=1024, tools=[weather_tool], messages=messages, ) for block in response.content: if block.type == "text": print(block.text) Walking through what matters here: get_weather itself is ordinary Python — nothing agent-specific about it — it geocodes a city name and pulls both the current temperature and the week’s daily highs in one request. The weather_tool dictionary is the schema Claude actually sees, and the description matters more than it looks — a vague description is one of the most common causes of a model calling a tool with the wrong arguments. The while response.stop_reason == “tool_use” loop is the real mechanical heart of standard tool calling: every time Claude requests the tool, your code has to actually run it, wrap the result as a tool_result block, append it to the conversation, and call the API again — and this repeats for as many tool calls as the task needs. For a single lookup like this one, that’s one pass through the loop and done, which is exactly why tool calling fits this case well: one call, one result, and a result small and relevant enough that Claude genuinely benefits from seeing it directly before writing a natural-language answer. Code Execution for Fan-Out and Aggregation Now change the question, using the exact same get_weather function — completely unchanged: “given these fifteen cities, which one will have the coldest high temperature this week, and what’s the average weekly high across all of them?” Run that through the tool-calling loop above and you’d get fifteen separate tool calls, fifteen full JSON payloads of daily temperatures pushed into Claude’s context, and Claude would then have to manually compare and average them in natural language — slow, token-expensive, and exactly the kind of arithmetic a model is more error-prone at than a for-loop is. This is precisely the case Programmatic Tool Calling was built for. 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 import json from anthropic import Anthropic client = Anthropic() # Same get_weather function from the previous example, unchanged weather_tool = { "name": "get_weather", "description": ( "Get the current temperature and this week's daily high " "temperatures for a city. Returns JSON with city, " "current_temp_c, week_high_temps_c (7 daily highs), and unit." ), "input_schema": { "type": "obj [truncated for AI cost control]
Tool Calling vs. Code Execution for AI Agents: Choosing the Right Action Primitive
Summary
Theory is easier to trust once it's running against a real API, so both examples in this article use the same tool — a get_weather function backed by <a href="https://open-meteo.
Tool Calling vs. Code Execution for AI Agents: Choosing the Right Action Primitive
Report an error
The correction channel is not available yet. You can copy the article reference below for later.
Correction instructions