Function Calling
Chapter of Language Models in Practice, part of AI & LLM Foundations.
What you will understand at the end
- The request/response contract that turns “the model wants to call a function” into code your application actually executes — and why the model never runs anything itself
- How a model selects which function to call and populates its arguments from natural language, and where that selection and population process breaks
- The specific defenses against parameter hallucination and ambiguous tool selection, before this Part extends the pattern to multi-tool agent design
The model proposes; your code disposes
The single most important mental model for function calling: the LLM never executes anything. It
emits a structured description of a function call it would like made — a name and a set of arguments
— and your application code is entirely responsible for deciding whether to run it, running it, and
feeding the result back. This is true regardless of provider or SDK dressing; “function calling” and
“tool use” name the same underlying mechanism (Anthropic’s Messages API calls it tool_use;
OpenAI’s Chat Completions API calls it function_call/tool_calls), and every framework built on
top of either is implementing the same four-step loop:
sequenceDiagram
participant App as Your application
participant Model as LLM
App->>Model: Request + function/tool definitions
Model->>App: "I want to call get_weather(city='Paris')"
Note over App: Model has NOT executed anything yet
App->>App: Validate, then execute get_weather("Paris")
App->>Model: Function result: "18°C, cloudy"
Model->>App: Final natural-language answer
The gap between “the model requested a call” and “the call happened” is exactly where the
application-level responsibility sits: validating arguments, enforcing permissions, handling
execution errors, and deciding what a failure should look like from the model’s point of view. None
of that is the model’s job, and code that treats a tool_use block as already-executed is the most
common source of confused, hard-to-debug agent behavior.
The function-schema contract
A function definition given to the model is a JSON Schema describing the function’s name, purpose, and parameters — the same schema mechanism covered from the extraction side in Structured Outputs, applied here to describing an action instead of a data shape:
tools = [{
"name": "get_weather",
"description": "Get the current weather for a given city. Use this whenever "
"the user asks about current conditions, not historical or forecast data.",
"input_schema": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "City name, e.g. 'Paris'"},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"], "description": "Temperature unit"},
},
"required": ["city"],
},
}]
response = client.messages.create(
model="claude-opus-4-8", max_tokens=1024,
tools=tools,
messages=[{"role": "user", "content": "What's it like in Paris right now?"}],
)
for block in response.content:
if block.type == "tool_use":
print(block.name, block.input) # "get_weather" {"city": "Paris"}
Two fields do almost all the work of correct selection and population, and both are frequently under-invested in relative to their leverage:
descriptionis the primary selection signal. The model decides whether to call a function, and which of several candidates, almost entirely from the description text — not the function name. A description that states only what the function does (“Gets weather data”) gives weaker selection signal than one that also states when to call it (“Use this whenever the user asks about current conditions, not historical or forecast data”) — being prescriptive about the trigger condition, not just the capability, is the single highest-leverage edit to a tool description.- Per-parameter
descriptionandenumdo the population work. Acityparameter with no description leaves the model guessing at format (full name? airport code? “Paris, France” vs. “Paris”?); anenum-constrainedunitparameter cannot drift to a value your code doesn’t handle.
Parameter hallucination
The failure: the model calls a real function with a plausible-looking argument that has no basis
in anything the user said or any data the model was given — inventing a city value, fabricating an
order_id, or supplying a default that happens to look reasonable rather than surfacing that the
information is missing.
This is a special case of hallucination (treated generally in Hallucination Management), and it’s particularly dangerous in function calling specifically because a hallucinated argument doesn’t look like an error — it’s a syntactically perfect call to a real function with a wrong value, which will execute successfully and produce a wrong-but-plausible result rather than failing loudly.
Concrete defenses:
- Make genuinely required parameters
requiredin the schema, and give the model permission to ask instead of guess. A system prompt line like “If a required parameter isn’t clear from the conversation, ask the user rather than guessing a value” measurably reduces confident invention — the model defaults to guessing when guessing is implicitly the only offered path forward. - Constrain everything constrainable. Free-text fields invite invention;
enum, date-format patterns, and numeric bounds close off entire classes of plausible-but-wrong values. - Validate before executing, and reject with a specific reason. If
order_iddoesn’t match your system’s ID format, don’t execute the call and hope — return atool_resultwithis_error: trueand a message the model can act on (“order_id must match format ORD-XXXXX”), which lets the model self-correct on the next turn instead of silently succeeding against garbage. - For destructive or high-stakes calls, require the model to restate what it’s about to do before you execute it, and gate execution behind that restatement matching the actual arguments — this catches cases where the model’s stated intent and its populated arguments have quietly diverged.
Ambiguous selection
The failure: with several tools available, the model picks the wrong one — not because it
hallucinated, but because two tool descriptions overlap in a way that makes selection genuinely
underdetermined from the model’s point of view. A search_orders and get_order_by_id pair with
similar descriptions will get confused on requests that could plausibly go to either.
Concrete defenses:
- Make descriptions mutually exclusive, not just individually accurate. “Search orders by customer name or date range” and “Retrieve a single order by its exact order ID” don’t overlap; “Look up order information” and “Get order details” do.
- Reduce the tool count actually in context for a given turn. Every function definition consumes context and, more importantly, competes for selection attention — a system that always exposes all 40 tools to every request will make more selection errors than one that narrows to the 5 relevant tools per request. This scales into the tool-discovery and tool-search patterns covered in Part 04 of Agentic AI Engineering — Tool Discovery and Tool Selection Strategies.
- Force the tool when the request genuinely has only one right answer.
tool_choice: {"type": "tool", "name": "..."}removes selection ambiguity entirely for requests you can classify upstream — don’t leave the model to choose between two tools when your application logic already knows which one applies.
From one function to many
This chapter deliberately stayed in single-function territory: one function offered, one call made, one result returned. Real agents almost always offer several tools, choose among them per-turn, and sometimes call more than one in the same turn — that’s a distinct set of design problems (tool registries, parallel vs. sequential invocation, tool-choice strategy) covered next in Tool Calling.
Metadata
| Author | Amit Singh |
| Scope | ai-foundations |
Local graph
Linked from 8 notes
1. Tool Calling Architecture
Covers the mechanics of function/tool calling in modern LLM APIs -- schema definition, the model's structured-call output, execution, and result injection back into the conversation -- as the foundational primitive every agent framework builds on.
11. Probability, Sampling & Decoding
The math intuition underneath every model call — how a raw logit vector becomes a probability distribution, why temperature and top-p reshape that distribution differently, why beam search lost to sampling for chat and agent models, and how entropy and KL divergence turn 'the model is uncertain' and 'alignment training' into something you can actually reason about.
2. Prompt Design Patterns
Catalogs reusable prompt patterns — chain-of-thought, ReAct, self-consistency, and role/persona framing — with guidance on when each pattern earns its added token cost over a plain instruction.
3. Structured Outputs
Covers forcing an LLM into a validated schema — JSON mode, function-calling-style schemas, and grammar-constrained decoding — and the failure modes, like schema drift and hallucinated fields, that break naive implementations.
5. Tool Calling
Extends function calling into multi-tool agent design — tool registries, tool-choice strategies, parallel vs. sequential invocation — and how tool descriptions themselves become part of the prompt-engineering surface.
9. AI Failure Modes
Surveys production failure modes beyond hallucination — prompt injection, context poisoning, tool-call loops, silent schema violations, and cascading errors in multi-agent chains — as the taxonomy a staff engineer defends against.
10. Building Reliable LLM Applications
Covers the engineering practices that turn a probabilistic model call into a reliable system component — retries with validation, evals as CI gates, circuit breakers, and observability for non-deterministic outputs.
AI & LLM Foundations
A book-shaped table of contents for AI & LLM Foundations: the pre-agentic substrate — symbolic AI through transformers, tokens, embeddings, attention, foundation models, and turning a raw LLM API into a dependable application component. Book 1 of the AI Systems Engineering series.
Related notes
9. AI Failure Modes
Surveys production failure modes beyond hallucination — prompt injection, context poisoning, tool-call loops, silent schema violations, and cascading errors in multi-agent chains — as the taxonomy a staff engineer defends against.
10. Building Reliable LLM Applications
Covers the engineering practices that turn a probabilistic model call into a reliable system component — retries with validation, evals as CI gates, circuit breakers, and observability for non-deterministic outputs.
8. Hallucination Management
Covers the mechanisms behind LLM hallucination — parametric knowledge gaps, exposure bias, overconfident sampling — and mitigation strategies like grounding via RAG, citation requirements, and confidence-calibrated refusal.
7. Model Selection & Routing
Covers building a model router that picks among providers and tiers by task complexity, latency SLA, and cost — the pattern that replaces always calling the biggest model once traffic reaches production scale.