Tool Use and Function Calling

Topics Covered

Function Calling Internals

What Function Calling Really Is

How Models Learn to Call Functions

The Function Schema

Inside the Model's Decision Process

Parsing Tool Calls

Constrained Decoding

The Problem Constrained Decoding Solves

How Constrained Decoding Works

Grammar Representations

The Mechanics in Detail

Trade-offs

When to Use Constrained Decoding

Tool Selection Strategies

Why Scale Matters

Strategy 1: Flat Tool Lists (Small Scale)

Strategy 2: Category-Based Grouping

Strategy 3: Retrieval-Based Tool Selection

Strategy 4: Hierarchical Planning

Strategy 5: Routing to Specialized Models

Combining Strategies

Error Recovery

Categories of Failures

Strategy 1: Retry with Clarification

Strategy 2: Error Propagation to the LLM

Strategy 3: Structured Error Responses

Strategy 4: Tool Alternative Suggestions

Strategy 5: Graceful Degradation

Strategy 6: Hallucination Guards

Error Budgets and Circuit Breakers

Putting It Together

Paper Study: Toolformer

Key Takeaways for Practitioners

Coding Practice

Function calling is how LLMs interact with the outside world. A pure language model can only produce text. It cannot query a database, call an API, run code, or read a file. Function calling bridges this gap by teaching the model to produce structured output that a runtime interprets as a request to invoke external tools. This lesson covers how function calling actually works under the hood, how to make it reliable, and how to recover when things go wrong.

What Function Calling Really Is

The name "function calling" is slightly misleading. The model does not actually call functions. It produces text that looks like a function call, and your application code parses that text and invokes the real function. The model is essentially trained (or prompted) to follow a specific output format when it wants to use a tool.

A function call interaction looks like this:

  1. Your application sends a prompt plus a list of available tools with their schemas.
  2. The model generates either a text response (if it can answer directly) or a structured output indicating which tool to call with which arguments.
  3. Your application parses the structured output, calls the actual function, gets a result.
  4. The application sends the result back to the model in the conversation history.
  5. The model generates a natural-language response incorporating the tool result.

This loop, prompt, tool call, result, final response, is the foundation of every tool-using agent. Understanding each step is essential to debugging when it fails.

The five steps of a tool call laid across the model and application boundary, with every step on the model side shown as plain text generation.
Key Insight

The model never actually calls functions. It produces text that looks like a function call, and your application code does the real invocation. This distinction matters because it means the model can hallucinate function calls that do not exist, produce malformed arguments, or decide not to call a function at all.

How Models Learn to Call Functions

Modern function-calling models like GPT-4, Claude, and Gemini are explicitly trained to emit structured output when tools are available. During fine-tuning, they see thousands of examples of (conversation with tool list) → (structured tool call). The model learns to pattern-match: when it sees a tool list in the prompt and the user request matches a tool's purpose, emit a tool call in the expected format.

The structured output format varies by model. OpenAI uses a specific JSON schema in its API. Anthropic uses XML-tagged tool_use blocks. Open-source models might use special tokens or a JSON format defined by their template. The API abstracts these differences, but underneath, it is all just text generation conditioned on a format the model was trained to produce.

The Function Schema

When you register a function with an LLM, you provide a schema that tells the model:

  1. The function name
  2. A description of what the function does and when to use it
  3. The parameters it accepts, with types and descriptions
  4. Which parameters are required

Here is an example schema in OpenAI format:

json
1{
2  "name": "get_weather",
3  "description": "Get current weather for a city. Use when the user asks about weather conditions.",
4  "parameters": {
5    "type": "object",
6    "properties": {
7      "city": {
8        "type": "string",
9        "description": "The city name, for example 'San Francisco'"
10      },
11      "units": {
12        "type": "string",
13        "enum": ["celsius", "fahrenheit"],
14        "description": "Temperature units to return"
15      }
16    },
17    "required": ["city"]
18  }
19}

The description fields matter enormously. The model uses them to decide whether to call the function and what to pass as arguments. Vague descriptions lead to inappropriate calls and wrong arguments. A good description explains what the function does, when it should be called, and gives examples if the behavior is non-obvious.

Inside the Model's Decision Process

When the model receives a prompt with tool schemas, it makes two decisions:

  1. Should I use a tool at all? If the user asks "what is 2+2", the model should answer directly rather than calling a calculator. If the user asks "what is the weather in Paris right now", the model should call the weather tool because it has no way to know current weather from its weights.
  2. Which tool and with what arguments? If multiple tools could apply, the model picks based on the descriptions and matches arguments from the user query to the schema fields.

These decisions happen during generation. The model is predicting tokens one at a time, and the early tokens commit it to a path. Once the model starts emitting the JSON structure for a function call, it is committed to making a call. Once it commits to a specific function name, it cannot easily switch. This is why prompt engineering matters: the system prompt and tool descriptions shape these early-token decisions.

Whether to call a tool and which tool to call, both decided from the registered description strings, with a vague description absorbing the wrong query.
Common Pitfall

Vague tool descriptions are the number one cause of wrong tool selection. A description like 'gets data' forces the model to guess. A description like 'gets the current stock price for a single ticker symbol, use when the user asks about current prices, not historical trends' gives the model clear criteria for selection.

Parsing Tool Calls

Parsing the model's tool call output is where production systems often fail. Even with structured outputs, the model can produce malformed JSON, wrong types, missing fields, or hallucinated field names. Your parser must handle:

  1. Malformed JSON: Missing commas, unbalanced brackets, extra text around the JSON. Use a forgiving JSON parser or extract the JSON substring before parsing.
  2. Type coercion: The schema says "integer" but the model produced "42" (a string). Decide whether to auto-coerce or reject.
  3. Missing required fields: The model omitted a field marked required. Either prompt again asking for it, or use a default if the field has one.
  4. Extra fields: The model added fields not in the schema. Ignore them or reject the call.
  5. Hallucinated functions: The model called a function name that does not exist. This usually means the tool descriptions were confusing or the user query did not match any tool well.

Robust parsing is a core part of a reliable function-calling system. Think of the LLM's output as untrusted input that needs validation, not as a reliable API response.