Beyond Text In, Text Out

The biggest upgrade in LLM capabilities in the past year isn't better reasoning — it's function calling. Instead of generating text about what it would do, the model generates structured JSON that calls your functions. The LLM becomes a reasoning layer that decides what to do, while your code handles how to do it.

This is the foundation of every useful AI agent. Without function calling, an LLM is just a chatbot. With it, the LLM can search databases, call APIs, send emails, create records, and orchestrate multi-step workflows.

How Function Calling Works

  1. You define tools with JSON schemas describing what each function does and what parameters it takes
  2. The model receives a user message plus the available tools
  3. If a tool is relevant, the model returns a tool call (function name + arguments) instead of a text response
  4. Your code executes the function and returns the result
  5. The model incorporates the result and either calls another tool or generates a final text response

Designing Good Tool Schemas

The quality of your tool definitions directly determines how well the model uses them. This is where most implementations fail.

Be Specific in Descriptions

// Bad
{"name": "search", "description": "Search for things"}

// Good
{"name": "search_products",
 "description": "Search the product catalog by name, category, or price range. Returns up to 10 matching products with name, price, and availability.",
 "parameters": {
   "query": {"type": "string", "description": "Search query - product name or category"},
   "max_price": {"type": "number", "description": "Maximum price in USD. Optional."},
   "in_stock_only": {"type": "boolean", "description": "If true, only return available products"}
 }}

Use Enums for Constrained Choices

Don't let the model freeform where you need specific values. Use enums:

"status": {"type": "string", "enum": ["open", "in_progress", "resolved", "closed"]}

Required vs Optional Parameters

Mark parameters as required only when the function literally can't work without them. Optional parameters with good defaults give the model flexibility to call tools with minimal information.

Multi-Step Tool Chains

Real-world tasks require multiple tool calls in sequence. A customer service agent might:

  1. Call lookup_customer(email) to find the account
  2. Call get_recent_orders(customer_id) to see their history
  3. Call check_refund_eligibility(order_id) to see if a refund applies
  4. Call process_refund(order_id, amount) to execute the refund
  5. Generate a confirmation message to the customer

The model orchestrates this sequence autonomously. You provide the tools and guardrails; the model figures out the right order and passes data between steps.

Error Handling Is Everything

In production, tools fail. APIs time out, records don't exist, permissions are denied. How you handle tool errors determines whether your AI agent is useful or frustrating.

  • Return structured errors, not stack traces. The model needs to understand what went wrong to decide what to do next
  • Include recovery hints. "Customer not found. Try searching by phone number instead." gives the model a next step
  • Set max iterations. Without a limit, the model can loop forever retrying a broken tool. 5-10 iterations is a reasonable ceiling
  • Log every tool call with input, output, latency, and success/failure. This is your observability layer

Security: The Hard Part

Function calling creates a new attack surface. The model decides which functions to call based on user input. This means:

  • Never trust model-generated parameters blindly. Validate and sanitize all inputs before executing
  • Use permissions. A customer-facing agent should not have access to delete_account() or update_pricing()
  • Rate limit tool calls. A prompt injection attack could make the model call expensive APIs in a loop
  • Separate read and write tools. Read-only tools are inherently safer. Require confirmation for write operations

Parallel Tool Calls

Claude and GPT-4 both support parallel tool calls — the model can request multiple function calls in a single turn. This is a massive performance win for independent lookups (e.g., fetching customer info and order history simultaneously).

Support it in your implementation: check for multiple tool calls in the response and execute them concurrently with asyncio.gather() or Promise.all().

Structured Outputs vs Tool Calls

Both force the model to generate structured JSON. The difference:

  • Structured outputs: the model's response IS the structured data. Use for classification, extraction, formatting
  • Tool calls: the model requests an action that your code performs. Use when the model needs external data or side effects

In practice, most production systems use both: tool calls for actions and structured outputs for the final response format.

Related Articles

Multi-AgentArchitecture

AI Multi-Agent Systems

Architecture patterns for multi-agent systems with structured communication.

WorkflowOrchestration

AI Workflow Orchestration

Building reliable multi-step AI workflows with state management.

AI AgentsChatbots

AI Agents vs Chatbots

The 4-level spectrum from rule-based chatbot to autonomous operator.