Designing tools your agent won't misuse
Your agent is only as reliable as its worst tool. Field notes on tool ergonomics — the highest-leverage, least glamorous work in agentic AI.
When an agent misbehaves, everyone blames the model or the prompt. In my experience the culprit is usually neither: it's a tool with bad ergonomics. Tool design is API design where the consumer is a language model — brilliant, literal-minded, and unable to read your source code. Some field notes.
Name tools for intents, not endpoints
POST /v2/records means nothing to a model. create_customer_note means everything. If your tool list reads like an OpenAPI spec, you've delegated interface design to whoever built the backend, and the model pays for it in wrong guesses.
The test: could a new teammate pick the right tool using only the names? Models choose tools the same way — by name and description, under time pressure.
The description is load-bearing
Every tool description is injected into the model's context on every turn. It's the most-read documentation you will ever write, so write it like documentation:
def search_orders(customer_email: str, status: str = "any") -> list[dict]:
"""Search a customer's orders by email.
Use this FIRST when the user mentions an order without an ID.
status: one of "pending", "shipped", "delivered", "any".
Returns at most 20 orders, newest first, each with an
order_id usable with get_order_details.
"""
Notice what's in there: when to use it, valid enum values, result limits, and which tool it chains into. That last one matters most — models plan by reading descriptions, and telling them how tools compose eliminates whole classes of flailing.
Return what the next step needs, nothing else
A raw API response is written for a frontend, not for reasoning. 40KB of JSON with updatedAtTimestampMs and nested locale metadata makes every later decision in the conversation worse, because it's all still sitting in context. Project the result down to the fields a human expert would jot on a sticky note — and always include the IDs needed for follow-up calls.
Make errors instructive
The difference between an agent that recovers and one that spirals is almost always error message quality:
- Bad:
KeyError: 'cust_9f3a' - Good:
Customer 'cust_9f3a' not found. Emails, not IDs, are accepted here — call search_customers(email) to get a valid ID.
Bake the recovery path into the failure. Every error message is a free prompt.
Guard the dangerous verbs
Any tool that mutates the world gets three things from me: an explicit confirmation parameter for destructive variants, idempotency where possible (so a retried call doesn't double-charge), and tight scoping — delete_draft(draft_id), never delete(resource_type, id). Models under pressure reach for the general tool; don't stock one.
The payoff
None of this is clever. All of it compounds: swapping a generic tool set for intent-shaped ones cut tool-call errors by more than half in one production agent I work on — with the same model and prompt untouched. Before reaching for a bigger model, spend an afternoon on your tools. It's the cheapest capability upgrade in the field.