The short version
Do not teach the model your booking rules. Give it tools that call the same services your own booking page already calls, validate every argument in code, and gate every change behind a confirmation the model cannot fake. The assistant then cannot disagree with the product because it has no logic of its own.
Where chatbots go wrong
The tempting version of an AI booking assistant puts the rules in the prompt: opening hours, how long a haircut takes, which staff member does which service. It demos beautifully. Then the business changes a rule in the real booking system, nobody updates the prompt, and the assistant keeps offering slots that do not exist.
Any time two places know the same rule, they will eventually disagree. With a chatbot the disagreement ends in a customer standing at a closed door.
The model chooses, the product decides
Aria is the conversational layer on top of Prime Coworking’s booking product. Customers write things like “need an appointment tomorrow after 5” in any language. Aria owns no availability or booking logic at all. It has 15 tools, and each one wraps the exact domain service that the product’s own booking widget already calls.
The model’s whole job is to work out what the customer wants and pick the right tool with the right arguments. What counts as an open slot is decided by the booking engine, the same way for a chat customer and a widget customer. They cannot drift apart because there is only one copy.
// A tool is a thin, validated wrapper. It contains no booking rules.
const findSlots = {
name: 'find_open_slots',
description: 'List real open appointment slots for a service on a given date.',
parameters: z.object({
serviceId: z.string(),
staffId: z.string().optional(),
date: z.string().regex(/^\d{4}-\d{2}-\d{2}$/),
}),
async run(args, ctx) {
const input = this.parameters.parse(args)
return bookingService.getAvailability(ctx.tenantId, input)
},
}The parse call is doing real work. Models produce arguments that look right and are not: a date in the wrong format, a service id invented from a name. Validate on the way in, and return an error message the model can read, so it can correct itself and try again.
A public endpoint still needs teeth
The booking API behind a customer chat is unauthenticated by design, because customers are not logged in. That means a stranger can talk to it, and a stranger can try to talk it into cancelling someone else’s appointment.
So every change to a booking is gated behind an emailed one time code, with a name and phone match as a fallback. Requests are rate limited per IP, the flow locks after repeated bad attempts, and every mutation is written to an audit log. The system prompt also has an integrity block, but treat that as a courtesy to the model and not as security. The code enforces the rules even if the model is talked out of them.
The question to ask of any tool that changes something is: if the model were completely compromised, what is the worst call it could make, and does the code stop it?
Conversation memory has sharp edges
Aria keeps a bounded history of 40 messages in MongoDB. Bounded history is sensible, since long chats cost money and eventually overflow the context window. But there is a trap in how you cut it.
A tool call and its result are a pair, and model providers reject a history where a call has no matching result, or where a result appears with no call before it. If you simply slice the last 40 messages, you can cut a pair in half. The same thing happens when a request crashes after the model asked for a tool and before the result was saved. Here is a sketch, using a simple internal format with the roles user, assistant and tool. Aria’s conversation memory has orphan turn protection for exactly this class of problem.
export function trimHistory(messages, max = 40) {
let history = messages.slice(-max)
// Never start in the middle of a pair: drop leading tool results
// and assistant tool calls until a plain user message begins the history.
while (history.length && history[0].role !== 'user') history.shift()
// Never end with a tool call that has no result (a crashed turn).
const last = history[history.length - 1]
if (last && last.role === 'assistant' && last.toolCalls?.length) history.pop()
return history
}Plan for the provider to fail and to change
Model calls fail in boring ways: timeouts, overloaded servers, rate limits. Aria retries up to four times with backoff before giving up politely. And the tool loop never imports a vendor SDK directly. It talks to a small provider interface, so Gemini 2.5 Flash today can be replaced tomorrow by editing one adapter. Prices and models change every few months, and that is not the time to rewrite your agent.
The chat widget is a React 19 iframe with a focus trap and abort on unmount, and the server owns the conversation state. The internal product is private, so there is no repository to link, but the Aria case study has the full picture.