• LLM chatbots
  • Next.js
  • Cost control

Put the Model Last: A Cheaper, Safer Chatbot

The short version

Treat the model as the last stage of your pipeline, not the first. Cheap code should answer what it can, refuse what it must, and cache what repeats. Whatever is left goes to the model, and its reply is checked before anyone sees it.

Why a public chatbot is a different animal

A chatbot behind a login has known users and a bounded audience. A chatbot on a public website is an open endpoint that costs money every time someone sends a message. People will ask it the same five questions all day, try to make it say something embarrassing, and now and then paste a novel into it.

The Copper Larder is a demo front of house host for a British bistro. It answers menu questions, takes callback requests and streams its replies. The interesting part is everything wrapped around the model. The request goes through 8 stages, and only stage 7 costs a token.

The cheap stages come first

Ahead of the model sit session caps, a handoff for complaints, scripted intercepts and an exact match cache. The cheapest and most certain checks go first.

  • Session caps stop one visitor, or one script, from running up the bill.
  • Complaint handoff means an angry customer reaches a human instead of getting a chirpy paragraph.
  • Intercepts are plain pattern matches for the questions everyone asks, such as opening hours or whether there is parking. The bistro has about 19 of them. They cost no tokens and give the same correct answer every time.
  • An exact match cache returns a stored answer when the same normalised question has been answered before.
JavaScript
const intercepts = [
  { test: /\b(opening|open)\s+(hours|times)\b|\bwhen are you open\b/i, answer: () => hoursCard(restaurant) },
  { test: /\bparking\b/i, answer: () => infoCard(restaurant.parking) },
]

export function tryIntercept(message) {
  const hit = intercepts.find((rule) => rule.test.test(message))
  return hit ? hit.answer() : null
}

const normalise = (text) => text.toLowerCase().replace(/[^a-z0-9 ]/g, '').replace(/\s+/g, ' ').trim()

Notice that the answers come from restaurant data and are not typed by hand into the rules. The dish and info cards are built only from typed restaurant data, so they can only show what that data says.

Remember what the guest told you

If someone says “I’m vegan” in message 2, that must still hold in message 20. Models forget, and long histories get trimmed. The Copper Larder treats it as a conversation wide dietary lock. A sturdy way to build one is to store the constraint in session state, inject it into every turn and filter the menu cards in code, so the model is never trusted to remember a constraint that matters.

Check the reply before it ships

The model is the one stage that can say anything, so its output is checked. In this bot the important rule is that it must never confirm availability, because it cannot know it. If a reply promises that a table is free, a guardrail catches it, rewrites it with a correction, and, importantly, blocks that reply from entering the cache. A simple version looks like this. A cached mistake is a mistake you serve for free to every future visitor.

JavaScript
const promisesAvailability = /\b(we have|there is|i can confirm|you are booked|table is available)\b/i

export function checkReply(reply) {
  if (promisesAvailability.test(reply)) {
    return {
      reply: reply + '\n\nI can’t see live availability, so please leave a callback request and the team will confirm.',
      cacheable: false,
    }
  }
  return { reply, cacheable: true }
}

Always have a graceful failure

The model will be down, or rate limited, or the key will be missing in some environment. Every one of those paths returns a warm, on brand message and a callback card, so the visitor can still leave their number. A broken demo is a bad impression. A polite “I’ll have someone call you” is just a slightly different product.

Two smaller habits round it out. Replies stream over server sent events, with aria live regions so screen readers announce text as it arrives. And the app stores no raw IP addresses, so there is nothing sensitive to leak. See the Copper Larder case study for the rest.

From the projectThe Copper LarderHannah — an AI front-of-house concierge for hospitality sites.Read the case study →

Keep reading