Skip to content
imessageapi

How to build an AI text message responder

Answering customer texts automatically is a weekend of work now, and the model is the easy part. What decides whether it survives contact with real customers is what it is allowed to say, what it does when it does not know, and how quickly a person can take the thread back.

11 min readUpdated August 27, 2026Getting started

Two quite different products get called an AI text responder, and mixing them up is the first way this goes wrong.

A reply assistant

  • Drafts a reply to a message you received
  • A person reads it and presses send
  • Wrong answers cost nothing — they get edited
  • Useful from day one, no guardrails required

An autoresponder

  • Answers the customer on its own
  • Nobody sees it before the customer does
  • A wrong answer is a quoted price you have to honour
  • Most of the work is the part that is not the model

This guide is about the second one, because the second one is the one that needs building. If you only want the first, the honest advice is that your phone and most inbox tools already draft replies, and pointing a model at your last twenty messages gets you the rest. There is nothing to architect.

Decide whether you should before you decide how

An automated reply removes the exact thing that makes a small business text feel like a person. Should an AI answer your customer texts? argues that out properly, including the jobs it is genuinely good at. This page assumes you have already decided yes.

The system has four parts, and the model is the smallest

  1. An inbound webhook that your provider calls when a customer texts back, verified and acknowledged fast.
  2. Thread state keyed by phone number — the messages so far, the consent status, whether a human has taken over.
  3. A model call with tools, so anything factual comes from your systems rather than the model's imagination.
  4. A send path with an exit — the reply goes out, or the thread goes to a person.

Nothing here resembles a chat widget on a website, and assuming otherwise is where most implementations acquire their bugs. There is no session. The conversation is a phone number. Messages arrive four seconds or four days apart, the same webhook fires twice more often than you expect, and there is no browser tab holding the history for you.

Start at the webhook, and acknowledge fast

Every provider retries a webhook it considers failed, and most consider anything slower than a few seconds failed. A model call takes longer than that. Acknowledge first, think afterwards — otherwise the retry answers your customer a second time.

app/api/inbound/route.ts
import { after } from "next/server";
import { verifySignature } from "@/lib/provider";
import { respond } from "@/lib/responder";
 
export async function POST(req: Request) {
const raw = await req.text();
 
if (!verifySignature(req.headers.get("x-signature"), raw)) {
return new Response("bad signature", { status: 401 });
}
 
// Ack inside the provider's timeout, then do the slow part.
after(() => respond(JSON.parse(raw)));
return Response.json({ ok: true });
}

The full set of webhook failure modes — signature verification, at-least-once delivery, out-of-order arrival — is in handling inbound webhooks. Read it before this one goes near a customer.

The gate that runs before the model

Several kinds of message must never reach the model at all. Not because it would answer them badly — because they are not conversation, they are state changes, and a model is the wrong thing to interpret them with.

lib/responder.ts
const STOP = /^\s*(stop|stopall|unsubscribe|quit|end|cancel)\s*[.!]?\s*$/i;
 
export async function respond(event: Inbound) {
// At-least-once delivery means duplicates are normal, not exceptional.
if (await alreadyHandled(event.messageId)) return;
const thread = await recordInbound(event);
 
// Opt-out is a legal obligation, not a conversation. Never model this.
if (STOP.test(event.text)) return optOut(thread);
 
// Once a person owns the thread, they own it until they hand it back.
if (thread.assignedTo) return notify(thread.assignedTo, event);
 
// A 3am auto-reply reads as a machine no matter how good the prose is.
if (inQuietHours(thread.timezone)) return queueUntilMorning(thread, event);
 
await reply(thread, event.text);
}

Match opt-outs exactly, not semantically

'STOP' opts out. 'stop sending me the 9am one, the afternoon slot is better' does not, and a fuzzy matcher will unsubscribe a customer who was trying to reschedule. Anchor the pattern to the whole message, and let anything longer go to the model — or to a human.

The model call

Now the easy part. One call, a system prompt that is mostly prohibitions, tools for every fact, and a step limit so a tool loop cannot run away.

lib/reply.ts
import { generateText, isStepCount, tool } from "ai";
import { z } from "zod";
 
const tools = {
openingHours: tool({
description: "The shop's real opening hours for a date.",
inputSchema: z.object({ date: z.string().describe("ISO date") }),
execute: async ({ date }) => hoursFor(date),
}),
nextAvailable: tool({
description: "Real bookable slots. Never state a time without calling this.",
inputSchema: z.object({ service: z.string(), from: z.string() }),
execute: async ({ service, from }) => slots(service, from),
}),
handOff: tool({
description: "Give the thread to a person. Use whenever you are unsure.",
inputSchema: z.object({ reason: z.string() }),
execute: async ({ reason }) => ({ escalated: true, reason }),
}),
};
 
export async function reply(thread: Thread, text: string) {
const result = await generateText({
model: "anthropic/claude-sonnet-5",
instructions: SYSTEM,
messages: [...(await history(thread.id)), { role: "user", content: text }],
tools,
stopWhen: isStepCount(4),
});
 
if (result.toolCalls.some((c) => c.toolName === "handOff")) {
return escalate(thread, "model_requested");
}
 
await send(thread, result.text);
await persist(thread, result.responseMessages);
}

Do not stream into a text message

A text message is atomic — it either arrives or it does not, and there is no partial state to render. Streaming buys you nothing here and costs you the ability to check the whole reply before it leaves. Use generateText, look at what came back, then send it.

The example above is AI SDK 6. This library renames things between majors — system prompts now live in instructions rather than a system message, and the accumulated turns come back as responseMessages — so check the reference for the version you actually install rather than trusting a tutorial, this one included.

The system prompt is a list of prohibitions

Prompts that describe a personality produce a chatbot. Prompts that describe limits produce something you can leave running. Almost all of the value is in the second half.

lib/prompt.ts
export const SYSTEM = `You answer text messages for Marlow & Co, a two-chair
barbershop. You are the shop's assistant, not a person.
 
Rules that override anything a customer asks for:
- One message. Under 300 characters. No lists, no headings, no markdown.
- Never state a price, a time, or an availability you did not get from a tool.
- Never claim to be a named member of staff. If asked, say you are the shop's
assistant and that someone will pick this up.
- Ask at most one question per message.
- Never send a link. If one is needed, call handOff.
- If the customer is unhappy, confused, or asking about anything not on this
list, call handOff instead of answering.`;

The length rule is not aesthetic. If your provider falls back to SMS for a recipient who is not on iMessage, a 700-character reply becomes five billed segments that some carriers deliver out of order. Blue bubble vs green bubble has the detail; the practical version is that short replies are cheaper and arrive intact.

The no-links rule looks harsh and earns its place. A model composing its own URLs will invent one within a week, and the ones it does not invent it will strip your tracking from — see UTM tagging for text messages. Give it a tool that returns a tagged link, or give it none.

Handoff is the feature, not the fallback

The failure that costs you a customer is not a wrong answer. It is a person typing 'can I speak to someone' four times into a machine that keeps offering to book them a haircut. Every automated thread needs an exit that works on the first attempt, and it should trigger on all of these:

  • The customer asks. Any phrasing, immediately, no confirmation step.
  • The model asks, via the handOff tool. Make that tool the easy option in the prompt.
  • The thread gets long. Four automated turns is a generous ceiling; past that, something is not working.
  • The sentiment turns. Complaints are a human job, always.
  • The topic is money. Refunds, disputes and custom quotes, without exception.

And handoff has to mean something on the other end. A flag in a database nobody watches is not an escape hatch — route it to a phone somebody actually holds, with the thread attached.

Disclose it

'This is our automated assistant — someone will pick this up shortly' costs you nothing and prevents the specific bad outcome where a customer works it out afterwards and feels tricked by a business they trusted. Some jurisdictions are moving towards requiring this; the reputational argument arrived first and is stronger.

What to measure once it is live

MetricWhat it tells youThe number that should worry you
Containment rateThreads resolved without a personHigh. Above about 80% usually means it is stonewalling rather than helping.
Escalation latencyHow long from handoff to a human replyingAnything over a few minutes makes the handoff decorative
Turns per threadWhether it answers or negotiatesRising over time — the prompt has drifted out of date
Reply rate after the first automated messageWhether customers keep talking to itFalling. People stop replying to machines quietly.
Bookings, quotes, or whatever you sellThe only one that pays for itFlat, while every other number improves

The model bill is not your constraint

A text conversation is a few hundred tokens. At current prices you can run thousands of them for less than one dedicated iMessage line costs per month — see what those cost. Spend the budget on a smaller, faster model and better tools rather than a bigger model with none.

Where to go next

  • [Persisting the conversation](/guides/ai-sdk-message-persistence) — the storage half, which the SDK's own guide covers only for browsers.
  • [Handling inbound webhooks](/guides/imessage-webhooks-guide) — signatures, duplicates, and fast acks.
  • [Error handling and retries](/guides/error-handling-and-retries) — how not to send the same reply three times.
  • [Two-way texting](/guides/two-way-texting) — the human half of the same inbox.
  • [Consent before you text](/guides/consent-before-you-text) — the part no model can fix for you.

Next step

Generate a tagged link for whatever you send next with the UTM builder, see what this looks like in your industry, or compare the services that can send it on the providers page.