In a web chat, the client holds the conversation. useChat keeps the message list in React state and posts it back on every turn, so persistence is a convenience — it survives a refresh, and that is mostly it.
In a texting agent there is no client. A webhook arrives carrying one message and a phone number, and your database is the only copy of the conversation that exists. Persistence stops being a nicety and becomes the thing that makes the agent coherent at all.
Check the version before you copy anything
This SDK renames things between majors, and the search results for it are full of code from two versions ago. Everything here is AI SDK 6: system prompts go in instructions, the accumulated model turns come back as responseMessages, and the stream helpers are toUIMessageStream and createUIMessageStreamResponse. Check the reference for what you installed.
Pick the message format first
There are two, they are not interchangeable, and choosing the wrong one costs you a migration later.
| UIMessage | ModelMessage | |
|---|---|---|
| What it is for | Rendering a conversation to a human | Sending a conversation to a model |
| Carries | Ids, parts, tool state, metadata, custom data parts | Roles and content, and nothing you do not need |
| Produced by | useChat, and the UI message stream helpers | convertToModelMessages, and responseMessages on a result |
| Store it when | A browser will render this thread | The only reader is the model |
The SDK's documentation recommends storing UIMessage, and for a web chat that is right — it is the richer format, and converting down is lossless in the direction you need. For a phone channel the calculus flips: there is no useChat anywhere in your system, nothing produces a UIMessage, and inventing one so you can convert it back is ceremony.
The rule
If a browser will ever render the thread from these rows, store UIMessage. If the only reader is the model and your web inbox renders from plain text columns instead, store ModelMessage — it is exactly what responseMessages hands you.
The table
One row per message, and the vendor's own message id as a unique column. That column is what makes the whole thing idempotent, which matters more here than anywhere else in the design.
create table threads ( id text primary key, line text not null, -- the business number this belongs to contact text not null, -- E.164, normalised on the way in assigned_to text, -- set once a human takes over opted_out_at timestamptz, unique (line, contact)); create table messages ( id text primary key, -- generated before insert, not after thread_id text not null references threads(id), role text not null, -- 'user' | 'assistant' | 'tool' content jsonb not null, -- the ModelMessage content body text, -- flat text, for the inbox and search provider_id text unique, -- the vendor's id. The dedupe key. created_at timestamptz not null default now()); create index on messages (thread_id, created_at);provider_id being unique is doing real work. Webhook delivery is at-least-once, so the same inbound message will arrive twice sooner or later. An insert that violates that constraint is the cheapest possible duplicate check, and it is atomic in a way an if (await seen(id)) check is not.
Loading a thread
import type { ModelMessage } from "ai"; const WINDOW = 40; // turns, not messages — see compaction below export async function history(threadId: string): Promise<ModelMessage[]> { const rows = await sql` select role, content from messages where thread_id = ${threadId} order by created_at desc limit ${WINDOW} `; // Newest-first for the limit, oldest-first for the model. return rows.reverse().map((r) => ({ role: r.role, content: r.content }));}Note what is not here: no system message. In AI SDK 6 system prompts belong in instructions, and system messages inside messages are rejected by default. If you stored one in a previous version, that is a migration, not a runtime flag to flip.
Saving the turn
generateText returns responseMessages — the assistant and tool messages the model produced, already in ModelMessage shape and already in order. Append them and you are done.
import { generateText, createIdGenerator } from "ai"; const nextId = createIdGenerator({ prefix: "msg", size: 16 }); export async function reply(thread: Thread, text: string) { const result = await generateText({ model: "anthropic/claude-sonnet-5", instructions: SYSTEM, messages: [...(await history(thread.id)), { role: "user", content: text }], tools, }); // One transaction: the reply is stored and sent, or neither happened. await db.transaction(async (tx) => { for (const m of result.responseMessages) { await tx.insert("messages", { id: nextId(), thread_id: thread.id, role: m.role, content: m.content, body: m.role === "assistant" ? result.text : null, }); } await enqueueSend(tx, thread, result.text); });}Generate ids before you store, not after
Ids that appear only once a message has been persisted cannot be referenced by anything that happens in between — a delivery receipt, a retry, an escalation. createIdGenerator gives you a stable id up front. The SDK offers the same thing server-side for streamed responses through generateMessageId.
Compaction, because these threads never end
A web chat is a session. A customer thread is a relationship — the same phone number, on and off, for years. Left alone it grows past any context window and takes your latency and your bill with it.
The useful shape is a rolling window plus a durable summary, refreshed when the window slides:
- Keep the last N turns verbatim. Recent exchanges are where almost all the relevant detail lives.
- Keep a summary row per thread — who this is, what they buy, what went wrong last time. Regenerate it when messages fall out of the window, not on every turn.
- Keep facts as facts, not prose. Their usual appointment, their car, their dog's name: columns, not a paragraph the model has to re-read and can misread.
- Never summarise consent. Opt-in and opt-out are columns with timestamps. A summary that says 'the customer seemed happy to be contacted' is not a legal record.
For compaction *inside* a single multi-step run — a tool loop that gets long — the SDK has prepareStep and a pruneMessages helper. That is a different problem from the one above, which happens between conversations rather than within them.
What else belongs on the thread
The transcript is not the state. These fields are, and every one of them is something you will otherwise try to infer from the transcript at three in the morning:
- Consent status and the timestamp it changed — see consent before you text.
- Assignment. Who owns this thread right now, human or machine.
- Escalation history. How many times, and why.
- The provider's delivery status per outbound message, which arrives later than the send and belongs next to it.
- A retention date. Messages are personal data and they age badly — data retention for messages covers what to keep and for how long.
If you do also have a web chat
Plenty of teams run both: an agent that answers texts and a chat widget on the site, sharing one brain. On the web side the SDK's own persistence path applies, and it is worth using rather than reinventing.
- Save inside the stream, not after it.
toUIMessageStreamtakes anonEndcallback that receives the complete messages including the response — that is the hook, and it fires even though the response is streaming. - Send only the last message.
prepareSendMessagesRequeston the transport lets the client post one message and an id; the server loads the rest. On a long thread this is a large saving. - Validate what you load.
validateUIMessageschecks stored messages against your current tool and metadata schemas. Rows written by last quarter's tool definitions will not always match this quarter's, and failing loudly at load time beats a confusing model error.
Both halves can share one storage layer if the web side stores UIMessage and the phone side stores ModelMessage in the same table with a discriminator. What they must share is the thread and the consent state — two systems texting the same customer with two different ideas about whether they opted out is a genuinely bad afternoon.
The agent that sits on top of all this is in how to build an AI text message responder.
Next step
Generate a tagged link for whatever you send next with the UTM builder, see what this looks like in your industry, or compare the services that can send it on the providers page.