00.0

Home  /  AI and conversation design

AI and conversation design

The model is not the hard part. Being wrong in public is.

Every voice and chat product eventually meets the same moment: it is not sure, and it has to do something anyway. Most of them hide it. They pick the likeliest reading, act on it with full confidence, and let the person discover the mistake later, when it is expensive.

Designing that moment is the job. Four AI products, one of them a customer facing agent for a company you have heard of, taken from problem definition to production.

listening
In plain terms

Conversation design is the design of what a voice agent or chatbot says and does. The AI model is bought. The design is everything around it: what the thing is allowed to do, how sure it has to be before it acts, what it says when it is not sure, and how a person gets to a human when it goes wrong.

That last part is the whole subject of this page. A product that is confidently wrong in front of a customer costs more than one that admits it needs to check.

What the discipline actually is

Conversation design is not writing what the bot says.

That part takes an afternoon. The work is everything around it: what the product is allowed to conclude, from how little, and what it does at each level of doubt.

A conversation is a sequence of commitments. Every turn either narrows what the system believes or widens it, and a system that never widens is a system that cannot be corrected. Most products are built to sound confident because confidence demos well, and then they meet a real caller with an accent, a bad line, and a problem that does not fit any of the six intents anyone thought of.

So the artefacts are not scripts. They are decision boundaries, escalation rules, correction paths, and an evaluation set that can prove a change made things better rather than just different.

Not this"Write me a prompt"

A prompt is one artefact of many, and the one most likely to be rewritten in month two. Designing around the prompt alone produces a product that works in the demo and fails on the fourth call.

ThisWhat happens at 0.61 confidence

Does it guess, ask, or hand over? Who does it hand to, how fast, and what does the person receiving it already know? Those three answers decide whether the product is trusted, and none of them is a wording problem.

The mechanic

One sentence, several readings, and a visible way out.

This is a real intent model from a live product. Pick something a caller might actually say and watch what the system considers, how sure it is, and what it does about the readings it rejected. The escalation lane is always drawn, because an escape hatch that only appears in the failure case is one nobody trusts.

What the system considers

What the job actually is

Eight tasks, roughly in the order they happen.

Not one of them is writing what the product says. Each one carries the failure you get if you skip it, and every failure listed was found by shipping something and reading what came back rather than by reasoning about it beforehand. That is the only method that finds them.

Where this comes from

One product for a company you have heard of, three of my own.

Verizon Connect

2023 to 2024  ·  Dublin  ·  fleet software

A customer facing agent, sole designer, problem definition through to production. The brief was not "add AI". It was that the support team was carrying more contacts than it could absorb, and a large share of them were the same small set of questions arriving over and over. The target was the load, not the technology.

So the first move was to find out what people actually ask, at volume, rather than what the organisation assumed they ask. The highest-traffic flows are not usually the most interesting ones, and that is exactly why they are the right ones to take off a human.

Then I tested the conversation before anything could hold one. Wizard of Oz: a person behind the curtain answering as the agent, while the customer believes they are talking to a system. It is the cheapest way to find out whether a designed conversation survives contact with a real one, and it costs a room and an afternoon rather than a sprint. Everything I learned that way was learned before a single model decision was made.

What came out of it was a prototype of the main use flow, taken through to production, and the artefact that mattered more than the prototype: a written boundary of what the agent would and would not attempt.

In scopeWhat it was allowed to answer
  • The highest-volume repeated questions, the ones where a human adds nothing but availability
  • Anything answerable from information the system already holds and can show its working for
  • The first, structured part of a longer request, so a human picks it up already knowing what it is about
Deliberately outWhat it always handed over
  • Anything with a money, contract or account consequence. Being nearly right there is not a service, it is a dispute
  • Anything urgent enough that a second attempt costs the customer their day
  • Long-tail questions. High variety and low volume is the worst possible trade: maximum design cost, minimum load relieved
  • Anything a frustrated customer said twice. Repetition is a state, not a query
Wizard of Oz

The conversation tested with real people before the agent existed. A person behind the curtain, a customer who does not know.

Sole designer

Problem definition, research, prototype and production. No design team to hand to and none to hide behind.

The out list

The boundary of what it refuses is the artefact that protected the project. It is also the one nobody asks for.

ParrotB

Live  ·  own product  ·  four languages

A voice receptionist for one-person trades. A plumber under a sink cannot answer the phone, and every missed call is a job that goes to whoever did pick up. An answering service that books badly is worse than no answering service, because now the diary is wrong as well as the day.

The intent model above is this product's. The failure it taught me was not the model misunderstanding the caller, which turned out to be rare. It was bookings that fit the sentence but not the working day: the right job, at a time that ignored the ninety minutes it takes plus twenty minutes of travel, landing on top of a buffer that already existed.

The rule that fixed it

Never quote an exact minute. Offer hour windows out loud, and store the precise start time underneath.

Checked against the day

Job length plus travel, both inside one working slot, after the buffer on the booking before it. Only then may the agent say yes.

Measured, not asserted

An evaluation set per behaviour, so a change to the intake can be shown to be an improvement rather than argued to be one.

If you are building one of these

The question to ask a candidate is not about models.

Ask what their product does at 0.61 confidence. Ask who it escalates to and what that person receives. Ask how they would know a change made it better. Anybody who has actually shipped one of these has an answer ready, because those three questions are what the work turned out to be.

I take six to twelve month contracts, remote, part time or full time, as a senior designer or a product owner. AI features are the half of the work I would choose if somebody made me choose.

NextThe match block

Role, length, hours, right to work and contracting, all on one screen.

Services →
OrTen seconds to a straight answer

Four questions and a verdict that is willing to be no.

Contact →
The full case study

ParrotB: the office a small trade never had

Portfolio →