Design products around useful work, uncertainty, and user control.
Most AI features that fail do not fail because the model is weak. They fail because they were aimed at the wrong job, they hid their uncertainty, or they asked users to trust output they had no easy way to check. AI product design is the discipline of choosing where a probabilistic component earns its place, deciding how much autonomy it gets, and shaping the screens, evidence and controls around it so people can get useful work done when the model is right and recover cheaply when it is wrong.
The mental model is a cost of error curve. On one axis is how often the model is wrong for this task; on the other is what a wrong output costs and how easily a person can spot it. Generation of messy, varied content with cheap verification sits in the sweet spot. Exact calculations, irreversible actions and anything a user cannot check sit outside it, and belong to code, rules or a human. Every design decision in this chapter, from copilot versus automation to how uncertainty is displayed, is a way of moving a feature to a better spot on that curve.
The questions build in three layers. The first asks whether to use GenAI at all and what kind of product you are building (AI-native, wrapper, copilot). The second covers trust at the interface: showing uncertainty, abstaining, human review, and interfaces beyond chat. The third covers the business: where defensibility comes from, whether data is really a moat, when to automate, and how to measure value and learn from user behaviour so the product improves after launch.
Use GenAI when the input is unstructured or varied, the output is language, code or media that a rule cannot produce, and a person or a check can verify the result more cheaply than producing it from scratch.
GenAI is good at a specific class of work: reading messy input (free text, documents, images, transcripts) and producing flexible output (summaries, drafts, classifications with explanations, extracted fields, code). Traditional software is good at the opposite: exact rules over structured data. The question is not "could a model do this" but "is this the kind of work where flexibility is worth occasional error".
Three conditions together make a strong case. First, the input varies too much for rules: support tickets in twelve languages, contracts from hundreds of counterparties, scanned forms with no fixed layout. Second, a good-enough answer has value: a first draft that is 80% right saves real time if a person finishes it. Third, verification is cheaper than creation: a reviewer can check a summary against the source in a minute, while writing it took twenty. This asymmetry is the economic core of most successful GenAI features.
The volume and frequency of the task matter too. A model that saves ten minutes on a task done twice a year will not change behaviour; the same saving on a task done forty times a day will. Look for work that is frequent, tedious, and currently done by skilled people who would rather do something else.
Finally, check that you can measure success. If you cannot say what a correct output looks like for a sample of fifty real inputs, you cannot evaluate the feature, and you will ship on the strength of a demo. Teams that start by collecting real examples and labelling a small set of expected outcomes usually discover in the first week whether the use case is viable.
Avoid it when a rule, a calculation or a database lookup gives the exact answer, when errors are costly and hard to detect, or when the result must be identical every time. In those cases code is cheaper, faster and auditable.
A language model produces plausible output, not guaranteed output. That is a feature for drafting and a defect for anything with one correct answer. Tax calculations, pricing, account balances, eligibility rules and date arithmetic all have exact answers that code computes perfectly in microseconds. Routing them through a model adds latency, cost and a small but real error rate for no gain.
The second warning sign is costly, invisible errors. If a wrong output causes harm and the user cannot tell it is wrong, a model is a liability. A dosage calculation, a legal deadline or a bank transfer amount fall here. The model might be right 99% of the time, but the 1% is not caught, and the product will be judged on that 1%.
Third, some requirements demand determinism and auditability: the same input must give the same output, and you must be able to explain exactly why. Regulated decisions such as credit approval often fall under rules that require explainable reasons. Even with temperature set to zero, model outputs can vary across runs and across provider updates, and the reasoning is not a reliable audit trail.
The right pattern is usually a split, not a rejection. Let the model do the fuzzy part (extract the invoice fields from a scanned PDF) and let code do the exact part (validate totals, apply tax rules, post to the ledger). The model handles variety; code handles correctness. Many good AI features are 80% ordinary software with a model at one narrow step.
An AI-native product is one whose core workflow only makes sense because a model can read, generate or act on unstructured material. Remove the AI and the product does not degrade gracefully; it stops being the product.
Most existing software is AI-enhanced: a CRM adds a summary button, an email client adds a draft reply. The product works without the AI and the feature is a convenience. An AI-native product is designed the other way round. Its main loop assumes that a model does real work in every session, so the screens, data model and pricing are built around that assumption.
The difference shows up in design decisions. An AI-native research tool might have no blank document state at all: you start by describing what you need, it gathers sources, drafts, and you edit. An AI-native support desk might have no ticket queue in the traditional sense; most tickets are resolved before a person sees them, and the human interface is an exceptions queue. The unit of work changes from "a record I edit" to "a proposal I review".
AI-native design also changes what the product stores. Because the model's proposals and the user's edits are the core interaction, the product naturally records what was suggested, what was accepted and what was changed. That record is the raw material for evaluation and improvement, which is why AI-native products can get better with use faster than enhanced ones.
Being AI-native is not automatically better. It brings model cost into every session, makes the product's quality depend on a component you do not fully control, and requires strong handling of failure because there is no non-AI path to fall back on. It is the right choice when the AI capability removes a step users hated, not when it is added for positioning.
A wrapper is a product whose value is mostly a prompt and an interface around someone else's model, with little proprietary workflow, data or integration. It can be a fine starting point, but it is easy to copy and exposed to the model provider adding the same feature.
The term is usually used as criticism, but it describes a real spectrum. At one end is a thin wrapper: a branded form that sends user text through a fixed prompt to a model and shows the result. At the other end is a product that also calls a model but adds substantial layers around it: connections to the customer's systems, retrieval over their documents, domain-specific evaluation, review flows, audit logs, and a workflow that fits how a team actually works.
Thin wrappers are vulnerable for two reasons. First, anyone can rebuild them in days, because the prompt is the only asset and prompts are easy to reverse-engineer from outputs. Second, the model provider can absorb them: general assistants keep gaining features like document upload, web search and templates, and a wrapper whose only advantage is convenience loses it when the general product catches up.
That does not make wrappers worthless. A thin wrapper is a cheap way to test whether users want an outcome at all. Many durable products started as wrappers and then thickened: they learned which tasks users repeated, connected to the systems where those tasks lived, and built the review and data layers that made switching costly. The mistake is staying thin while believing the prompt is the moat.
The practical question is: what would a competitor need besides model access to match you? If the honest answer is "our prompt", you have a wrapper. If it is "two years of integrations, a labelled outcome dataset, and a workflow embedded in the customer's approval process", you have a product that happens to use a model.
A copilot is an AI assistant embedded in someone's work that suggests, drafts or explains while the person keeps decision authority. Nothing consequential happens until the user accepts, edits or rejects the proposal.
The copilot pattern puts the model beside the user, not in their place. It produces candidate work: a code completion, a draft email, a summary of a candidate's experience, a proposed category for an expense. The user reviews and decides. Because the human is the final check, the model can be useful at accuracy levels that would be unacceptable for automation; a draft that is right 70% of the time and easy to fix still saves effort.
Good copilots share a few traits. They work inside the existing tool, where the context already lives, rather than in a separate window the user copies text into. They are fast enough to stay in flow: a suggestion that takes fifteen seconds breaks concentration, so copilots often use smaller models, streaming, or precomputed suggestions. They make acceptance and editing cheap, with one key to accept and direct editing of the draft. And they show their basis: which documents or fields the suggestion came from.
The main risk is automation bias, the well-studied tendency of people to accept machine suggestions without checking, especially when tired or busy. A copilot that is usually right trains users to stop reviewing, and then the rare error passes through. Designs that counter it include highlighting uncertain parts, requiring explicit confirmation for high-stakes fields, and measuring how often users accept suggestions unchanged (a very high rate can signal rubber-stamping rather than quality).
A copilot is often the right first step even when full automation is the goal, because every accept, edit and reject is labelled data about where the model is reliable.
Show uncertainty through evidence and specific limits, not invented percentages: cite the sources, say what was not found or conflicts, and offer the next step. Users act well on "two sources disagree on the date" and badly on "87% confident".
A model's fluent tone is the same whether it is right or wrong, so the interface has to supply the signal the prose does not. The tempting approach is a confidence score, but numbers a model writes about itself are generally not calibrated: a stated 90% does not mean it is right nine times in ten. Even token-level probabilities, where available, measure confidence in wording rather than in facts. Showing an uncalibrated number gives users false precision.
Better signals are observable and specific. Citations let the user check a claim in one click. Explicit gaps ("the contract does not mention a termination fee") tell them what the system looked for and did not find. Conflicts ("the policy page says 30 days; the FAQ says 14") surface real ambiguity instead of resolving it silently. Scope statements ("based on documents updated before March") explain what the answer covers. Each of these is something the system can actually know, unlike a felt confidence.
Uncertainty display should also be proportional to stakes. A low-stakes autocomplete needs no caveat. A medical or financial answer needs sources inline and a clear statement of limits. Blanket disclaimers at the bottom of every answer ("AI can make mistakes") are ignored within a week and do nothing; targeted warnings on the specific uncertain part get read.
If you do want a numeric or tiered signal, derive it from things you can measure and calibrate: retrieval scores, agreement across several samples, whether a verifier check passed, or a classifier trained on labelled past errors. Then check calibration against real outcomes before showing it. A three-level label (verified, partial evidence, no evidence) that is honestly calibrated beats a precise number that is not.
When the evidence it needs is missing or contradictory, or when the cost of a wrong guess exceeds the cost of no answer. Abstaining must be designed in, with a threshold, a measured rate and a useful next step, because models default to answering.
Language models are trained to produce helpful completions, and a confident answer usually looks more helpful than a refusal. Left to themselves they will fill gaps with plausible content, which is how hallucinations reach users. Saying "I don't know" is therefore a product decision you implement, not a behaviour you can assume.
The decision depends on evidence and stakes. In a grounded system (one that answers from retrieved documents), the clearest trigger is that no retrieved passage supports the answer. Other triggers include conflicting sources, a question outside the product's scope, missing required inputs ("which account?"), and high-stakes topics where a partial answer could mislead. The higher the cost of a wrong answer, the lower the bar for abstaining.
Implementation usually combines several layers. The prompt explicitly permits and describes abstention. A retrieval check refuses to answer when no passage passes a relevance threshold. A post-generation check verifies that each claim is supported by a cited passage and abstains or trims unsupported claims. Each layer catches failures the others miss.
Abstention has its own cost. A system that refuses too often is useless and users stop asking. So treat it as a measured trade-off: on an evaluation set with known answerable and unanswerable questions, track both the wrong-answer rate and the false-abstain rate, and tune thresholds to the product's stakes. And never abstain with a dead end; say what was searched, what was missing, and who or what can help.
From what competitors cannot copy by calling the same model: deep workflow fit, integrations into systems of record, distribution, rights to proprietary outcome data, and the trust and compliance needed to execute real actions. The model itself is rarely a moat.
Frontier model capability is available to everyone with an API key, and it improves for everyone at once. So an AI product's advantage cannot come from the model's raw ability, and gains from prompting are usually copied within months. Defensibility has to come from things around the model that take time, relationships or permission to build.
Workflow depth is the most common source. A product that fits a team's actual process, including their approval steps, exception handling, roles and reporting, is costly to replace even if a rival's model output is similar. Switching means retraining people and re-mapping processes, not just changing a vendor. Integrations into systems of record (ERP, EHR, CRM, core banking) compound this: each connector involves security reviews, data mapping and edge cases, and the product that owns the write path into the system becomes part of the infrastructure.
Distribution matters more than founders like to admit. An incumbent with millions of users can ship a merely adequate AI feature and win because it is already where work happens. A startup needs a wedge, usually a narrow job done far better, to earn a place. Proprietary data can be a moat, but only under specific conditions (covered in the next question): exclusive access, rights to use it, and a measurable effect on outcomes.
Finally, trusted execution is underrated. Once an AI system takes actions such as filing claims, changing records or sending payments, customers need audit trails, permission models, certifications and a track record. These take years to accumulate and are what large buyers check first. A rival with a better demo but no audit history often cannot get through procurement.
None of these are permanent. Assess them as rates: how quickly could a funded competitor or the model provider match each layer? Spend effort on the layers where that time is longest.
Only when three things hold: your access to it is exclusive and durable, you have the legal right to use it for improving the product, and it measurably improves outcomes that a general model plus public data cannot match. Most data fails at least one test.
"We have data" is the most common claimed moat and the least examined. General models are trained on vast public corpora, so data that is merely large or domain-flavoured often adds little. What matters is whether your data changes outcomes in ways a competitor cannot reproduce.
The first test is exclusivity and durability. Is the data available only to you, and will it stay that way? Public filings, scraped websites and licensed datasets that anyone can buy are not exclusive. Data generated by your product's own use (which drafts users accepted, which predictions turned out correct, which claims were later disputed) is much more likely to be, because it only exists where your workflow runs.
The second test is rights. Enterprise contracts commonly restrict using a customer's data to train models that serve other customers, and privacy laws limit use of personal data beyond its original purpose. If the data cannot legally flow into evaluation, retrieval or training across customers, it improves one account at a time and does not compound. Check contracts before building a roadmap on it.
The third test is measurable lift. Run the experiment: does adding this data (as retrieval context, few-shot examples, evaluation sets or fine-tuning data) improve the target metric on held-out cases compared with a strong general model alone? Often the most valuable data is outcome-labelled failure data: rare cases with known correct resolutions, such as confirmed fraud or equipment faults that preceded breakdowns. That is scarce, hard to collect and hard to fake.
Even data that passes all three tests erodes as models improve and public data grows, so the moat is the data-generating process, not a static dataset. A product whose daily use keeps producing new labelled outcomes can keep its lead; a one-time data purchase cannot.
Yes, and often it should. Chat suits open-ended exploration, but most work is better served by AI placed directly in existing screens: inline suggestions, pre-filled forms, smart defaults, ranked queues and one-click actions, where context is already known.
Chat became the default AI interface because general assistants popularised it, not because it fits every task. A chat box has real weaknesses for structured work: users must know what to ask, must describe context the application already has, and must parse a block of prose to find the part they need. It also hides what the AI can do; an empty text box offers no clue about capabilities.
Most high-adoption AI features are embedded. Examples include inline completions in an editor; a support console that pre-drafts the reply and highlights the relevant order; a form that pre-fills fields from an uploaded document with each value linked to where it was found; a queue sorted by predicted urgency; a spreadsheet column that classifies each row. In each case the application supplies the context, the model's output lands where the user already works, and accepting it is one action.
Embedded design also makes output structured, which makes it checkable. A pre-filled field can be validated by code; a paragraph of chat cannot easily be. Structure lets you show per-field evidence, mark low-certainty fields for attention, and log precisely which values users changed. This is much better data for improvement than a chat transcript.
Chat still earns its place for open-ended questions ("why did revenue dip in March?"), for exploring unfamiliar material, and as a fallback when the embedded features do not cover a need. A common pattern is both: embedded features for the frequent, predictable tasks, plus a scoped assistant panel that already knows the current record. Watch what users type into that panel; frequent requests are candidates for new embedded features.
When the task is repeatable, errors are cheap or reversible, measured accuracy on real traffic is high and stable, and human review adds little beyond rubber-stamping. Even then, automate the confident majority and route exceptions to people.
Copilot and automation are points on an autonomy dial, not separate products. At one end the model suggests and a person decides everything. At the other the model acts and nobody looks. Most well-run systems sit in between: the model acts alone on cases that meet strict conditions and sends the rest to a human queue.
Four conditions favour moving toward automation. Repeatability: the task has a stable definition and a large volume, such as tagging tickets or extracting fields from standard documents. Cheap or reversible errors: a wrong tag can be fixed later at low cost; a wrong payment cannot. Measured accuracy: you have real-traffic data, often from a copilot phase, showing the model's decisions match human decisions at a rate the business accepts, broken down by segment. Review adds little: if reviewers accept 98% of suggestions unchanged in under three seconds each, they are not really reviewing, and the human step costs money while providing false assurance.
The usual mechanism is selective automation. A rule or calibrated score decides, per item, whether to act automatically. High-certainty items in low-risk categories go straight through; the rest go to a person. This captures most of the savings while keeping humans on the cases where they add value. The threshold should come from data: pick the point where automated items meet your accuracy target on a held-out set.
Automation needs controls a copilot does not. Sample a percentage of automated decisions for human audit each week so drift is caught. Track error rate by category, since an overall rate can hide one failing segment. Make every action logged and reversible where possible, and keep a kill switch that drops back to copilot mode. When the model or prompt changes, re-validate before restoring full automation.
Measure completed user outcomes and their quality against a baseline without the feature: tasks finished, time to finish including fixing AI mistakes, error rates downstream, and retention. Usage of the AI button and volume of generated text are not value.
AI features are unusually easy to measure badly. Clicks on a "generate" button, number of messages, tokens produced and thumbs-up rates all rise when a feature is novel, and none of them show that work got done better. A support assistant that generates long replies agents then rewrite will look busy on every activity metric while making agents slower.
Start from the outcome the user wanted: a ticket resolved, a contract signed off, a report published, code merged. Then measure the feature's effect on that outcome with a comparison. The most reliable comparison is an A/B test: randomly give some users or accounts the feature and compare outcome rates and times. Where randomisation is impractical, compare the same team before and after with care for seasonality, or compare accounts that adopted against similar ones that did not, knowing that adopters self-select.
Time saved must be net time. Include the time to read, verify and correct the AI output. A useful proxy is edit distance between the AI draft and what the user finally kept: if users rewrite most of every draft, the feature may be costing time even when they keep using it. Pair this with downstream quality: reopened tickets, escalations, contract disputes, bugs traced to AI-written code. A feature that speeds work up and pushes errors downstream is not saving money.
Finally, track retention of the feature (do users still use it after the first month?) and unit economics (model and infrastructure cost per completed outcome, against the value of that outcome). Many AI features show a spike of curiosity and then decay; a flat or rising week-eight usage curve among active users is a stronger signal than any launch-week number.
Good review design puts the proposed change, its evidence and its risk in one view, so a reviewer can verify rather than redo the work, and makes accept, edit and reject equally easy. It also guards against rubber-stamping and records every decision.
Adding "a human in the loop" does not make a system safe by itself. If the reviewer sees a wall of AI output with no context, they either redo the work (no savings) or approve it without checking (no safety). The design goal is to make real verification fast: the reviewer should be able to confirm a correct proposal in seconds and spot a wrong one just as quickly.
Four elements do most of the work. Show the change, not just the result: a diff of old and new contract clauses, highlighted field changes in a record, a before-and-after of an email. Attach evidence to each claim: the source passage, the transaction, the policy rule, one click away and ideally inline. Rank attention by risk: highlight fields the system is least certain about, values outside normal ranges, and changes with high consequence, so the reviewer's limited attention goes where errors are likely. Make all three outcomes cheap: accept, edit in place and reject with a reason, with editing as easy as accepting.
Review quality degrades with volume and fatigue. Automation bias means reviewers increasingly approve whatever the machine proposes. Counter-measures include batching reviews into manageable sessions, requiring active confirmation on high-stakes fields rather than one "approve all" button, tracking per-reviewer time per item and acceptance rate, and occasionally inserting known-bad items (seeded errors) to measure whether reviewers catch them.
Every review decision should be recorded with structure: what was proposed, what the reviewer changed, the reason code on rejections, the time taken and who decided. That record supports audits, and it is the best evaluation and training signal the product generates.
Capture implicit signals (accepts, edits, rejections, retries, abandonment) in structured form alongside each output, add lightweight explicit feedback with reason codes, and route it into error analysis, eval sets and prioritisation. Thumbs alone are too sparse and biased to steer a product.
Most teams add a thumbs-up and thumbs-down button and call it a feedback loop. In practice only a small share of users click either, and those who do are skewed toward strong reactions, so the data is sparse and biased. A thumbs-down also says nothing about why. Thumbs are worth having, but they should be the smallest part of the loop.
The richest signals are implicit, produced by normal use. Did the user accept the suggestion, edit it, or discard it? How much did they change (edit distance, which fields)? Did they regenerate, rephrase the request, or copy the output elsewhere? Did they abandon the flow? Did the downstream outcome succeed (ticket stayed closed, code passed review)? These are collected on every interaction, not just when someone bothers to click.
To use them, log each AI output with an output id that ties together the inputs, prompt version, model, retrieved context and every user action that followed. Without that join key you cannot answer "which prompt version produced the drafts people rewrote most". Respect privacy terms when storing content; often you can store diffs, field names and reason codes rather than full text.
Then close the loop deliberately. Review a sample of negative signals each week, cluster them into failure types (wrong fact, wrong tone, missing context, ignored instruction), and turn representative cases into regression eval items so fixes stay fixed. Use frequency and cost of each failure type to prioritise. Be cautious about training directly on raw feedback: it encodes user preferences that may conflict with correctness, and a model tuned to maximise thumbs-up can learn to flatter rather than be right.
Match the interaction to the wait: stream text for reading tasks, show progress and partial results for multi-step work, move long jobs to the background with notification, and precompute where you can. Perceived speed depends on time to first useful output, not total time.
Model calls are slow compared with ordinary software. A simple answer may take a second or two; a long generation, a retrieval pipeline or an agent running several tool calls can take tens of seconds or minutes. Users judge responsiveness by time to first useful output and by whether they understand what is happening, so the design question is how to make waiting feel purposeful or remove it from the user's path.
For text that people read, streaming (showing tokens as they are generated) makes a ten-second answer feel responsive, because reading starts after a fraction of a second. It works less well for structured output that cannot be used until complete, such as JSON that drives a form; there, show a skeleton of the form and fill fields as they arrive, or stream a short human-readable status.
For multi-step work, show progress truthfully: "searching 3 sources", "reading contract (12 pages)", "drafting". Real steps build trust and let users cancel early if the system is heading the wrong way. Fake progress bars that stall at 90% do the opposite. Partial results (the first findings while the rest continue) let users start working sooner.
For anything over roughly half a minute, consider an asynchronous pattern: the user submits, continues other work and is notified when the result is ready, with the job running on a durable queue so a closed tab does not lose it. Finally, remove waits entirely where possible: precompute summaries when a document is uploaded rather than when it is opened, prefetch likely suggestions, cache repeated results, and use a smaller, faster model for interactive steps while reserving larger ones for background work.