heytappr.book a demo

your in-app ai shouldn't write its own answers

a chatbot that invents a settings page is worse than no chatbot at all. we think the ai should pick the answer and a person should write it. this is the reasoning, and what it looks like in practice.

by shritupdated 6 min read

still from Whisper of the Heart: Shizuku and Seiji at a library table stacked with books
the library, from Whisper of the Heart (Studio Ghibli, 1995). still shared by studio ghibli for free use.

tl;dr (too long, didn't read)

in-app ai makes stuff up when you let the model write the answer. in product help that means confident directions to menus that don't exist.

the fix is boring and it works: let the ai figure out what the user means and pick the answer, but only from answers your team wrote and approved. no match? it says so. no guessing.

that's how heytappr is built. the ai matches the question and the current page to an approved guide, then walks the user through the real screen. anything it can't answer gets logged, so you know what to write next.

the menu that doesn't exist

imagine a user asks the new ai help widget in your product: "how do i turn off email notifications for comments?"

it answers instantly. go to settings, then notifications, then comments, and switch off the email toggle. the reply is friendly and specific, and it sounds exactly like something your own team would write.

but your notifications page has no comments section. it never did. the model has seen thousands of other apps with that menu, and it filled in the blank with the most likely answer. the user clicks around for a couple of minutes, gets annoyed, and files a ticket that opens with "your bot lied to me".

that's a hallucination, and in product help it does a particular kind of damage. a wrong fact in an essay is embarrassing. a wrong instruction inside your product makes people trust the product less, because the mistake came from the product itself.

Pokémon, via GIPHY

why product help is a hard case for generative ai

language models are very good at sounding right. product help happens to be a place where sounding right and being right come apart all the time, for reasons that have nothing to do with how smart the model is.

start with the interface itself. the model has read millions of settings pages, and yours is one it has probably never seen. even if it read your docs last month, you may have shipped a redesign last week. then there are plans and permissions. "click export" is a wrong answer for someone on a plan without export, or for a viewer who doesn't have admin rights.

and the user has no easy way to check. if a chatbot tells you something about history, you can look it up. inside an app, the only way to test an instruction is to try it and fail, which is precisely the experience you were trying to prevent.

retrieval helps a lot here. feeding the model your help articles before it answers cuts down the invention considerably. it doesn't close the gap, though, because the model still writes the final sentence. it can blend two articles, drop a step, or describe the version of the page from before the redesign. the error rate goes down. it doesn't go to zero, and one confident wrong answer is enough to undo a lot of goodwill.

let the ai pick, and let people write

the design we settled on for heytappr splits the job in two, and we think it's the right default for anything that tells users where to click.

one half of the job is understanding messy human questions, and ai is genuinely excellent at it. "invoices going to the wrong person", "change billing email" and "my accountant isn't getting receipts" are the same request in three disguises. recognizing that is a perfect task for a model, and much harder for keyword search.

the other half is writing the instructions, and that's the part a model shouldn't do. only your team knows where things actually are in your product. so the work divides roughly like this:

jobwho does itwhy
understand the questionaipeople phrase the same need a hundred ways
read the current pageaithe same question means different things on different screens
choose which guide fitsai, from approved guides onlythe choice is flexible, the options aren't
write the stepsyour teamonly you know where things are in your product
approve and publishyour teamnothing reaches users without a person saying yes
no good matchnobody guessessay so, and log the question as a gap

what you end up with is an assistant that's flexible about what users say and strict about what it says back. that asymmetry is the whole point.

what "i don't know" should look like

most chatbots are built to always produce something, because an empty reply feels like a failure. in product help we'd argue the opposite. the real failure is a confident answer that sends someone to the wrong place.

when heytappr has no approved guide that fits a question, it tells the user it doesn't have a walkthrough for that yet. the question then goes into the team's analytics as unanswered demand. after a few weeks that list becomes an unusually honest to-do list: the exact questions users asked, in their own words, that nobody has written a guide for.

seen that way, every "i don't know" is a small piece of free user research. a hallucinated answer throws that signal away and costs you some trust on top of it.

One Piece, via GIPHY

show the step instead of describing it

there's a second, quieter reason text answers go wrong. even a perfectly correct description, like "click the gear icon, then billing", still asks the user to turn words into a location on a screen. which gear? the one at the top right, or the one in the sidebar?

so heytappr doesn't describe steps in a chat bubble. each step of a guide points at the real control on the live page through a CSS selector. heytappr spotlights that element, says the instruction out loud, and moves to the next step when the user actually clicks it. if a release breaks a selector, the step is flagged in the team's needs-attention list instead of quietly pointing at nothing.

put the pieces together and there's very little room left for anything to be made up. the ai understands the question, a person wrote the answer, and the answer is shown on the real interface.

when a generative chatbot is still the right call

this approach is deliberately less magical than a chatbot. a generative bot can take a swing at anything, including questions nobody prepared for. an approved-answers assistant can only answer what your team has written. that's a real trade-off, and we don't think one side wins everywhere.

our rough rule is that a generative chatbot is fine for open questions where a slightly wrong answer is cheap: explaining a concept, drafting something, brainstorming. approved answers belong wherever the answer is an instruction inside your product, like navigation, settings, billing or permissions. plenty of teams run both, with a docs chatbot for concepts and guided walkthroughs for the "where do i click" questions. if you work in fintech, healthcare or insurance, where a wrong instruction has consequences, the approved version is usually the only one compliance will sign off on.

before you ship any ai into your product, it's worth being able to answer a few plain questions:

  1. can it say something your team never wrote, and if so, what stops it from inventing a menu?
  2. when it doesn't know, does it guess or say so?
  3. can you see every question it couldn't answer?
  4. does it know which page the user is on?
  5. after a redesign, how do you find out which answers broke?
  6. can you pause or edit a single answer without redeploying anything?

if some of those don't have good answers yet, you'll learn which ones matter eventually. it's cheaper to learn it from a list than from a support ticket.

see it in your product

fifteen minutes: heytappr answering a real question by walking a user through a real interface, out loud.

questions people ask

Why do AI chatbots hallucinate in product support?

Because a language model writes the final answer from patterns it has seen, and it has seen far more generic settings pages than yours. Without strict limits it fills gaps with plausible-sounding menus and steps that don't exist in your product.

Does retrieval-augmented generation (RAG) stop hallucinations?

It reduces them a lot, because the model reads your real docs first. But the model still writes the final sentence, so it can merge articles, skip steps, or describe an outdated screen. For step-by-step instructions, choosing from pre-approved content is safer.

How does HeyTappr prevent hallucinations?

HeyTappr's AI only chooses which of your team's approved, published guides matches the user's question and current page. It never writes the steps itself. With no match, it says so and logs the question for your team.

Is an approved-answers assistant good for regulated industries?

Yes. Every instruction a user sees was written and approved by your team, which is what compliance reviews in fintech, healthcare, and insurance usually require.

go deeper

keep reading

spotted something out of date? tell us and it gets fixed.