The short version
- I built a support chat that answers the moment someone asks, using two sources: a set of saved FAQs, and the page the person is already looking at.
- The FAQs are found by meaning, not keywords. That look-up-then-answer step is called RAG. The page adds what FAQs can't know: who's asking, and what this record says right now.
- Slack is the control room. Every chat lands in a thread. From there I can jump in and answer, or add, edit and delete FAQs by typing a message.
A help center is the project everyone agrees is important and nobody wants to write. The painful part isn't the writing. It's accuracy. An article has to know what the product does, including the rules that never show up on the happy path.
So I skipped the articles. The answers that stay true live as short question-and-answer pairs. The facts that change per screen travel with the question. The assistant may speak from those two piles and nothing else.
This is the honest build log: how it works, the exact prompt, the parts that went wrong first, and a workflow you can copy.
What I actually built
Not a docs site with a table of contents. A chat inside the product. Someone types "Can I get my money back?" and gets a short answer, then a few bullets if they help.
Behind it there are three things:
FAQs
Stored in a vector index, found by meaning.
Live context
From the page: role, prices, what's switched on.
System prompt
Sets the voice and the limits.
Next to it sits one Slack channel. It gets a copy of every conversation, it's where a teammate steps in, and it's where the FAQ library is edited. There's no admin screen. If you can post in the channel, you can change what the assistant knows.
The anatomy of it
A classic help center is search on top, a section per topic, and one template so every page sounds like the same writer. This does the same three jobs:
- Search is an embedding. The question becomes a list of numbers that captures its meaning (
text-embedding-3-small, 1,536 dimensions). That's how "money back" finds the FAQ titled "Refunds" even though they share no words. - Sections are FAQ vectors. Each saved Q&A is embedded the same way and stored in Pinecone, with its text beside it. At question time, the closest few come back.
- The template is the system prompt. Direct answer first, then two to four bullets. If a fact isn't in either source, it says so and gives the support email.
Found by meaning, not keywords — try it
The visitor asks…
- Refundsclosest match
- How to publish
- What each role can do
Once that shape is fixed, adding knowledge is a fill-in: write the question, write the answer, and the next similar question can find it.
One question, start to finish
Five steps. Tap through them in your head:
- Gather the screen. The page collects what it already knows: signed in or not, role, and the record being viewed (names, prices, enabled features).
- Search the FAQs. The question is embedded and the index returns the nearest matches. Embed, search, hand the text to a model: that's RAG.
- Build one prompt. Rules, live record, matched FAQs and recent messages go in together. The record is pasted in, so the model never has to remember it.
- Call the model. LangChain sends it to OpenAI, Anthropic or Gemini. It's just the adapter, so switching providers changes nothing else.
- Mirror it to Slack. The visitor already has their answer. The same turn is posted to a thread with the question, the reply, the page, and every FAQ used, each with its id and score.
One question, start to finish
- 1Gather the screenSigned in or not, role, and the record being viewed: names, prices, enabled features.
- 2Search the FAQsThe question is embedded; the index returns the nearest matches. Embed, search, hand the text to a model — that's RAG.
- 3Build one promptRules, live record, matched FAQs and recent messages go in together. The record is pasted in, never remembered.
- 4Call the modelLangChain sends it to OpenAI, Anthropic or Gemini. Switching providers changes nothing else.
- 5Mirror it to Slacknot on the critical pathThe visitor already has their answer. The thread gets the question, the reply, the page, and every FAQ used with its id and score.
Slack as the back office
Watch everything. A new chat is a new message. Follow-ups stay in its thread. If someone asks for a human, the channel gets a ping with the page and user.
Take over. Reply inside the thread:
reply Happy to help. Refunds for this one close 24 hours before it starts.The model steps aside and the visitor sees a named person has joined. When you're done:
exit You're back with the assistant. Ask it anything else.Fix the library. Same channel, three commands:
faq add
Can I get a refund?
Refunds follow the policy on the record. If none is set, say so and give the support address. Do not guess a deadline.
faq edit abc123
Can I get a refund?
Refunds are available until 24 hours before the start time, unless this record says otherwise.
faq delete abc123The next question uses the updated text. No redeploy and no overnight reindex.
The loop is simple. The assistant fumbles something, you open the thread, see which FAQ it used and its id, and either reply for that one person or faq edit for everyone after.
One thread, the whole back office
assistant mirror
New chat · /event/summer-market
Q: “Can I get a refund?”
A: Refunds close 24 hours before the start time.
used FAQ abc123 · Refunds
teammate, typing in the thread
reply Happy to help. Refunds for this one close 24 hours before it starts.
The model steps aside — the visitor sees a named person has joined.
teammate, typing in the thread
faq edit abc123
Can I get a refund?
Refunds are available until 24 hours before the start time, unless this record says otherwise.
The next question uses the updated text. No redeploy, no overnight reindex.
The two-source method
Accurate support needs two kinds of truth: what's true of the product in general, and what's true of this screen. Most bots capture one and quietly get the other wrong.
1. FAQs: the product, written by a person. Refund rules, how to publish, what each role can do. They hold across every record.
2. The live record: this screen, sent by the page. Prices, enabled features, who's asking. An FAQ can explain refunds in general, but it can't know this event's price.
Why both. FAQs alone are almost right and wrong about what's on screen. The live record alone knows the prices but will invent a policy nobody wrote. Together, the reply quotes the price because the record said so and states the refund rule because an FAQ said so. If they conflict, the record wins.
Two kinds of truth, one prompt
the FAQs
The product, written by a person. Refund rules, how to publish, what each role can do. True across every record.
the live record
This screen, sent by the page. Prices, enabled features, who's asking. An FAQ can't know this event's price.
one prompt
The model answers only from these two piles — and if they conflict,the record wins
The whole prompt fits in one place:
You are a support assistant.
Answer only from EVENT CONTEXT and matched FAQs.
- Start with a direct answer. Add 2–4 bullets when they help.
- Reuse FAQ wording when it applies.
- If the event and an FAQ disagree, prefer the event.
- Never invent a policy, price, deadline, link, or feature.
- If the fact is missing, say what you know and point to support@yourdomain.com.
Do not mention "context", "knowledge base", or the model.What took longer than expected
If I pretended it worked first try, this would be an ad. The honest speed bumps:
- Teaching it to say "I don't know." Early versions were helpful to a fault. A missing deadline became a plausible one. The fix was naming the forbidden inventions and giving it the exact sentence to use instead. It took several rounds.
- FAQs that contradict each other. Two chunks can score well and disagree. The Slack post shows both ids, so you spot it. Someone still has to delete one.
- Command wording. Symbol prefixes were a trap, because Slack turns a leading
>into a quote. Plain words (reply,exit) worked. Commands also only work inside a thread, and a private "go into the thread" message saved a lot of confusion.
The rule I didn't break: a person owns the FAQ text. Retrieval will faithfully serve a stale answer, and the model will repeat it confidently. Slack makes the correction a one-message job. It doesn't make a wrong FAQ true.
What you write once vs. what runs every time
| You, in Slack | The assistant, every question | |
|---|---|---|
| Product truth | faq add / edit / delete | Retrieves the closest matches |
| This screen | Step in only if the record isn't enough | Attaches role, prices, flags |
| A live chat | reply to join, exit to hand back | Answers until a person takes over |
| A missing fact | Put the support address in the prompt | Says what it knows, then stops |
The model is the last step. The FAQs and the live record are the authority.
Steal my workflow
- Turn real questions into FAQs. One pair per question support already answers.
- Embed them and retrieve a few. Only the closest chunks go in the prompt.
- Attach the live record. Label it as data and say which source wins.
- Write the refusal into the prompt. Name what the model must not invent.
- Mirror every turn into one Slack thread. Include the FAQ ids. That's your feed and your debug log.
- Take over with
reply, give back withexit. - Edit FAQs in the same channel. Embed on write so changes are instant.
Try it on one cluster. Ten FAQs, one page of context, one channel. Ask a question, read the thread, fix the FAQ from the id, ask again. Once that reply is right, the rest is more FAQs.
FAQ
Same purpose, different moment. A help center makes people search. This builds one answer from the FAQs and the live screen. You can still publish articles too.

Written by
Rushit Jivani
Senior Software Architect & AI Application Engineer. Architecture that shows up in crash rate, performance and how fast the team ships.