Deploying an AI chatbot without defining its KPIs is like running a paid ad campaign without checking conversion data. You know the tool is running β but is it performing? Is it actually deflecting tickets? Are users satisfied with the answers they get? Without measurement, you cannot optimize, and you cannot justify the budget to anyone who asks.
Most chatbot dashboards show session counts and message volume. Those numbers feel like progress. They are not performance. This guide covers the 15 chatbot KPIs that actually matter β with precise definitions, calculation formulas, 2026 benchmarks by industry sourced from Gartner, Forrester, and Salesforce State of Service research, and the measurement mistakes that make these numbers misleading. Use the RAG chatbot evaluation template to score 50 pre-launch answers and refusals, or browse the free AI chatbot tools hub for the complete toolset. Whether you are running a RAG-based AI agent or evaluating whether your current chatbot is delivering value, this is the framework.
Chatbot KPI definition: A chatbot KPI is a quantified indicator that measures how well a conversational assistant delivers a business outcome β resolving a request without a human, satisfying the user, or lowering the cost of an interaction. Unlike raw usage statistics such as session counts, a KPI is attached to a formula and a target, so it can be tracked over time and acted on. The core set combines deflection, quality, and cost metrics.
The 15 KPIs covered in this guide, with what each one measures and how to calculate it:
| KPI | What it measures | Formula |
|---|---|---|
| 1. Containment rate | Share of conversations the bot resolves with no human involved. | (Bot-resolved conversations / Total conversations) × 100 |
| 2. Deflection rate | Reduction in tickets that would otherwise have reached support. | (Tickets avoided / Expected tickets without bot) × 100 |
| 3. CSAT | User satisfaction with the answer received. | (Positive ratings / Total CSAT responses) × 100 |
| 4. First response time | How fast the bot replies to the first message. | Time from first user message to first bot reply |
| 5. Escalation rate | Share of conversations handed over to a human agent. | (Escalated conversations / Total conversations) × 100 |
| 6. Abandonment rate | Share of conversations dropped before any resolution. | (Abandoned conversations / Total initiated) × 100 |
| 7. Intent coverage rate | Share of real question types with an adequate answer. | (Question types covered / Top question types reviewed) × 100 |
| 8. Cost per resolved conversation | What one autonomous resolution actually costs. | (Platform cost + maintenance) / Bot-resolved conversations |
| 9. Engagement rate | Share of visitors who start a conversation. | (Conversations initiated / Unique visitors) × 100 |
| 10. Lead conversion rate | Share of conversations producing a commercial action. | (Leads captured / Total conversations) × 100 |
| 11. Return user rate | Share of users who come back to the assistant. | (Users with 2+ sessions / Total users) × 100 |
| 12. NPS (chatbot-specific) | Willingness to recommend the assistant. | % Promoters (9β10) minus % Detractors (0β6) |
| 13. First contact resolution | Issues solved in a single session, no second channel. | (Issues resolved in 1 session / Total issues) × 100 |
| 14. Average handling time | How long a resolved conversation takes end to end. | Total handling time / Resolved conversations |
| 15. Ticket reopen rate | Bot resolutions that did not hold. | (Reopened after bot resolution / Bot resolutions) × 100 |
TL;DR
- Start with three metrics: containment rate, CSAT, and cost per resolved conversation β they tell you whether your chatbot is working, whether users like it, and whether it pays for itself.
- Industry benchmark for containment rate on a well-configured RAG chatbot: 40β65% (Gartner, 2025).
- The four KPI categories are: engagement, deflection, quality, and business impact.
- Most chatbot deployments measure session volume and ignore deflection rate β that is the single most common measurement mistake.
- Heeya surfaces containment rate, CSAT, escalation rate, and conversation history natively in the dashboard.
Table of Contents
- Why Most Chatbot Dashboards Lie
- The 4 KPI Categories
- The 15 Metrics: Definitions, Formulas, and Benchmarks
- User Experience KPIs for a Chatbot
- 2026 Benchmarks by Industry
- Chatbot KPIs by Industry
- How to Instrument These Metrics
- Dashboard Framework
- Common Measurement Mistakes
- How Heeya Surfaces These Metrics
- Further Reading
- FAQ
Why Most Chatbot Dashboards Lie
The default analytics in most chatbot platforms are built around engagement, not outcomes. You see total conversations, average session length, and message counts. These numbers grow as traffic grows β they do not tell you whether the chatbot is solving problems or frustrating users into abandoning the conversation.
A chatbot that starts 10,000 conversations and resolves 2,000 autonomously is performing worse than one that starts 3,000 conversations and resolves 2,400. Session volume as a headline metric hides this. The Salesforce State of Service report (2025) found that only 34% of customer service teams track deflection rate as a primary chatbot KPI β the metric that most directly measures the tool's impact on support volume.
The fix is not more data. It is measuring the right outcomes across four categories.
The 4 KPI Categories
Every meaningful chatbot metric falls into one of four categories:
- Engagement β are users finding and using the chatbot?
- Deflection β is it reducing the volume of tickets, calls, and human agent interactions?
- Quality β are the answers accurate and satisfying?
- Business impact β is it generating leads, saving costs, or improving conversion?
A healthy chatbot program tracks at least two metrics per category. A mature one tracks all 15 on a cadence that matches their rate of change: daily for operational metrics, monthly for financial ones, quarterly for NPS and return rate. For enterprise teams building the business case for AI investment, our guide on generative AI enterprise ROI and use cases in 2026 provides the financial frameworks and benchmarks to quantify impact at scale.
The 15 Metrics: Definitions, Formulas, and Benchmarks
| # Metric | Category | Formula | 2026 Benchmark |
|---|---|---|---|
| 1. Containment Rate | Deflection | (Conversations resolved by bot / Total conversations) × 100 | 40β65% (RAG), 20β35% (rule-based) |
| 2. Deflection Rate | Deflection | (Tickets avoided / Expected tickets without bot) × 100 | 35β55% (e-commerce), 25β45% (B2B SaaS) |
| 3. CSAT | Quality | (Positive ratings / Total CSAT responses) × 100 | 70β82% (RAG bot), 55β65% (rule-based) |
| 4. First Response Time | Engagement | Time from user's first message to bot's first reply | <2 seconds (modern SaaS chatbot) |
| 5. Escalation Rate | Deflection | (Escalated conversations / Total conversations) × 100 | 15β30% (varies by domain complexity) |
| 6. Abandonment Rate | Engagement | (Abandoned conversations / Total initiated) × 100 | 20β40% (alert if >50% in first 3 exchanges) |
| 7. Intent Coverage Rate | Quality | % of top question types with an adequate bot answer | Target: cover 80%+ of top-20 FAQ |
| 8. Cost per Resolved Conversation | Business Impact | (Monthly platform cost + maintenance) / Bot-resolved conversations | $0.10β$0.50 (vs. $8β$15 for human agent) |
| 9. Engagement Rate | Engagement | (Conversations initiated / Unique page visitors) × 100 | 2β8% depending on page type and widget placement |
| 10. Lead Conversion Rate | Business Impact | (Leads captured via chatbot / Total conversations) × 100 | 3β12% (B2B), 1β5% (e-commerce support) |
| 11. Return User Rate | Engagement | % of users with 2+ sessions in a defined period | 15β35% (internal/HR bots), lower for e-commerce |
| 12. NPS (chatbot-specific) | Quality | % Promoters (9β10) minus % Detractors (0β6) | +10 to +35 for RAG-based AI agents |
| 13. FCR (First Contact Resolution) | Quality | (Issues resolved in 1 session / Total issues reported) × 100 | 50β70% for well-tuned RAG chatbots |
| 14. AHT (Avg. Handling Time) | Business Impact | Avg. time from conversation start to resolution | <3 min for bot-resolved; baseline vs. human AHT |
| 15. Ticket Reopen Rate | Quality | (Tickets reopened after bot resolution / Total bot resolutions) × 100 | <8% (alert if >15%) |
Benchmarks sourced from Gartner Customer Service Technology Survey 2025, Forrester The Total Economic Impact of AI Customer Service 2025, and Salesforce State of Service 2025. RAG = Retrieval-Augmented Generation.
Here is what each metric actually measures and when to act on it:
1. Containment Rate
The master metric. Containment rate measures the percentage of conversations the chatbot resolves entirely without human intervention. If this number is low, every downstream financial KPI suffers. A containment rate below 30% after 30 days signals an incomplete knowledge base or poor intent coverage β not a technology failure. According to Gartner's 2025 Customer Service Technology Survey, organizations with mature RAG deployments average 55β65% containment rate; those running rule-based bots average 20β35%.
When to act: below 30% at 30 days β audit your knowledge base coverage. Above 60% β focus on maintaining quality as you scale.
2. Deflection Rate
Deflection rate measures impact at the ticket volume level, not the conversation level. A user who found their answer via the chatbot and never opened a ticket counts toward deflection but not containment. Forrester's 2025 Total Economic Impact study found that AI chatbot deployments with strong RAG architectures achieved 40β55% ticket deflection within 90 days of launch in e-commerce and SaaS contexts. For e-commerce teams specifically, see our guide on how to reduce e-commerce support tickets with an AI chatbot for concrete implementation patterns.
Measure by comparing ticket volume in equivalent periods before and after deployment, or by running A/B tests on pages with and without the chatbot widget. This is distinct from containment: deflection rate is always higher because it captures self-service that happens before a ticket is opened.
3. CSAT
Collect CSAT at conversation end using a 1β5 scale or thumbs up/down. Expect a 15β30% response rate β that is normal. The important number is the percentage of responses that are positive. Segment by topic to identify which subject areas are underperforming. A CSAT below 60% on bot-resolved conversations is a signal to review those conversations directly, not to adjust the model.
Important: measure CSAT separately for bot-resolved and human-escalated conversations. Mixing them masks the chatbot's actual quality signal.
4. First Response Time (FRT)
For AI chatbots, FRT should be under 2 seconds. This is an infrastructure metric, not a content metric β it does not fluctuate based on knowledge base quality. Latency spikes above 5 seconds measurably degrade CSAT. The broader impact of customer service response time on conversion rates is significant β slow first responses correlate directly with lost revenue, not just lower satisfaction scores. Monitor for degradation during high-context loads (large knowledge bases, long conversation histories). If FRT degrades, contact your platform provider β you cannot fix this in the UI.
5. Escalation Rate
The percentage of conversations transferred to a human agent. A healthy range is 15β30% depending on domain complexity. Two warning patterns: an escalation rate above 40% means the bot is not covering the real question set. An escalation rate below 5% with low CSAT means users are abandoning instead of escalating β a worse outcome than escalation.
6. Abandonment Rate
Not all abandonment is bad β a user who found their answer and closed the window is a success. The signal to investigate is high abandonment in the first 3 exchanges. That pattern indicates friction in the opening of the conversation: a poorly phrased greeting, a misunderstood first question, or an overly long initial response. Analyze drop-off points in the conversation flow, not the aggregate abandonment number.
7. Intent Coverage Rate
The percentage of question types in your real conversation logs that have an adequate answer in the knowledge base. You cannot automate this metric β it requires a human to review a sample of 20β30 conversations monthly and flag gaps. Target: cover 80% of your top-20 FAQ topics. Anything below that is a knowledge base problem, not a model problem.
8. Cost per Resolved Conversation
The most persuasive ROI metric for internal reporting. Formula: (monthly platform subscription + maintenance hours at loaded cost) divided by the number of conversations the bot resolved autonomously that month. A realistic example: $50/month platform + 1 hour maintenance at $60/hr = $110/month. If the bot resolves 500 conversations: $0.22 per conversation. The Forrester benchmark for human agent cost per interaction in 2025 is $8.01 for chat and $12.31 for phone β making the comparison straightforward. See our AI chatbot ROI calculator guide for a full cost model.
9. Engagement Rate
The percentage of page visitors who initiate a conversation. Typical range: 2β8%. A rate below 2% usually points to widget placement (buried in a corner, below the fold) or a call-to-action label that does not match visitor intent. A rate above 10% on high-traffic pages is excellent β focus on containment rate, not engagement rate, as the primary optimization target.
10. Lead Conversion Rate
The percentage of conversations that produce a measurable commercial action β a form submission, email captured, demo booked. B2B service firms typically see 3β12% when the chatbot is configured with an explicit lead capture step. E-commerce support bots typically see 1β5%. Track this separately from customer support conversations to avoid mixing intent signals.
11. Return User Rate
An indirect proxy for answer quality: users return when they found value. Benchmark: 15β35% for internal bots (HR, IT helpdesk) where the same users interact repeatedly. Lower for external customer support bots where many interactions are one-time (order questions, account issues). A low return rate on an internal bot is a red flag worth investigating.
12. NPS (Chatbot-Specific)
Ask: "On a scale of 0 to 10, how likely are you to recommend our AI assistant?" Formula: % Promoters (9β10) minus % Detractors (0β6). Benchmark for RAG-based agents: NPS of +10 to +35. Collect quarterly rather than monthly β you need sufficient sample size to make it meaningful. NPS is most useful for board-level reporting; CSAT is more actionable for operational improvement.
13. FCR (First Contact Resolution)
Measures whether the user's issue was resolved in a single session without them needing to return through another channel. FCR is the metric that bridges chatbot quality with overall support effectiveness. Salesforce State of Service (2025) found that teams with high chatbot FCR rates (60%+) reported 23% lower overall support costs compared to teams where chatbots primarily triaged rather than resolved. High FCR also directly supports efforts to reduce customer churn with AI β unresolved first contacts are one of the leading drivers of churn in subscription and service businesses. Track FCR by conversation category to find where the bot is triaging versus resolving.
14. AHT (Average Handling Time)
For bot-resolved conversations, AHT should be under 3 minutes. The more interesting measurement is the delta: how does bot AHT compare to human agent AHT for the same question categories? When the bot handles a question type in 90 seconds that takes a human agent 8 minutes, that is the number to put in your ROI report. Track AHT by question category, not in aggregate.
15. Ticket Reopen Rate
The percentage of bot-resolved conversations where the user returned with the same issue through another channel. A reopen rate above 8% indicates the bot is marking conversations as resolved when the user's actual problem was not solved β a knowledge base accuracy issue, not a volume issue. This metric catches the difference between "conversation ended" and "user satisfied."
User Experience KPIs for a Chatbot
Financial KPIs tell you whether the chatbot pays for itself. User experience KPIs tell you whether people can actually get what they came for β and they move first. A degrading experience shows up in fallback and abandonment weeks before it shows up in cost per resolved conversation. Seven metrics cover the experience layer:
- CSAT after chat. Asked at the end of the conversation, on bot-resolved sessions only. It is the only direct statement of user opinion you get; everything else is inference.
- Containment / self-service rate. The share of users who complete their task without a human. From the user's point of view this is not a cost metric β it is the difference between getting an answer now and waiting for a reply.
- Fallback rate. The share of user messages that trigger a "I don't have that information" response or a generic non-answer. It is the earliest signal of a knowledge base gap, and it is diagnostic rather than a headline KPI: track it per topic, not in aggregate.
- Average handling time. How long a resolved conversation takes end to end. A rising AHT with stable containment usually means the bot is getting there, but through too many turns.
- Abandonment rate. Users who leave mid-conversation. Look at where they leave: drop-off inside the first three exchanges is a friction problem, drop-off after a correct answer is often a success.
- Escalation rate. A UX metric as much as a deflection one. Users need a visible, working route to a human; an escalation path that is hard to find pushes people into abandonment instead.
- Sentiment. Tag conversations where the user expresses frustration β repeated rephrasing, "that's not what I asked," all-caps, requests for a human in the first turn. Sentiment catches the failures that a 15β30% CSAT response rate never samples.
Read them as a set, not individually. High containment with a high fallback rate means the bot is closing conversations it did not answer. Low escalation with high abandonment means users gave up rather than asked for help. Low AHT with low CSAT means the bot is fast at being unhelpful. The pairings are where the diagnosis lives β and all seven are visible in conversation logs, which is why a monthly read of 20β30 real transcripts remains the highest-value UX review you can run.
2026 Benchmarks by Industry
| Industry | Containment Rate | Deflection Rate | CSAT | Escalation Rate | Lead CVR |
|---|---|---|---|---|---|
| E-commerce / Retail | 55β70% | 40β60% | 72β80% | 15β25% | 1β4% |
| B2B SaaS | 40β60% | 30β50% | 70β82% | 20β35% | 5β12% |
| Professional Services (legal, finance, accounting) | 30β50% | 25β45% | 65β78% | 25β40% | 8β15% |
| Healthcare | 30β45% | 20β40% | 68β76% | 30β45% | 2β6% |
| HR / Internal IT Helpdesk | 50β70% | 40β65% | 74β84% | 15β25% | N/A |
| Real Estate | 35β55% | 30β50% | 68β78% | 25β40% | 10β18% |
| Education / eLearning | 45β65% | 35β55% | 70β80% | 20β30% | 4β10% |
Benchmarks reflect RAG-based AI chatbot deployments. Rule-based chatbots typically perform 15β25 percentage points lower on containment and deflection. Sources: Gartner Customer Service Technology Survey 2025, Forrester TEI of AI Customer Service 2025, Salesforce State of Service 2025.
Professional services and healthcare industries show lower containment rates not because the technology is less effective, but because the questions are genuinely more complex and regulatory caution appropriately routes more conversations to human agents. A 35% containment rate in a legal services context can represent excellent performance. For a broader view of where these figures stand relative to the market, the 2026 customer support automation benchmark provides cross-industry comparisons across automation rate, resolution time, and cost metrics.
Chatbot KPIs by Industry
The 15 KPIs apply everywhere, but the two or three you report to leadership should reflect what the chatbot is actually there to do. Three common contexts:
Finance and banking
Regulated environments add a compliance dimension to quality measurement, and the most useful KPIs are the ones that expose it:
- Authentication drop-off. The share of users who abandon at the identity-verification step before reaching an answer. This is often the single largest source of unresolved conversations in banking β and it is a flow problem, not a knowledge base problem.
- Compliant-answer rate. The share of sampled answers that stay inside approved wording, include the required disclaimers, and decline to give personalized advice when they should. Scored manually on a monthly sample; there is no automated substitute.
- Ticket deflection on billing and account questions. Fees, statements, card status, and payment dates are high-volume and highly repetitive β measure deflection on that subset rather than on all conversations, where complex advisory questions drag the average down.
Expect lower containment than in retail: the benchmark table above puts professional services, including finance, at 30β50%. A refusal to answer a question that requires a licensed advisor is correct behavior and should not be counted as a failure.
E-commerce
- Order-status containment. "Where is my order?" is usually the highest-volume intent and the most automatable one. Track containment on it separately β it is where the deflection budget is won or lost.
- Assisted conversion rate. The share of conversations followed by a purchase in the same session. Tag chatbot sessions in your analytics so pre-sale conversations are not credited to the last-click channel.
- Return and refund policy deflection. Policy questions are answerable from a single well-maintained document; persistent escalations here point to an outdated knowledge base.
SaaS support
- Deflection against baseline. Ticket volume per active account before and after deployment, not raw ticket counts β account growth otherwise hides the gain.
- Time to first response. The clearest KPI for justifying the chatbot to a support team measured on SLAs: the bot answers in seconds, at any hour, in every time zone.
- FCR and ticket reopen rate on how-to questions. Documentation-answerable questions should resolve in one session and stay resolved. A rising reopen rate on this category means the docs, not the model, need work.
How to Instrument These Metrics
You do not need a custom analytics stack to track these KPIs. Here is how to instrument each category:
Deflection and containment
Your chatbot platform should expose conversation status at the end of each session: resolved by bot, escalated to human, or abandoned. If it does not, that is a platform capability gap worth addressing. For deflection rate, pull ticket volume from your helpdesk (Zendesk, Freshdesk, Linear, or email) for equivalent periods before and after deployment. The difference is your baseline deflection estimate.
CSAT and NPS
Trigger CSAT collection on conversation end for resolved sessions only. Use a 5-star rating or a single binary (thumbs up/down) β longer surveys have lower completion rates without meaningfully better data. For NPS, trigger a separate survey via email on a quarterly sample of users who had at least one bot interaction in the period.
Cost per resolved conversation
Calculate monthly: (platform subscription + (maintenance hours × loaded hourly rate)) / bot-resolved conversations. Pull bot-resolved conversation count from your platform's dashboard. Compare to your human agent cost per interaction, which you can estimate from (total support headcount cost / total human-handled conversations per month).
Lead conversion
Tag chatbot-originated leads in your CRM using a source field. If your chatbot has a built-in form tool, leads are automatically labeled. If not, use UTM parameters or a dedicated form endpoint to separate chatbot-sourced submissions from other channels.
Intent coverage
This one requires a human. Export a random sample of 20β30 conversations monthly. Read them and tag each bot response as adequate, partial, or no-answer. The percentage of adequate responses across your top question categories is your intent coverage rate. Schedule a 30-minute monthly review β it is the most actionable quality signal you have.
Dashboard Framework
Not all KPIs need daily attention. Match monitoring cadence to the rate of change and business impact of each metric:
| Cadence | KPIs to Track | Owner |
|---|---|---|
| Weekly | Containment rate, escalation rate, abandonment rate, FRT | Support lead / Ops |
| Monthly | CSAT, FCR, cost per resolved conversation, lead conversion rate, intent coverage review | Project manager / Customer success |
| Quarterly | NPS, return user rate, deflection rate (vs. baseline), full ROI analysis | Director / Executive sponsor |
Start lean: track containment rate, CSAT, and cost per resolved conversation in week one. Add the remaining metrics as your measurement infrastructure matures. A dashboard with three well-tracked metrics beats one with fifteen poorly instrumented ones.
Common Measurement Mistakes
Tracking session volume as a success metric
Session volume is a reach metric, not an outcome metric. A chatbot that starts more conversations but resolves fewer is regressing, not growing. Always pair session volume with containment rate.
Mixing bot and human CSAT scores
Human agent conversations typically score higher on CSAT than bot-resolved ones for complex issues. If you average them together, you get a number that neither accurately reflects bot performance nor human agent quality. Segment always.
Treating low escalation rate as success
An escalation rate below 5% combined with high abandonment is a failure mode, not a win. It means users are giving up rather than asking for human help. Watch the abandonment and escalation metrics together.
Measuring deflection without a baseline
You cannot calculate deflection rate without pre-deployment ticket volume data. If you are launching a new chatbot, pull 60β90 days of historical ticket data before launch so you have a valid comparison point.
Ignoring the ticket reopen rate
A chatbot can achieve a high containment rate by marking conversations as resolved prematurely. Ticket reopen rate is the check on containment rate quality. If containment is high but reopen rate is above 15%, your bot is closing conversations, not solving problems.
How Heeya Surfaces These Metrics
Heeya's analytics dashboard natively surfaces the metrics that matter most for operational monitoring: containment rate, escalation rate, conversation history with full message logs, and CSAT collection built into the widget. You do not need a third-party BI tool or custom integration to track your core KPIs. For SMBs looking to connect these KPIs to a broader transformation roadmap, our guide on transforming SMB customer support with AI ties measurement to operational change management.
The conversation history view lets you review individual sessions β which is how you run the monthly intent coverage audit without exporting to a spreadsheet. Filter by unresolved or escalated conversations to focus your review on the gaps that matter.
For teams that want to go further on ROI measurement, the Heeya chatbot platform integrates with CRM tools via form submissions (for lead conversion tracking) and exposes conversation metadata via API for teams building custom dashboards. The RAG for customer service guide covers how the retrieval architecture directly impacts containment rate and FCR β the two quality metrics most affected by knowledge base structure.
On Heeya's Standard and Premium plans, you get the built-in contact form tool for lead capture, enabling you to track lead conversion rate without any additional configuration.
Further Reading
- AI Chatbot ROI Calculator 2026 β turn your KPIs into a financial business case
- Best AI Chatbot Platforms 2026 β compare platforms on analytics capability and pricing
- How Much Does an AI Chatbot Cost in 2026? β full cost breakdown for budgeting
- RAG for Customer Service 2026 β how retrieval architecture affects your KPIs
- AI Chatbot RFP Template 2026 β turn these KPIs into vendor requirements and a scoring matrix
- Heeya AI Chatbot Platform β built-in analytics, RAG-native, GDPR-compliant
FAQ
What is a chatbot KPI?
A chatbot KPI is a quantified indicator that measures how well a conversational assistant delivers a business outcome β resolving a request without a human, satisfying the user, or lowering the cost of an interaction. Unlike raw usage statistics such as session counts or message volume, a KPI has a formula and a target, so it can be tracked over time and acted on. The three KPIs most teams start with are containment rate, CSAT, and cost per resolved conversation.
What is the difference between containment rate and deflection rate?
Containment rate measures the percentage of conversations the chatbot resolves without any human intervention. Deflection rate measures the impact on incoming ticket volume β it includes users who found their answer via the chatbot and never opened a ticket. Deflection rate is always higher than containment rate. Both matter: containment tells you how capable your bot is; deflection tells you how much it reduces support workload.
How long does it take to reach a good containment rate?
With a RAG-based chatbot and a well-structured knowledge base, containment rate typically stabilizes in 2β4 weeks. Week one reveals knowledge gaps; week two you fill them. Rule-based chatbots take 2β3 months to reach comparable performance because gaps must be addressed one intent at a time rather than through document updates.
Should I collect CSAT on every conversation?
No. Collecting CSAT on every conversation increases friction and reduces completion rates without improving data quality. A better approach is to collect CSAT on a representative sample β every 3rd or 5th resolved conversation β or exclusively on bot-resolved conversations to isolate chatbot quality from human agent performance.
Is a 70% deflection rate achievable?
Yes, in high-repetition domains. E-commerce businesses handling order tracking, returns, and shipping questions regularly achieve 60β70% deflection with a well-configured RAG chatbot. In complex domains β legal, medical, or highly customized B2B β 30β45% deflection is strong performance. See our best AI chatbot platforms guide for deployment patterns by industry.
Which chatbot KPIs should I start tracking first?
Start with three: containment rate, CSAT, and cost per resolved conversation. Containment rate tells you whether your bot is working. CSAT tells you whether users find its answers useful. Cost per resolved conversation tells you whether it pays for itself. Once those three are stable and tracked consistently, add deflection rate, escalation rate, and lead conversion rate.
What are the top KPIs to measure chatbot success in finance?
Track containment rate on informational queries, authentication drop-off (users who abandon at the identity-verification step before reaching an answer), compliant-answer rate (the share of sampled answers that respect approved wording and required disclaimers), ticket deflection on billing and account questions, escalation rate, and CSAT measured on bot-resolved conversations only. Containment in financial services runs lower than in retail β 30β50% is a realistic range β because advisory questions should be routed to a licensed human. A correct refusal is not a failure and should not be counted as one.
What user experience KPIs should a chatbot have?
Seven metrics cover the experience layer: CSAT collected at the end of the chat, containment or self-service rate, fallback rate (messages that get a non-answer), average handling time, abandonment rate, escalation rate, and sentiment on frustration signals. Read them in pairs rather than individually: high containment with a high fallback rate means the bot is closing conversations it did not answer, and low escalation with high abandonment means users gave up instead of asking for a human.
Track the metrics that matter β without building a custom dashboard
Heeya surfaces containment rate, CSAT, escalation rate, and full conversation history natively. GDPR-native, RAG-powered, flat monthly pricing. No credit card required to start.