Skip to main content
AlonChat
Chatbot Analytics: 10 Metrics That Actually Matter
Back to Blog
Analytics
September 16, 202610 min read0 views

Chatbot Analytics: 10 Metrics That Actually Matter

Most chatbot dashboards show dozens of metrics. Only ten of them actually tell you whether your chatbot is working. Here is which ones to watch and what to do when the numbers look bad.

Metrics Without Context Are Just Numbers Your chatbot dashboard shows total messages, unique users, average session length, and fifteen other metrics. You glance at it, see the numbers going up, and assume everything is fine. But are those numbers actually telling you that your chatbot is performing well? Total messages going up could mean your chatbot is helping more customers. Or it could mean customers are sending more messages because the chatbot is not understanding them and they keep rephrasing. The difference between a useful metric and a vanity metric is whether it drives a decision. If a metric goes up, do you know what to do? If it goes down, do you know what to fix? If the answer to both is no, the metric is decoration. Here are ten metrics that actually matter, what they tell you, and what to do when they look wrong. The 10 Metrics 1. Resolution Rate What it measures: The percentage of conversations that the chatbot resolves without human intervention. A "resolved" conversation is one where the customer got their answer and did not need to be transferred to a human agent. What good looks like: For a well-trained chatbot with a solid knowledge base, 70-85% is a strong resolution rate. Below 60% suggests significant knowledge gaps. Above 90% is excellent but might indicate that complex questions are not being escalated when they should be. What to do when it is low: Review the conversations that were not resolved. Look for patterns. Are customers asking about topics not covered in your knowledge base? Add that content. Are they asking questions the chatbot misunderstands? Add Q&A pairs for those specific phrasings. The resolution rate is a direct reflection of your knowledge base completeness. 2. Handover Rate What it measures: The percentage of conversations that get transferred from the chatbot to a human agent. This is roughly the inverse of resolution rate, but not exactly, because some conversations end without resolution and without handover (the customer just leaves). What good looks like: 15-30% is typical for a chatbot handling customer support. Below 10% is exceptional. Above 40% means the chatbot is not handling enough on its own and is essentially an expensive routing system. What to do when it is high: Analyze why handovers happen. Is the chatbot escalating because it lacks information? Fill the knowledge gaps. Is it escalating because it is configured too conservatively (low confidence thresholds)? Adjust the threshold. Are customers requesting human agents even when the chatbot could help? Improve the chatbot's first response quality so customers trust it. 3. Average Response Time What it measures: How long it takes the chatbot to respond to a customer message. For AI chatbots, this includes the time to retrieve relevant context, generate a response, and deliver it. What good looks like: Under 3 seconds feels instant. 3-5 seconds feels responsive. 5-10 seconds feels slow but acceptable. Above 10 seconds and customers start wondering if anyone is there. For comparison, the average human agent response time is 1-2 minutes. What to do when it is slow: Response time is usually a platform issue, not something you control directly. But if you notice specific types of questions taking longer (complex queries that require more context retrieval), it might indicate that the knowledge base needs better organization. Cleaner, well-structured content retrieves faster. 4. Customer Satisfaction Score (CSAT) What it measures: Direct customer feedback on the chatbot interaction, typically collected via a post-conversation survey (thumbs up/down, star rating, or short survey). What good looks like: CSAT above 80% is strong for chatbot interactions. 70-80% is acceptable and indicates room for improvement. Below 70% is a red flag. Keep in mind that CSAT for chatbots tends to run 5-10 points lower than for human agents, partly due to inherent bias (some customers rate chatbots lower simply because they prefer humans). What to do when it is low: Read the negative feedback. If customers are saying the chatbot was unhelpful, check whether it had the right information. If they say it was confusing, look at conversation flow. If they say it felt robotic, adjust the tone settings. CSAT is a lagging indicator, meaning it tells you there is a problem but not exactly what the problem is. You need to dig into the conversations behind low scores. 5. Containment Rate What it measures: The percentage of conversations that stay within the chatbot from start to finish, regardless of whether the question was fully answered. This differs from resolution rate because a contained conversation is not necessarily resolved. The customer might have given up. What good looks like: High containment with high satisfaction means the chatbot is handling things well. High containment with low satisfaction means customers are trapped. Always pair containment rate with satisfaction scores. What to do when it is misaligned: If containment is high but satisfaction is low, make the escalation path to human agents clearer. If containment is low but satisfaction is decent, your chatbot might be escalating appropriately. 6. First Response Resolution Rate What it measures: The percentage of customer questions answered correctly in the chatbot's first response, without the customer needing to rephrase or follow up. What good looks like: Above 60% is solid. Approximate this by looking at conversations where the customer sends one message, gets a response, and either says thanks or leaves without further questions. What to do when it is low: Review conversations where the first response missed. Did the chatbot have the right information but present it poorly? Or did it not retrieve the right information at all? The fix depends on which failure mode you see. 7. Conversation Volume by Channel What it measures: How many conversations happen on each connected channel (Messenger, Instagram, website widget, WhatsApp). What good looks like: No universal benchmark. What matters is knowing where your customers are. If 80% of conversations happen on Messenger, that is where to focus optimization. What to do with this data: Prioritize knowledge base improvements for the highest-volume channels. If a channel is growing rapidly, ensure your chatbot performs well there before volume gets too high. 8. Fallback Rate What it measures: The percentage of messages where the chatbot could not find relevant information and fell back to a generic response like "I am not sure about that" or "Let me connect you with our team." What good looks like: Below 15% is excellent. 15-25% is acceptable for a chatbot in its first few months. Above 30% means your knowledge base has significant gaps or your retrieval system is underperforming. What to do when it is high: Export the messages that triggered fallback responses. Group them by topic. You will quickly see which areas of your business are under-documented. These are your highest-priority content additions. Every Q&A pair you add for a common fallback topic directly reduces the fallback rate. 9. Conversation Completion Rate What it measures: The percentage of conversations where the customer reaches a natural conclusion (question answered, order placed, appointment booked) versus abandoning mid-conversation. What good looks like: Above 75% is healthy. Below 60% suggests customers are leaving frustrated. High abandonment at specific points in the conversation often indicates a UX problem at that step. What to do when it is low: Identify where in the conversation customers drop off. If they leave after the welcome message, the message might not be engaging enough. If they leave after the first chatbot response, the response quality might be poor. If they leave mid-flow (for example, during an order process), there might be a confusing step. Fix the specific drop-off point. 10. Knowledge Base Coverage What it measures: The percentage of customer questions that your knowledge base has content for, regardless of whether the chatbot retrieves it correctly. This is a leading indicator. It tells you how complete your content is before customers even ask. What good looks like: You want coverage for at least your top 50 customer questions. If you can identify the 50 most common questions your business receives and your knowledge base has clear answers for at least 45 of them, you are in good shape. What to do to improve it: Review your customer conversation history (from before you had a chatbot, or from escalated conversations). List every unique question topic. Check whether your knowledge base covers it. Fill the gaps systematically, starting with the most frequently asked topics. Building a Metrics Habit The biggest mistake with analytics is checking them once during setup and then ignoring them. Spend 15 minutes every Monday reviewing your key metrics. If resolution rate dropped, investigate. If fallback rate increased, check your knowledge base. The goal is not perfect numbers. It is continuous improvement. A chatbot that resolves 70% today and 80% in three months is delivering real, measurable value. Track the metrics, do the work, and the numbers take care of themselves. Related AlonChat resources Analytics Best AI chatbot in the Philippines AI chatbot training Deployment options
analyticsmetricschatbot-performanceresolution-ratecustomer-satisfaction
AlonChat Team

Written by

AlonChat Team

Ready to Build Your AI Agent?

Start your free trial today and deploy an AI agent in under 10 minutes.

Start Free Trial