Standard product A/B testing metrics — click-through, session length — don't capture whether an AI feature's answer was actually correct, complete, or useful. A prompt change that increases engagement could just as easily be making answers longer and less useful.
What to measure alongside the usual metrics
- Explicit user feedback (thumbs up/down) rate and ratio, tied to the specific variant shown.
- A sample of outputs from each variant, periodically reviewed by a human against a quality rubric.
- Task completion, if applicable — did the user's underlying goal get resolved, not just did they engage with the response.
The trap to avoid
Optimising purely for engagement metrics can reward a model that hedges less, sounds more confident, and is wrong more often — engagement and correctness aren't the same axis, and conflating them is how AI features quietly get worse while metrics improve.
See testing LLM applications, a practical guide.
— Pranjul Rathour, GenAI Engineer from Kanpur, India. Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at any campus: pranjulrathour41@gmail.com.
