Conversational AI Advisor
Designing a conversational AI triage and quality framework for an online technology provider — from quantitative log analysis through to the behavioural specification used as the engineering build brief.
Designing a conversational AI triage and quality framework — from quantitative log analysis through to a behavioural specification the client's engineering team built from.
The deliverable from this project — a behavioural specification covering conversational flow, tone, personalisation, ethical guardrails, and interaction patterns — was adopted as the client’s engineering build brief and taken into implementation by their product and engineering teams. I led research and design end-to-end and was the primary author of both the triage framework and the specification.
The brief asked for better routing. The research found something more interesting: the routing wasn’t the problem.
Client: Online technology provider (NDA) · Domain: Conversational UX / AI Products
01 — The Brief
An online technology provider was absorbing avoidable load on its human support agents. Customers were contacting agents for queries the conversational AI should have handled: password management, product comparisons, standard how-to questions.
The brief had two layers: design a triage model that correctly routes queries between self-service, AI-handled, and human agent channels; and raise the quality of AI-handled interactions so that deflected queries actually resolve.
Framed that way, this is a routing problem — point each query at the right channel and the load resolves itself. That framing didn’t survive the research.
02 — My Role
I led the research and design process end-to-end from establishing the quantitative baseline through to the final specification and wireframes. My work spanned conversation log analysis, interview design, persona development, conversation design theory, and specification writing. I was the primary author of the triage framework and the AI behavioural specification that structured the final build brief.
03 — Building the Evidence Base
Research Baseline
Historic customer conversations were analysed and categorised by CSAT score, establishing a quality baseline across the interaction set. Query distribution revealed a significant proportion of agent-handled conversations falling within clear self-service categories.
Low-performing sessions shared a consistent structure: short exchanges, no meaningful resolution, early disengagement. High-performing sessions read like dialogue — substantive, tailored, with natural repair.
1. Log Analysis
We reviewed hundreds of historic customer conversations. Given the sensitivity of the data, analysis ran on a lightweight Python pipeline using a locally-hosted LLM — keeping all content off cloud infrastructure. To guard against false pattern recognition, the model was required to cite session IDs for each finding, which I cross-checked manually against the source transcripts. This established a quantitative baseline across utterance characteristics, query types, and failure modes, and surfaced which query categories were consistently misrouted to human agents.
2. Live Conversation Transcripts
Alongside the log data, we reviewed live transcripts to understand the real texture of customer language: how users phrase questions, where they disengage, and what signals a conversation heading towards resolution versus abandonment.
3. User Interviews
We designed a structured interview guide grounded in the Technology Acceptance Model (TAM) and the Unified Theory of Acceptance and Use of Technology (UTAUT), targeting two primary user cohorts. The guide explored platform navigation behaviour, expectations of conversational AI, and barriers to adoption.
4. Persona Creation
Research findings informed six detailed personas, each representing a distinct attitude and behaviour pattern towards AI tools. These became the primary lens for all subsequent design decisions.
5. Conversation Design & Specification
Drawing on conversation analysis theory and competitive benchmarking of existing AI products, we developed a comprehensive specification document defining conversational flow, response guidelines, tone of voice, ethical guardrails, and technical interaction patterns.
6. Wireframing
Three iterations of wireframes explored different approaches to invoking the AI, presenting responses, and progressive disclosure of live human support — each refined based on internal feedback and client review.
04 — The Problem Behind the Problem
The brief assumed queries were reaching the wrong channel. The data showed something different: many sessions were failing inside the right channel, because the interaction itself was mismatched from the first message. The real problem wasn’t where queries went — it was what users believed they were talking to when they got there.
What the data told us
The conversation analysis revealed clear structural patterns. Customer queries were short and direct, often technically specific, rarely conversational. Agent responses varied significantly by query type, with complex product topics producing substantially longer replies. The gap between these two registers was itself a signal: many sessions were mismatched from the start, with customers asking brief, bounded questions into an interface configured for open-ended conversation.
Low-performing sessions shared a common profile: exchanges that ended without resolution, customers disengaging before a clear outcome, and repetitive loops that failed to advance the query. High-performing sessions were characterised by contextually tailored responses, active clarification before answering, and a rhythm that felt closer to human dialogue.
What users told us
Interviews surfaced six themes that shaped the design direction. Three concerned the relationship users had with AI as a category: trust had to be earned through transparency and accuracy; AI was broadly accepted for routine queries but human intervention was strongly preferred for anything complex or high-stakes; and attitudes ranged widely enough — from enthusiasm to deep scepticism — that designing for one end of that spectrum would actively alienate the other.
The remaining three were more specific to the product context: users couldn’t access capabilities they didn’t know existed, making discoverability a design problem as much as a support one; tailored, low-effort interaction was the baseline expectation, not a differentiator; and adoption decisions were framed around security and efficiency — users evaluated the product like a professional tool, not a consumer feature.
Together, the data and the interviews converged on the same reframe: years of tree-branch chatbots had trained users to treat conversational AI as a bounded menu, so they selected rather than asked — and underused a system capable of considerably more.
Routing queries more accurately would have optimised a broken interaction. The design problem was the mental model.
05 — Designing the Conversation, Not the Interface
Conversation Design Theory
Rather than designing the AI’s behaviour by instinct, we grounded it in established conversation analysis theory — specifically the Natural Conversation Framework (NCF), which structures interactions around three nested layers:
Three Principles of Conversation Design
Recipient design shaped how we defined the AI’s persona. Everything the AI says — its register, level of detail, tone — should reflect who it’s speaking to. We used historic transcripts and brand guidelines to calibrate this: human enough to feel responsive, precise enough that users were never misled about what they were talking to.
Minimisation governed response length. In natural conversation, short phrases are the norm and elaboration is offered rather than assumed. We defined explicit response length tiers — a concise default for routine queries, an expanded register for complex ones — with clear guidance on which applied when.
Repair was treated as inevitable, not exceptional. The specification anticipated conversation breakdown at every stage: the AI paraphrases unclear queries, offers worked examples, and escalates to a human agent before the user reaches frustration rather than after.
UX Principles Applied
Jakob’s Law shaped the invocation model: users arrive with expectations formed by other chatbot experiences, and fighting those expectations carries a friction cost the design couldn’t afford. Where those expectations were useful, we met them. Where they weren’t, we reframed them deliberately.
The Doherty Threshold governed latency handling. Conversational AI generates responses at variable speeds — without loading states and in-progress indicators, users read silence as failure. We specified these explicitly rather than leaving them to engineering discretion.
Two further principles shaped how information was delivered. Cognitive load theory drove the decision to chunk longer answers progressively, keeping the user in control of depth. And the risk of cognitive dissonance — the discomfort of an interface that behaves differently from how it presents — meant predictability was treated as a hard requirement.
In conversational UI, surprises don’t create delight; they create abandonment.
Competitive Benchmarking
Benchmarking focused on how users approach conversational AI agents — specifically what expectations they bring and where those expectations create friction. The clearest pattern across existing products was the legacy of tree-branch chatbots: years of interfaces built on branching selections had conditioned users to treat conversational AI as a bounded tool. They would select rather than ask, and underuse systems that were capable of considerably more.
The design implication ran counter to a straightforward application of Jakob’s Law. Rather than meeting that established mental model to reduce onboarding friction, we proposed disrupting it deliberately — placing the chat invocation point outside the location users would predict from prior experience. The intent was to interrupt the tree-branch assumption before the first interaction: signalling through placement and context that this was a different kind of tool before the user had typed anything. Meeting users inside their existing mental model would have reproduced the very limitation the product was trying to remove.
06 — The AI Specification
The research and design thinking culminated in a comprehensive specification the client’s engineering and product teams could use to build and evaluate the conversational AI. This included:
- the AI role & purpose
- conversational flow
- tone & language
- personalisation
- compliance & ethics
- guardrails
07 — Outcome
The specification was adopted as the client’s engineering build brief and taken into implementation by their product and engineering teams. That adoption is the result I value most from this project: a research-led behavioural document precise enough for engineers to build from, without flattening the conversation design thinking behind it.
It also validated the reframe. The client commissioned a routing fix; what they implemented was a redesign of what their AI is, how it speaks, and when it hands over to a human — because that’s what the evidence said the problem was.
Specific client details, product names, and visual assets have been omitted under NDA obligations.
More Projects
Ecotravel - Sustainable Travel Mobile App
A collaborative UX project developing a digital solution for sustainable travel, showcasing research synthesis and remote team collaboration across three time zones.
View Case Study → 02Green Grow Precision Planting - Digital Growing Companion
A comprehensive thesis project developing an AI-augmented mobile app with smart sensor integration to help home gardeners achieve consistent yields through data-driven insights.
View Case Study →