Quick answer
Define the purpose and sample before opening a transcript. Score only observable evidence, allow “not scorable,” separate critical failures from weighted coaching items, and preserve the source, reviewer and decision trail. Calibrate reviewers on the same cases. Analyze QA, CSAT and operational speed as different signals, then assign corrective work and verify whether it improved the customer-visible outcome.
A scoring method that reviewers can defend
Define before scoring
Version each criterion, evidence rule, weight, critical-failure rule, scorable population and review period. Do not tune the rubric after seeing which team or agent scores poorly.
Use four outcomes
Use pass, fail, partial and not scorable, with criterion-specific anchors. A missing transcript segment is not the same as observed failure. Preserve reviewer confidence and evidence links.
Separate gates from coaching
Security, privacy, unauthorized financial action and fabricated information may be mandatory failures even when the weighted score is high. Coaching items can identify improvement without falsely declaring an unsafe interaction acceptable.
Minimum review record
review_id · purpose · population · sampling_method · conversation_id · channel · locale · human_or_ai · rubric_version · criterion_id · outcome · evidence_pointer · reviewer · confidence · critical_failure · disagreement · calibration_set · appeal · corrective_action · owner · due_date · effectiveness_check
The 18-point customer support QA checklist
Each control below states what to inspect, what evidence can support a score, and when not to pass. Adapt the anchors to the service; do not convert every sentence into an equally weighted checkbox.
Scope and sampling are valid
Confirm channel, queue, language, product, customer tier, human/AI participation, review window, eligibility rule and sampling method before scoring. A convenient set of escalations is not a team-quality sample.
- Score from
- Reproducible population and ticket IDs; exclusions; random or stratified selection; reviewer assignment; comparison dimensions.
- Do not pass when
- Fail when the sample is cherry-picked, duplicates are hidden, or results are generalized beyond the defined population.
Monitoring is transparent and data is minimized
Use a documented quality purpose, tell workers and customers where required, and expose reviewers or models only to fields necessary for that purpose. Separate credentials, payment data, health data, secrets and restricted investigations.
- Score from
- Notice and purpose; lawful or contractual basis where applicable; field inventory; access roles; retention/deletion; redaction test.
- Do not pass when
- Fail when monitoring is covert, excessive, repurposed without review, or sends unnecessary sensitive narrative to an AI system.
The record preserves full interaction context
Review the customer message, prior thread, channel, timestamps, attachments, automation events, ownership changes and final state. Do not grade one isolated agent sentence as if it were the whole interaction.
- Score from
- Immutable transcript or recording reference; event timeline; actor type; edits; missing segments; channel constraints; final state.
- Do not pass when
- Fail when context is missing but the reviewer still infers intent, fault, response time or resolution. Mark not scorable instead.
The customer need and desired outcome are identified
State the customer’s task, affected object, expected result, observed result, urgency and requested outcome in evidence-linked language. Separate the stated request from the underlying issue when they differ.
- Score from
- Quoted or timestamped evidence; normalized issue; affected account/order/workspace; desired outcome; uncertainty; competing interpretations.
- Do not pass when
- Fail when the reviewer or agent invents a goal, collapses several issues into one, or treats sentiment as the problem definition.
Facts and instructions are accurate and supported
Check product behavior, account state, policy version, knowledge version and tool result that existed at response time. Distinguish verified facts, reasonable hypotheses and unknowns.
- Score from
- Source link/version; account or tool evidence; retrieval time; quoted policy; calculation inputs; confidence; correction trail.
- Do not pass when
- Fail for fabricated features, stale steps, unsupported causal claims, invented tool results, or confident answers where evidence is unavailable.
Clarifying questions precede consequential assumptions
When identity, product version, jurisdiction, account state, dates, scope or desired remedy changes the answer, ask the smallest necessary question before acting. Avoid interrogating customers for data already available.
- Score from
- Decision-changing ambiguity; question asked; available context checked; answer received; safe temporary path; no-response handling.
- Do not pass when
- Fail when the agent acts on a material assumption, or creates friction with irrelevant questions that do not alter routing or resolution.
Policy, eligibility and exceptions are applied correctly
Use the policy and product terms applicable to the customer, event date, region, plan and request. Explain which condition matters and route exceptions rather than silently denying or granting them.
- Score from
- Policy ID/version/effective date; eligibility inputs; exception path; decision owner; customer-facing explanation; conflict record.
- Do not pass when
- Fail when a general rule is applied to the wrong date or region, an exception is invented, or policy language is presented as legal advice.
Identity, permissions and sensitive actions are controlled
Verify identity only to the level required, use approved channels, respect role permissions and require additional authorization for refunds, deletion, credential changes or irreversible actions. Never request secrets in free text.
- Score from
- Verification method/result; actor and role; permission check; action preview; approval; idempotency key; audit event; rollback availability.
- Do not pass when
- Fail for over-verification, secret collection, privilege escalation, action without approval, duplicate execution, or exposing one customer’s data to another.
Severity, priority and routing match evidence
Classify impact, urgency, affected users, workaround availability, security/privacy signals, service state and contractual commitments. Preserve uncertainty and follow the current routing matrix.
- Score from
- Severity inputs; priority rule/version; affected scope; incident link; specialist queue; SLA target; override and approver.
- Do not pass when
- Fail when an angry tone alone raises severity, a calm report hides a serious incident, or AI confidence substitutes for required specialist review.
The response leads with a direct, usable next step
Open with the answer, current state or next required action. Put prerequisites before commands, order steps, state who performs each action and distinguish required from optional work.
- Score from
- First-response answer; ordered steps; prerequisites; owner; expected observable result; stop condition; alternative path.
- Do not pass when
- Fail when the customer must search a long preamble for the action, instructions arrive out of order, or success cannot be observed.
Tone is respectful, specific and non-performative
Acknowledge the concrete inconvenience or risk without inventing feelings, blame, certainty or intimacy. Match urgency and channel while remaining professional and inclusive.
- Score from
- Customer impact acknowledged; neutral ownership language; apology tied to a known failure; no blame; no manipulative reassurance.
- Do not pass when
- Fail for canned empathy that delays help, promises that cannot be kept, arguing with the customer, or language that stereotypes or shames.
Language is clear, accessible and correctly localized
Use short sentences, descriptive links, meaningful headings, text alternatives and terminology the audience knows. Translate the task and policy meaning—not just words—and preserve locale-specific dates, currencies and support routes.
- Score from
- Reading level appropriate to audience; glossary; locale reviewer; link purpose; attachment alternative; date/time zone; accessible channel.
- Do not pass when
- Fail when jargon obscures action, machine translation changes eligibility or safety meaning, or required information exists only in an inaccessible image or audio file.
Ownership and handoffs remain continuous
Name the current owner, destination team, reason, evidence package, expected response and fallback. Do not ask the customer to repeat information already transferred with authorization.
- Score from
- From/to owner; transfer time; reason; evidence and permissions; receiving acknowledgment; customer notice; fallback and escalation timer.
- Do not pass when
- Fail for blind transfers, orphaned tickets, circular routing, conflicting promises, or handoffs that expose data the receiving role does not need.
Time and status expectations are truthful
Give a measurable next-update time or event, specify the time zone and distinguish an internal target from a contractual commitment. Update before a deadline is missed and explain changed assumptions.
- Score from
- SLA or target source; due time/time zone; dependency; current state; next update; breach warning; revised commitment and owner.
- Do not pass when
- Fail for “soon” without a checkpoint, guaranteed resolution dates without authority, silent deadline misses, or stopping the SLA clock incorrectly.
Resolution is complete and verified against the original need
Confirm that the requested outcome or an agreed alternative was achieved, not merely that an agent replied or a tool returned success. Test the customer-visible state and cover every issue unit.
- Score from
- Action result; before/after state; customer-visible verification; unresolved issue list; acceptance evidence; reopen path.
- Do not pass when
- Fail when a backend success message is treated as customer success, one of several issues is ignored, or closure occurs before required propagation or confirmation.
Workarounds and escalations are safe and bounded
State what a workaround changes, who may use it, duration, side effects, monitoring and reversal. Escalate security, privacy, legal, financial, health, safety or irreversible cases to the approved specialist route.
- Score from
- Risk classification; allowed audience; expiry; side effects; monitoring; rollback; specialist acceptance; emergency stop.
- Do not pass when
- Fail when a workaround bypasses control, becomes permanent without review, hides an incident, or invites unsafe actions outside the agent’s authority.
AI participation, uncertainty and human authority are visible
Identify which replies, summaries, scores, tags or actions came from automation. Preserve model/version, inputs, evidence links, confidence or abstention, reviewer decision and override reason for consequential uses.
- Score from
- Actor type; model/prompt/workflow version where appropriate; retrieved sources; proposed score; human decision; override; incident and rollback.
- Do not pass when
- Fail when AI output is presented as verified evidence, the same model grades itself without independent review, or automation can take consequential action outside approved boundaries.
Closure, feedback and learning form a controlled loop
Summarize what changed, what remains, how to reopen or escalate, and when feedback may be requested. Send CSAT only to eligible interactions and analyze response rate and selection bias separately from QA. Convert recurring failures into owned coaching, content, product or control work.
- Score from
- Closure summary; unresolved items; reopen route; survey eligibility/time; response denominator; coaching/action owner; due date; effectiveness check.
- Do not pass when
- Fail when closure hides unresolved work, survey scores are treated as the entire truth, feedback is used punitively without context, or corrective actions have no owner and verification.
Calibration case: two reviewers disagree for valid reasons
This hypothetical example explains the method; it is not an OpenMax result. A customer asks for a duplicate subscription charge to be reversed. The agent acknowledges the problem, verifies the account through an approved method, finds two captured charges, submits one refund and says it should “arrive tomorrow.” The payment rail actually gives a multi-day estimate, and no follow-up owner is recorded.
- Reviewer A: mostly resolved. The requested financial action was authorized and completed once; identity, evidence and idempotency pass. Tone and directness also pass.
- Reviewer B: expectation failure. “Tomorrow” is unsupported, the rail estimate is omitted, and no owner or checkpoint exists. Time/status and closure fail; operational record is partial.
- Calibration decision. Both observations are retained. The interaction is not converted to one vague average. The scorecard records the successful refund, the false timing commitment, the missing follow-up control and the customer-visible verification still required.
- Corrective action. Update the response template to insert the rail-specific estimate, require a next-check date and owner, and sample later cases to verify the new control—not merely that training was assigned.
Operate the QA program, not just the checklist
Calibrate and measure disagreement
Use a stable set of ordinary, edge and high-risk cases. Review independently, compare criterion-level decisions, adjudicate disagreements, refine anchors prospectively and monitor both human-human and human-AI agreement. Agreement alone is not correctness, so retain expert ground truth and appeals.
Report denominators and uncertainty
Report eligible interactions, sampled, scorable, not scorable, pass/partial/fail by criterion, critical failures, reviewer disagreement, appeals, corrective actions and verified follow-up. Segment carefully by channel, locale, product and human/AI participation without ranking tiny groups.
Connect findings to owned change
Route defects to coaching, knowledge, policy, product, workflow, staffing, accessibility, localization, security or privacy owners. Set acceptance evidence and a due date. Re-sample comparable work after release and distinguish correlation from causal improvement.
How OpenMax can coordinate support QA
OpenMax can freeze an eligible sampling frame, create a minimized review view, assemble transcript and system evidence, propose criterion outcomes with evidence pointers, abstain when context is insufficient, route critical findings to humans, record overrides and appeals, assign corrective work, and monitor comparable follow-up samples. People retain rubric ownership, worker/privacy decisions, specialist judgment, employment consequences, customer remedies and final acceptance.
Privacy, fairness and interpretation boundaries
- Do not use QA as undisclosed surveillance. Define necessity and proportionality, provide required notices, consult affected workers where appropriate, minimize access and set retention before monitoring begins.
- Do not infer protected traits, emotion, honesty or intent from voice, accent, writing style or sentiment. Evaluate observable service behavior and validated operational evidence.
- Do not rank agents, vendors, locales or AI systems without comparable populations, adequate denominators, calibration, uncertainty and an appeal route. A score is a decision input, not ground truth.
- CSAT response rate, satisfaction score, QA pass rate, first response time, resolution time, reopen rate and business outcome measure different things. Never merge them into a causal story without appropriate design.
Sources, editorial method, and limitations
OpenMax editors reviewed Intercom’s current conversation-rating eligibility and reporting definitions, the NIST AI RMF Core outcomes for roles, measurement, human oversight and monitoring, the UK Information Commissioner’s guidance and call-center example for worker monitoring, and W3C WCAG 2.2 accessibility requirements. We then synthesized the original 18-control rubric, evidence anchors, failure conditions and calibration case. Sources were rechecked September 3, 2026.
- Intercom — Measure customer satisfaction with conversation ratings
- Intercom — Reporting metrics and attributes
- NIST — AI Risk Management Framework Core
- UK ICO — Specific data protection considerations for worker monitoring
- W3C — Web Content Accessibility Guidelines 2.2
Frequently asked questions
Is CSAT the same as quality assurance?
No. CSAT reflects feedback from customers who were eligible and chose to respond. QA applies a defined rubric to sampled interaction evidence. Report the survey response denominator, QA sample and scorable population separately.
Should all 18 items have equal weight?
Usually not. Set criterion-specific anchors and weights prospectively. Treat selected security, privacy, accuracy or unauthorized-action failures as gates rather than allowing high tone scores to average them away.
Can AI score every conversation?
It can propose outcomes for cases with sufficient permitted evidence. It must abstain on missing or restricted context, surface evidence and uncertainty, and route consequential findings to calibrated human review with override and appeal records.
How many conversations should be reviewed?
There is no universal number. It depends on the decision, volume, variation, risk, desired precision, strata and reviewer capacity. Publish the eligible population, selection method, sample, not-scorable count and limitations instead of presenting a convenient number as representative.
How should QA findings improve service?
Assign each recurring failure to the relevant coaching, knowledge, policy, product, workflow, staffing, localization, accessibility, security or privacy owner. Define acceptance evidence, due date and a comparable follow-up sample.

