Quick answer
Analyze 1,000 CSAT comments as a governed evidence workflow: define the decision and eligible corpus; quantify request, response, and text coverage; minimize data; preserve original language and rating target; build a versioned codebook; calibrate two human reviewers on a stratified sample; require AI labels to cite exact spans or abstain; audit errors on a locked holdout; aggregate with honest denominators; and turn only reviewed themes into owned corrective work.
Start with a CSAT metric and corpus contract
Unit and rating target
Choose comment, completed rating, conversation, or unique respondent as the unit. Record whether the rating concerns a teammate, chatbot, AI agent, product, or overall service. Never merge unlike rating objects silently.
Population and timestamp
Declare who was eligible to receive a request, who received one, who responded, who added text, and which requested, responded, started, or updated timestamp controls the window.
Denominators and multi-label math
Publish counts for requests, responses, text comments, eligible comments, and unique respondents. A comment can carry several themes, so theme shares may exceed 100%; say so rather than forcing false exclusivity.
Missingness and inference limit
Blank comments, survey nonresponse, inaccessible channels, deleted records, language exclusions, and failed joins are data—not zero dissatisfaction. Limit findings to the observed corpus unless a reviewed sampling design supports broader inference.
Ten-step workflow for 1,000 comments
Run the steps in order. Each creates an auditable output and a stop condition; the next step must not hide failed eligibility, privacy, calibration, or evaluation work.
Freeze the analysis question before touching the comments
Write one decision the review must support: for example, which verified service failures deserve corrective work next quarter. Define owner, deadline, allowed uses, prohibited uses, and what evidence would change the decision. Do not begin with “find insights”; an unbounded prompt makes attractive but unauditable themes.
- Required evidence and output
- Approved question, decision owner, intended audience, exclusions, review date, and a statement that comment analysis does not represent nonrespondents.
- Human checkpoint
- A qualified reviewer confirms scope, evidence, uncertainty, privacy, and whether the proposed action is authorized before the record advances.
Define the eligible 1,000-comment corpus
Specify survey instrument and version, rating object, requested/responded timestamps, date window, products, regions, channels, languages, actor types, duplicates, edits, deleted records, and whether the latest completed response or every response is selected. Preserve the source ID and extraction query.
- Required evidence and output
- A reproducible manifest with input count, exclusions by reason, missing comments, blank text, duplicate policy, extraction time, and immutable snapshot hash.
- Human checkpoint
- A qualified reviewer confirms scope, evidence, uncertainty, privacy, and whether the proposed action is authorized before the record advances.
Measure who is missing before reading who responded
Calculate request and response denominators using the selected timestamp and population. Compare response coverage by permitted operational cohorts such as channel, language, product, issue type, accessibility route, and agent type. Treat differences as coverage warnings, not customer traits or weights invented after seeing results.
- Required evidence and output
- Request rate, response rate, text-comment rate, nonresponse table, unknown values, excluded cohorts, and a documented decision on whether conclusions are descriptive or generalizable.
- Human checkpoint
- A qualified reviewer confirms scope, evidence, uncertainty, privacy, and whether the proposed action is authorized before the record advances.
Minimize and protect the text
Remove fields not needed for the stated purpose; redact or tokenize direct identifiers; restrict raw-text access; separate customer-visible comments from internal notes; define retention and deletion; and prevent prompts, logs, exports, or screenshots from leaking credentials, health data, payment data, secrets, or third-party information.
- Required evidence and output
- Field-level data inventory, lawful basis or consent review where applicable, access list, processor/model route, retention clock, deletion test, incident path, and redaction exceptions reviewed by a qualified person.
- Human checkpoint
- A qualified reviewer confirms scope, evidence, uncertainty, privacy, and whether the proposed action is authorized before the record advances.
Preserve language, context, and rating object
Store original text beside any translation, detected language, translation method and version, confidence, conversation excerpt allowed for review, rating target, numeric/ordinal score, channel, product, and issue state. Do not normalize away negation, sarcasm, accessibility-related phrasing, mixed language, or product names.
- Required evidence and output
- Original-to-translation linkage, terminology glossary, low-confidence queue, no-translation path, reviewer language capability, and explicit distinction among teammate, chatbot, AI-agent, product, and overall-service ratings.
- Human checkpoint
- A qualified reviewer confirms scope, evidence, uncertainty, privacy, and whether the proposed action is authorized before the record advances.
Build a codebook with an explicit “other” path
Create operationally distinct themes with definition, inclusion, exclusion, positive and negative examples, parent-child rules, multi-label policy, and an “other/uncertain” code. Separate issue topic from sentiment, severity, resolution evidence, request type, and proposed action so one label does not pretend to answer six questions.
- Required evidence and output
- Versioned codebook, change log, example provenance, maximum labels per comment, precedence rules, uncertain code, and owner approval before the full run.
- Human checkpoint
- A qualified reviewer confirms scope, evidence, uncertainty, privacy, and whether the proposed action is authorized before the record advances.
Calibrate on a blinded stratified sample
Before processing all 1,000 comments, select a reproducible sample across rating, language, channel, product, issue state, length, and time. At least two qualified reviewers label independently, reconcile disagreements, revise the codebook, and reserve an untouched holdout. Agreement is diagnostic, not proof that the categories are true or fair.
- Required evidence and output
- Sample seed and strata, independent labels, disagreements, adjudicator, codebook revisions, per-label agreement, rare-class review, holdout lock, and stop/go criteria.
- Human checkpoint
- A qualified reviewer confirms scope, evidence, uncertainty, privacy, and whether the proposed action is authorized before the record advances.
Run AI coding with evidence spans and abstention
For each comment, require comment ID, codebook version, proposed labels, exact supporting span, confidence that has been calibrated for this task, contradiction or missing-context flag, and abstention reason. Validate structured output; quarantine parse failures; make retries idempotent; never invent a quote or silently replace an unsupported label.
- Required evidence and output
- Prompt/model/version, parameters, schema validation, source span offsets, abstentions, retries, parse failures, token/cost log, access log, and deterministic link from every output to the immutable input.
- Human checkpoint
- A qualified reviewer confirms scope, evidence, uncertainty, privacy, and whether the proposed action is authorized before the record advances.
Audit errors by label and cohort, not only overall accuracy
Review the locked holdout and targeted slices. Report precision and recall per label, confusion pairs, unsupported evidence spans, missed negation, translation error, abstention quality, multi-label omissions, and error differences across permitted cohorts. Low-volume or sensitive findings go to human review rather than optimistic aggregation.
- Required evidence and output
- Holdout results with denominators and intervals, false-positive/negative examples, subgroup minimum sizes, reviewer corrections, threshold rationale, residual risks, and a rollback decision.
- Human checkpoint
- A qualified reviewer confirms scope, evidence, uncertainty, privacy, and whether the proposed action is authorized before the record advances.
Aggregate themes into decisions without losing the evidence
Count eligible comments and unique respondents separately; permit multiple labels without making percentages sum falsely to 100; publish denominators, unknowns, intervals, and example-selection rules. For each priority, link representative and contradictory comments, operational owner, corrective hypothesis, due date, verification metric, customer follow-up, and decision log.
- Required evidence and output
- Reproducible table, no hidden deduplication, no cherry-picked quote, base-rate context, contradiction set, action owner, acceptance criteria, and a scheduled remeasurement using the same metric contract.
- Human checkpoint
- A qualified reviewer confirms scope, evidence, uncertainty, privacy, and whether the proposed action is authorized before the record advances.
Worked example: 1,000 exported rows become 742 eligible comments
This hypothetical walkthrough illustrates arithmetic and review, not OpenMax customer data or a benchmark. An export contains 1,000 rows. The manifest removes 96 unanswered survey records with no comment, 54 duplicate snapshots under the predeclared latest-completed rule, 31 test records, 22 internal-note leaks, 18 records outside the date window, and 37 rows whose rating object cannot be reconciled. The retained corpus is 742 comments; every exclusion remains counted by reason.
- Coverage before themes. The analyst reports request, response, and text-comment rates from their own denominators and flags underrepresented language and phone cohorts.
- Calibration before scale. Two reviewers independently code a stratified sample, find that “slow response” and “unresolved outcome” are often confused, revise definitions, and lock a holdout.
- Evidence before a theme count. AI proposes labels only with exact text spans. Unsupported labels and uncertain translations abstain; reviewers correct high-impact and sampled records.
- Action without causal overclaim. A reviewed billing-clarity theme becomes an owned documentation and invoice-message experiment. The team measures task completion, repeat contact, new-comment coverage, and harms; it does not claim the theme caused low CSAT or that the change will raise it.
Measure model quality, research quality, and service outcomes separately
Coding quality
Per-label precision and recall, confusion, unsupported spans, abstention, translation error, reviewer override, and drift. Overall accuracy alone can hide a failed rare theme.
Research quality
Coverage, missingness, sampling, duplicate rate, codebook stability, reviewer agreement, contradictory evidence, example-selection integrity, and reproducibility.
Service outcome
Customer-confirmed task completion, repeat contact, reopen, complaint, accessibility, safety, time, cost, and distribution across cohorts. Keep these separate from rating response and model quality.
How OpenMax can coordinate the analysis
OpenMax can coordinate approved extracts, immutable manifests, redaction, language routes, versioned codebooks, blinded reviewer tasks, evidence-linked AI proposals, abstentions, adjudication, holdout evaluation, correction queues, action ownership, deadlines, and remeasurement. Humans retain the research question, data purpose and authority, codebook approval, sensitive interpretation, threshold choice, publication, and consequential service decisions.
Privacy, fairness, and interpretation boundaries
- Do not upload raw customer text to an unapproved model, retain it indefinitely, expose internal notes, or reuse it for a new purpose without authority, notice, access controls, deletion, security, and processor review.
- Do not infer protected traits, health, disability, identity, honesty, intent, emotion, employee performance, or customer value from wording, grammar, name, language, channel, or sentiment alone.
- Do not publish a quote merely because it is vivid. Verify consent or authority, remove identifiers, preserve meaning and context, represent contradictory evidence, and prevent search or linkage from re-identifying the person.
- Do not claim that AI discovered the root cause, that a theme represents all customers, or that an action improved CSAT without an explicit causal design, comparable population, stable metric, follow-up window, uncertainty, missing-data analysis, and harm review.
Sources, editorial method, and limitations
OpenMax editors reviewed Intercom’s current conversation-rating setup and remarks view, conversation-rating dataset and metric definitions, conversation reporting population and timestamp behavior, NIST AI RMF 1.0, and the NIST Generative AI Profile. We synthesized the ten-step workflow, metric contract, and hypothetical 1,000-row case. Sources were rechecked September 3, 2026.
- Intercom — Measure customer satisfaction with conversation ratings
- Intercom — Reporting metrics and attributes
- Intercom — Conversations reporting
- NIST — AI Risk Management Framework 1.0
- NIST — Generative AI Profile
Frequently asked questions
Can AI analyze all 1,000 comments without human review?
It can process records, but a governed result still needs human corpus approval, codebook calibration, holdout error review, sensitive-case review, and action authorization.
Should low ratings and negative comments be analyzed together?
Keep score, text, rating target, timestamp, and evidence separate. Their relationship can be analyzed, but disagreement is useful data rather than an error to erase.
How many themes should the codebook have?
There is no universal number. Use the smallest operationally distinct set that reviewers can apply reliably, retain other/uncertain, and split or merge only through versioned evidence.
Does a large theme reveal the root cause?
No. Frequency describes coded observations in the eligible corpus. Root cause needs corroborating operational evidence and a tested causal explanation.
What can OpenMax automate?
It can coordinate authorized extraction, manifests, redaction, coding proposals, evidence links, abstention, review, evaluation, action routing, and remeasurement while people retain purpose, approval, interpretation, and consequential decisions.

