快速答案

把1,000条CSAT评论作为Governed Evidence Workflow分析:定义Decision与Eligible Corpus;量化Request、Response和Text Coverage;最小化Data;保留Original Language与Rating Target;建立Versioned Codebook;让两名Human Reviewer在Stratified Sample上校准;要求AI Label引用Exact Span或Abstain;在Locked Holdout审计错误;用诚实Denominator汇总;仅把已复核Theme转为有Owner的Corrective Work。

先建立CSAT指标与语料合同

单位与评分对象

选择Comment、Completed Rating、Conversation或Unique Respondent作为Unit。记录Rating针对Teammate、Chatbot、AI Agent、Product还是Overall Service;不得静默合并不同Rating Object。

人群与时间字段

声明谁Eligible Receive Request、谁Received、谁Responded、谁Added Text,以及Requested、Responded、Started或Updated中哪个Timestamp控制Window。

分母与多标签计算

发布Request、Response、Text Comment、Eligible Comment与Unique Respondent Count。一个Comment可有多个Theme,因此Theme Share可能超过100%;应说明而非强制虚假互斥。

缺失与推断边界

Blank Comment、Survey Nonresponse、Inaccessible Channel、Deleted Record、Language Exclusion与Failed Join都是Data,不是Zero Dissatisfaction。除非经复核Sampling Design支持更广推断,否则结论仅限Observed Corpus。

1,000条评论的十步工作流

按顺序执行。每步都产生可审计Output与Stop Condition;后续步骤不能掩盖Eligibility、Privacy、Calibration或Evaluation失败。

01

先冻结分析问题,再接触评论

先写清本次复核支持的一个Decision,例如下季度哪些已验证Service Failure值得优先修复。定义Owner、Deadline、Allowed Use、Prohibited Use以及什么Evidence会改变决定。不要只写“寻找洞察”;无边界Prompt会产生好看但不可审计的主题。

必要证据与输出
获批Question、Decision Owner、Audience、Exclusion、Review Date,并明确评论分析不代表未回复人群。
人工检查点
Qualified Reviewer确认Scope、Evidence、Uncertainty、Privacy及Proposed Action是否获授权后,Record才能前进。
02

定义合格的1,000条评论语料

明确Survey Instrument/Version、Rating Object、Requested/Responded Timestamp、Date Window、Product、Region、Channel、Language、Actor Type、Duplicate、Edit、Deleted Record,以及选择Latest Completed Response还是Every Response。保留Source ID与Extraction Query。

必要证据与输出
可复现Manifest:Input Count、按原因Exclusion、Missing Comment、Blank Text、Duplicate Policy、Extraction Time与Immutable Snapshot Hash。
人工检查点
Qualified Reviewer确认Scope、Evidence、Uncertainty、Privacy及Proposed Action是否获授权后,Record才能前进。
03

阅读回复者之前,先衡量谁缺失

用所选Timestamp与Population计算Request/Response Denominator。按获准Operational Cohort比较Coverage,如Channel、Language、Product、Issue Type、Accessibility Route与Agent Type。差异是Coverage Warning,不是Customer Trait,也不能看完结果后随意造Weight。

必要证据与输出
Request Rate、Response Rate、Text-Comment Rate、Nonresponse Table、Unknown Value、Excluded Cohort,以及结论仅为Descriptive还是可Generalize的书面决定。
人工检查点
Qualified Reviewer确认Scope、Evidence、Uncertainty、Privacy及Proposed Action是否获授权后,Record才能前进。
04

最小化并保护文本

删除与Purpose无关Field;Redact/Tokenize直接Identifier;限制Raw Text访问;区分Customer-Visible Comment与Internal Note;定义Retention/Deletion;防止Prompt、Log、Export、Screenshot泄漏Credential、Health/Payment Data、Secret或Third-Party Information。

必要证据与输出
Field-Level Data Inventory、适用时Lawful Basis/Consent Review、Access List、Processor/Model Route、Retention Clock、Deletion Test、Incident Path及由Qualified Person复核的Redaction Exception。
人工检查点
Qualified Reviewer确认Scope、Evidence、Uncertainty、Privacy及Proposed Action是否获授权后,Record才能前进。
05

保留语言、语境与评分对象

Original Text与Translation并存,并记录Detected Language、Translation Method/Version、Confidence、获准Review的Conversation Excerpt、Rating Target、Numeric/Ordinal Score、Channel、Product与Issue State。不得清洗掉Negation、Sarcasm、Accessibility-Related Phrasing、Mixed Language或Product Name。

必要证据与输出
Original-Translation Link、Terminology Glossary、Low-Confidence Queue、No-Translation Path、Reviewer Language Capability,并明确区分Teammate、Chatbot、AI Agent、Product与Overall Service Rating。
人工检查点
Qualified Reviewer确认Scope、Evidence、Uncertainty、Privacy及Proposed Action是否获授权后,Record才能前进。
06

建立带明确“其他”路径的编码本

创建Operationally Distinct Theme,含Definition、Inclusion、Exclusion、Positive/Negative Example、Parent-Child Rule、Multi-Label Policy与Other/Uncertain Code。Issue Topic、Sentiment、Severity、Resolution Evidence、Request Type和Proposed Action必须分开,不能一个Label回答六个问题。

必要证据与输出
Versioned Codebook、Change Log、Example Provenance、每条Maximum Label、Precedence Rule、Uncertain Code及Full Run前Owner Approval。
人工检查点
Qualified Reviewer确认Scope、Evidence、Uncertainty、Privacy及Proposed Action是否获授权后,Record才能前进。
07

在盲化分层样本上校准

处理全部1,000条前,按Rating、Language、Channel、Product、Issue State、Length与Time抽取可复现Sample。至少两名Qualified Reviewer独立标注、对账Disagreement、修订Codebook并保留Untouched Holdout。Agreement只是诊断,不证明Category真实或公平。

必要证据与输出
Sample Seed/Strata、Independent Label、Disagreement、Adjudicator、Codebook Revision、Per-Label Agreement、Rare-Class Review、Holdout Lock与Stop/Go Criteria。
人工检查点
Qualified Reviewer确认Scope、Evidence、Uncertainty、Privacy及Proposed Action是否获授权后,Record才能前进。
08

用证据片段与弃权机制运行AI编码

每条Comment必须输出Comment ID、Codebook Version、Proposed Label、Exact Supporting Span、经本任务校准的Confidence、Contradiction/Missing-Context Flag与Abstention Reason。验证Structured Output;隔离Parse Failure;Retry须Idempotent;不得编造Quote或静默替换无证据Label。

必要证据与输出
Prompt/Model/Version、Parameter、Schema Validation、Source Span Offset、Abstention、Retry、Parse Failure、Token/Cost Log、Access Log,以及每个Output到Immutable Input的确定链接。
人工检查点
Qualified Reviewer确认Scope、Evidence、Uncertainty、Privacy及Proposed Action是否获授权后,Record才能前进。
09

按标签与人群审计错误,而非只看总体准确率

复核Locked Holdout与Targeted Slice。报告Per-Label Precision/Recall、Confusion Pair、Unsupported Evidence Span、Missed Negation、Translation Error、Abstention Quality、Multi-Label Omission,以及获准Cohort间Error Difference。低量或Sensitive Finding转人工,不做乐观聚合。

必要证据与输出
含Denominator/Interval的Holdout Result、False Positive/Negative Example、Subgroup Minimum Size、Reviewer Correction、Threshold Rationale、Residual Risk与Rollback Decision。
人工检查点
Qualified Reviewer确认Scope、Evidence、Uncertainty、Privacy及Proposed Action是否获授权后,Record才能前进。
10

把主题汇总为决策,同时不丢证据

分别统计Eligible Comment与Unique Respondent;允许Multi-Label但不能让百分比虚假相加为100;发布Denominator、Unknown、Interval与Example Selection Rule。每个Priority关联Representative/Contradictory Comment、Operational Owner、Corrective Hypothesis、Due Date、Verification Metric、Customer Follow-Up与Decision Log。

必要证据与输出
可复现Table、无Hidden Deduplication、无Cherry-Picked Quote、Base-Rate Context、Contradiction Set、Action Owner、Acceptance Criteria,以及按同一Metric Contract计划Re-measurement。
人工检查点
Qualified Reviewer确认Scope、Evidence、Uncertainty、Privacy及Proposed Action是否获授权后,Record才能前进。

实操案例:1,000行导出变为742条合格评论

以下是假设演示,说明Arithmetic与Review,不是OpenMax客户数据或Benchmark。导出有1,000行;Manifest按预先声明规则排除96条无Comment的未回答Survey、54条Duplicate Snapshot、31条Test Record、22条Internal-Note Leak、18条Window外记录及37条无法对账Rating Object的行,保留742条Comment;每项Exclusion仍按Reason计数。

  1. 先看覆盖,再看主题。 分析师分别按自己的Denominator报告Request、Response与Text-Comment Rate,并标记Language/Phone Cohort代表不足。
  2. 先校准,再扩展。 两名Reviewer独立编码Stratified Sample,发现“Slow Response”与“Unresolved Outcome”常混淆,修订Definition并锁定Holdout。
  3. 主题计数前先看证据。 AI仅在有Exact Text Span时建议Label;Unsupported Label与Uncertain Translation选择Abstain,Reviewer纠正High-Impact与Sampled Record。
  4. 行动但不夸大因果。 经复核Billing-Clarity Theme转为有Owner的Documentation/Invoice-Message Experiment。团队衡量Task Completion、Repeat Contact、New-Comment Coverage与Harm,但不声称Theme导致低CSAT或改动一定提高CSAT。

分开衡量模型质量、研究质量与服务结果

编码质量

Per-Label Precision/Recall、Confusion、Unsupported Span、Abstention、Translation Error、Reviewer Override与Drift。Overall Accuracy会掩盖Rare Theme失败。

研究质量

Coverage、Missingness、Sampling、Duplicate Rate、Codebook Stability、Reviewer Agreement、Contradictory Evidence、Example Selection Integrity与Reproducibility。

服务结果

Customer-Confirmed Task Completion、Repeat Contact、Reopen、Complaint、Accessibility、Safety、Time、Cost及Cohort Distribution;与Rating Response及Model Quality分开。

OpenMax如何协调分析

OpenMax可协调Approved Extract、Immutable Manifest、Redaction、Language Route、Versioned Codebook、Blinded Reviewer Task、Evidence-Linked AI Proposal、Abstention、Adjudication、Holdout Evaluation、Correction Queue、Action Ownership、Deadline与Re-measurement。Research Question、Data Purpose/Authority、Codebook Approval、Sensitive Interpretation、Threshold Choice、Publication与Consequential Service Decision由人负责。

1 · 界定Question、Population、Rating Object、Allowed Use
2 · 准备Manifest、Minimization、Language、Codebook
3 · 校准Blind Label、Disagreement、Revision、Holdout
4 · 分析Evidence Span、Abstention、Validation、Audit
5 · 行动与学习Owner、Correction、Outcome、Re-measurement

隐私、公平与解读边界

  • 不得把Raw Customer Text上传未批准Model、无限期保留、暴露Internal Note,或在没有Authority、Notice、Access Control、Deletion、Security与Processor Review时改作新用途。
  • 不得仅从Wording、Grammar、Name、Language、Channel或Sentiment推断Protected Trait、Health、Disability、Identity、Honesty、Intent、Emotion、Employee Performance或Customer Value。
  • 不能仅因Quote生动就发布。验证Consent/Authority,删除Identifier,保留Meaning/Context,呈现Contradictory Evidence,并防止Search/Linkage重新识别人。
  • 没有Explicit Causal Design、Comparable Population、Stable Metric、Follow-Up Window、Uncertainty、Missing-Data Analysis与Harm Review时,不得声称AI发现Root Cause、Theme代表所有客户或Action提高了CSAT。

来源、编辑方法与限制

OpenMax编辑复核Intercom当前Conversation Rating设置与Remarks View、Conversation Rating Dataset/Metric Definition、Conversation Reporting Population/Timestamp行为、NIST AI RMF 1.0与NIST Generative AI Profile,再原创综合十步工作流、指标合同与假设1,000行案例。资料于2026年9月3日复核。

范围说明 Vendor文档描述其当前自有Product Dataset且可能变化;NIST框架为自愿采用。来源不验证本工作流、不提供OpenMax Customer Data、不证明Representativeness,也不保证Accuracy或CSAT Improvement。必须测试真实Instrument、Population、Language、Model、Reviewer与Decision。

常见问题

AI能在没有人工复核时分析全部1,000条吗?

AI可处理Record,但可治理结果仍需Human Corpus Approval、Codebook Calibration、Holdout Error Review、Sensitive-Case Review与Action Authorization。

低评分与负面评论应该合并分析吗?

Score、Text、Rating Target、Timestamp与Evidence应分开;可分析关系,但不一致是有价值Data,不应被抹掉。

编码本应该有多少主题?

没有通用数字。使用Reviewer能可靠应用的最小Operationally Distinct集合,保留Other/Uncertain,并仅通过Versioned Evidence拆分或合并。

大主题能揭示根因吗?

不能。Frequency只描述Eligible Corpus中的Coded Observation;Root Cause需要Operational Evidence佐证与经测试的Causal Explanation。

OpenMax可自动化什么?

可协调Authorized Extraction、Manifest、Redaction、Coding Proposal、Evidence Link、Abstention、Review、Evaluation、Action Routing与Re-measurement;Purpose、Approval、Interpretation与Consequential Decision由人负责。