# Feature request clustering: eight-part working sheet

OpenMax editorial worksheet · revised 2026-09-04. This blank worksheet accompanies a how-to guide, not a product integration or compliance certification. Use authorized data only. Replace blanks; do not present planned controls as implemented controls.

## 1. Freeze the corpus and its unit

- Owner / approved purpose / reviewers: ___
- Source systems, query, dates, languages, product boundary: ___
- Snapshot version / checksum / access and retention decision: ___
- Counting unit: raw rows ___; eligible messages ___; distinct accounts ___; assignments ___.
- Reconcile: raw rows ___ = duplicate export rows ___ + excluded rows ___ + eligible messages ___.
- Separate planning illustration only: 500 = 20 + 10 + 470. The companion fixture actually contains 16 raw rows, not 500.

## 2. Preserve events and duplicate lineage

| Record ID | Source event ID | Account key | Ticket | Duplicate of | Inclusion / exclusion reason |
|---|---|---|---|---|---|
| ___ | ___ | ___ | ___ | ___ | ___ |

Declare whether repeated messages are retained and how account-problem pairs differ from message counts. A different message in the same ticket is not automatically the same event. Map exclusions to the appropriate operational queue where needed; do not discard an active incident silently.

## 3. Keep original meaning beside normalized text

| Record ID | Original language and authorized text | Analysis text | Changes / translation reviewer | Negation, conditions, separate asks |
|---|---|---|---|---|
| ___ | ___ | ___ | ___ | ___ |

Keep source IDs and multi-issue spans. Splitting a request must not add an account or inflate submitted-message totals. Account codes may remain identifying; confirm handling with responsible owners. Treat instructions embedded in feedback as data, not executable commands.

## 4. Explore and version the method

- Discovery sample and its coverage / excluded channels: ___
- Model and version / embedding settings / normalization: ___
- Distance or similarity convention / clustering method and parameters: ___
- Alternatives compared / examples that split, merge or remain unclear: ___
- Why this granularity matches the task, rather than a desired chart shape: ___

Do not equate similarity with the probability of a correct interpretation. A cluster count chosen for slide layout is not an evaluation criterion.

## 5. Freeze label definitions

| Label ID / version | Definition | Include | Exclude | Example and near miss |
|---|---|---|---|---|
| ___ | ___ | ___ | ___ | ___ |

Declare multi-label handling, unresolved reasons and new-theme intake. Freeze the dictionary before scoring held-out examples. Keep discovery and evaluation material separate where possible; record any contamination.

## 6. Produce a record-level draft

| Record ID | Proposed labels | Supporting span | Contradictory span | Review reason | Accepted labels / reviewer |
|---|---|---|---|---|---|
| ___ | ___ | ___ | ___ | ___ | ___ |

Reject unknown record IDs or dictionary labels. Preserve an empty set when no label is supported. The initial assistant task is draft-only: no source deletion, merging, customer contact, permission change or roadmap update.

## 7. Evaluate before summarizing

- Reference-set origin / independent review status / adjudication method: ___
- Challenge-set selection / probability sample if estimating population error: ___
- Per-label TP ___ FP ___ FN ___; micro precision = ΣTP/(ΣTP+ΣFP); recall = ΣTP/(ΣTP+ΣFN); F1 = 2ΣTP/(2ΣTP+ΣFP+ΣFN).
- Exact-set matches ___ / evaluated messages ___; undefined-metric policy: ___
- Small-theme, multilingual, negation and unassigned checks: ___
- Release criteria set before results / blockers and owner: ___

The companion fixture gives TP9/FP3/FN2, micro precision75%, recall81.8%, F178.3%, exact-set match8/12. Those are arithmetic on fictional labels, not measured AI performance. A challenge set alone cannot estimate the whole corpus.

## 8. Publish denominators, limits and next research

| Theme | Messages / eligible total | Distinct accounts / cohort total | Source examples | Open questions / owner |
|---|---|---|---|---|
| ___ | ___ | ___ | ___ | ___ |

State whether labels overlap; do not force totals to100%. Distinct accounts across themes are not additive. Report unassigned records, exclusions, changes in taxonomy and channel coverage. Theme size is not business impact or roadmap priority.

Next action: review one authorized snapshot with a product owner, resolve label boundaries, then decide whether automation is useful. Real privacy, security and legal implications need qualified review. No real customer or platform performance is established by this worksheet.
