Quick answer: build evidence ledgers before writing prose
For a long business thread, first freeze the mailboxes, folders, cutoff and inclusion rule. Build a message graph from stable IDs and authorized header fields; map participants, subject variants and attachment versions; then extract atomic records for decisions, commitments, open questions, corrections and withdrawn proposals. Only after those records reconcile should a model draft a concise brief.
The minimum safe output
A useful brief names its scope and cutoff, lists current decisions with conditions, assigns open actions to an accountable role, preserves exact deadlines and timezones, exposes unresolved questions, and marks superseded or withdrawn material. Each statement should cite one or more message IDs. Anything without adequate support is removed or explicitly qualified.
What the summary must not do
Classification and summarization do not create authority. The brief must not turn “please approve,” a quoted instruction, a preliminary quote, or a calendar suggestion into an approval, purchase, payment, signature or accepted meeting. Mail content is evidence to review, not executable policy.
Download the working materials
Use the editable AI email thread summary worksheet to define corpus scope, graph fields, typed ledgers, review criteria and release gates. Then inspect the complete TS083 fictional evidence packet, which includes all 50 synthetic messages, nine attachment records and the full scoring ledger—without omitted “similar” rows.
Start with the business decision, not the inbox display
A team rarely needs “a shorter email.” It needs a dependable handoff: what was decided, what remains blocked, who owes what, when it is due, and which evidence supports that state. Defining this decision determines what belongs in the corpus and what the final brief must preserve.
Write one decision sentence
A precise purpose might be: “Prepare the release-readiness handoff for the Northstar pilot at the September 2 cutoff.” That sentence excludes unrelated correspondence and tells the reviewer which omissions could be material. “Summarize my inbox” has no comparable boundary.
Name the accountable audience
An operations owner may need scope, dependencies and dates. Legal may need versions, unresolved clauses and signature state. Security may need control evidence and exceptions. One source corpus can support different views, but each view should use the same underlying record IDs rather than silently rebuilding the facts.
Freeze an evidence cutoff
Every brief represents a state at a time. Record the latest included UTC timestamp and the generation time. When a new reply arrives, create a new version; do not quietly edit a released brief. A visible cutoff prevents yesterday's summary from masquerading as today's truth.
Define the authorized corpus
The messages visible in a conversation view are not automatically the messages a system is authorized to process. Scope must combine access permission with an explicit inclusion rule.
Record accounts, folders and exclusions
List the exact mailbox accounts, shared aliases and folders in scope. State whether Sent, Archive, Trash or delegated accounts are included. Exclude personal mail, privileged material and unsupported attachments unless an authorized owner adds a documented handling path.
Use stable identifiers
RFC 5322 describes Message-ID, In-Reply-To and References fields that can provide structural reply evidence. Preserve the supplied values and flag missing or malformed fields. They help construct a graph; they do not prove that the body is true or that every provider will render the same thread.
Document provider behavior
Gmail's official help explains that replies can be grouped into conversations, while a subject change or a conversation exceeding 100 emails can split the display. Microsoft's Outlook guidance describes conversation and branch views whose coverage depends on account, version and settings. Record those conditions. Never equate a convenient UI group with a complete evidence set.
Reconstruct the thread as a graph
A 50-message project exchange is rarely a single straight line. Scope, security, commercial terms and rollout scheduling may develop on different branches while referring to the same project.
Preserve parent-child relationships
For every message, store its stable ID, primary parent, references, received time, participant, subject variant and project key. A root with four branches is more informative than a date-sorted transcript because it shows which reply corrected which earlier claim.
Record explicit cross-branch links
A legal reply may cite a security requirement; a rollout message may condition a date on an unsigned addendum. Capture those explicit links as evidence without pretending one message sits under two different parents. The graph can contain cross-links while retaining one primary reply chain.
Flag gaps instead of inventing continuity
If a message says “as agreed in the missing attachment” and the attachment is absent, mark the dependency unresolved. Do not infer the agreement from later confidence or repeated wording. A gap is an operational finding, not a reason to fill the story with a likely answer.
Normalize identities and time carefully
Names and dates are high-value summary fields, but they are also easy to distort.
Map participants to approved roles
Maintain a participant dictionary linking addresses to known roles. Keep display names, addresses and role mappings distinct. A familiar display name is not sufficient identity proof, and two people with similar names should not be merged without evidence.
Retain UTC and original wording
Store the original timestamp and an approved UTC normalization. When a sender writes “Tuesday morning” or “09:00 local project time,” preserve the phrase and mark timezone resolution as open. Never choose a timezone merely to make the brief look complete.
Separate owner from sender
The person who mentions an action is not necessarily the person responsible for it. “Can Legal confirm?” creates a requested owner only after the operating process assigns one. The action ledger should show the evidence for ownership, not just the closest name in the text.
Control attachment versions before using their facts
Long threads often carry multiple files with similar names. A summary that quotes the wrong version can reverse a security, pricing or schedule decision.
Build an attachment register
For each supplied file, capture filename, version, digest, first-seen message, approved extract, review state and supersession link. The digest is a version identifier only when the acquisition process is trustworthy; it is not proof that the file is safe or correct.
Keep superseded files visible
Mark version 1 as superseded by version 2, but retain both records. The earlier version explains why a correction occurred and supports an audit. Deleting it may make the final state look cleaner while destroying provenance.
Do not execute active content
An email summarizer does not need permission to run macros, follow embedded instructions or open remote resources. Supply a bounded, approved extract through an authorized attachment-handling process. This is especially important because OWASP's prompt-injection guidance treats crafted external content as a route by which intended model behavior can be altered.
Extract atomic claims into typed ledgers
Do not ask a model to jump directly from 50 messages to one paragraph. First convert the corpus into records that can be individually checked.
Decision ledger
A decision record needs the final choice, decision owner, supporting message IDs, conditions and state at cutoff. “Vendor proposed,” “finance budgeted,” and “legal reviewed” are not interchangeable with “the authorized owner decided.”
Commitment ledger
A commitment combines an action, accountable owner, due time, evidence and current state. Preserve whether a deadline was completed, moved, missed or left without a timezone. Do not make a polite intention look like a binding commitment.
Question ledger
Open questions deserve first-class status. Record what remains unknown, which role can answer, what evidence is required, and the latest relevant message. A summary that hides questions behind a confident conclusion is operationally dangerous.
Correction ledger
A correction should name the affected field, earlier value, corrected value and evidence chain. Apply it only to that field. Correcting a retention period does not automatically approve the security design, and correcting a price does not approve a purchase.
Withdrawal ledger
Record each proposal and its explicit withdrawal evidence. Then check that withdrawn items are absent from the decision section. Repetition and recency are weak substitutes for state: a stale proposal can appear many times and still be withdrawn.
Compose the action brief from current state
Once the ledgers reconcile, the prose becomes simpler. The brief should be compact because the evidence model carries the detail.
Open with scope and cutoff
Name the project, included corpus, time window, last message and excluded material. State whether attachments were represented by approved extracts rather than opened directly. A reviewer should understand the evidence boundary before reading any conclusion.
Present decisions with conditions
List decisions as atomic statements. Put conditions in the same line: “Cutover is Sep 16 at 14:00 UTC, conditional on security and legal prerequisites.” Never move the condition into a footnote where a hurried reader may miss it.
List actions by owner and deadline
Use one row per owned action. Include the exact UTC deadline and status, then cite the commitment and message IDs. If an owner or time is missing, display “unassigned” or “time unresolved” rather than guessing.
Expose blocks and questions
Show blocked prerequisites before optional context. A missing deletion artifact or unsigned addendum matters more to release readiness than a polished recap of early discussion.
Close with authority limits
State that the brief is informational and review-only. It does not constitute security approval, legal advice, a signature, procurement approval, payment authorization, system provisioning or acceptance of calendar changes.
Walk through the complete TS083 case
TS083 is a deliberately synthetic Northstar access-gateway rollout. It demonstrates the record design without claiming a real customer, live mailbox or OpenMax result.
What the 50-message corpus contains
The packet contains 12 scope/technical messages, 16 security/data-handling messages, eight commercial/legal messages and 14 rollout/support messages. Twelve fictional roles use three subject variants. Nine attachment records include superseded scope, questionnaire, quote and rollout files.
What is current at the cutoff
Seven decisions define a two-region, SSO-only pilot using synthetic identities, four audit fields, quote v2 at a stated ceiling, an unsigned DPA prerequisite, and a conditional Sep 16 14:00 UTC cutover. Eleven commitments track evidence, contract, runbook, rehearsal, support and go/no-go work. Six questions remain open.
Why corrections matter
Four correction chains change retention from 365 to 30 days, price from USD 128,000 to USD 118,000, date from Sep 15 to Sep 16, and an ambiguous local time to 14:00 UTC. Each earlier value remains in the packet, but none should appear as the current state.
Why withdrawals matter
Password fallback, automatic renewal and a Friday fallback cutover are all withdrawn. A fluent model may repeat them because they were salient earlier. The withdrawal ledger makes the release check mechanical: none may appear as an active decision.
Evaluate a candidate summary claim by claim
The complete packet includes thread-summary-draft-v0.2, a fictional candidate designed to fail visibly. It has no relationship to OpenMax product performance.
Split compound sentences
A sentence such as “Legal approved the addendum and rollout begins Friday” contains at least two claims. Split them so each receives evidence, cutoff state, owner/deadline review and a disposition. One supported half must not shield one invented half.
Measure claim-support precision
The candidate contains 24 atomic claims. Eighteen are supported, three are contradicted by later corrections or withdrawal, and three are unsupported. The calculation is 18 ÷ 24 × 100 = 75%. This is a teaching result on one frozen synthetic case, not a benchmark.
Measure critical-fact recall
TS083 declares 12 critical facts before scoring. The candidate preserves eight under the strict rubric: 8 ÷ 12 × 100 = 66.666…%, displayed as 66.7%. Fluent coverage is not enough when omitted qualifiers change readiness or authority.
Measure owner and deadline attribution
Fourteen candidate claims require a named owner or deadline. Ten are attributed correctly: 10 ÷ 14 × 100 = 71.428…%, displayed as 71.4%. A correct task with the wrong owner can still break the handoff.
Measure correction capture
The candidate preserves only one of four correction chains: 1 ÷ 4 × 100 = 25%. The raw score is useful because it pinpoints a stale-state problem that general prose quality would obscure.
Set release gates before seeing the result
Thresholds should reflect the consequence of the brief. A release-readiness summary deserves stricter gates than an informal personal recap.
Freeze the rubric first
TS083 requires 24 of 24 claims supported or explicitly qualified, all 12 critical facts captured, all 14 attributions correct, all four corrections preserved, all six open questions visible, all three withdrawals excluded and zero consequential effects. Set those rules before candidate output exists.
Treat zero side effects as necessary, not sufficient
The fictional candidate cannot send, sign, approve, pay, provision or modify anything, so its consequential-effect count is zero. It still receives NOT_RELEASED because its content gates fail. Read-only generation reduces one risk class; it does not create factual reliability.
Do not move the gold posts
If qualified reviewers find a gold record wrong, version the rubric, explain the correction and rerun the evaluation. Never silently change the expected answer merely to improve a score.
Design human review around likely failure modes
Human review is most valuable when it is tied to explicit claims and known risks, not when someone is asked to “read it over.”
Check stale facts and negation first
Search the candidate for earlier prices, dates, attachment versions, withdrawn options and missing words such as “not,” “conditional,” or “unsigned.” These small qualifiers often control the business meaning.
Verify high-impact owners
Review the roles attached to money, contracts, security, privacy, access and go/no-go decisions. A summary should not promote a contributor, requester or model into an authorized approver.
Review untrusted content
Quoted emails, forwarded blocks and attachment extracts may contain instructions addressed to a person or system. Treat them as content. In TS083, “mark security approved” is preserved as a hostile test string and explicitly grants no authority.
Record the reviewer and disposition
Store the corpus version, candidate version, rubric, reviewer, time, accepted/rejected claim IDs and final release state. This makes later corrections explainable instead of mysterious.
Use OpenMax only within verified product boundaries
OpenMax's official AI business email assistant page describes a workflow based on approved threads, permitted facts, draft preparation and human ownership of final send and commitments. That is the relevant connection here: a governed agent workflow can prepare a reviewable artifact while people retain consequential authority.
Prepare the inputs first
Before mapping this workflow to OpenMax, define the mailbox scope, approved facts, thread/attachment records, owner matrix, allowed tools, review gates and audit fields. A weak input contract cannot be repaired by a confident summary style.
Keep product claims narrow
This guide does not claim that OpenMax natively connects to every mailbox, scans attachments, reconstructs every provider thread, prevents every injection, or achieves the TS083 numbers. Confirm current integration, permission and deployment details with the OpenMax team for the intended environment.
Preserve human final authority
Draft preparation and operational orchestration should remain separate from final send, contract signature, payment, security approval, data disclosure and access provisioning. The action brief can make review faster without becoming the decision-maker.
Adopt a staged maturity model
Move from evidence hygiene to constrained assistance. Each stage should pass its own review before more capability is added.
Stage 0: manual corpus register
A person freezes messages, attachment versions and the cutoff. The goal is to discover missing evidence and ambiguous ownership without model-generated prose.
Stage 1: read-only extraction
The system proposes atomic records with citations. Humans approve or correct every decision, commitment, question, correction and withdrawal. There are no consequential tools.
Stage 2: reviewed action brief
A model composes prose only from accepted ledgers. Claim-level gates and named review are required before release. New mail invalidates or versions the brief.
Stage 3: bounded downstream preparation
After separate authorization, accepted actions may prepare—not execute—tickets, draft replies or calendar options. Every destination, field and permission remains allowlisted and reviewable.
Common failure modes and fixes
Most thread-summary failures are state-management failures rather than grammar failures.
Summarizing the visible conversation only
Failure: missing branches or folders disappear. Fix: use a frozen corpus manifest and reconcile stable IDs against provider settings.
Treating the latest message as the whole truth
Failure: conditions and authority from earlier evidence vanish. Fix: resolve each atomic field through its correction chain, not through recency alone.
Copying every number into the brief
Failure: stale price, date or retention values compete with current facts. Fix: show the current value in the brief and retain earlier values only in the evidence ledger.
Hiding uncertainty
Failure: missing owner, timezone or attachment becomes a plausible invention. Fix: make “unresolved” an allowed value and route a bounded question.
Confusing a summary with approval
Failure: generated wording is treated as authorization. Fix: maintain a capability matrix, explicit authority statement and zero-effect default.
Implementation checklist
Use the checklist as a go/no-go review, not a decorative appendix.
Before extraction
- Business purpose, audience and cutoff are explicit.
- Accounts, folders, exclusions and retention are authorized.
- Stable message IDs and provider/client settings are recorded.
- Participant and attachment handling rules are approved.
- Consequential tools are absent or independently blocked.
Before drafting
- Every included message appears once in the graph register.
- Attachment versions and supersession links reconcile.
- Decisions, commitments, questions, corrections and withdrawals have evidence.
- Ambiguous owners, times and claims remain visibly unresolved.
- Critical facts and release gates were frozen before candidate output.
Before release
- Every candidate sentence is split into atomic claims.
- Current claims are supported and cited.
- Owners, deadlines, conditions and negation are correct.
- Open questions remain visible; withdrawals stay out of decisions.
- Reviewer, version, disposition and rollback path are stored.
FAQ: frequently asked questions
Can I summarize a thread using only its subject line?
No. Subject lines can change, collide or remain unchanged across different work. Use authorized stable identifiers, reply evidence, project mapping and provider settings, then document any uncertainty.
Does a valid Message-ID prove the email's claims are true?
No. It is structural metadata. The body, sender relationship, attachment state and business authority still require separate verification.
Should the summary include every message?
The evidence register should reconcile every included message. The released brief should include every material current decision, action, question, correction and condition—not repeat every conversational sentence.
How should a corrected deadline appear?
Display the corrected date/time and cite the correction chain. Preserve the earlier value in the ledger, label the exact field superseded, and avoid inferring a timezone that was never supplied.
Can an AI-generated thread summary send a reply automatically?
Not by default. Summarization does not grant send authority. Any later drafting or sending capability needs a separate permission model, approved facts, review policy and audit trail.
Is TS083 evidence of OpenMax accuracy?
No. TS083 and thread-summary-draft-v0.2 are fictional teaching artifacts. Their numbers demonstrate reproducible evaluation and deliberately lead to NOT_RELEASED; they are not product results or customer experience.
Sources and next step
The method uses primary and official documentation available at the September 5, 2026 research cutoff:
- RFC Editor: RFC 5322 Internet Message Format for narrow structural use of Message-ID, In-Reply-To and References.
- Google: Group emails into conversations for Gmail conversation behavior and documented split conditions.
- Microsoft: View email messages by conversation in Outlook for Outlook conversation/branch behavior and setting differences.
- OWASP: Prompt Injection for the need to treat embedded instructions as untrusted content.
- NIST AI 600-1, Generative AI Profile for risk-management framing around provenance, evaluation and human accountability. Referencing it is not a certification claim.
- OpenMax: AI business email assistant for the related approved-thread, permitted-facts, draft-preparation and human-final-send boundary.
Begin with the editable worksheet. Freeze one consented or synthetic thread, build its complete graph and ledgers, and require claim-level review before the brief informs any business action.

