A worked example: where the task becomes difficult
A fictional software team has two hundred feedback messages about onboarding. Some customers dislike the setup form, while others fail at invitation acceptance. Combining both as “onboarding is confusing” hides different fixes. Remove unnecessary personal identifiers, retain stable message references and ask the agent to distinguish reported problems, feature requests and praise. Review the categories before running the complete dataset.
Decisions to make before implementation
Define whether one message can belong to several themes and how duplicates are handled. Separate a customer count from a message count when one person writes repeatedly. Keep severity, frequency and commercial impact distinct. A rare account-access failure may matter more than many color preferences. Decide how ambiguous or sarcastic feedback is flagged for human review rather than forced into a confident category.
A practical sequence for the work
Use the sequence below as a task boundary, not as a claim that the example has been executed. Work with approved inputs and the project’s actual architecture. If a required integration or permission is unavailable, keep that stage visibly incomplete rather than generating a plausible substitute result.
- Prepare a permitted dataset with stable references and minimized personal data.
- Develop categories on a sample and inspect disagreements with human coding.
- Classify the complete set with explicit uncertainty and duplicate handling.
- Connect proposed improvements to representative examples and measured counts.
A detailed brief you can adapt for your agent
Replace the illustrative context with your approved facts and controlled inputs. Keep the stated boundaries when adapting the brief. The expected deliverable matters more than a particular tool name: ask for an explanation grounded in the inspected material and evidence for the requested outcome.
Analyze this redacted onboarding-feedback dataset. Propose a coding scheme on the first sample, then classify with stable message references. Separate invitation failures from form confusion. Preserve uncertain and opposing examples. Do not invent quotes or calculate unsupported percentages. Return theme counts, denominator definitions and an evidence-backed improvement queue.Failure modes that an attractive preview can hide
A model can invent quotes or report percentages without counting. Require exact source references for excerpts and calculate totals independently. Do not erase positive feedback when investigating complaints, and do not equate silence with satisfaction. Feedback collected through a support channel may overrepresent frustrated customers compared with the whole user base.
Technical references: NIST: AI Risk Management Framework
Acceptance checks and the evidence to retain
Provide the category definitions, traceable classifications, counts and unresolved examples. Product decisions should reference the specific journey affected rather than a generic sentiment score.
| Controlled case | Expected evidence |
|---|---|
| One person submits several similar messages | Customer and message counts remain distinguishable. |
| Message fits two themes | The documented multi-label rule is applied consistently. |
| Summary includes a quotation | The wording and reference match an authorized source message. |
Specific answers
Common questions
Should I rely on an overall sentiment score?
Use it cautiously; task-specific themes and supporting examples are often more actionable.
Can AI quote customer messages?
Only reproduce authorized material accurately and minimize private information in the published output.
What is the practical completion criterion?
Provide the category definitions, traceable classifications, counts and unresolved examples. Product decisions should reference the specific journey affected rather than a generic sentiment score.
Sources and editorial method
These references support the indicated technical facts. Workflows, examples and decision tables are original Roseram analysis. Illustrative costs are not vendor prices. No search volume, organic difficulty, ranking result or product endorsement is implied.
- NIST: AI Risk Management Framework ↗
A reference for managing AI risk; these task examples are original editorial workflows, not a NIST endorsement or a compliance assessment.
Roseram offers AI software and may compete with tools discussed here. Sources checked 2026-10-11. Send a sourced correction.
Your next step
Keep a practical checklist.
Mark your progress. This checklist and helpfulness choice are saved on this device only; they are not public reviews.
0 of 5 complete
Share your experience in the community or submit a sourced correction. Public experiences remain separate from editorial claims.
Bring your next idea
Keep learning. Build with context.
Get Roseram model and workspace reopening updates. The guide remains available whether or not you subscribe.
Explore the workspace guide →