AI-assisted practical guide. Examples are hypothetical; proposed workflows are editorial suggestions.
Borderline examples reveal how a classifier behaves when labels are not obvious. Build them deliberately instead of testing only cases whose answers are already easy to guess.
Define the decision first
Write the allowed labels, their meanings, and what should happen when information is insufficient. Use synthetic examples for an initial trial so private data is unnecessary. Keep the expected handling and its rationale separate from the model's output.
Include cases where two labels seem plausible, a key fact is absent, or the wording resembles a familiar category but the intent differs.
Review a hypothetical disagreement
Imagine a support classifier must distinguish a request for information from a request to change an account. A message asks how a change would work but does not authorize it. A useful expected behavior may be to classify the information request and avoid triggering the change.
If reviewers disagree about the label, resolve the definition before treating the model's answer as a failure. Ambiguous evaluation criteria produce ambiguous results.
Report the scope honestly
Keep examples, outputs, and decisions in a small evaluation record. A synthetic trial does not establish real-world accuracy or safety. Use it to refine definitions and identify escalation needs. The goal is a clearer decision boundary and observable handling of uncertainty, not an impressive score built from easy or selectively chosen examples.