VOICE AGENT EVALUATION BRIEF 1. DECISION THIS SET MUST SUPPORT Model, provider, or release being evaluated: Decision to make from the results: Production failure or risk being investigated: Release cadence and required repeatability: 2. DEPLOYMENT CONDITIONS Languages and markets: Caller accents and demographic slices that matter to the deployment: Channel: wideband, 8 kHz telephony, separate channels, or mixed: Devices and call path: Noise and SNR bands: Conversation behavior: interruption, overlap, backchannels, repair, silence: Domain scenarios and critical vocabulary: 3. HELD-OUT DESIGN Target usable hours and unique speakers: Speaker-disjoint rule: Prompt or scenario separation rule: Training-data non-reuse commitment: Access controls and split manifest owner: Refresh or release policy: 4. REFERENCE DATA Transcript style: verbatim, normalized, or both: Human review and adjudication method: Speaker, turn, and timestamp fields: Critical entities or intents: Condition metadata required for slicing: 5. METRICS AND SLICES Primary metric and calculation method: Secondary metrics: Minimum slices: language, accent, channel, noise, scenario, speaker group: Critical entity or task-success checks: Thresholds for release, warning, and failure: Treatment of missing or low-volume slices: 6. ACCEPTANCE AND DELIVERY Audio format validation: Reference-transcript quality check: Quota tolerance: Metadata completeness threshold: Duplicate and leakage checks: Replacement and re-review rules: Delivery contents and naming convention: 7. RIGHTS AND REVIEW Permitted model and evaluation uses: License and exclusivity: Sensitive-data or biometric review needed: Retention, deletion, and access requirements: Buyer legal or safety reviewer: 8. OPEN QUESTIONS Template note: This worksheet is an evaluation-scoping aid, not a claim that a fixed evaluation pack is available and not legal advice. A real collection requires an approved scope, consent and license terms, sample, acceptance criteria, and fulfilment owner.