SPEAKER DIARIZATION EVALUATION AND ACCEPTANCE TEMPLATE Version: 1.0 (2026-08-02) Purpose ------- Use this template to compare a diarization system, annotation delivery, or speaker-labeled speech-data proposal against one written brief. Complete the fields with the supplier and the people who own the downstream product. This is an ungated scoping template, not a benchmark result, legal opinion, or claim that a fixed Spirelight dataset, sample, service configuration, accuracy level, or evaluation pack is available. Project scope, reference data, metrics, thresholds, reviewer coverage, rights, security, schedule, sample availability, and price must be confirmed in writing. 1. SCOPE AND DEPLOYMENT ----------------------- Project / system: Decision this evaluation must support: System or delivery being evaluated: System version, model version, or delivery version: Evaluation owner: Supplier owner: Reviewers and approval date: Deployment job: [ ] Meeting transcription [ ] Call-center or telephony transcription [ ] Voice-agent turn taking [ ] Broadcast / interview transcription [ ] Multi-party conversational training data [ ] Speaker-attributed ASR evaluation [ ] Other: Languages, locales, and code-switching requirements: Expected speaker-count range per recording: Known or unknown speaker count at inference: Audio channel layout (mixed, stereo, separate channels, other): Sample rate, codec, bit depth, and container: Microphone / device / call path: Environment, distance, reverberation, and noise conditions: Expected overlap, interruptions, and short backchannels: Expected recurring speakers across recordings: Latency or streaming constraints: Downstream consequence of a speaker-label error: Out of scope: 2. EVALUATION POPULATION AND SLICES ----------------------------------- Define the deployment population before selecting a headline metric. Every material slice should contain enough independently sourced speech time and speakers to support the decision. Record thin slices as limitations. Total recordings: Total scored speech time: Total distinct speakers: Sessions per speaker: Minimum / median / maximum recording duration: Source and collection period: Required slice table (add rows as needed): Slice | Recordings | Speech time | Speakers | Why it matters | Thin-slice flag ------|------------|-------------|----------|----------------|---------------- All scored audio | | | | Headline only; never the sole decision | Speaker-count band | | | | Cluster-count sensitivity | Overlap band | | | | Simultaneous speech | Backchannel / short-turn band | | | | Brief segments | Channel / codec | | | | Telephony or compression shift | Noise / SNR band | | | | Acoustic robustness | Distance / room | | | | Far-field conditions | Language / locale | | | | Linguistic and accent variation | Device / microphone | | | | Capture-path variation | Recurring vs new speakers | | | | Cross-session behavior | Other product-critical slice | | | | | Excluded audio and reason: Known population gaps: 3. REFERENCE ANNOTATION ----------------------- Reference labels define the answer against which the system or delivery is scored. A metric is not interpretable without the convention used to create those labels. Reference-annotation guideline version: Label schema and speaker-ID format: Turn-boundary convention: Minimum speech / silence duration rules: Overlap representation (multiple active speakers, dominant speaker, other): Backchannel convention: Non-speech and vocal-event convention: Unknown / uncertain speaker convention: Cross-recording speaker identity rule: Timestamp precision: Annotator language and domain requirements: Annotator training and qualification method: Independent annotation or review design: Adjudication process: Reference-set quality metric and result: Disagreement retained for audit: Guideline changes during the project: Reference artifacts to deliver: [ ] Human-readable guideline [ ] Machine-readable schema [ ] Annotated example with difficult cases [ ] Adjudication / disagreement record [ ] Reference manifest and checksums [ ] Version history [ ] Other: 4. METRICS AND SCORING ----------------------- Define the scoring implementation before results are seen. Do not compare two reported diarization error rates unless the overlap, collar, reference, and speaker-mapping rules are the same. Primary metric: [ ] Diarization error rate (DER) [ ] Jaccard error rate (JER) [ ] Speaker-attributed word error rate [ ] Turn-boundary / segmentation metric [ ] Product-specific task metric [ ] Other: DER components to report separately: [ ] Missed speech [ ] False-alarm speech [ ] Speaker confusion Scoring implementation / package and version: Collar around reference boundaries: Overlap included or excluded: Scored and unscored regions: Speaker mapping / permutation rule: Oracle or estimated number of speakers: Treatment of unknown speakers: Micro- or macro-averaging: Confidence interval or resampling method: Statistical comparison method, if used: Always report: [ ] Headline metric [ ] Metric by every product-critical slice [ ] DER component breakdown where DER is used [ ] Speaker-count estimation error [ ] Result with overlap included when overlap matters in deployment [ ] Recording-level distribution, not only an aggregate [ ] Reference-annotation quality result [ ] Failed / unscorable files and reasons [ ] Model, data, code, and configuration versions Do not insert universal thresholds here. Agree thresholds from product risk, reference-label reliability, a baseline, pilot evidence, and the cost of each error type. 5. HELD-OUT DESIGN AND LEAKAGE CONTROL -------------------------------------- Evaluation-set owner: Creation date and freeze date: Access-control list: Storage location and security classification: Separation rules: [ ] No evaluation recording appears in training, tuning, or annotation examples [ ] No segment from the same source recording crosses a split [ ] Speaker-disjoint split where the deployment expects unseen speakers [ ] Session-disjoint split [ ] Prompt / scenario separation where relevant [ ] Source and license separation documented [ ] Deduplication method documented [ ] Evaluation labels hidden from the delivery team until scoring [ ] Repeated evaluation and overfitting policy documented Leakage checks performed: Known overlap with public or vendor data: How a suspected leak is investigated: Rules for replacing contaminated items: 6. PILOT AND BASELINE --------------------- Baseline system or current production result: Baseline configuration and scoring rules: Pilot sample provenance and selection method: Pilot size and slice coverage: Supplier access before acceptance is frozen: Changes allowed after pilot review: Regression tests for previously working conditions: Required error review: [ ] Inspect overlap failures [ ] Inspect short turns and backchannels [ ] Inspect similar-sounding speakers [ ] Inspect far-field / reverberant audio [ ] Inspect channel and codec changes [ ] Inspect speaker-count over- and under-estimation [ ] Inspect recurring-speaker consistency where relevant [ ] Record product impact, not only metric movement 7. ACCEPTANCE AND DELIVERY -------------------------- Complete one row for every requirement. “TBD” is a blocker, not a pass. Requirement | Slice | Metric / check | Threshold | Evidence | Rework rule | Owner ------------|-------|----------------|-----------|----------|-------------|------ Reference labels | All | | | | | Headline diarization | All | | | | | Overlap | | | | | | Short turns | | | | | | Speaker-count estimate | | | | | | Product-critical language / locale | | | | | | Product-critical channel | | | | | | Delivery completeness | All | | | | | Other | | | | | | Delivery contents: [ ] Original or agreed audio files [ ] Reference annotations [ ] System predictions or delivered labels [ ] File and segment manifest [ ] Speaker / session / condition metadata [ ] Guideline and schema versions [ ] Scoring code, configuration, and dependency versions [ ] Slice definitions and result table [ ] Error examples and limitations [ ] Rights / provenance references [ ] Checksums and transfer record Accepted delivery formats: Replacement / rework window: Change-control process: Final acceptance owner and sign-off: 8. RIGHTS, PRIVACY, SECURITY, AND REVIEW --------------------------------------- Intended model and evaluation uses: Source, permission, license, and provenance records required: Applicable notice and lawful-basis record: Consent evidence required when consent is relied on: Voice, likeness, biometric, or cloning permissions required for the use: Restrictions on reuse, redistribution, or model development: Controller / processor roles to be reviewed: Subprocessors and transfer locations: Access controls, encryption, and audit logging: Retention, deletion, withdrawal, and objection handling: Incident and rights-request contacts: Legal and security reviewers: Open evidence gaps: Decision if a required record is missing: 9. FINAL DECISION RECORD ------------------------ Decision: [ ] Accept [ ] Accept with documented limitations [ ] Rework and re-evaluate [ ] Reject Evidence reviewed: Material limitations: Conditions on deployment or reuse: Next evaluation date or trigger: Approvers and date: Related Spirelight buyer resources ---------------------------------- Speaker diarization guide: https://www.spirelight.ai/guides/speaker-diarization Audio annotation service: https://www.spirelight.ai/services/audio-annotation Speech data buyer tools: https://www.spirelight.ai/resources/buyer-tools Custom conversational speech collection: https://www.spirelight.ai/services/full-duplex-conversational-speech-data