Custom annotation layer
Emotion labeling for speech, tagged per turn on your taxonomy
Label affect in speech at a resolution a training run can actually use: discrete classes on every turn, dimensional ratings where continuous targets work better, and agreement figures so you know how far to trust each label.
Free sample: Name a language and domain. The team labels a representative excerpt against your taxonomy and sends it back manually within 48 hours, free of charge.
Collection specification
- Discrete labels
- Anger, disgust, fear, happiness, sadness, surprise, and neutral, tagged per utterance or per speaker turn
- Dimensional labels
- Valence and arousal ratings layered on the same segments when continuous targets suit the objective better
- Intent
- Kept separate from affect, delivered as dialogue acts or mapped to a domain taxonomy you define
- Agreement
- Several annotators per batch, with inter-annotator agreement reported per class rather than averaged away
- Source audio
- Your recordings, or spontaneous conversational corpora recorded for the project
- Output
- JSON or CSV keyed to segment timestamps and speaker turns, in a schema agreed before the first batch
- Pricing
- Custom, scoped to your conditions
Emotion is a graded label, so the agreement number ships with it
Words have a ground truth. A transcript is right or it is not, and two good annotators agree almost every time. Affect has no such anchor: one clipped sentence reads as irritation to one listener, fatigue to another, and nothing at all to a third, and none of them is wrong.
That does not make the signal useless, it makes the reporting matter. Every batch is labeled by several annotators and shipped with per-class agreement, so you can weight training on the classes that held up and treat the contested ones with the caution they deserve.
Discrete classes, dimensional ratings, or both on the same segments
Most teams start with discrete classes per utterance or turn, because the categories map cleanly onto a classifier head. Where a model needs something continuous, valence and arousal ratings are layered onto exactly the same segments, so one pass produces both views of the audio and the two never drift apart.
Intent is handled separately. It can arrive as dialogue acts, or mapped into whatever domain taxonomy you define, because teams almost always want intent expressed in their own categories rather than ours.
This offer is not the right fit when
- You need facial or physiological emotion recognition rather than speech.
- You want one emotion label for an entire recording rather than per turn.
- You want labels at the lowest possible price with no agreement reporting.
Get a free labeled sample
- Your details
- Your project
- Verify
Three short steps. The team then follows up manually within 48 hours, and confirms volume, rights, QA, and delivery if you want the scope priced.
Would rather talk it through first? Talk through your label taxonomy
Prepare the brief
Bring a specification your vendors can price consistently
Use the free worksheet to define coverage, rights, delivery fields, held-out rules, and acceptance tests before requesting a collection plan.
Buyer documentation
Specimen data card and provenance structure, evaluation brief template, and acceptance and held-out questions.
Related services
Guides for this decision
Want to talk through your taxonomy first? Category sets, dimensional scales, and agreement thresholds are usually a conversation, not a form field. Book a short call to work through the taxonomy and conditions the labels need to hold up under.