Platform architecture and the engineering behind speech data.
Gustav leads the platform architecture behind Spirelight. He builds the systems used to manage recording, transcription, QA, metadata, contributor workflows, and dataset delivery. He writes about the engineering behind speech data: pipelines, data quality, evaluation, and the practical details of turning raw audio into training-ready datasets.
Scripted vs spontaneous speech data isn't a style preference, it's a decision your target audio should make for you. Here's the actual test, the measurable gap between the two, and what training on the wrong register breaks downstream.
Diarization looks solid in the demo, then breaks on real audio: overlapping speech, similar voices, short turns, channel changes, crosstalk. Here's what each failure mode actually does downstream, and how to catch it early.
Teams ask how many hours of call history they need before speech analytics is reliable. The honest answer: hours are the wrong axis. Here's what actually determines sufficiency, and how to measure it yourself.
A working engineer's tour of speaker diarization: segmentation, embeddings, EEND, overlap, and how Diarization Error Rate behaves once you push it through a real pipeline.
A hands-on look at the augmentation techniques that make speech models more robust, which ones risk corrupting your transcripts, and a default recipe that holds up at scale.
Word error rate is the default ASR accuracy metric, but the number means nothing without context. How WER is calculated, what counts as good, and what really moves it.