Platform architecture and the engineering behind speech data.
Gustav leads the platform architecture behind Spirelight. He builds the systems used to manage recording, transcription, QA, metadata, contributor workflows, and dataset delivery. He writes about the engineering behind speech data: pipelines, data quality, evaluation, and the practical details of turning raw audio into training-ready datasets.
A working engineer's tour of speaker diarization: segmentation, embeddings, EEND, overlap, and how Diarization Error Rate behaves once you push it through a real pipeline.
A hands-on look at the augmentation techniques that make speech models more robust, which ones risk corrupting your transcripts, and a default recipe that holds up at scale.
Word error rate is the default ASR accuracy metric, but the number means nothing without context. How WER is calculated, what counts as good, and what really moves it.