Gustav Aggeboe
Writer

Gustav Aggeboe

Platform architecture and the engineering behind speech data.

Gustav leads the platform architecture behind Spirelight. He builds the systems used to manage recording, transcription, QA, metadata, contributor workflows, and dataset delivery. He writes about the engineering behind speech data: pipelines, data quality, evaluation, and the practical details of turning raw audio into training-ready datasets.

๐Ÿ“ Copenhagen, DenmarkWebsite

Posts by Gustav Aggeboe

For Developers

Scripted vs Spontaneous Speech: Which One Your Model Needs

Scripted vs spontaneous speech data isn't a style preference, it's a decision your target audio should make for you. Here's the actual test, the measurable gap between the two, and what training on the wrong register breaks downstream.

July 30, 2026
For Developers

Where Speaker Labels Break: Diarization Failure Cases to Plan For

Diarization looks solid in the demo, then breaks on real audio: overlapping speech, similar voices, short turns, channel changes, crosstalk. Here's what each failure mode actually does downstream, and how to catch it early.

July 30, 2026
For Developers

How Much Call History You Need Before Speech Analytics Is Reliable

Teams ask how many hours of call history they need before speech analytics is reliable. The honest answer: hours are the wrong axis. Here's what actually determines sufficiency, and how to measure it yourself.

July 30, 2026
For Developers

Speaker Diarization: Splitting Audio by Speaker at Scale

A working engineer's tour of speaker diarization: segmentation, embeddings, EEND, overlap, and how Diarization Error Rate behaves once you push it through a real pipeline.

June 21, 2026
For Developers

Audio Data Augmentation for Speech Models: A Practical Guide

A hands-on look at the augmentation techniques that make speech models more robust, which ones risk corrupting your transcripts, and a default recipe that holds up at scale.

June 12, 2026
For Developers

Word error rate: what WER really measures, and how to cut it

Word error rate is the default ASR accuracy metric, but the number means nothing without context. How WER is calculated, what counts as good, and what really moves it.

June 9, 2026