If a speech AI model performs poorly, more often than not, the fault lies at an earlier stage, with the data.

We have seen cases of voice assistants being trained on speakers from just one city and then being rolled out across the country. We have had multilingual customer support projects where the transcription guidelines were altered midway through the audio data collection process, necessitating a full-scale review of thousands of audios. We have even encountered datasets which initially seemed sufficient but which proved to lack any variation in terms of recording device or environment. There is nothing wrong with the machine learning process in these cases. The projects failed simply because the data was incomplete.

Our experience has led us to change how we approach speech AI data collection projects. Even before the recording hours, number of contributors, or delivery discussions come up, we take time to understand the product in question. A speech recognition model for a bank's customer support platform will require a different type of audio dataset from a voice assistant for a delivery vehicle or from a multilingual healthcare app.

Sometimes, a single sentence will help to pinpoint what needs to be done. Rather than asking for "1,000 hours of speech", it is better to make sure that the client defines the needed result. For example, "We need conversational mobile-phone recordings from adult Danish and Swedish speakers to improve speech recognition for customer support across the Nordic region" provides us with much more clarity than a mere amount of recordings.

After the product has been specified, the following conversation is always related to the people behind it. Probably, the most frequent mistake we face relates to the confusion between the specification of language and the specification of the speakers. A speech recognition model doesn’t just learn "Danish". It learns how people speak in Danish with various backgrounds, accents, and experience. And the speakers should be considered depending on their background. This includes where they come from, what other languages they speak, and what environments they speak in regularly. If your users live in Copenhagen, Aarhus, Odense, Malmö, or Hamburg, those accents must be taken into account. The more similar speakers are to your users, the more helpful the final speech dataset will be.

This approach applies to the recording conditions too. In most cases, clients want clean audio since they believe it to be a safer option. Sometimes, it is; however, in most cases, it is not. Your users do not use your product in studios. They will most likely interact with it from kitchens, offices, warehouses, train stations, or vehicles. All those places must be included in the list of recording environments. A model trained on studio-quality recordings will produce impressive results in benchmarks and fail to understand the conversations it was supposed to understand. The goal of the recording is not to gather perfect recordings. It is to record the conversations used in real life.

Recording devices are no exception either. Smartphones, laptops, wireless headsets, or call centre headsets record the speech differently. And those differences affect the performance of the model, which is why we consider devices as part of the specification rather than the operation. Also, we often gather metadata about the devices and environments, as those factors explain why the model performs well in case of some groups of users and badly in case of others. This data becomes extremely valuable when working with refinement of ASR training data or troubleshooting production problems.

We also believe that quality assurance should not only be performed after collecting the data. In any case, each custom speech data collection should include a pilot dataset. Evaluating the samples will help to refine recording instructions, contributor recruitment, and transcription guidelines before starting the full collection. By the time the final dataset is delivered, there should be almost no surprises, as the specification will be verified with real recordings.

When choosing a speech data collection service provider, many companies compare their competitors based on recording volume, timeline, and pricing. These aspects are important; however, they are not usually what determines the success of the speech AI model. What really counts is whether the provider understands the product well enough to collect the right data.

This is what we spend most of our time at Spirelight doing. We cooperate with AI teams to define the users, accents, recording environment, devices, transcript standards, metadata, and quality controls before the actual data collection begins. After those decisions are made, the speech data collection turns into a structured process rather than assumptions. See how Spirelight builds custom speech datasets.