The EU AI Act does not regulate speech data directly. It regulates the systems trained on it, and the obligations flow backward: emotion inference from voice has been banned in workplaces and schools since February 2025, general-purpose model providers have had to summarize their training data publicly since August 2, 2025, and from August 2, 2026 voice bots must disclose they are AI while high-risk systems must document where every hour of training audio came from.

This page is the practical version for teams that buy or commission speech data: the three dates that matter, the penalties, the GDPR layer underneath, the emotion recognition rules for cars and contact centers, and the five documents to demand from any vendor before you sign.

This guide is practical orientation for engineering and procurement teams. It is not legal advice; for decisions with legal consequences, talk to counsel.

What the AI Act actually asks of voice teams

The AI Act regulates AI systems, not datasets. No article tells you how to record speech. But three of its mechanisms reach the training data anyway, and each lands on a different part of a voice product.

Article 5 bans a short list of practices outright, including emotion inference from biometric data in workplace and education contexts. That is a product constraint: some voice features may not exist at all in those settings. The general-purpose AI rules make training data a public matter, because GPAI providers must publish a summary of their training content by source category. And Article 10 requires providers of high-risk systems to run documented data governance: how the data was collected, where it came from, and how it was annotated.

Article 10 is the one that lands on your desk as a data buyer. When your system is high-risk, you have to be able to show the data's paperwork. If your vendor cannot produce it, that gap is now yours.

The timeline that matters for voice

The Act phases in over several years. Three dates carry nearly all the weight for speech products.

DateWhat took effectWhat it means for voice
February 2025Prohibited practices (Article 5)Emotion inference from biometric data banned in workplace and education contexts, with an explicit exception for safety uses such as driver fatigue detection.
August 2, 2025General-purpose AI obligationsGPAI providers must publicly summarize training content by source category: licensed, scraped, user-provided, synthetic.
August 2, 2026Article 50 transparency and the bulk of high-risk obligationsVoice bots must clearly disclose they are AI at the start of an interaction. Article 10 data governance applies to high-risk systems: documented collection processes, data origin, and annotation procedures.

The February 2025 tier came first because it covers the uses the legislators considered unacceptable rather than merely risky. For voice, the headline item is the Article 5 ban on emotion inference from biometric data in workplace and education contexts. The safety carve-out is explicit and matters for automotive: driver fatigue detection is named as the kind of use the ban does not touch.

The August 2, 2025 milestone changed the economics of scraped audio. The training-content summary template forces general-purpose model providers to state publicly which of their training data was licensed, which was scraped, which came from users, and which is synthetic. Provenance stopped being an internal detail on that date.

August 2, 2026 is the date most voice teams feel directly. Article 50 disclosure becomes enforceable: a voice bot has to say it is AI, clearly, at the start of the interaction, not in a terms page. And the bulk of high-risk obligations under Article 10 apply, which is where the documentation duties concentrate. The full text of both is on the AI Act Explorer: Article 50 and Article 10.

Penalties, briefly

Prohibited-practice violations carry fines of up to EUR 35 million or 7% of global turnover. Most other obligations, including the Article 50 transparency duties, carry up to EUR 15 million or 3%.

In practice the fine is rarely the first cost. The first cost is diligence: enterprise customers and their counsel now ask where your training data came from, and a missing answer stalls a deal long before a regulator calls.

The GDPR layer underneath

The AI Act did not replace GDPR; it sits on top of it. Voice recordings are personal data in essentially every case, because a voice can identify a person. When a system actually uses voice to identify someone, the recording becomes biometric special-category data under GDPR Article 9, and processing it requires explicit consent.

Purpose limitation is the rule that catches teams by surprise. A call recording made for quality assurance was collected for quality assurance, and it cannot simply be repurposed as training data just because the tape exists. The lawful basis has to cover training, which in practice means consent that names AI training. The same logic is why many open corpora are unsafe for commercial models: the terms under which they were recorded never mentioned training at all. Our guide to free speech datasets for commercial use covers which corpora actually clear that bar.

Emotion recognition: two verticals to get right

In-cabin voice. Driver fatigue detection from voice is the textbook safety exception: it is explicitly the kind of use the Article 5 ban does not cover. Empathetic-companion features are the careful case. An assistant that adapts its tone to the driver's mood is not a safety system, and for professional drivers the cabin is a workplace, so the feature needs a harder look than the fatigue monitor sitting next to it. The broader collection questions for that vertical live in our automotive voice data guide.

Contact centers. The line runs between the two sides of the call. Emotion inference on customers is high-risk: permitted, but with the full Article 10 documentation load attached. Emotion inference on employees is prohibited, because a contact center is a workplace. A dashboard that scores agents' emotional state crosses the Article 5 line; a system that scores customer sentiment for routing lives in high-risk territory with paperwork attached.

The five documents to ask any speech data vendor for

Article 10 compliance is mostly a documentation problem, and the documentation has to originate where the data did. You cannot reconstruct consent records after the fact. Before buying, ask for:

  1. License chain. Who granted what rights, from the speaker to the vendor to you. Our speech data licensing guide covers the terms worth checking.
  2. Consent records. What the speakers agreed to, and whether AI training was named explicitly rather than folded into general terms.
  3. Collection process documentation. How and where the audio was recorded, and under what instructions.
  4. Annotation provenance. Who labeled the data, and under what quality assurance process.
  5. Per-dataset statistical properties. Languages, speaker demographics, and recording conditions, so you can assess whether the data fits your deployment population.

A vendor that cannot produce these is not selling you data at a discount. It is transferring the risk to you, because under Article 10 the documentation duty sits with the provider of the high-risk system, which is you.

How Spirelight handles this

Spirelight collects speech data in the EU, under GDPR, from day one: the company is based in Denmark, and consent, contracting, and storage were built for European law rather than adapted to it afterwards.

Documentation varies by dataset and custom scope. Before buying, ask which license, consent, collection, annotation, QA, and statistical records are available for the specific delivery; for custom work, put the required artifacts into the written brief and quote. See our speech data collection process or browse the catalogue, then verify the records for the item you are evaluating.

When this is not your problem

Honesty requires the reverse case. If you are building an internal transcription tool that never identifies, scores, or decides anything about a person, your system is probably not high-risk, and most of the Article 10 machinery is not aimed at you. Buying a documentation pack you do not need is spending money on comfort.

Two things apply regardless. GDPR covers voice recordings of EU speakers no matter what the system does with them. And Article 50 attaches the moment your bot talks to users, because disclosure follows the interaction, not the risk class.