An AI training data company is the supplier that stands between your model and the real-world data it needs to learn from. Some sell finished datasets off a shelf. Others recruit people, then record or label fresh data to your spec and hand back something built for your deployment. Most sit somewhere in between, and the label covers a wider range of businesses than the phrase suggests.
This guide opens with a public-evidence comparison of speech data providers, maps the wider market by category, then explains what these companies actually do, the main types of AI data services on the market, and how to evaluate one before you sign. The examples lean toward speech and voice. Any provider name is a shortlist lead only; verify current capabilities, ownership, evidence, and commercial terms directly.
Buyers comparing AI training data services usually want the list first, so this guide opens with the evidence: a comparison of five speech data providers built only from what their own public pages state, with every claim linked to its source.
Speech data providers compared: public evidence for buyers (2026)
Disclosure: Spirelight publishes this comparison and appears in it. To keep it useful anyway, every cell records only what each provider's own public pages stated when we checked them on August 2, 2026, with the page linked so you can verify it yourself. Providers are listed alphabetically, and there is no ranking. "Not publicly documented" means we did not find the evidence on the pages we checked; it does not mean the capability is missing, and the right response is to ask that vendor for the artifact. All scale, certification, and coverage figures are vendor-stated, not independently audited.
| Public evidence | DataForce | Defined.ai | LXT | Shaip | Spirelight |
|---|---|---|---|---|---|
| Custom speech collection | Documented: scripted and conversational speech, wake-words, multi-speaker dialogues; studios and in-the-wild consumer devices | Documented: conversational dialogues, IVR interactions, studio emotional recordings | Documented: scripted, conversational, expressive, and environmental audio in remote, on-site, and studio settings | Documented: monologue, dialogue, multi-party, wake-word, acoustic, ASR, TTS, and call-center collection | Documented: remote, moderated, on-site, and studio recording scoped per project brief |
| Ready-made audio catalogue | Not publicly documented on the pages checked; the positioning is custom collection | Public marketplace listing 666 audio datasets with item-level specs; no prices shown | Off-the-shelf section advertised, but the category pages describe built-to-spec datasets; no item-level inventory shown | Browsable catalogue with per-language detail pages; stated size varies by page (55k+ vs 70,000+ hours, vendor-stated) | Collection configurations for about 60 languages, explicitly labeled as custom targets rather than finished inventory |
| Audio annotation and transcription | Documented: transcription with noise, speaker, and event tags, speaker tracking, timestamps | Documented: ASR transcript annotation, human transcription, diarization, intent and sentiment tagging | Documented: diarization, event timestamping, classification, linguistic annotation, transcription | Documented: transcription with word-level timestamps, labeling, diarization, phonetic transcription | Documented: transcription, anonymous speaker turns, timestamps, event, intent, and emotion labels to a buyer schema |
| Public license terms | Not publicly documented | Full data license agreement published: a non-exclusive, non-sublicensable internal-use grant covering training, testing, and benchmarking | Not publicly documented | Not publicly documented; licensing described as flexible, terms set per purchase | Not publicly documented; license terms are defined in the project contract |
| Consent and provenance documentation | Vendor statement: contributors opt in and sign data-use agreements; automated PII scrubbing | Governance page: vendor-stated 100% contributor consent and per-dataset provenance records | Brief vendor statements: protocols vetted for consent and privacy; rights-cleared supplier network | Vendor statements: consent documentation and provenance metadata included with licensed sets | Per project: the written scope defines consent language, permitted uses, and provenance records before collection begins |
| Public pricing signals | Not publicly documented | Not publicly documented; the route is quote and sample requests | Worked range published: vendor-stated $15K for focused 50-200 hour sets up to $150K+ beyond 1,000 hours | Not publicly documented | No numeric rates were visible on the pages checked; planning figures appear only where a dated evidence record exists, and the final price is quoted in writing |
| Sample access | Not publicly documented on the pages checked | On request: "Request free samples and quotes"; no public playback found | Not publicly documented | On request via the licensing page; no public playback found | On request: an admin confirms per request whether an applicable example may be shared |
| Named speech case studies | Verint (contact-center transcription, 30 projects since 2021, vendor-stated), iGenius (TTS voice recordings), and Speechmatics (ASR text normalization) | Omilia and Theseus AI named; further speech cases anonymized | Dubber named (voice AI data across 10 languages, vendor-stated); other speech cases anonymized | Not publicly documented: speech case studies exist but customers are anonymized | Not publicly documented yet |
| Security and compliance claims (vendor-stated) | SSAE 16 SOC 2, ISO 27001, ISO 9001, HIPAA, and GDPR statements; certificate scope not broken out on the pages checked | ISO 27001, ISO 27701, and ISO 42001 plus GDPR and HIPAA statements | Five ISO 27001-certified facilities plus GDPR and HIPAA statements | ISO 9001, ISO 27001, SOC 2, and HIPAA badges plus GDPR statements | No certification claims published; EU data-protection footing described on the privacy page |
| Language coverage claim (vendor-stated) | "Over 250 languages and dialects" on one page, "over 200 languages" on another | "500+ languages, dialects, and locales" | "1,000+ language locales" sitewide; "100+ languages" for the speech dataset offer | Varies by page, from "50+ languages/100+ dialects" to "150+ languages" | Contributor recruitment across 50+ languages and 30+ markets; per-locale feasibility confirmed per project |
| EU footing and GDPR positioning | GDPR compliance stated; part of TransPerfect; no EU residency claim found on the pages checked | US company founded in Seattle, offices in Lisbon and Porto; GDPR and EU AI Act readiness statements | Canada headquarters with presence in Germany and Romania; GDPR statements | Part of Ubiquity Global Services (New York) since February 2026; UK office; claim-level GDPR statements | Danish company (Spirelight ApS, CVR DK44262509) in Valby, Denmark; complaints route to the Danish Data Protection Agency |
How to read this table: an empty public record is a screening prompt, not a verdict. Where a vendor's own pages disagree, for example on catalogue hours or language totals, ask which figure applies to your brief and get it in writing. Where everything is documented, the next question is whether the artifact survives contact with your specification: request the license text, a matched sample, and the consent wording rather than the marketing page. The same eleven fields map onto the free speech data provider scorecard, which turns them into a weighted, evidence-scored comparison against one brief.
Every row was checked on August 2, 2026 against the linked official pages. Provider pages change; if a row is out of date, tell us what changed and we will recheck the source and correct it.
The wider market: AI training data providers by category
The comparison above goes deep on five speech data providers. The wider market of AI training data providers is much larger, and many buyers run a speech vendor search alongside a broader review of AI data services and AI data solutions such as annotation, evaluation, and human feedback. The list below is alphabetical, is not a ranking or a recommendation, and describes each company in terms drawn from its own public self-description, checked on August 8, 2026. Spirelight publishes this page and appears in the list.
- Appen: a training data provider that describes its offer as expert-validated data for training frontier AI models.
- Cogito Tech: describes itself as providing custom data curation and labeling services for computer vision, NLP, and generative AI models.
- DataForce: a TransPerfect company that describes itself as delivering multimodal training data and services across language, voice, image, and video.
- Defined.ai: describes itself as an AI data marketplace where enterprises buy or commission training data across modalities.
- FutureBee AI: describes itself as offering pre-labeled datasets and custom AI data solutions across speech, text, image, and video.
- iMerit: describes itself as providing expert-led model evaluation and training data solutions.
- Prolific: a participant platform that describes itself as a way to collect high-quality data from real people.
- Sama: describes itself as delivering human-verified annotation, validation, and evaluation at scale.
- Scale AI: describes itself as working across the AI stack, from the data that trains models to the systems that put them to work.
- Shaip: describes itself as collecting, annotating, licensing, and evaluating multimodal data, including conversational AI.
- Sigma.AI: describes itself as providing independent evaluation, high-fidelity data, and human intelligence for teams building AI systems.
- Spirelight: a Danish speech data company offering custom speech data collection, audio annotation, and per-project speech dataset configuration. Spirelight publishes this page.
- Surge AI: describes itself as providing human intelligence for AGI development.
- TELUS Digital: describes itself as offering frontier AI training data, engineering, and intelligent automation alongside broader digital services.
- Toloka: describes itself as building data solutions that integrate human expertise and technology to accelerate AI development.
Inclusion in this list is not a capability, quality, or fit claim, and no provider here is ranked or recommended. If a description is out of date, tell us what changed and we will recheck the source and correct it.
What AI training data companies actually do
Strip away the marketing and an AI training data company does some slice of one pipeline: find the right data, collect or create it, label it, check it, and deliver it under a license you can build on. Where a vendor sits on that pipeline is the first thing worth understanding, because two businesses that both call themselves a training data service can do almost no overlapping work.
At the front of the pipeline is sourcing and collection. For text and images that can mean licensing existing material or filtering what already exists. For speech it means recruiting real people across languages and accents and recording them under defined conditions, which is closer to field operations than to a catalogue lookup. Next comes annotation: transcription, labeling, segmentation, and whatever schema your model needs. Then quality control, where reviewers catch errors, measure agreement, and decide what ships. Finally delivery and licensing, which fix the formats you receive and the rights you get to train, evaluate, and ship models on the data. If the terms are new to you, our primer on what speech data is sets the baseline.
The main types of AI data services
Companies in this space differ along three axes. Knowing where a vendor lands on each tells you more than any capability list.
- Generalist crowd versus domain specialist. A generalist handles many data types, text, image, video, audio, and search relevance, through a large and broad crowd. A specialist concentrates on one domain, such as speech, and builds its crowd, tooling, and reviewers around it. Provider category does not predict fit; compare evidence for the exact brief.
- Marketplace versus commissioned collection. A marketplace sells existing datasets you license as-is, which is fast when something fits. Commissioned collection means the vendor recruits, records, and labels to your spec, so you get data shaped for your deployment rather than someone else's.
- Annotation-only versus end-to-end. Some companies only label data you already have. Others run the whole chain, from finding speakers to delivering a licensed, quality-checked corpus. The right choice depends on how much of the pipeline you want to own.
Most real engagements mix these. You might license a ready-made set for the common part of your problem and commission a specialist for the part no catalogue covers. Our guide to buying AI training data walks through that build, buy, or commission decision in more detail.
How to evaluate an AI training data company
Once you have a shortlist, the differences that matter rarely appear in a sales deck. Press on these.
Data quality and QA. Ask who checks the work, what the annotation guidelines are, how inter-annotator agreement is measured, and what the error rate looks like after review. If a supplier cannot describe its QA process and provide evidence, treat the quality claim as unverified. For speech specifically, our notes on speech data quality list the metrics worth requesting.
Consent and licensing. Request source and permission records, the license for the intended use, applicable privacy notices and lawful basis, and consent records when consent is relied on. Compare redistribution, model-related rights, retention, transfer, exclusivity, and risk allocation in the contract. Our guide to speech data licensing covers what those clauses mean in practice.
Language and dialect coverage. A large dataset can still miss the speakers and accents you deploy into. Ask for the distribution, not just the headline hours: how many speakers, which dialects, what age and gender spread, what recording environments. Do not infer coverage from supplier category; require speaker-cell distributions, screening method, feasibility assumptions, and sample evidence for the target brief.
Security and data handling. Confirm how the vendor stores and transfers data, who can access it, whether personal data is minimized or de-identified where required, and how the arrangement maps to the regulations you answer to. For regulated buyers this is not optional.
Ability to collect custom data. The clearest divide among AI data services is whether a company can only sell what it already has or can go out and build what you need. For a narrow use case, ask whether a representative pilot can be scoped, which success threshold controls expansion, and what capacity and schedule assumptions apply.
Build a current supplier map
Provider ownership, services, and delivery models change. Build the shortlist from current official materials and dated evidence, then send one specification to every candidate. Names buyers may encounter include Appen, Defined.ai, Shaip, TELUS Digital, Sama, Sigma.ai, Summa Linguae, Way With Words, and Spirelight; inclusion here is not a capability, quality, or fit claim.
| Evidence field | What to verify |
|---|---|
| Delivery model | Finished license, custom collection, annotation service, software platform, or a combination |
| Modality and language | Named sourcing method, screening, sample evidence, current capacity, and relevant delivery history |
| QA | Guidelines, reviewer qualifications, agreement metrics, sampling, rework, and acceptance thresholds |
| Rights and privacy | Source chain, permissions, license scope, applicable lawful basis and notices, retention, transfers, and consent records when relied on |
| Commercial scope | Inventory status, minimum, schedule, included work, sample availability, price, and change control |
Where a speech-focused proposal may fit
Supplier category does not establish capability. Compare each shortlisted provider's documented language access, recording method, QA, capacity, sample evidence, subcontractors, and delivery history against the brief.
Spirelight can assess recording, metadata capture, transcription, annotation, and quality checks as a proposed speech-data workflow. Collection configuration pages describe possible targets and evidence-gated planning inputs; they do not establish finished inventory, capacity, sample availability, rights, schedule, or final price.
For a custom brief, provide the language, dialect, recording conditions, speaker profiles, volume, intended use, and acceptance criteria. Feasibility, pilot design, capacity, deliverables, rights, schedule, sample availability, and price are confirmed in writing on the data services page.