An AI training data company is the supplier that stands between your model and the real-world data it needs to learn from. Some sell finished datasets off a shelf. Others recruit people, then record or label fresh data to your spec and hand back something built for your deployment. Most sit somewhere in between, and the label covers a wider range of businesses than the phrase suggests.

This guide opens with a public-evidence comparison of speech data providers, maps the wider market by category, then explains what these companies actually do, the main types of AI data services on the market, and how to evaluate one before you sign. The examples lean toward speech and voice. Any provider name is a shortlist lead only; verify current capabilities, ownership, evidence, and commercial terms directly.

Buyers comparing AI training data services usually want the list first, so this guide opens with the evidence: a comparison of five speech data providers built only from what their own public pages state, with every claim linked to its source.

Speech data providers compared: public evidence for buyers (2026)

Disclosure: Spirelight publishes this comparison and appears in it. To keep it useful anyway, every cell records only what each provider's own public pages stated when we checked them on August 2, 2026, with the page linked so you can verify it yourself. Providers are listed alphabetically, and there is no ranking. "Not publicly documented" means we did not find the evidence on the pages we checked; it does not mean the capability is missing, and the right response is to ask that vendor for the artifact. All scale, certification, and coverage figures are vendor-stated, not independently audited.

Public evidenceDataForceDefined.aiLXTShaipSpirelight
Custom speech collectionDocumented: scripted and conversational speech, wake-words, multi-speaker dialogues; studios and in-the-wild consumer devicesDocumented: conversational dialogues, IVR interactions, studio emotional recordingsDocumented: scripted, conversational, expressive, and environmental audio in remote, on-site, and studio settingsDocumented: monologue, dialogue, multi-party, wake-word, acoustic, ASR, TTS, and call-center collectionDocumented: remote, moderated, on-site, and studio recording scoped per project brief
Ready-made audio catalogueNot publicly documented on the pages checked; the positioning is custom collectionPublic marketplace listing 666 audio datasets with item-level specs; no prices shownOff-the-shelf section advertised, but the category pages describe built-to-spec datasets; no item-level inventory shownBrowsable catalogue with per-language detail pages; stated size varies by page (55k+ vs 70,000+ hours, vendor-stated)Collection configurations for about 60 languages, explicitly labeled as custom targets rather than finished inventory
Audio annotation and transcriptionDocumented: transcription with noise, speaker, and event tags, speaker tracking, timestampsDocumented: ASR transcript annotation, human transcription, diarization, intent and sentiment taggingDocumented: diarization, event timestamping, classification, linguistic annotation, transcriptionDocumented: transcription with word-level timestamps, labeling, diarization, phonetic transcriptionDocumented: transcription, anonymous speaker turns, timestamps, event, intent, and emotion labels to a buyer schema
Public license termsNot publicly documentedFull data license agreement published: a non-exclusive, non-sublicensable internal-use grant covering training, testing, and benchmarkingNot publicly documentedNot publicly documented; licensing described as flexible, terms set per purchaseNot publicly documented; license terms are defined in the project contract
Consent and provenance documentationVendor statement: contributors opt in and sign data-use agreements; automated PII scrubbingGovernance page: vendor-stated 100% contributor consent and per-dataset provenance recordsBrief vendor statements: protocols vetted for consent and privacy; rights-cleared supplier networkVendor statements: consent documentation and provenance metadata included with licensed setsPer project: the written scope defines consent language, permitted uses, and provenance records before collection begins
Public pricing signalsNot publicly documentedNot publicly documented; the route is quote and sample requestsWorked range published: vendor-stated $15K for focused 50-200 hour sets up to $150K+ beyond 1,000 hoursNot publicly documentedNo numeric rates were visible on the pages checked; planning figures appear only where a dated evidence record exists, and the final price is quoted in writing
Sample accessNot publicly documented on the pages checkedOn request: "Request free samples and quotes"; no public playback foundNot publicly documentedOn request via the licensing page; no public playback foundOn request: an admin confirms per request whether an applicable example may be shared
Named speech case studiesVerint (contact-center transcription, 30 projects since 2021, vendor-stated), iGenius (TTS voice recordings), and Speechmatics (ASR text normalization)Omilia and Theseus AI named; further speech cases anonymizedDubber named (voice AI data across 10 languages, vendor-stated); other speech cases anonymizedNot publicly documented: speech case studies exist but customers are anonymizedNot publicly documented yet
Security and compliance claims (vendor-stated)SSAE 16 SOC 2, ISO 27001, ISO 9001, HIPAA, and GDPR statements; certificate scope not broken out on the pages checkedISO 27001, ISO 27701, and ISO 42001 plus GDPR and HIPAA statementsFive ISO 27001-certified facilities plus GDPR and HIPAA statementsISO 9001, ISO 27001, SOC 2, and HIPAA badges plus GDPR statementsNo certification claims published; EU data-protection footing described on the privacy page
Language coverage claim (vendor-stated)"Over 250 languages and dialects" on one page, "over 200 languages" on another"500+ languages, dialects, and locales""1,000+ language locales" sitewide; "100+ languages" for the speech dataset offerVaries by page, from "50+ languages/100+ dialects" to "150+ languages"Contributor recruitment across 50+ languages and 30+ markets; per-locale feasibility confirmed per project
EU footing and GDPR positioningGDPR compliance stated; part of TransPerfect; no EU residency claim found on the pages checkedUS company founded in Seattle, offices in Lisbon and Porto; GDPR and EU AI Act readiness statementsCanada headquarters with presence in Germany and Romania; GDPR statementsPart of Ubiquity Global Services (New York) since February 2026; UK office; claim-level GDPR statementsDanish company (Spirelight ApS, CVR DK44262509) in Valby, Denmark; complaints route to the Danish Data Protection Agency

How to read this table: an empty public record is a screening prompt, not a verdict. Where a vendor's own pages disagree, for example on catalogue hours or language totals, ask which figure applies to your brief and get it in writing. Where everything is documented, the next question is whether the artifact survives contact with your specification: request the license text, a matched sample, and the consent wording rather than the marketing page. The same eleven fields map onto the free speech data provider scorecard, which turns them into a weighted, evidence-scored comparison against one brief.

Every row was checked on August 2, 2026 against the linked official pages. Provider pages change; if a row is out of date, tell us what changed and we will recheck the source and correct it.

The wider market: AI training data providers by category

The comparison above goes deep on five speech data providers. The wider market of AI training data providers is much larger, and many buyers run a speech vendor search alongside a broader review of AI data services and AI data solutions such as annotation, evaluation, and human feedback. The list below is alphabetical, is not a ranking or a recommendation, and describes each company in terms drawn from its own public self-description, checked on August 8, 2026. Spirelight publishes this page and appears in the list.

  • Appen: a training data provider that describes its offer as expert-validated data for training frontier AI models.
  • Cogito Tech: describes itself as providing custom data curation and labeling services for computer vision, NLP, and generative AI models.
  • DataForce: a TransPerfect company that describes itself as delivering multimodal training data and services across language, voice, image, and video.
  • Defined.ai: describes itself as an AI data marketplace where enterprises buy or commission training data across modalities.
  • FutureBee AI: describes itself as offering pre-labeled datasets and custom AI data solutions across speech, text, image, and video.
  • iMerit: describes itself as providing expert-led model evaluation and training data solutions.
  • Prolific: a participant platform that describes itself as a way to collect high-quality data from real people.
  • Sama: describes itself as delivering human-verified annotation, validation, and evaluation at scale.
  • Scale AI: describes itself as working across the AI stack, from the data that trains models to the systems that put them to work.
  • Shaip: describes itself as collecting, annotating, licensing, and evaluating multimodal data, including conversational AI.
  • Sigma.AI: describes itself as providing independent evaluation, high-fidelity data, and human intelligence for teams building AI systems.
  • Spirelight: a Danish speech data company offering custom speech data collection, audio annotation, and per-project speech dataset configuration. Spirelight publishes this page.
  • Surge AI: describes itself as providing human intelligence for AGI development.
  • TELUS Digital: describes itself as offering frontier AI training data, engineering, and intelligent automation alongside broader digital services.
  • Toloka: describes itself as building data solutions that integrate human expertise and technology to accelerate AI development.

Inclusion in this list is not a capability, quality, or fit claim, and no provider here is ranked or recommended. If a description is out of date, tell us what changed and we will recheck the source and correct it.

What AI training data companies actually do

Strip away the marketing and an AI training data company does some slice of one pipeline: find the right data, collect or create it, label it, check it, and deliver it under a license you can build on. Where a vendor sits on that pipeline is the first thing worth understanding, because two businesses that both call themselves a training data service can do almost no overlapping work.

At the front of the pipeline is sourcing and collection. For text and images that can mean licensing existing material or filtering what already exists. For speech it means recruiting real people across languages and accents and recording them under defined conditions, which is closer to field operations than to a catalogue lookup. Next comes annotation: transcription, labeling, segmentation, and whatever schema your model needs. Then quality control, where reviewers catch errors, measure agreement, and decide what ships. Finally delivery and licensing, which fix the formats you receive and the rights you get to train, evaluate, and ship models on the data. If the terms are new to you, our primer on what speech data is sets the baseline.

The main types of AI data services

Companies in this space differ along three axes. Knowing where a vendor lands on each tells you more than any capability list.

  • Generalist crowd versus domain specialist. A generalist handles many data types, text, image, video, audio, and search relevance, through a large and broad crowd. A specialist concentrates on one domain, such as speech, and builds its crowd, tooling, and reviewers around it. Provider category does not predict fit; compare evidence for the exact brief.
  • Marketplace versus commissioned collection. A marketplace sells existing datasets you license as-is, which is fast when something fits. Commissioned collection means the vendor recruits, records, and labels to your spec, so you get data shaped for your deployment rather than someone else's.
  • Annotation-only versus end-to-end. Some companies only label data you already have. Others run the whole chain, from finding speakers to delivering a licensed, quality-checked corpus. The right choice depends on how much of the pipeline you want to own.

Most real engagements mix these. You might license a ready-made set for the common part of your problem and commission a specialist for the part no catalogue covers. Our guide to buying AI training data walks through that build, buy, or commission decision in more detail.

How to evaluate an AI training data company

Once you have a shortlist, the differences that matter rarely appear in a sales deck. Press on these.

Data quality and QA. Ask who checks the work, what the annotation guidelines are, how inter-annotator agreement is measured, and what the error rate looks like after review. If a supplier cannot describe its QA process and provide evidence, treat the quality claim as unverified. For speech specifically, our notes on speech data quality list the metrics worth requesting.

Consent and licensing. Request source and permission records, the license for the intended use, applicable privacy notices and lawful basis, and consent records when consent is relied on. Compare redistribution, model-related rights, retention, transfer, exclusivity, and risk allocation in the contract. Our guide to speech data licensing covers what those clauses mean in practice.

Language and dialect coverage. A large dataset can still miss the speakers and accents you deploy into. Ask for the distribution, not just the headline hours: how many speakers, which dialects, what age and gender spread, what recording environments. Do not infer coverage from supplier category; require speaker-cell distributions, screening method, feasibility assumptions, and sample evidence for the target brief.

Security and data handling. Confirm how the vendor stores and transfers data, who can access it, whether personal data is minimized or de-identified where required, and how the arrangement maps to the regulations you answer to. For regulated buyers this is not optional.

Ability to collect custom data. The clearest divide among AI data services is whether a company can only sell what it already has or can go out and build what you need. For a narrow use case, ask whether a representative pilot can be scoped, which success threshold controls expansion, and what capacity and schedule assumptions apply.

Build a current supplier map

Provider ownership, services, and delivery models change. Build the shortlist from current official materials and dated evidence, then send one specification to every candidate. Names buyers may encounter include Appen, Defined.ai, Shaip, TELUS Digital, Sama, Sigma.ai, Summa Linguae, Way With Words, and Spirelight; inclusion here is not a capability, quality, or fit claim.

Evidence fieldWhat to verify
Delivery modelFinished license, custom collection, annotation service, software platform, or a combination
Modality and languageNamed sourcing method, screening, sample evidence, current capacity, and relevant delivery history
QAGuidelines, reviewer qualifications, agreement metrics, sampling, rework, and acceptance thresholds
Rights and privacySource chain, permissions, license scope, applicable lawful basis and notices, retention, transfers, and consent records when relied on
Commercial scopeInventory status, minimum, schedule, included work, sample availability, price, and change control

Where a speech-focused proposal may fit

Supplier category does not establish capability. Compare each shortlisted provider's documented language access, recording method, QA, capacity, sample evidence, subcontractors, and delivery history against the brief.

Spirelight can assess recording, metadata capture, transcription, annotation, and quality checks as a proposed speech-data workflow. Collection configuration pages describe possible targets and evidence-gated planning inputs; they do not establish finished inventory, capacity, sample availability, rights, schedule, or final price.

For a custom brief, provide the language, dialect, recording conditions, speaker profiles, volume, intended use, and acceptance criteria. Feasibility, pilot design, capacity, deliverables, rights, schedule, sample availability, and price are confirmed in writing on the data services page.