Datasets
Spirelight builds speech datasets to order rather than reselling a fixed inventory. Each of the 60 pages below is a collection configuration for one language, with recruitment, recording, transcripts, QA, consent records, and commercial training rights confirmed in writing against your brief. Every page carries a planning rate and a free sample.
Showing 61 of 61 datasets
Javanese · Indonesia
Javanese speech dataset. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Sinhala · Sri Lanka
Sinhala speech dataset. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Igbo · Nigeria
Igbo speech dataset. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Zulu · South Africa
Zulu speech dataset. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Pashto · Afghanistan
Pashto speech dataset. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Burmese · Myanmar
Burmese speech dataset. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Nepali · Nepal
Nepali speech dataset. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Afrikaans · South Africa
Afrikaans speech dataset. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Cantonese · Hong Kong
Cantonese speech dataset. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Khmer · Cambodia
Khmer speech dataset. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Gujarati · India
Gujarati speech dataset: up to 500 hours, 50 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Norwegian · Norway
Norwegian speech dataset: up to 500 hours, 50 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Dutch (Western) · Netherlands
Dutch (Western) speech dataset: up to 1,000 hours, 100 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Mandarin Chinese · Taiwan
Mandarin Chinese speech dataset: up to 1,500 hours, 150 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Marathi · India
Marathi speech dataset: up to 500 hours, 50 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Indonesian · Indonesia
Indonesian speech dataset: up to 2,000 hours, 200 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Malayalam · India
Malayalam speech dataset: up to 500 hours, 50 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Romanian · Romania
Romanian speech dataset: up to 1,000 hours, 100 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
English (Western) · United States
English (Western) speech dataset: up to 1,500 hours, 150 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Urdu · Pakistan
Urdu speech dataset: up to 1,000 hours, 100 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
English (African) · South Africa
English (African) speech dataset: up to 2,000 hours, 200 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
English (Asian) · Singapore
English (Asian) speech dataset: up to 2,000 hours, 200 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Persian (Farsi) · Iran
Persian (Farsi) speech dataset: up to 500 hours, 50 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Finnish · Finland
Finnish speech dataset: up to 500 hours, 50 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
German · Germany
German speech dataset: up to 1,000 hours, 50 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Kannada · India
Kannada speech dataset: up to 500 hours, 50 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Greek · Greece
Greek speech dataset: up to 1,000 hours, 100 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Hindi · India
Hindi speech dataset: up to 2,000 hours, 200 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Thai · Thailand
Thai speech dataset: up to 2,000 hours, 200 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Russian · Russia
Russian speech dataset: up to 1,000 hours, 100 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Turkish · Turkey
Turkish speech dataset: up to 2,000 hours, 200 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Malay · Malaysia
Malay speech dataset: up to 2,000 hours, 200 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Spanish (Western) · Spain
Spanish (Western) speech dataset: up to 1,000 hours, 100 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Portuguese (Western) · Portugal
Portuguese (Western) speech dataset: up to 1,000 hours, 100 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Portuguese (LatAm) · Brazil
Portuguese (LatAm) speech dataset: up to 2,000 hours, 200 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Spanish (LatAm) · Mexico
Spanish (LatAm) speech dataset: up to 2,000 hours, 200 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Hausa · Nigeria
Hausa speech dataset: up to 2,000 hours, 200 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Ukrainian · Ukraine
Ukrainian speech dataset: up to 1,500 hours, 150 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Swedish · Sweden
Swedish speech dataset: up to 500 hours, 50 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
French (Western) · France
French (Western) speech dataset: up to 1,000 hours, 100 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Vietnamese · Vietnam
Vietnamese speech dataset: up to 2,000 hours, 200 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Tagalog · Philippines
Tagalog speech dataset: up to 2,000 hours, 200 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Bengali · Bangladesh
Bengali speech dataset: up to 2,000 hours, 200 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
French (African) · DR Congo
French (African) speech dataset: up to 2,000 hours, 200 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Arabic MSA (Modern) · Saudi Arabia
Arabic MSA (Modern) speech dataset: up to 2,000 hours, 200 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Hebrew · Israel
Hebrew speech dataset: up to 500 hours, 50 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Korean · South Korea
Korean speech dataset: up to 500 hours, 50 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Yoruba · Nigeria
Yoruba speech dataset: up to 2,000 hours, 200 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Polish · Poland
Polish speech dataset: up to 1,000 hours, 100 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Tamil · India
Tamil speech dataset: up to 500 hours, 50 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Swahili · Kenya
Swahili speech dataset: up to 2,000 hours, 200 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Punjabi · India
Punjabi speech dataset: up to 500 hours, 50 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Telugu · India
Telugu speech dataset: up to 500 hours, 50 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Italian · Italy
Italian speech dataset: up to 1,000 hours, 100 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Japanese · Japan
Japanese speech dataset: up to 500 hours, 50 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Arabic (Levantine) · Lebanon
Arabic (Levantine) speech dataset: up to 500 hours, 50 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Arabic (Gulf) · Saudi Arabia
Arabic (Gulf) speech dataset: up to 500 hours, 50 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Arabic (Egyptian) · Egypt
Arabic (Egyptian) speech dataset: up to 2,000 hours, 200 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Arabic (Darija) · Morocco
Arabic (Darija) speech dataset: up to 1,500 hours, 150 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Amharic · Ethiopia
Amharic speech dataset: up to 2,000 hours, 200 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
Danish · Denmark
Danish speech dataset: up to 500 hours, 50 speakers. Recorded to your specification. Speaker profile, formats, rights, schedule, and price are confirmed in the written proposal.
We build custom speech datasets to spec: your language, dialect, recording conditions, and volume. Tell us what you need and we will scope it and send pricing.
If the catalogue does not have the language, dialect, or recording conditions you need, we collect it. Spirelight runs custom speech data collection with a vetted contributor network: scripted or spontaneous speech, your target demographics, devices, and acoustic environments, delivered with verified transcripts and full consent documentation. Bespoke datasets are licensed to you alone by default.
Typical projects include conversational speech for ASR, scripted prompts for wake words and commands, and multilingual corpora for low-resource languages. Tell us your spec and we will scope hours, timeline, and pricing.
Spirelight builds speech datasets to order rather than reselling a fixed inventory. Each catalogue page is a custom collection configuration for one language: recruitment, recording, transcripts, QA, consent records, and commercial training rights are confirmed in writing against your brief. If you need a speech recognition dataset, an ASR evaluation set, or conversational audio in a specific language, request a quote on the matching page and get a free sample.
Where an evidence-verified per-hour rate is published, it is a planning reference for custom collection. The final quote defines volume, speaker profile, recording conditions, QA, delivery format, usage rights, timeline, and any exclusivity requirements.
Exclusivity and restricted-use terms can be scoped per project. Tell us the intended use, required rights, and exclusivity window so they can be documented in the quote and contract.
Yes. The catalogue describes custom collection configurations and planned formats. Any published rate, speaker target, or monthly capacity is shown only with a dated evidence record. Availability, speaker mix, schedule, rights, and delivery criteria are confirmed against your project brief before work starts.
Where a sample is available, use it to review the documented audio format, transcript structure, and annotation approach. The final project specification and acceptance criteria are agreed separately.
The catalogue lists public collection configurations. Language, dialect, location, and speaker feasibility are confirmed for each brief; contact us if your target is not listed.
Each page shows planned formats for the collection configuration. The written proposal confirms the audio container, sample rate, bit depth, channels, transcript or annotation schema, and delivery method for the project.
Use "Request quote" on a collection page and include the language, speaker profile, volume, recording conditions, intended use, and required rights. Spirelight will confirm feasibility, scope, timeline, pricing, and acceptance criteria in writing.
Spirelight sells commissioned collection, not an off-the-shelf download library. Where a relevant sample from consented prior work is available, the team can share it after a verified request so you can inspect audio, transcript, and metadata quality before commissioning. The catalogue pages state availability honestly instead of presenting planned collection as finished stock.
Have a question that is not on this list? Book a call and tell us what you are building.