Gig
15
40
Apr 30, 2026
**About the role**
We are hiring an Audio AI Specialist to help us build high-quality training data for next-generation voice AI models. This is a hands-on role for someone with direct experience in audio AI training and evaluation across areas like text-to-speech, transcription, speech-to-speech, ASR, and conversational voice systems.
You will work across three core areas:
1. defining and applying audio quality standards,
2. recording high-quality speech on demand, and
3. performing annotation and QA tasks across speech datasets.
This is not a generic audio production role. We are looking for someone who understands what makes audio actually useful for model training.
**What you’ll do**
* Build, refine, and consistently apply audio quality guidelines for speech and voice datasets
* Review audio files against technical, linguistic, and task-specific standards, and make clear approval, rejection, or revision decisions
* Identify issues such as background noise, clipping, distortion, plosives, echo, low signal, channel imbalance, unnatural silence, overlap, segmentation errors, transcript mismatches, and speaker-label errors
* Perform annotation and QA tasks including transcription, timestamping, VAD/segmentation, diarization, pronunciation checks, and metadata validation
* Record speech on demand based on provided scripts and performance guidelines, delivering natural, clean, spec-compliant audio
* Document edge cases, update review rubrics, and improve internal SOPs and examples
* Partner with research, ML, and operations teams to turn model needs into data specifications, quality thresholds, and evaluation criteria
* Help maintain consistency and integrity across raw audio, transcripts, annotations, and metadata
**Must-have qualifications**
* Direct experience in audio AI training is required
* Hands-on experience working on datasets or evaluation workflows for TTS, transcription, ASR, speech-to-speech, or related voice AI systems
* Proven experience building or applying audio quality guidelines in a production workflow
* Proven experience with speech annotation tasks such as transcription, timestamp QA, VAD/segmentation, and diarization
* Strong ear for audio quality and naturalness, with the ability to catch subtle issues consistently
* Ability to produce high-quality recordings from a quiet environment using professional or near-professional recording equipment
* Strong written communication skills and ability to write precise, actionable feedback
* High attention to detail and strong judgment when reviewing edge cases
* Comfortable working with structured data and metadata in spreadsheets, CSV, or JSON workflows
**Nice-to-have qualifications**
* Basic Python, Bash, or SQL for QA checks and dataset review
* Background in linguistics, phonetics, speech research, or voiceover
* Experience reviewing both real and synthetic audio
* Multilingual experience or familiarity with accent and dialect variation
* Familiarity with consented, licensed, and compliant handling of voice data
* Familiarity with languages other than English
**What success looks like**
* You improve the consistency and reliability of our audio review decisions
* You raise the quality of training data before it reaches customers or models
* You become a trusted owner of annotation standards and audio QA workflows
* You can move fluidly between reviewing, annotating, recording, and improving guidelines
**Why join us**
You’ll play a direct role in improving the quality of the datasets that power modern voice AI. Your work will shape how training data is collected, reviewed, and delivered for real-world model development.