Correct English transcripts and align word-level timing to spoken audio with millisecond precision. This part-time contractor role offers 20+ hours per week and pays $10-$35 per hour.
About OpenTrain
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. The platform helps contributors discover projects, build a professional profile, and apply to opportunities that match their skills. Creating an OpenTrain account is free.
About AI Training Work
AI training is the human side of building modern artificial intelligence. Contributors prepare and review examples that help speech and language systems understand real-world audio, including varied accents, dialects, speaking styles, and challenging recording conditions.
This work offers a way to contribute to cutting-edge AI while developing specialized experience in transcription, annotation, and quality review. Many AI training projects are flexible and remote, making them suitable for people balancing work, studies, or family commitments.
The Role
OpenTrain is seeking an Audio Transcription and Alignment Specialist to improve speech data used in AI training. You will review and correct word-level transcripts, align word and segment timing boundaries to spoken audio with millisecond precision, and edit text and timing together in a specialized web-based audio labeling tool.
The role requires careful judgment when working with overlapping speakers, filler words, false starts, background noise, accents, dialects, mumbled words, and unclear speech. You will work independently and asynchronously while maintaining dependable turnaround across a large volume of audio tasks.
- Part-time contractor position
- 20+ hours per week
- Entry-level opportunity
- Based in India
- English-language work
- Pay range of $10-$35 USD per hour
What You'll Do
You will combine accurate transcription with precise audio annotation. Success in this role depends on consistently applying detailed rules while matching written transcripts and timing boundaries closely to the underlying speech.
- Review transcripts for spelling, grammar, and accuracy
- Follow detailed transcription style and formatting conventions
- Align word-level timing boundaries to spoken audio
- Align segment timing boundaries with millisecond-level precision
- Edit transcript text and timing in a web-based audio labeling tool
- Resolve overlapping speech, filler words, false starts, and unclear audio
- Account for accents, dialects, background noise, and mumbled words
- Maintain consistent quality across large volumes of audio tasks
- Work independently and asynchronously while meeting task deadlines
Requirements
Native or near-native English fluency and excellent listening comprehension are required. You should be able to distinguish subtle speech differences and make reliable decisions when audio is difficult to understand.
A methodical, self-directed approach is important. You must be comfortable following strict, rule-based conventions and applying style guides consistently while working at millisecond-level timing precision.
- Native or near-native English listening comprehension
- Ability to produce accurate speech transcriptions
- Ability to align word and segment boundaries with millisecond precision
- Strong attention to detail and consistency
- Good judgment with overlapping speech, filler words, false starts, accents, and unclear audio
- Ability to follow transcription style guides and formatting rules
- Reliable, independent, and self-directed task management
Helpful Background
Previous experience is helpful but not required for this entry-level opportunity. Relevant backgrounds include professional transcription, closed captioning, court reporting, linguistics, phonetics, audio data annotation, speech-to-text quality assurance, and audio post-production.
- Experience with transcription or closed captioning
- Background in linguistics or phonetics
- Audio annotation or speech-to-text quality assurance experience
- Audio post-production experience
- Familiarity with Praat, ELAN, or Subtitle Edit
- Fast and accurate typing skills
Who Should Apply
This role may suit careful listeners who enjoy structured, detail-focused work and can make consistent decisions across many audio examples. It is a strong fit for contributors interested in building practical experience in speech data annotation and AI training.
Apply if you can commit to 20+ hours per week, work reliably without constant supervision, and maintain accuracy when audio includes noise, multiple speakers, accents, dialects, or incomplete speech.
- Native or near-native English speakers in India
- Detail-oriented listeners comfortable with repetitive precision work
- Transcription, captioning, linguistics, phonetics, or audio specialists
- Contributors seeking flexible part-time AI training work
How to Get Started
Create a free OpenTrain account to build your profile and apply for opportunities that match your experience. OpenTrain helps people start and grow careers in AI training and data labeling by bringing specialized projects and professional development into one place.
- Create your free OpenTrain account
- Highlight relevant transcription, audio, or language experience
- Review the role details and submit your application
- Prepare to demonstrate careful listening and consistent annotation judgment