For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
J
Josia L.

Josia L.

AI Voice & Data Research Specialist (Freelance) — recorded data, labeled/localized audio datasets, and performed TTS/STT

Kenya flagMombasa, Kenya

Key Skills

Software

Other

Top Subject Matter

Voice AI training dataset labeling
accent/dialect variation capture
and speech quality evaluation for East African English and Swahili.

Top Data Types

AudioAudio
TextText
DocumentDocument

Top Task Types

TranscriptionTranscription
RLHFRLHF

Freelancer Overview

AI Voice & Data Research Specialist (Freelance) — recorded data, labeled/localized audio datasets, and performed TTS/STT. Brings 4+ years of professional experience across legal operations, contract review, compliance, and structured analysis. Core strengths include CrowdGen and Other. Education includes Craft Certificate, Bandari Maritime Academy and Kenya Certificate of Secondary Education, Secondary Education Credentials. AI-training focus includes data types such as Audio and labeling workflows including Evaluation, Rating, and Transcription.

Labeling Experience

Voice AI Research Specialist - Freelance

AudioAudioRLHFRLHF

Provided voice and language research support for Voice AI projects by delivering high-quality recordings and ensuring dataset usability for model development. Conducted TTS and STT evaluation audits to assess fluency, pronunciation accuracy, and cultural relevance for East African speech. Applied strict project guidelines and metadata schema requirements to support reliable RLHF voice research and reduce bias in multilingual speech systems. • Recorded and submitted prompted speech and natural conversational audio per technical dataset criteria. • Audited TTS/STT outputs and graded AI responses across fluency, pronunciation, and naturalness. • Reviewed and labeled localized datasets to identify accent variations and regional dialect markers. • Supported semantic labeling and prompt engineering for RLHF-based voice research workflows.

2024 - Present

AI Voice & Data Research Specialist (Freelance) — recorded data, labeled/localized audio datasets, and performed TTS/STT evaluation audits.

AudioAudio

Recorded and submitted prompted speech and conversational audio for Voice AI training datasets according to project-specific technical criteria. Evaluated and rated TTS/STT outputs by assessing fluency, pronunciation accuracy, syntactic naturalness, and cultural relevance for East African English and Swahili. • Graded AI-generated voice responses against provided guidelines. • Labeled localized audio samples to reduce algorithmic bias from accent/dialect variation. • Conducted multi-turn voice evaluation with QA targets above 98%. • Applied metadata schemas and quality benchmarks during review.

2024 - Present

Linguistic Validator (Voice Data QA) - Freelance

AudioAudioTranscriptionTranscription

Performed voice-related research and linguistic validation work for multilingual AI datasets with an emphasis on East African English and Swahili. Produced clear, noise-free audio samples and ensured accurate transcription and time alignment for speech recognition readiness. Collaborated remotely with international teams while maintaining consistent annotation schema usage and high-quality checks. • Validated complex multilingual datasets focused on East African English and Swahili structures. • Transcribed and time-stamped spoken audio, including hesitations, overlaps, and code-switching. • Flagged mislabels and low-quality entries to improve dataset integrity and reduce downstream errors. • Evaluated AI-generated transcriptions against source audio to correct phonetic and dialect-specific mismatches.

2023 - 2024

AI Data Annotator & Linguistic Validator (Freelance) — transcription, time-stamping, schema-based annotation, and dataset quality validation.

OtherAudioAudioTranscriptionTranscription

Annotated and validated multilingual datasets focused on East African English and Swahili linguistic structures for speech recognition and NLP pipelines. Transcribed and time-stamped spoken audio with accurate capture of natural speech phenomena such as hesitations, overlaps, and code-switching. • Applied consistent annotation schemas across large volumes of text and audio. • Corrected phonetic mismatches and dialect-specific errors in model outputs. • Identified and flagged mislabelled or low-quality data entries. • Maintained uniformity needed for machine learning pipeline compatibility.

2023 - 2024

Education

S

Secondary Education Credentials

Kenya Certificate of Secondary Education, Secondary Education

Kenya Certificate of Secondary Education
Not specified
B

Bandari Maritime Academy

Craft Certificate, Marine Engineering

Craft Certificate
Not specified

Work History

F

Freelance

Voice AI Research Specialist

Mombasa
2024 - Present
F

Freelance

Linguistic Validator (Voice Data QA)

N/A
2023 - 2024