For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
S
Shadrack G.

Shadrack G.

AI Data Specialist – LLM Evaluation, RLHF & Multilingual Annotation

Kenya flagNairobi, Kenya

Key Skills

Software

Data Annotation TechData Annotation Tech
RemotasksRemotasks
MercorMercor
AppenAppen
MindriftMindrift
TelusTelus

Top Subject Matter

Healthcare – Clinical Language & Medical NLP
Linguistics – Multilingual NLP & Swahili-English Annotation
AI Safety – Content Moderation & Red Teaming

Top Data Types

ImageImage
TextText
DocumentDocument

Top Task Types

Text GenerationText Generation
RLHFRLHF
TranscriptionTranscription
Prompt + Response Writing (SFT)Prompt + Response Writing (SFT)
Text SummarizationText Summarization
Red TeamingRed Teaming
Evaluation/RatingEvaluation/Rating
ClassificationClassification
Question AnsweringQuestion Answering

Freelancer Overview

I've been working in AI data and language services for over three years, primarily focused on LLM evaluation, RLHF preference ranking, hallucination detection, and multilingual annotation. Most of my work has come through platforms like Outlier, DataAnnotation, and Remotasks, where I've completed thousands of tasks across text labeling, response comparison, content safety review, and adversarial red-teaming. I consistently rank in the top tier for quality, which has given me early access to higher-complexity specialist projects over time. What sets me apart is the combination of language and domain expertise I bring. I'm a native Swahili speaker with strong English proficiency, which means I can handle bilingual annotation tasks that most contributors can't. On top of that, I'm currently completing postgraduate nursing studies, so I'm genuinely comfortable working with clinical and healthcare language, not just flagging it as sensitive, but actually understanding it. I tend to be precise and detail-oriented by nature, and I take guideline interpretation seriously, which shows up in my accuracy scores and inter-rater agreement rates.

Labeling Experience

AI Red Teaming & Content Safety Review

TextTextRed TeamingRed Teaming

This project focused on adversarial testing and content safety review for large language models on Outlier AI. Tasks included writing red-teaming prompts designed to elicit harmful, biased, or policy-violating model responses, then documenting the model's failure modes in structured reports. I also reviewed AI-generated content for toxicity, misinformation, hate speech, and safety guideline violations as part of content moderation workstreams. Monthly output averaged 300–500 evaluated items. Quality was maintained through strict guideline adherence and zero escalations raised by the client team over the full duration of the project. Turnaround SLAs were consistently met across all batches.

2023 - Present

Hallucination Detection & Factual Accuracy Review

TextTextEvaluation/RatingEvaluation/Rating

This project involved reviewing LLM-generated text to identify factual inaccuracies, unsupported claims, and hallucinated content across general knowledge and healthcare domains. For each task I read the model output against the source prompt, flagged any incorrect or fabricated information, and wrote a structured justification explaining the error type and severity. Project size ranged from 200–400 reviewed outputs per month depending on the platform. Quality was tracked through accuracy scores against gold-standard labels. I achieved 97% accuracy on healthcare-specific tasks, supported by my background in postgraduate nursing studies which gave me confidence with clinical terminology and medical facts.

2022 - Present

LLM Response Evaluation & RLHF Preference Ranking

TextTextRLHFRLHF

The scope of this project involved evaluating and ranking AI-generated responses for large language model training across multiple ongoing engagements on Outlier and DataAnnotation. On a typical week I would work through batches of 50–150 prompt-response pairs, comparing two or more model outputs and selecting the better response based on structured rubrics covering accuracy, helpfulness, instruction-following, and harmlessness. Over the course of three years I completed upwards of 5,000 preference ranking tasks. Quality was measured through platform-side accuracy scores and inter-rater agreement checks, I consistently maintained a 95th-percentile rating, which qualified me for higher-complexity specialist task pools.

2022 - Present

Bilingual English-Swahili Annotation

TextTextEntity (NER) ClassificationEntity (NER) Classification

The scope of this project covered building and expanding bilingual English-Swahili training datasets for NLP models. Specific tasks included named entity recognition tagging, intent classification labeling, sentiment annotation, and post-editing of machine-translated Swahili text to correct unnatural phrasing and cultural errors. I also completed transcription tasks for Swahili audio recordings. Project volume averaged 800–1,000 annotated items per month across Appen and Remotasks. Quality was measured through inter-rater agreement scores, where I consistently exceeded 92%, and I was one of fewer than 5% of active contributors selected for a specialist Swahili linguistic review project based on annotation precision and native fluency.

2021 - Present

Education

C

Coventry University

Postgraduate Diploma, Nursing Studies (Adult Nursing)

Postgraduate Diploma
2023 - 2026
U

University of Nairobi

Bachelor of Science (BSc), Linguistics & Communication

Bachelor of Science (BSc)
2016 - 2020

Work History

C

Coventry University

Postgraduate Nursing Student (Research & Clinical Studies)

Coventry
2023 - 2026
F

Freelance Language Services

Translator & Transcriptionist (English-Swahili)

Nairobi
2020 - 2021