For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
M
Mark

Mark

AI Data Quality Specialist

Kenya flagWashington, DC, Kenya

Key Skills

Software

AppenAppen
Other

Top Subject Matter

Natural Language Processing
Llms Domain Expertise
Rlhf Domain Expertise

Top Data Types

TextText
ImageImage
DocumentDocument

Top Task Types

Entity (NER) ClassificationEntity (NER) Classification
ClassificationClassification
Red TeamingRed Teaming

Freelancer Overview

AI Data Quality Specialist. Brings 3+ years of professional experience across legal operations, contract review, compliance, and structured analysis. Core strengths include Internal, Proprietary Tooling, and Appen. Education includes Bachelor of Science, Georgetown University (2021). AI-training focus includes data types such as Text and labeling workflows including Evaluation, Rating, and Entity (NER) Classification.

Labeling Experience

AI Data Quality Specialist

TextText

As an AI Data Quality Specialist, I evaluated large volumes of LLM-generated text for accuracy, coherence, and alignment with human values. I contributed to RLHF prompt dataset creation and developed annotation guidelines to improve consistency among annotators. My work involved curating safety datasets, identifying problematic content, and mentoring junior team members. • Evaluated over 500 LLM outputs weekly using rubric-based assessments. • Created prompt–response pairs for RLHF data in STEM, law, and creative writing. • Improved inter-annotator agreement by 18% via refined guidelines. • Documented model hallucinations and edge cases for safety data curation.

2023 - Present
Appen

NLP Data Annotator

AppenAppenTextTextEntity (NER) ClassificationEntity (NER) Classification

As an NLP Data Annotator, I performed entity extraction, text classification, and coreference resolution on text datasets for major enterprise clients. The projects included multilingual annotation, with tasks in both English and Spanish. I maintained high annotation accuracy and offered feedback to refine guidelines. • Completed text classification and entity recognition on diverse NLP projects. • Participated in multilingual annotation efforts. • Maintained a 98.5% quality score over 12 project cycles. • Provided feedback that improved company-wide style guides.

2021 - 2022

Research Assistant – Computational Linguistics

TextTextClassificationClassification

As a Research Assistant in Computational Linguistics, I contributed to the annotation of linguistic corpora for discourse and syntactic analysis. This role involved both manual and script-assisted labeling of raw text for academic research. I also supported the development of annotation tools and presented findings to the research community. • Annotated text corpora for discourse analysis and syntactic parsing. • Built scripts to clean and preprocess linguistic datasets. • Co-authored a poster on automated argument mining. • Supported faculty in research-related labeling activities.

2019 - 2021

Developer – Argument Quality Annotator Tool (Open Source Project)

OtherTextTextClassificationClassification

I developed an open-source GUI annotation tool for assessing the quality of arguments in news articles. This project included the creation and labeling of datasets for argument strength and evidence metrics. The tool was adopted by university labs as part of their annotation workflow. • Engineered a Python tool for argument quality annotation. • Labeled news article arguments for strength and evidence. • Supported academic research in computational argumentation. • Tool adopted by multiple university research groups.

Not specified

Project Lead – LLM Bias Detection Dataset (Personal Project)

TextTextRed TeamingRed Teaming

For the LLM Bias Detection Dataset project, I curated a dataset of adversarial prompts to uncover demographic and factual biases in open-source language models. The process combined prompt engineering and targeted data labeling techniques. Findings were shared publicly, supporting ongoing AI safety work. • Designed and labeled 2,000 adversarial prompt entries. • Focused on surfacing bias in open-source LLMs. • Shared outcomes with the AI safety community. • Work recognized by 3,000+ readers and experts online.

Not specified

Education

G

Georgetown University

Bachelor of Science, Computer Science and Cognitive Science

Bachelor of Science
2017 - 2021

Work History

G

Georgetown University

Research Assistant – Computational Linguistics

Washington, DC
2019 - 2021