Data Annotator — Text, Audio & Multimodal
In this ongoing role, I label AI response pairs using multi-dimensional rubrics and perform pairwise preference ranking to generate RLHF training data for large language models. I correct machine-generated transcripts verbatim and apply rich transcription conventions. I also conduct speaker count voting, language identification annotation, and categorize user prompts by task type. • Labeled for factuality, instruction following, helpfulness, and style. • Flagged model errors and non-compliance in AI outputs. • Worked with both text and audio datasets, including 25+ languages. • Used proprietary annotation platforms for large-scale labeling tasks.