User Testing Evaluator (Freelance/Ongoing) — One Forma
Conducted structured AI content evaluation and comparative preference judgments to rank model-generated responses. Applied accuracy, fluency, coherence, relevance, and helpfulness rating rubrics to generate reliable training signals. Produced written feedback used to refine RLHF pipelines for improved model behavior. • Evaluated outputs across text, image, and audio modalities when required by the task • Interpreted and followed complex, evolving guideline formats including search relevance and instruction-following • Maintained high consistency across large high-volume annotation batches to meet daily targets • Delivered pairwise/comparative judgments and preference feedback for training iterations