Outlier – Data Annotation & AI Evaluation Agent (Remote)
Evaluated LLM-generated responses for accuracy, helpfulness, factuality, tone, and instruction-following quality as part of RLHF feedback pipelines. Performed comparative ranking of AI outputs to support model fine-tuning and alignment improvements. Conducted quality checks and data verification to maintain consistent, high-accuracy annotations. • Evaluated generative AI expressiveness, coherence, and contextual appropriateness across prompt types • Annotated text datasets for classification, sentiment, and linguistic quality evaluation • Supported instruction-following and prompt-response quality assessment • Worked independently in a fully remote environment while meeting productivity targets