The LRL Humans Project - AI Data Annotator & Model Evaluator (Freelance)
Evaluated AI-generated text responses for correctness, relevance, helpfulness, factual accuracy, and alignment with user intent. Performed prompt adherence and multi-dimensional safety/quality assessments to support large-scale model training and improvement. Conducted preference-based evaluations by comparing outputs from multiple AI models and rating selected responses. • Rated instruction following, coherence, completeness, and safety • Checked relevance to user intent and task requirements • Supported RLHF-style preference learning via pairwise model comparisons • Applied detailed project annotation guidelines to evaluation tasks