Data Labeling and Annotation
Conducted Comparative Analysis through rigorous Side by Side (SxS) evaluations of model generated responses, ranking them based on dimensions of truthfulness, helpfulness, and safety. Prompt Engineering by working through developed complex prompts to stress test model constraints and identify edge case failures in logic. Also conducted Data Labeling & Justification by writing detailed technical justifications for rankings, ensuring all feedback met strict quality benchmarks and 150+ word count requirements under tight deadlines. Conducted deep dive research to validate the technical accuracy of model outputs across diverse domains, including STEM, humanities, and creative writing.