Data Annotation Specialist & AI Trainer | DataAnnotation.Tech
Performed LLM evaluation and comparative rating using strict guidelines for truthfulness, helpfulness, and formatting. Conducted RLHF-style comparative analysis by providing detailed written justifications to rank preferred completions. Built adversarial prompts for edge-case and jailbreak testing to stress-test model safety and constraints. • Rated and graded LLM responses against truthfulness/helpfulness criteria • Wrote high-quality justifications for preference ranking • Created adversarial prompts for safety boundary evaluation • Checked and fixed code generation outputs for syntax and structure