Freelance Content Moderator & Data Annotator (AI Data Evaluator)
Performed RLHF annotation by rating and ranking AI-generated responses for quality, helpfulness, harmlessness, and honesty to improve LLM outputs. Reviewed and classified large volumes of user-generated text content against platform safety policies, flagging policy-violating material. Carried out adversarial and prompt-injection testing to identify model vulnerabilities and edge cases in AI safety guardrails. • Rated and ranked response quality using RLHF-style guidelines • Flagged and categorized harmful categories such as hate speech, misinformation, graphic violence, and NSFW content • Ensured inter-rater reliability by meeting strict quality benchmarks • Escalated ambiguous or high-severity cases with detailed written rationales