AI Trainer, RLHF Annotator & Reasoning Evaluator (Freelance)
Reviewed and rated 5,000+ AI-generated responses for accuracy, reasoning quality, helpfulness, and safety using structured multi-criteria rubrics to support RLHF workflows. Performed comparative preference ranking and side-by-side evaluation to produce training signals for reward model improvement cycles. Engineered adversarial and boundary-testing prompts and documented findings in structured reports for evaluation leads. • Preference ranking and comparative rating of model outputs • Safety, helpfulness, and correctness evaluation against rubrics • Reward-model feedback contribution based on side-by-side comparisons • Adversarial/boundary testing to surface reasoning and safety gaps