AI Data Annotator & LLM Evaluator
As an AI Data Annotator & LLM Evaluator at OpenTrain AI, I evaluated and ranked LLM-generated responses across eight quality dimensions. I authored high-quality prompts and responses for instruction tuning and flagged unsafe or biased content to improve model alignment. I consistently exceeded project quality thresholds with RLHF comparison tasks and contributed to safety and alignment improvements. • Evaluated 3,500+ RLHF comparison tasks with a 97.8% inter-annotator agreement rate. • Authored 500+ prompts and gold responses for LLM fine-tuning. • Identified harmful, biased, or policy-violating content for major AI clients. • Used Labelbox, Label Studio, and Amazon SageMaker Ground Truth for annotation.