AI Response Evaluator & Prompt Specialist | Self-Directed / Freelance Platforms
Evaluated thousands of AI-generated responses for accuracy, helpfulness, safety, and instruction-following quality across multiple platforms. Performed side-by-side preference labeling by comparing two AI outputs per prompt and selecting/ranking responses using detailed rubrics. Produced structured prompts and conducted factual verification against primary/authoritative sources, while applying platform-specific safety escalation when needed. • Preference labeling and response ranking based on rubric criteria. • Instruction-following and quality evaluation (coherence, helpfulness, correctness). • Safety labeling including toxicity, harmful content, hate/policy violation detection. • Factual correctness checks via cross-referencing credible references.