AI Prompt & Response Evaluator | Remote Projects
Evaluated AI-generated responses for factual accuracy, relevance, clarity, safety, and instruction following. Applied consistent rubric-based judgment across multiple candidate outputs and flagged issues such as hallucinations and logical errors. Produced structured written justifications to support evaluation decisions and guide model improvement. • Compared multiple AI responses per prompt and selected the best output • Assessed grammar, reasoning quality, helpfulness, and user experience • Identified unsafe content and instruction-following failures • Provided detailed human feedback for performance tuning