AI Response Evaluator and Dataset Validator
Evaluated AI-generated textual responses for logical consistency, factual correctness, and guideline alignment. Compared responses across different AI models such as ChatGPT and Claude to identify and document reasoning differences. Revised and improved outputs for clarity, structure, and compliance with annotation instructions. • Produced structured summaries from engineering and technical datasets. • Verified dataset accuracy by cross-referencing multiple technical documentation sources. • Worked within annotation-style workflows and instruction-guided evaluation tasks. • Ensured consistency across multi-source documentation during dataset validation.