I have practical experience evaluating AI-generated outputs across text, image, and video modalities — assessing respons
I have practical experience evaluating AI-generated outputs across text, image, and video modalities — assessing responses for accuracy, coherence, instruction-following, and alignment with user intent. Through cross-model benchmarking using tools like Midjourney, DALL-E, Runway, Sora, and Veo, I developed structured prompt engineering workflows and a sharp eye for identifying subtle quality gaps between model outputs. I'm comfortable articulating detailed, rubric-style feedback that goes beyond surface-level critique. My academic background in Information Technology sharpened my ability to handle classification and data tasks with precision and consistency. I understand the importance of applying evaluation criteria uniformly across large volumes of outputs, staying objective, and flagging edge cases clearly. These skills map directly to what AI training and RLHF pipelines require from human evaluators.