Data Analyst — Training different AI models (model response rating, prompt evaluation/redrafting, and adversarial prompt editing)
Trained and evaluated multiple AI models by assessing their responses against defined criteria. Performed rating and revision work on model outputs using various rubrics and provided justifications for assessments. Participated in prompt engineering by creating, evaluating, and rewriting prompts, then using edited prompts to test model behavior. • Evaluated one or two model responses and rated them across different dimensions. • Wrote justification for ratings and revised responses to improve quality. • Created prompts, evaluated prompt quality, and rewrote prompts for better performance. • Executed “make the model fail” red-team style tests by adding hard but natural constraints and specific requests.