AI Systems Tester & Prompt Evaluator (Freelance)
As an AI Systems Tester & Prompt Evaluator, I designed, evaluated, and refined language model workflows by annotating bot responses and reviewing prompt effectiveness. I performed RLHF-style preference ranking, error identification, and iterative prompt/annotation cycles to improve system output for healthcare and customer service conversation data. My evaluations included assessment of fluency, factual accuracy, instruction-following, and safety in LLM outputs. • Assessed LLM outputs for instruction-following, accuracy, helpfulness, and tone • Conducted iterative annotation and prompt refinement cycles, reducing error rates • Performed RLHF preference ranking by evaluating conversational AI outputs • Identified failure patterns such as hallucinations and documented findings