AI Response Evaluator & Workflow Developer (Self-Directed Project, Remote)
Evaluated AI-generated text outputs for accuracy, coherence, tone, and instruction-following. Ranked multiple model responses against defined quality criteria to identify hallucinations, factual errors, and formatting violations. Stress-tested prompt behavior across task types such as summarization, Q&A, and creative writing to ensure reliable model performance. • Reviewed responses against a rubric and guideline set • Logged evaluation decisions in a structured manner • Detected hallucinations, unsafe outputs, and logical inconsistencies • Improved prompt/query quality by probing model behavior