AI Response Evaluator (Remote Contract Work)
Evaluated AI-generated responses for reasoning quality, factual consistency, clarity, and safety. Reviewed prompts and model outputs across technical and general knowledge tasks, identifying hallucinations and logical or formatting issues. Performed comparative ranking and conducted edge-case testing, then provided structured feedback to improve alignment and response quality. • Assessed reasoning quality and factual consistency of model outputs. • Detected hallucinations, logical flaws, and formatting inconsistencies. • Ranked responses comparatively by usefulness and accuracy. • Ran edge-case tests and supplied structured improvement feedback.