AI Research & Evaluation Contributor (Remote | Freelance)
Evaluated AI-generated responses for correctness, clarity, logical consistency, and factual accuracy using structured prompts and benchmark tasks. Reviewed technical and mathematical outputs against detailed evaluation guidelines to detect hallucinations, formatting issues, and reasoning flaws. Provided human feedback intended to improve model alignment and response quality. • Assessed reasoning quality and instruction adherence using scoring rubrics. • Performed quality assurance on outputs for factuality and logical coherence. • Identified and reported errors such as hallucinations and reasoning inconsistencies. • Supported LLM evaluation workflow development through benchmark prompt design.