AI Content Evaluator & Reasoning Analyst (Self-Directed / Contract)
Evaluated AI-generated responses across analytical, creative, and instructional tasks for accuracy, logical consistency, depth of reasoning, and adherence to instructions. Applied structured rubric-based scoring to judge response quality, including factual correctness, clarity, helpfulness, and format appropriateness, similar to RLHF evaluation workflows. Rewrote and optimized underperforming outputs to improve instruction-following, reduce ambiguity, and increase utility. • Designed and stress-tested structured prompts to probe output control and identify failure modes. • Performed preference ranking and binary quality judgments within multi-criteria scoring frameworks. • Evaluated failure patterns such as instruction drift, hallucination markers, and inappropriate tone calibration. • Focused domain evaluation strengths in Web3/DeFi, growth marketing, content strategy, financial writing, and technical explainers.