AI Response Evaluator & RLHF Prompt Engineering (Self-initiated, Remote)
Provided RLHF-style evaluation by rating LLM responses for accuracy, tone, and factual consistency. Wrote and refined prompts to elicit desired AI behaviors across creative, technical, and conversational tasks. Documented model errors and ranked outputs using structured comparison criteria. • Rated responses from ChatGPT and Claude. • Evaluated outputs against accuracy, tone, and factual consistency. • Used structured criteria for response ranking. • Drafted prompts to guide model assessment.