AI Output Evaluation & Prompt Engineering — Senior Algorithm Engineer
Performed AI output evaluation for LLM-generated content including relevance, factual accuracy, tone, and business appropriateness. Iterated on prompt strategies and conducted large-scale output ranking and benchmarking across production LLM tasks. Focused the feedback and labels on quality and safety signals suitable for RLHF-style reward modeling. • LLM output scoring and comparative evaluation (helpfulness, accuracy, coherence/tone, safety) • Prompt refinement and output quality evaluation across production tasks • Internal prompt benchmarking framework to compare model versions • Feedback collection aligned to RLHF evaluation needs