AI Trainer & PromptEngineer | Outlier AI (Remote)
Provided reinforcement learning from human feedback (RLHF) by rating and comparing AI model responses to guide fine-tuning cycles. Evaluated LLM outputs for factual accuracy, reasoning quality, coherence, and instruction-following across STEM, science, and general knowledge domains. Identified and reported factual errors and inconsistencies while maintaining high inter-annotator agreement with updated annotation guidelines. • Rated and compared prompt-response outputs for quality • Conducted model response accuracy QA (97% personal accuracy rate) • Maintained inter-annotator agreement above 90% • Completed 500+ prompt-response evaluation tasks per month across concurrent projects