Data Expert (LLM Evaluation) - Outlier AI
Data Expert responsible for evaluating and refining large language model outputs to ensure accuracy, coherence, and compliance with prompt requirements. The position required advanced prompt engineering and expertise in reinforcement learning from human feedback concepts to improve model alignment and safety. Work also included fact-checking generated claims and contributing to high-quality training data pipelines through large-scale text review and organization. • Evaluated and refined LLM responses across diverse domains for factual accuracy, coherence, and adherence to prompt guidelines. • Designed multi-turn prompts to stress-test reasoning and uncover edge-case vulnerabilities. • Performed RLHF-style output ranking, scoring, and editing to align with human values and safety protocols. • Conducted fact-checking against authoritative sources and contributed to training data preparation through large dataset categorization and annotation.