Freelance RLHF Engineer, Air Dawg Remote (Nov 2025 – Present)
Performed RLHF-based evaluation and ranking of large language model outputs to improve response quality and alignment. Annotated model outputs for reasoning quality, factual accuracy, and safety constraints to guide preference learning. Supported prompt engineering and red-teaming workflows aimed at reducing algorithmic bias and alignment errors. • Evaluated LLM responses for reasoning, factual accuracy, and safety • Ranked and annotated outputs for RLHF optimization • Participated in prompt engineering iterations • Contributed to red-teaming pipeline improvements