GenAI Model Trainer & AI Evaluation Specialist | OneForma (Remote)
Trained and evaluated generative AI model outputs using RLHF-oriented preference and comparative ranking workflows. Assessed responses for helpfulness, factual accuracy, safety, instruction adherence, reasoning quality, and bias presence according to provided guidelines and benchmarks. Provided evaluator feedback that was used to improve alignment and future model training and optimization. • Ranked LLMs by multi-criteria evaluations (helpfulness, factual accuracy, safety, instruction adherence, reasoning, tone). • Performed preference ranking and comparative analysis of response pairs for reinforcement learning tasks. • Detected hallucinations, logical fallacies, edge cases, and biases using advanced reasoning checks. • Delivered structured evaluator feedback to support iterative model training and optimization.