AI/ML/NLP/RLHF Data Annotator & Machine Learning Trainer
Executed RLHF tasks including ranking and rating model outputs to guide preference learning. Performed rewriting of responses to improve alignment and safety for downstream AI use. Supported evaluation and iterative improvement loops for training pipelines. • Preference ranking and rating of model outputs. • Rewriting outputs to improve alignment and safety. • Reinforcement learning dataset support for training. • Iterative evaluation aligned to project requirements.