Independent Contractor — AI Data Annotator & Analyst (RLHF, dialogue annotation, LLM evaluation, red-teaming)
Performed reinforcement learning from human feedback (RLHF) by ranking multiple AI-generated responses using strict truthfulness, utility, and harmlessness metrics. Annotated complex multi-turn dialogue datasets by applying linguistic and logical analysis to guide model outputs toward natural, helpful, and accurate conversation. Evaluated and iteratively refined large language model prompts and responses for quality, accuracy, alignment, and reduced hallucinations via external source cross-referencing. • Truthfulness/utility/harmlessness response ranking • Multi-turn dialogue annotation with linguistic/logical reasoning • Prompt and response grading plus revision for safety and factual precision • Adversarial prompt creation for red-teaming and robustness testing