Data Annotator
Project Scope Description The client was fine-tuning a proprietary LLM to improve its performance in professional, high-stakes communication (specifically legal and medical advisory domains). The model frequently suffered from "hallucinated" citations and tone inconsistencies when responding to complex, multi-layered queries. My objective was to evaluate, critique, and rewrite model outputs to ensure factual accuracy, logical flow, and adherence to specific brand-voice guidelines. Specific Data Labeling Tasks Performed Ranking & Preference Labeling: Performed pairwise comparisons (A/B testing) on model-generated responses, ranking them based on factual correctness, conciseness, and neutrality. Fact-Checking & Citation Verification: Cross-referenced model assertions against verified primary sources to identify and flag inaccuracies, providing corrected citations where the model failed. Safety & Bias Mitigation: Evaluated responses for subtle harmful biases or inflammatory language, rewriting responses to ensure they adhered to strict safety and "helpful assistant" guardrails. Tone/Style Calibration: Adjusted the output style to shift from a generic robotic tone to a precise, professional persona, ensuring the model's vocabulary aligned with industry-standard terminology. Project Size Volume: Evaluated and audited over 2,500 individual dialogue turns and long-form responses. Duration: Ongoing freelance contract since Q3 2024. Tooling: Worked directly within the client’s custom internal evaluation dashboard and integrated Notion/Slack workflows for documenting edge-case patterns. Quality Measure Consistency Rating: Maintained a 97% alignment score with the client’s "Gold Standard" evaluation set (a set of benchmark responses vetted by the project lead). Feedback Loop: Acted as a senior contributor in documenting "failure modes," which directly influenced the refinement of the prompt-engineering instructions for the wider annotation team. Weekly quality audits focused on reasoning depth, ensuring that my rationale for flagging an error was logically sound and defensible to the engineers.