Senior AI Data Trainer / Specialist (Contract & Independent Projects)
Evaluated and ranked multi-turn LLM responses using strict dimensions of truthfulness, helpfulness, harmlessness, and structural logic for RLHF-style training. Authored large volumes of complex prompt sets to test frontier-model reasoning, mathematical logic, and coding performance. Identified, documented, and categorized edge-case anomalies to reduce downstream hallucinations and bias. • Conducted RLHF prompt/response evaluation and ranking • Produced 500+ complex multi-step prompts for model capability testing • Performed edge-case auditing and anomaly categorization • Developed localized annotation guidelines and schemas that improved labeling consistency and reduced quality variance