Freelance AI Data Labeler & Response Evaluator — Luei.ai (2025 – Present)
Conducted rubric-based evaluation of AI responses for instruction adherence, factual correctness, logical coherence, and overall response quality. Performed comparative ranking tasks (preference labeling) on multiple model responses and provided written rationale for each decision. Labeled multi-turn conversational data by tracking dialogue context and assessing response relevance and persona consistency. • Instruction-following evaluation • Comparative / preference (RLHF-style) labeling • Hallucination and policy violation detection • Safety and quality rubric scoring