AI Training & Data Annotation — Outlier AI
Evaluated and compared responses generated by AI models (LLMs) to assess quality and usefulness. Assessed clarity, logic, factual accuracy, and response safety, and flagged errors, inconsistencies, and hallucinations. Provided human feedback to support language model improvement and reviewed technical outputs related to programming and problem-solving. • Created advanced prompts for reasoning, programming, and logic testing • Performed analysis of response quality dimensions (clarity, logic, factuality, safety) • Identified hallucinations and other output defects • Contributed feedback loops for iterative LLM improvement