AI Evaluation & Data Annotation Specialist
No description provided.
Hire this AI Trainer
Sign in or create an account to invite AI Trainers to your job.
AI Evaluation & Data Annotation Specialist (Outlier / DataAnnotation / Alignerr). – Evaluated outputs from frontier language models — including Gemini, ChatGPT, and Grok — assessing reasoning quality, factual accuracy, and alignment with safety and policy guidelines. – Conducted structured data annotation across a wide range of technical and analytical task categories, consistently meeting quality benchmarks under high-volume conditions. – Identified and documented prompt inconsistencies, logical errors, and model performance gaps; produced detailed, structured feedback that fed directly into model improvement pipelines. – Operated under strict project confidentiality requirements, demonstrating disciplined adherence to annotation guidelines and quality protocols throughout all engagements. – Evaluated visual question-answering (VQA) tasks and model responses at Alignerr, verifying accuracy and correctness of AI outputs against image-based queries; identified errors and supplied corrected responses to support multimodal model improvement. Graduating with a bachelor from Social Work & Social Sciences – Core modules: Sociology, Social Economics, Descriptive Social Statistics, Political Sociology, and Human Rights. – Fieldwork: Conducted structured client intake interviews, needs assessments, and case documentation within local social institutions.
No description provided.
Evaluated visual question-answering (VQA) tasks and model responses against image-based queries. Checked correctness and accuracy of multimodal outputs and reported issues found during review. • Assessed answer validity relative to image content • Logged errors and inconsistencies for follow-up • Supplied corrected responses when needed • Contributed to multimodal model improvement efforts
Assessed outputs from frontier language models (Gemini, ChatGPT, and Grok) for quality and compliance. Focus areas included reasoning quality, factual accuracy, and alignment with safety and policy guidelines. • Measured response quality against established benchmarks • Identified prompt inconsistencies and logical errors • Documented model performance gaps for improvement • Supplied structured feedback for model improvement pipelines
Bachelor of Social Work and Social Sciences, Social Work and Social Sciences
1.5 Years of AI evaluation and data labeling experience