Outlier AI – Freelance Human-in-the-Loop (HITL) AI Specialist & RLHF Contributor
Applied human feedback to train and improve large language models using RLHF-style workflows. Ranked and evaluated AI-generated responses on correctness, relevance, reasoning quality, safety, and alignment with user intent. Performed prompt and output review to detect factual inaccuracies and policy/compliance issues. • Response evaluation and ranking for accuracy, relevance, reasoning, safety, and intent. • RLHF contribution for LLM training and improvement. • Prompt and model output auditing for factual and policy issues. • QA reviews and dataset contribution via validation and annotation pipelines.