AI Evaluator & Prompt Engineer (Ongoing) — Outlier, Remote
Performed evaluation and refinement of an OpenClaw agent by judging logical reasoning and decision-making quality. Provided high-quality human feedback to support RLHF and improve model alignment and accuracy. Assessed agent responses for safety, functional integrity, and compliance with intended instructions. • Evaluated tool-use execution and workflow decision making • Reviewed outputs to identify issues affecting alignment • Supplied human feedback to guide model improvement • Conducted safety and integrity checks on agent responses