AI Trainer & LLM Evaluator (Freelance) - Remote AI Training Platforms
Evaluate and rank AI-generated responses across diverse domains using precise side-by-side RLHF criteria, focusing on truthfulness, safety, conciseness, and formatting. Analyze complex logical prompts and programming snippets in Python and SQL to verify code validity and execution efficiency. Author golden-standard responses and instruction-following prompts for SFT fine-tuning and benchmark evaluations. Identify model errors, logic hallucinations, and formatting flaws, documenting feedback to improve training datasets.