Freelance AI Evaluation Consultant - Self-Employed
You generate complex programmatic prompts and multi-turn conversational scenarios to test large language models and evaluate their behavior. You review, rank, and validate AI-generated code and dataset content using strict logical safety and quality guidelines. The role requires strong computer science foundations, analytical reasoning, and precise technical feedback skills. • LLM prompt and scenario creation for testing • Code output evaluation and ranking across Python, JavaScript, and SQL • Dataset review using quality-ranking metrics • Identification of hallucinations, subtle flaws, and logic edge cases