Independent AI Evaluator and Prompt Specialist - Self-Directed / Freelance
This role involved evaluating and benchmarking major large language models to understand their behavior under different constraints. You designed and refined custom prompts to reduce hallucinations and improve output accuracy. You also reviewed results for factual precision, logical consistency, and safety alignment while maintaining organized performance records. • Benchmarked LLMs including ChatGPT, Gemini, and Claude • Built and iterated role and few-shot prompts • Performed red-teaming style evaluation and RLHF simulation • Tracked model performance variations in cloud spreadsheets