AI Evaluation Specialist (Independent / Freelance)
Conducted systematic, safety-focused evaluation of seven major LLM platforms (Pi, Grok, Claude, Gemini, ChatGPT, Copilot, and Meta AI/Muse Spark). Labeled model behaviors by analyzing hallucinations, factual accuracy, policy drift, and consistency across prompts, and documented anomaly patterns. Performed hands-on RLHF-related process documentation and AI behavioral pattern analysis for safety and ethics. • Hallucination detection and factual accuracy verification • Prompt engineering, response evaluation, and response rewriting • Jailbreak/social-engineering recognition and red-teaming style checks • Policy-drift detection and safety anomaly identification