Independent IT Consultant & Technical Evaluator (AI model evaluation and prompt/repro test-case development)
Evaluated the performance and reasoning of large language models on advanced computer systems topics through repeatable test cases. Authored and refined prompts to simulate real-world scenarios across cloud architecture, cybersecurity, data systems, and enterprise integration. Documented edge cases and failure points with detailed reproduction steps to guide model training refinements. • Developed prompts and reproducible evaluation scenarios for LLM testing. • Assessed reasoning gaps and provided feedback on technical soundness. • Improved evaluation metrics and rubrics for model training. • Produced detailed documentation of failures, edge cases, and recommended fixes.