AI Model Evaluation & Prompt Engineering (Independent Projects)
Independently evaluated AI-generated Python code and LLM responses for correctness, logical soundness, efficiency, and edge-case robustness. Generated adversarial prompts to stress-test reasoning depth and identified hallucinated documentation references or fabricated technical claims. Rewrote flawed outputs into instruction-aligned “golden outputs” while ranking model responses for clarity, safety, and helpfulness.• Reviewed syntax errors, logical flaws, and inefficiencies in AI-generated code.• Assessed instruction-following compliance and safety/risk alignment.• Conducted red-teaming style adversarial prompt testing.• Completed multilingual response evaluations in English, Efik, and Ibibio.