Freelance LLM stress-testing / red-teaming (Independent AI Tester & Prompt Designer)
Conducted red-team style evaluations by intentionally crafting unusual questions and bilingual constraints to bypass or expose LLM weaknesses. Tested how models handle French dialogue and edge cases that could trigger language or safety failures. Used findings to adjust prompts and system instructions to steer model behavior toward safer and more helpful responses. • Forced and verified language-handling behavior in bilingual prompts • Probed limitations and safety boundaries using contrived scenarios • Iteratively refined instructions based on observed failures • Used evaluation results to develop stronger prompting intuition