3 years of daily intensive cross-platform LLM evaluation and adversarial testing (GPT-4, Claude, DeepSeek, Gemini, and other models).
Practiced structured adversarial evaluation of LLM behavior to detect subtle shifts, logical inconsistencies, hallucination patterns, and boundary failures across multiple model versions. Performed continuous model assessment to improve reliability and controllability of generated outputs using prompt and interaction strategies. Worked across leading open and closed models to validate performance and robustness under adversarial conditions. • Conducted RLHF-style preference ranking and hallucination audit activities as part of LLM evaluation. • Built and tested adversarial prompt crafting and cross-model prompt architectures. • Designed evaluation approaches to measure behavior shifts and version-to-version changes. • Assessed prompt quality and controllable generation behavior for workflow deployment.