Complex command annotation
Participated in a high-profile Large Language Model (LLM) evaluation project on the Appen platform, focusing on validating model behavior safety and persona consistency under complex constraints. Key Tasks Handled: Thoroughly analyzed multiple sets of complex prompts (including specific personas, personality traits, designated functions, and background constraints) to understand the model's operational boundaries. Independently designed and executed 3 targeted, adversarial testing queries for each prompt scenario. Conducted multi-turn dialogues to evaluate whether the model deviated from its designated persona or violated platform safety policies. Accurately identified and documented anomalous dialogue samples, performing compliance tagging and feedback reporting as required. Project Scale & Quality Achievements: Completed comprehensive adversarial dialogue testing across over 10 distinct complex prompt scenarios, designing and submitting dozens of high-quality evaluation queries. Strictly followed the project guidelines throughout the execution; all submitted evaluation and annotation data successfully passed review and achieved a 100% acceptance rate.