AI Red Teaming & Content Safety Review
This project focused on adversarial testing and content safety review for large language models on Outlier AI. Tasks included writing red-teaming prompts designed to elicit harmful, biased, or policy-violating model responses, then documenting the model's failure modes in structured reports. I also reviewed AI-generated content for toxicity, misinformation, hate speech, and safety guideline violations as part of content moderation workstreams. Monthly output averaged 300–500 evaluated items. Quality was maintained through strict guideline adherence and zero escalations raised by the client team over the full duration of the project. Turnaround SLAs were consistently met across all batches.