Assistant Editor & Content Strategist (LLM evaluation, ranking, and red-teaming)
Performed rubric-based and pairwise evaluation of LLM responses, focusing on reasoning quality, factual accuracy, and stylistic nuance. Produced concise, structured rationales to document decision justification for each evaluated output. Conducted safety-focused red-teaming activities by identifying bias, harmful content, and jailbreak vulnerabilities. • Pairwise ranking of LLM responses. • Rubric scoring and long-context reasoning assessment. • Factual accuracy and logical coherence checks. • Safety red-teaming: bias and harmful content identification.