Whitebeard
In this project, I served as an expert domain evaluator for the Whitebeard workflow, where my primary responsibility was to assess and refine AI model responses to complex technical prompts. I thoroughly analyzed paired model outputs across several advanced mathematical disciplines, including multivariable calculus, complex analysis (specifically Hadamard finite-part regularization), and combinatorics. For each task, I evaluated the prompts for ratability based on strict guardrails regarding safety, PII, and data completeness. I then performed a comparative analysis to select the preferred response, judging dimensions such as instruction following, technical accuracy, and structural layout. Finally, I remediated formatting failures—such as broken LaTeX delimiters or syntax errors in code blocks—to generate a polished, 'gold standard' final rewrite suitable for platform deployment.