Anthropic Post-Training Pipeline: RLHF & Code Alignment
In this role, I contributed to the post-training alignment pipelines for Anthropic’s LLM models via Revelo. My core responsibility was to evaluate and perform structured code audits on 50+ AI-generated pull requests and complex programming tasks. I scored outputs based on logical correctness, test coverage, and edge-case robustness. By comparing multiple model outputs, I produced highly curated, labeled feedback that was directly integrated into RLHF and SFT workflows to improve model performance and reliability in coding tasks. This work required deep technical analysis of Python codebases to ensure the model's generated logic adhered to industry best practices and security standards.