Agentic Coding Annotator | AI Coding Evaluation Trainee
Evaluated model-generated coding responses for correctness, clarity, instruction adherence, and edge case handling. Conducted pairwise rankings and wrote evidence-based rationales to support scoring decisions. Designed and applied rubrics specific to coding tasks, focusing on outcomes and code quality. • Reviewed Python code snippets for errors and crash cases • Graded LLM output based on task-specific evaluation rubrics • Investigated edge-case and test coverage in code responses • Recorded concise justifications for all evaluation tasks