AI Response Evaluation & Code Quality Assessment
Evaluating and rating AI-generated responses and code outputs for quality, correctness, and safety as part of LLM training workflows. Tasks include assessing Python and JavaScript code for logical accuracy, identifying edge cases, rating responses for helpfulness and factual grounding, and flagging unsafe or incorrect outputs. Applied domain expertise in full stack development and web application security to provide high-signal feedback on technical tasks.