Senior Software Engineer / AI Evaluation Specialist — Independent & Contract Roles
Performed large-scale evaluation and response ranking of AI-generated outputs across diverse domains to improve accuracy and coherence. Built structured annotation frameworks and conducted instruction-following checks to ensure alignment with task requirements and user intent. Contributed to RLHF-style feedback loops and prompt refinement using detailed error analysis to improve model reasoning performance. • Ranked AI responses for correctness, coherence, and reasoning quality • Designed training/annotation guidelines and quality control processes • Ran detailed error analysis to identify failure patterns and iterate • Refined prompts and applied RLHF-style feedback mechanisms