Software Engineer & AI Evaluation Specialist | NovaMind Technologies (2023–Present)
Evaluated over 5,000 AI-generated responses for factual accuracy, coherence, reasoning quality, tone, safety, and instruction-following across multiple LLM training projects. Designed and refined prompts and structured feedback workflows to improve model performance and engineering evaluation quality. Created SWE-bench-style coding tasks involving debugging, testing, code review, and implementation challenges across real-world software repositories. • Labeled/assessed responses across multiple rubric dimensions (factuality, coherence, safety, instruction-following) • Built structured feedback processes to guide model improvement • Authored coding tasks that function as training/evaluation data • Debugged, tested, and reviewed code to validate task and response quality