LLM Technical Accuracy Reviewer
I systematically evaluated and tested the technical accuracy of outputs generated by large language models (LLMs) including GPT-4 and Claude. My work involved identifying and documenting model errors in topics such as algorithm optimization and systems programming. I refined prompts to elicit precise, pedagogically useful responses from AI models. • Assessed LLM-generated responses for logical accuracy and technical correctness • Tracked recurring issues around concurrency, memory management, and abstraction • Developed and tested prompts for improved output quality • Documented model limitations and proposed improvements