LLM Fine-Tuning and Strategic Research Evaluation (SFT)
Engaged in human-in-the-loop evaluation and prompt testing for language models within a professional communications and research workflow. Responsibilities include testing generative AI models against real-world business and macroeconomic trend data to evaluate response accuracy and identify model hallucinations. Performed detailed text categorisation, grading model outputs based on logical structure, compliance with specific brand guidelines, and tone calibration. Authored comprehensive feedback justifications outlining structural errors, linguistic gaps, and contextual inaccuracies to assist in refining and training advanced language models.