Analyst – LLM Evaluation/Rating
As an Analyst at Turing, I evaluated outputs from Meta’s large language models by judging quality, accuracy, clarity, and overall usefulness. I provided detailed written explanations to justify evaluation decisions and highlighted areas for improvement in model responses. Using browser-based and PDF tools, I converted unstructured LLM-generated information into organized datasets for further analysis. • Judged side-by-side LLM outputs for quality, accuracy, and usefulness. • Wrote explanations and recommendations for model improvements. • Utilized browser-based tools and extensions for efficient evaluation management. • Structured LLM content from PDFs for downstream tasks.