AI Response Evaluator at Pareto AI (September 2025–December 2025)
Evaluated and compared large language model (LLM) outputs using provided evaluation guidelines to select the stronger response. Documented decision rationales to support dataset clarity and improve training signals. Performed quality checks to ensure evaluation reliability across different topics and language variations. • Compared multiple LLM responses and selected the best output per rubric • Wrote concise explanations supporting each evaluation decision • Conducted ongoing quality checks for factual accuracy and consistency • Provided feedback and observations to inform model updates and improvements