For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
T
Theo A.

Theo A.

Python-based LLM evaluation framework for customer feedback classification/summarization and LLM reliability/latency ben

Indonesia flagN/A, Indonesia

Key Skills

Software

Other

Top Subject Matter

Customer feedback processing
LLM evaluation
Language-to-SQL benchmarking

Top Data Types

TextText

Top Task Types

Text SummarizationText Summarization
Function CallingFunction Calling

Freelancer Overview

Python-based LLM evaluation framework for customer feedback classification/summarization and LLM reliability/latency ben. Brings 12+ years of professional experience across complex professional workflows, research, and quality-focused execution. Core strengths include Internal, Proprietary Tooling, and Other. Education includes Master of Science, National University of Singapore (2015) and Bachelor of Science, Nanyang Technological University (2013). AI-training focus includes data types such as Text and labeling workflows including Evaluation, Rating, and Text Summarization.

Labeling Experience

Python-based LLM evaluation framework for customer feedback classification/summarization and LLM reliability/latency benchmarking

TextText

Led development of an LLM evaluation framework to benchmark model outputs for customer-feedback classification and summarization tasks. Created reusable model-provider architecture with validated structured outputs to assess quality and reliability. Built automated evaluation pipelines comparing model accuracy, latency, and output reliability across multiple LLM families. • Benchmarked customer feedback classification • Implemented provider support for Ollama and OpenAI • Added structured output validation via Pydantic • Automated comparison of accuracy/latency/reliability

2024 - Present

Automated customer case summarization using GPT-3.5

OtherTextTextText SummarizationText Summarization

Developed an automated customer case summarization solution using GPT-3.5 to convert lengthy customer interaction histories into concise case notes. This enabled downstream Customer Service and Operations teams to more efficiently handle loan default prediction workflows. The summarization work relied on prompt-driven LLM generation and integration into business processes. • Built GPT-3.5-based case summarization • Condensed customer interaction histories into concise notes • Enabled faster handoffs for ops and service teams • Supported loan default prediction use cases

2022 - 2024

Automated experimentation design and statistical inference platform (R-Shiny)

OtherTextTextFunction CallingFunction Calling

Created A/B testing and experimentation infrastructure that supports data-driven statistical inference and model development decisions. Standardized interactive Frequentist and Bayesian experiment design via an R-Shiny platform to guide evaluation of model and operational changes. While not traditional data labeling, the work directly supported AI/ML training evaluation workflows. • Designed reusable experiment frameworks • Enabled Frequentist and Bayesian analysis in R-Shiny • Standardized automated model training/deployment workflows • Supported evaluation-driven decision making for ML initiatives

2019 - 2021

Education

N

National University of Singapore

Master of Science, Supply Chain Management

Master of Science
2014 - 2015
N

Nanyang Technological University

Bachelor of Science, Mathematical Sciences

Bachelor of Science
2010 - 2013

Work History

J

Julo Fintech

Head of Data Science & AI

N/A
2024 - Present
T

Traveloka

Senior Data Manager (Financial Services)

Singapore
2022 - 2024