For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
T
Tianyi L.

Tianyi L.

MSc Dissertation Data Labeling – COVID-19 Vaccination Prediction

United Kingdom flaglondon, United Kingdom

Key Skills

Software

No software listed

Top Subject Matter

Public Health
Covid-19 Domain Expertise
Machine Learning

Top Data Types

ImageImage
TextText
DocumentDocument

Top Task Types

ClassificationClassification
SegmentationSegmentation
Emotion RecognitionEmotion Recognition

Freelancer Overview

MSc Dissertation Data Labeling – COVID-19 Vaccination Prediction. Brings 3+ years of professional experience across complex professional workflows, research, and quality-focused execution. Core strengths include Python, Internal, and Proprietary Tooling. Education includes Master of Science, University College London (2023) and Bachelor of Science, University of Liverpool (2022). AI-training focus includes data types such as Medical, DICOM, and Image and labeling workflows including Classification, Segmentation, and Emotion Recognition.

Labeling Experience

Data Annotation – Private Equity Due Diligence AI Agent

TextTextClassificationClassification

I annotated key business logic, strategy styles, and risk factors from unstructured meeting minutes for an AI agent in private equity due diligence. The work involved integrating data from multiple sources via APIs and performing rule-based automated labeling. This ensured a high-quality prompt corpus for large language model (LLM) development. • Labeled unstructured text for entity and strategy recognition. • Built prompt corpora for AI due diligence applications. • Integrated and aligned multi-source financial data. • Employed rule-based methods for automated classification.

2024 - Present

Rule-based Data Labeling – Document Formatting Project

DocumentDocumentClassificationClassification

I translated complex formatting standards of official documents into rule-based, machine-readable label sets as part of an AI-powered formatting initiative. The labeling process included multi-round acceptance testing and suggestions for user experience optimization. These labels supported further automation of document formatting tasks in large organizations. • Developed labeling rules for document formatting. • Conducted classification for formatting compliance. • Managed acceptance testing for label accuracy. • Advised on user experience improvements based on labeling feedback.

2024 - 2024

Sentiment Analysis Data Labeling – Restaurant NLP Chatbot

TextTextEmotion RecognitionEmotion Recognition

I labeled over 10,000 restaurant user reviews for sentiment (positive/negative) and extracted keywords for a GPT-3.5-based NLP chatbot. Model performance was enhanced by iterative few-shot data labeling and adjusting prompt parameters. Accurate sentiment tagging directly improved the chatbot’s feedback understanding and classification. • Tagged sentiment classes for customer feedback. • Performed large-scale annotation for model fine-tuning. • Adjusted prompts for model optimization. • Extracted keywords for targeted NLP training.

2023 - 2023

Data Annotation – Wealth Management Knowledge Base Construction

DocumentDocumentSegmentationSegmentation

I was involved in the preprocessing and labeling of non-structured financial documents for constructing a knowledge base supporting a wealth management LLM. Tasks included segmentation labeling, metadata tagging, and generating QA pairs for reinforcement learning from human feedback (RLHF). My contributions improved LLM ingestion quality and response accuracy. • Segmented and tagged PDF and scanned documents. • Generated and screened question-answer pairs for RLHF. • Labeled model prediction errors for correction. • Enhanced document pipeline for large-scale LLM training.

2023 - 2023

MSc Dissertation Data Labeling – COVID-19 Vaccination Prediction

ClassificationClassification

I conducted data labeling as part of my MSc dissertation project on predicting COVID-19 vaccination rates in the UK. The process included managing large-scale medical datasets, treating outliers, and performing normalization to prepare the data for machine learning. These steps were integral to building reliable models and ensuring high data accuracy. • Labeled data for model training and evaluation. • Applied classification to medical records at the MSOA level. • Used feature extraction methods tailored to healthcare datasets. • Focused on end-to-end data pipeline management.

2022 - 2023

Education

U

University College London

Master of Science, Urban Spatial Science

Master of Science
2022 - 2023
U

University of Liverpool

Bachelor of Science, Applied Mathematics

Bachelor of Science
2018 - 2022

Work History

G

Guosen Securities

Product Operations Specialist

Shenzhen
2024 - Present