While I have not worked in a dedicated data labeling role, I have experience preparing, cleaning, validating, and labeli
While I have not worked in a dedicated data labeling role, I have experience preparing, cleaning, validating, and labeling datasets for machine learning and NLP projects. This includes collecting data from public sources such as GitHub, Reddit, and Kaggle, performing preprocessing and quality checks, creating training datasets, and evaluating model outputs. I have worked on projects involving DistilBERT fine-tuning for personality classification, LLM-based email generation, and RAG systems using LangChain and Qdrant. These projects required reviewing data quality, identifying incorrect predictions, refining prompts, and assessing model performance. Through this work, I developed skills that are closely related to AI training, annotation, and model evaluation workflows.