LSE Dissertation: Scalable scraping and dataset construction pipeline (object detection, embedding generation, zero-shot classification)
Built a scalable scraping and dataset construction pipeline combining object detection, embedding generation, and zero-shot classification for multimodal machine learning. Applied multimodal ML techniques using YOLOv8, CLIP, ResNet, and transformer-based models. Reported strong performance on vision tasks, indicating extensive model training and evaluation using constructed datasets. • Constructed datasets via scraping and automated annotation-like pipelines (detection, embeddings, zero-shot classification). • Trained and evaluated multimodal models including YOLOv8, CLIP, ResNet, and transformers. • Generated embeddings to support zero-shot classification workflows. • Achieved up to 91% macro-F1 across vision tasks through iterative evaluation.