Image Captioning Model Trainer and Data Curator
Developed and trained an image captioning model using multimodal datasets to generate accurate visual descriptions. Used Huggingface Visual Transformer and GPT-2 for feature extraction and language modeling, adjusting parameters for optimal fine-tuning outcomes. Built tools for data preparation, training, and performance evaluation of AI-generated captions. • Annotated datasets by pairing images with curated natural language captions for machine learning. • Applied data labeling and curation to improve model performance during iterative fine-tuning. • Ensured robust data validation and label consistency throughout the process. • Documented the workflow and results for future reproducibility and reference.