Multimodal Large Model Data Governance Specialist (Contractor) — OVD/VLM Ground Truth Curation
Curated multimodal datasets for open-vocabulary detection (OVD) and vision-language models requiring text-to-region ground truth mapping. Produced region-level supervision aligned with text queries to enable autonomous vision agent training. Automated pre-labeling steps for named entity recognition and relationship extraction to speed up data cleaning and reduce manual workload. • Text-to-region mapping ground truth for OVD/VLM training • Region supervision curation for multimodal alignment • NER and relationship extraction pre-labeling automation • Data cleaning cycle acceleration through local lightweight model deployment.