For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
W
William K.

William K.

Senior Multimodal AI Training Specialist, NovaSense AI Ltd.

Kenya flagN/A, Kenya

Key Skills

Software

AWS SageMakerAWS SageMaker

Top Subject Matter

Multimodal vision-language and cross-modal learning (image-text-audio) dataset curation and AI training
Multimodal instruction tuning dataset annotation and evaluation for image-text models
Video-caption and image-description dataset preparation

Top Data Types

ImageImage
VideoVideo

Top Task Types

Fine-tuningFine-tuning
ClassificationClassification
TrackingTracking
RLHFRLHF
Data CollectionData Collection

Freelancer Overview

Senior Multimodal AI Training Specialist, NovaSense AI Ltd.. Brings 9+ years of professional experience across complex professional workflows, research, and quality-focused execution. Core strengths include AWS SageMaker, Weights & Biases, and Notion. Education includes Master of Science, N/A (2018) and Bachelor of Science, N/A (2016). AI-training focus includes data types such as Image and Video and labeling workflows including Fine-tuning, Evaluation, and Rating.

Labeling Experience

AWS SageMaker

Senior Multimodal AI Training Specialist, NovaSense AI Ltd.

AWS SageMakerAWS SageMakerImageImageFine-tuningFine-tuning

Designed and maintained multimodal training pipelines across image, text, and audio modalities for production model families. Led dataset curation with QA protocols and inter-annotator agreement scoring to improve label consistency and downstream benchmark performance. Conducted systematic fine-tuning experiments on CLIP/BLIP-2 checkpoints using LoRA/QLoRA and evaluated retrieval accuracy for deployment-ready models. • Managed large-scale vision-language corpus curation (12M samples) across 18 languages • Implemented annotation QA and inter-annotator agreement scoring (81% to 96% consistency) • Fine-tuned CLIP/BLIP-2 with LoRA/QLoRA for cross-modal retrieval (0.87 Recall@5) • Mentored annotation specialists and junior ML engineers; established async review cadence

2022 - Present

AI Data Training Analyst, Cognitive Loop Technologies

ImageImageClassificationClassificationTrackingTracking

Managed the full annotation lifecycle for a multimodal instruction-tuning dataset of image-text pairs. Coordinated with external labeling vendors and applied active-learning sampling to reduce annotation cost while maintaining dataset diversity. Developed and validated evaluation frameworks for cross-modal understanding using hallucination rate, grounding accuracy, and caption quality metrics to guide model retraining. • Oversaw annotation for 800K image-text pairs with active-learning sampling • Defined and validated metrics including hallucination rate, grounding accuracy, and ROUGE-L • Measured improvements in COCO captioning CIDEr (112 to 134) from evaluation-informed retraining • Implemented prompt engineering templates for consistent zero-shot/few-shot evaluation across multiple model variants

2020 - 2021

Machine Learning Data Engineer, Praxis Annotate Inc.

VideoVideoData CollectionData CollectionClassificationClassification

Built automated data ingestion and pre-processing pipelines for video-caption and image-description datasets at multi-terabyte scale. Collaborated with research teams to define annotation schemas for novel visual reasoning tasks and translated ambiguous specifications into clear labeling guidelines. Authored an internal benchmark suite for multimodal model evaluation by implementing multiple metric modules and setting a standard QA gate for model releases. • Ingested and pre-processed ~4TB of video-caption and image-description data using Python and FFmpeg • Defined annotation schemas and labeling guidelines used by 120+ remote annotators • Implemented 14 Python evaluation metric modules with strong unit-test coverage • Supported multimodal benchmark suite adoption as a release QA gate

2019 - 2020

Junior AI Research Assistant, DeepVision Research Group

ImageImageFine-tuningFine-tuningClassificationClassification

Assisted with preparing training datasets for image classification and image-text alignment experiments. Performed data cleaning, deduplication, and metadata tagging for a large image corpus to support downstream training. Supported fine-tuning of ResNet and early CLIP-predecessor models with experiment tracking and documentation that informed research reporting. • Cleaned and deduplicated a 500K-image corpus and tagged metadata for alignment experiments • Supported fine-tuning of ResNet and early CLIP-predecessor models • Tracked experiments and documented hyperparameter sensitivity analyses • Contributed to training dataset readiness for published research

2018 - 2019

Education

N

N/A

Master of Science, Machine Learning

Master of Science
2018 - 2018
N

N/A

Bachelor of Science, Computer Science

Bachelor of Science
2016 - 2016

Work History

N

NovaSense AI Ltd.

Senior Multimodal AI Training Specialist

N/A
2022 - Present
C

Cognitive Loop Technologies

AI Data Training Analyst

N/A
2020 - 2021