For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
S
Surafel A.

Surafel A.

Machine Learning Engineer & Data Annotation Specialist

Ethiopia flagN/A, Ethiopia

Key Skills

Software

CVATCVAT
Label StudioLabel Studio

Top Subject Matter

Science and Technology (or Computer Science / Technology)
AI Safety Evaluation (or Model Training & Logic Validation)
Image Annotation QA (or Object Detection & Graphics)

Top Data Types

TextText
ImageImage
VideoVideo

Top Task Types

ClassificationClassification
Data CollectionData Collection
Object DetectionObject Detection
Fine-tuningFine-tuning
Point/Key PointPoint/Key Point

Freelancer Overview

Data Scientist and Machine Learning Specialist with a strong background in structuring high-quality datasets for AI models. Proficient in Python, data labeling workflows, and computer vision pipelines, with hands-on experience managing large-scale data engineering tasks. Notably executed an end-to-end computer vision data annotation pipeline involving over 14,000 images, utilizing advanced thresholding, Fourier transforms, and segmentation techniques to ensure impeccable label accuracy. Graduating with a Bachelor of Science in Data Science from Debre Berhan University (2026), I bring deep academic and practical knowledge in machine learning architectures, statistical modeling, and data classification. Experienced in data collection, text processing, and complex labeling workflows, I specialize in providing the high-fidelity validation and technical precision required to train and evaluate advanced AI models.

Labeling Experience

NLP Sentiment Analysis & Text AI Model Training — Commercial Bank of Ethiopia (CBE) | Sentiment analysis on customer feedback (2024–2025)

OtherTextTextClassificationClassification

Engineered and optimized Natural Language Processing (NLP) text processing pipelines to train, evaluate, and fine-tune machine learning classification models. Extracted, cleaned, and curated unstructured text datasets from Facebook and the Google Play Store, implementing advanced text preprocessing routines—including tokenization, text normalization, and noise reduction—to create high-fidelity training data. Built robust, NLP-driven analysis workflows to classify textual sentiments, validate model accuracy, and evaluate model performance metrics, translating raw conversational data into actionable strategic insights for banking and product systems. Gathered, structured, and processed high-volume social media and application store review text datasets. Applied NLP classification and sentiment analysis techniques to generate clean, model-ready training corpora. Conducted rigorous evaluation and logic validation of model outputs to ensure precision and remove algorithmic bias. Preprocessed and structured unstructured text feedback to optimize downstream supervised fine-tuning (SFT) and reporting.

2024 - Present

Computer Vision Sign Language Recognition AI

VideoVideoClassificationClassification

Engineered and curated a custom computer vision dataset for an end-to-end Sign Language Recognition AI system. Developed an automated data collection and keypoint annotation framework leveraging Google MediaPipe to extract high-fidelity spatial coordinates from real-time video streams. Directed the labeling and processing of multi-dimensional data points, tracking 21 landmarks per hand alongside face and holistic body posing mesh coordinates to establish an incredibly clean, model-ready training corpus. Built and optimized a Multi-Layer Perceptron (MLP) neural network architecture in Python to classify the extracted keypoint data streams into distinct sign language gestures with high precision. Managed dataset hygiene, feature normalization, and class-balancing protocols to eliminate tracking noise and optimize the model for real-time deployment. Managed automated and manual data collection workflows to curate sign language gestures. Utilized MediaPipe for real-time landmark tracking and Point Key Point annotation. Built and trained an MLP classification model in Python using structured coordinate data. Implemented dataset preprocessing, feature engineering, and cross-validation pipelines to maximize model accuracy.

2026 - 2026

Public Health Data Analyst — Academic & Independent Research (EPHI aligned) | Data cleaning & imputation for TB forecasting (2023–2024)

OtherTextTextData CollectionData Collection

Developed and optimized automated data preprocessing pipelines and feature engineering scripts in Python for complex epidemiological datasets. Engineered robust data cleaning routines to handle multi-dimensional data structures, ensuring high-fidelity, bias-free inputs ready for downstream machine learning model training and predictive time-series forecasting. Implemented meticulous data quality screening and missing-value analysis, applying rigorous statistical imputation strategies (such as mean, mode, and contextual logical constraints) to preserve dataset integrity across thousands of rows of global health data. Designed end-to-end data cleaning pipelines and automated preprocessing scripts in Python. Applied rigorous data quality screening and missing-value imputation to optimize training-ready datasets. Structured data features and key performance indicators (KPIs) to drive predictive modeling. Managed dataset hygiene and validation to ensure highly accurate, clean inputs for AI model evaluation.

2023 - 2024

Education

D

Debre Berhan University

Bachelor of Science, Data Science

Bachelor of Science
2022 - 2026

Work History

C

Commercial Bank of Ethiopia (CBE)

Data Scientist Intern

N/A
2024 - 2025
A

Academic & Independent Research

Public Health Data Analyst

N/A
2023 - 2024