For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
M
Mohammed B.

Mohammed B.

ConvoSense – LLM Output Evaluation & NLP Annotation Pipeline | Python, Transformers (2026 project)

United Arab Emirates flagAlain, Abu dhabi, United Arab Emirates

Key Skills

Software

Other

Top Subject Matter

LLM evaluation and NLP annotation
LLM red teaming and agent evaluation
Visual data annotation and AI output validation

Top Data Types

TextText
ImageImage
3D Sensor3D Sensor

Top Task Types

Red TeamingRed Teaming
ClassificationClassification
RLHFRLHF

Freelancer Overview

ConvoSense – LLM Output Evaluation & NLP Annotation Pipeline | Python, Transformers. Core strengths include Hugging Face, Python, and LLM APIs. Education includes Bachelor of Technology, Integral University (2026). AI-training focus includes data types such as Text, Image, and 3D Sensor and labeling workflows including RLHF, Red Teaming, and Classification.

Labeling Experience

DesignIQ – Visual Data Annotation & AI Output Validation | Python, Vision Transformers (2026 project)

OtherImageImageClassificationClassification

Annotated and labeled UI screenshot datasets across multiple visual categories including layout structure, spacing, hierarchy, and accessibility attributes. Reviewed and validated machine-generated design quality scores by correcting factual and formatting errors against ground truth benchmarks. Produced structured chain-of-thought reasoning labels for LLM-based design feedback training to reduce annotation ambiguity and iteration cycles.•Dense multi-page visual annotation with layout and accessibility attribute tags•Image quality score validation and correction workflow•Creation of structured reasoning labels for feedback training data•Ground-truth benchmark alignment to maintain labeled accuracy (85%)

2026 - Present

ConvoSense – LLM Output Evaluation & NLP Annotation Pipeline | Python, Transformers (2026 project)

TextTextRLHFRLHF

Built and evaluated transformer-based text classification models for sentiment, emotion, tone, and intent using RLHF-style response ranking across five label categories. Designed annotation protocols for ambiguous conversational inputs and resolved edge cases through structured logical reasoning. Conducted prompt-template iterations and assessed output quality against human preference benchmarks while correcting formatting, factual, and fluency errors.•Sentiment, emotion, tone, and intent labeling for conversation responses•Preference labeling and output quality assessment for model selection•Ambiguous-input protocol design to preserve consistent label semantics•Quality assurance documentation and audit-ready reporting to improve labeling consistency by 35%

2026 - Present

SmartGrow – Structured Data Labeling & Time-Series Annotation | ESP32, Python, MQTT (2025 project)

Other3D Sensor3D SensorClassificationClassification

Labeled and categorized continuous sensor data streams across eight measurement channels for ML training. Applied precise tagging protocols to classify normal versus anomalous readings for anomaly-detection datasets. Reviewed and corrected automated threshold-based classification outputs to improve annotation accuracy across a multi-day time-series window.•Time-series tagging across eight sensor channels•Normal vs anomalous classification label creation•Post-processing and correction of automated threshold logic outputs•Sustained annotation accuracy over 30 days of structured sensor data

2025 - Present

OmniAssist – Multimodal AI Red Teaming & Agent Evaluation | Python, LLM APIs, FastAPI (2025 project)

TextTextRed TeamingRed Teaming

Crafted adversarial and edge-case prompts to test LLM safety boundaries, bias resistance, and jailbreak limitations across text and voice-related inputs. Evaluated and ranked AI agent responses for accuracy, coherence, and task completion quality using structured preference labeling to guide model improvement cycles. Documented failure modes in contextual memory and reasoning workflows and produced correction reports based on observed quality gaps.•Adversarial prompt generation for safety and robustness testing•Preference labeling over ranked agent responses•Failure-mode documentation and correction reporting for improved outputs•Bias probing and edge-case coverage across modalities

2025 - Present

Education

I

Integral University

Bachelor of Technology, Computer Science Engineering (Data Science and AI)

Bachelor of Technology
2022 - 2026

Work History

C

Company not specified

Built and evaluated NLP and LLM pipelines through academic projects, including dataset annotation, model output ranking,

Location not specified
Not specified