For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
M
Mohamed H.

Mohamed H.

Mercor Intelligence - AI Engineer / Data Scientist / Technical Contributor/ Reviewer

USA flagNew York, Usa

Key Skills

Software

MercorMercor
Snorkel AISnorkel AI

Top Subject Matter

LLM benchmarking
AI evaluation and red-teaming
LLM benchmarking dataset authoring and evaluation

Top Data Types

TextText

Top Task Types

Computer Programming/CodingComputer Programming/Coding

Freelancer Overview

Mercor Intelligence - AI Engineer / Data Scientist / Technical Contributor/ Reviewer. Brings 19+ years of professional experience across complex professional workflows, research, and quality-focused execution. Core strengths include Mercor and Snorkel AI. Education includes Master of Data Science, University of Malaya (2019) and Bachelor of Computer Science, 6th October University (2009). AI-training focus includes data types such as Computer Code and Programming and labeling workflows including Evaluation and Rating.

Labeling Experience

Snorkel AI

Snorkel AI - AI Engineer / Task Author

Snorkel AISnorkel AI

Designed and implemented multi-step coding and reasoning tasks for the Terminal-Bench benchmark under the Terminus framework. Curated datasets and built automated evaluation pipelines to test complex problem-solving and reasoning capabilities in frontier GPT-level systems. Focused on dataset preparation and benchmark task authoring to enable consistent model scoring and comparison.• Authored multi-step coding/reasoning tasks for CLI/agentic evaluation.• Built automated evaluation pipelines for benchmark runs.• Curated high-difficulty datasets for LLM benchmarking.• Supported assessment of complex reasoning behaviors.

2025 - Present
Mercor

Mercor Intelligence - AI Engineer / Data Scientist / Technical Contributor/ Reviewer

MercorMercor

Worked as an AI Engineer / Data Scientist contributing to AI evaluation workflows for LLM benchmarking and red-teaming tasks. Built and refined Python-based evaluation pipelines, grading rubrics, and reasoning frameworks to assess model performance using structured scoring. Focused on data analytics and API-driven automation to support evaluation experiments for AI labs.• Contributed to SWE-Bench related evaluation activities.• Developed/maintained evaluation pipelines and assessment rubrics.• Supported LLM red-teaming and code-review automation.• Tracked model performance metrics for ongoing benchmarking.

2024 - Present

Education

6

6th October University

Bachelor of Computer Science, Computer Science

Bachelor of Computer Science
2007 - 2009
U

University of Malaya

Master of Data Science, Data Science

Master of Data Science
2019

Work History

V

Verisk Analytics

Lead Software Engineer

New York
2023 - Present
C

ContactCars.Com

Tech Lead, Backend & AI Engineering

Cairo
2020 - 2022