For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
H
Hadera Gebru G.

Hadera Gebru G.

Project Flamingo | RLHF, Pairwise Model Evaluation | Dataannotation.tech

Germany flagFrankfurt, Germany

Key Skills

Software

Data Annotation TechData Annotation Tech
MercorMercor

Top Subject Matter

LLM alignment and coding-focused instruction-following evaluation
Tool-use Domain Expertise
Rag Domain Expertise

Top Data Types

TextText
VideoVideo
DocumentDocument

Top Task Types

RLHFRLHF
TrackingTracking
Red TeamingRed Teaming

Freelancer Overview

Project Flamingo | RLHF, Pairwise Model Evaluation | Dataannotation.tech. Brings 5+ years of professional experience across complex professional workflows, research, and quality-focused execution. Core strengths include Data Annotation Tech and Mercor. Education includes Bachelor of Science, Addis Ababa Science and Technology University (2021). AI-training focus includes data types such as Computer Code, Programming, and Text and labeling workflows including RLHF, Evaluation, and Rating.

Labeling Experience

Mercor

Tool-use evaluation, Tool-augmented LLM evaluation, RAG, Multimodal comprehension Test | Micro1 & Mercor

MercorMercor

Evaluated whether the model can generate effective search queries and correctly use external sources under retrieval constraints. Performed retrieval-augmented generation (RAG-style) and multimodal comprehension evaluation using video understanding via transcripts or captions. Assessed the model’s ability to retrieve relevant information with search tools and produce accurate instruction-following summaries. • Tested tool use for search query generation • Verified retrieval constraint adherence and source usage • Evaluated multimodal/video comprehension from transcripts or captions • Rated summary correctness for instruction-following outputs

2025 - Present

RLHF and Pairwise Model Evaluation Engineer - Dataannotation.tech

TextTextTrackingTrackingRLHFRLHF

You created real-world coding prompts to evaluate instruction-following performance and model outputs across multiple programming languages. You assessed code correctness, constraint adherence, and reasoning quality while designing controlled multi-turn interaction flows. You applied RLHF-style pairwise evaluation concepts to surface failures, reasoning drift, and long-context consistency issues for downstream alignment use cases. • Designed and generated coding prompts in Python, JavaScript, and C++ • Built 15-turn multi-step interactions with follow-ups, memory checks, and deviation tracking • Conducted red-team style stress testing to expose constraint preservation failures • Supported preference optimization pipelines and coding-focused LLM evaluation benchmarks

2022 - 2024
Data Annotation Tech

Project Flamingo | RLHF, Pairwise Model Evaluation | Dataannotation.tech

Data Annotation TechData Annotation TechRLHFRLHF

Created coding prompts in Python, JavaScript, and C++ to evaluate instruction-following and model correctness. Conducted 15-turn multi-step interactions with follow-ups, memory checks, and deviation tracking to test long-context understanding and constraint preservation. Performed controlled multi-turn red-teaming style stress tests to expose reasoning drift and failures, contributing to alignment datasets and preference evaluation pipelines. • Evaluated code correctness, constraint adherence, and reasoning quality • Measured whether models preserve prior context and maintain coherence • Tracked deviations across turns to identify hallucination drift • Supported RLHF/pairwise model evaluation benchmark dataset creation

2022 - 2024

Education

A

Addis Ababa Science and Technology University

Bachelor of Science, Software Engineering

Bachelor of Science
2018 - 2021

Work History

M

Micro1 & Mercor

LLM Evaluation Engineer

N/A
2025 - 2026
M

MoEving

Full-Stack Developer

N/A
2022 - 2024