For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
S

Stephan O.

Project Coordinator (PMO) in Contract Review, Compliance, and Legal Research

United Kingdom flagWednesbury, United Kingdom

Key Skills

Software

Scale AIScale AI

Top Subject Matter

Legal Services & Contract Review
Regulatory Compliance & Risk Analysis
Legal Research & Document Analysis

Top Data Types

DocumentDocument
TextText
ImageImage

Top Task Types

RLHFRLHF
Evaluation/RatingEvaluation/Rating

Freelancer Overview

Over the past two months, I have been actively working as an AI Training Specialist on the Outlier platform, directly engaging in data labelling, reinforcement learning from human feedback (RLHF), and prompt engineering. In this role, I evaluate AI-generated outputs for factual accuracy, logical coherence, and alignment with strict project guidelines. This hands-on experience is strongly backed by my academic and professional background in data analysis; I hold an MSc in Business Analytics from Aston University and have extensive experience as a Project Coordinator and PMO Analyst. This combination allows me to approach AI training data not just as a labeller, but with a highly analytical, data-driven mindset that ensures the highest quality training inputs. What sets me apart is my proven ability to manage complex data structures and maintain rigorous quality control standards. In my previous roles such as my time at De-Light Solutions working on project tracking and documentation, and through my advanced proficiency in tools like Microsoft Excel and Jira , I have mastered the art of spotting data anomalies and maintaining 100% compliance with complex frameworks. I am highly adaptable to changing project methodologies, possessing a strong foundation in both Agile and Waterfall systems. My unique blend of direct Outlier experience, advanced data analytics education, and meticulous project governance skills enables me to deliver exceptionally precise, high-quality data labelling that directly enhances AI model performance.

Labeling Experience

AI Data Annotator

ImageImageRLHFRLHF

I served as an AI Training Specialist / Data Annotator executing Reinforcement Learning from Human Feedback (RLHF) for a prominent large language model (LLM) optimization project. The project focused on refining the model's capabilities in handling complex business logic, data structures, and enterprise management applications (such as advanced Excel workflows, SQL query interpretation, and Agile/Jira project frameworks). My day-to-day tasks involved evaluating pairs of AI-generated complex reasoning trajectories, grading them based on strict multi-dimensional rubrics (accuracy, logical flow, formatting constraints, and safety guidelines), and ranking them to train the reward model. A core part of my role was executing fact-checking and grounding procedures to verify technical data and eliminate hallucinations. I regularly drafted detailed analytical rationales (1-2 paragraphs) detailing precisely why a particular model output demonstrated superior logical sequence or constraint adherence, directly contributing to the model's alignment and advanced reasoning capabilities.

2026 - Present

Education

A

Aston University

Bachelor of Science, International Business and Management

Bachelor of Science
Not specified
A

Aston University

Master of Science, Business Analytics

Master of Science
Not specified

Work History

W

Wsec

Project Coordinator (PMO)

Birmingham
2022 - Present
D

De-Light Solutions

AI Assistant Project Manager

Remote
2021 - 2022