For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
O

Ojas J.

LLM Alignment via Direct Preference Optimization (DPO) - RLHF Data Curation & Training

India flagCoimbatore, India

Key Skills

Software

Label StudioLabel Studio
CVATCVAT

Top Subject Matter

Large Language Model Alignment and Fine-tuning
Multimodal Model Red-Teaming and Safety Assessment
Synthetic Data Generation for LLM Training

Top Data Types

TextText

Top Task Types

RLHFRLHF
Red TeamingRed Teaming
Text GenerationText Generation
ClassificationClassification

Freelancer Overview

Core Skills: - Languages: Python (primary), C, SQL, Bash - Data & ML: NumPy, Pandas, Matplotlib, Seaborn, Scikit-learn, PyTorch - Systems: Linux/Unix, Multi-threading, IPC (ZeroMQ), Git Key Projects: 1. HFT Simulation — Built a low-latency market data distribution system using ZeroMQ PUB/SUB and multi-threading, optimizing tick-to-trade latency with thread-safe queues and O(1) rolling statistics. 2. Custom UNIX Shell — Wrote a shell in C from scratch handling background execution, I/O redirection, pipes, and process management via POSIX calls (fork, execvp, waitpid) with leak-free memory handling.

Labeling Experience

Synthetic Data Generation & Model Distillation - Instruction Data Curation

TextTextText GenerationText Generation

I designed synthetic data generation pipelines leveraging GPT-4 and other open-source LLMs to produce instruction samples for legal and financial Q&A. I applied semantic filtering to refine generated data, ensuring high similarity and quality for downstream training. My work enabled efficient model distillation and boosted student model effectiveness. • Generated and filtered 50,000+ synthetic instruction data points • Automated sample selection using FAISS-based semantic similarity for consistency • Created datasets that substantially improved student-teacher knowledge transfer • Deployed and maintained data generation systems in Dockerized environments

2023 - 2024

Multimodal AI Red-Teaming & Safety Evaluation - Data Annotation & RLHF Dataset Creation

TextTextRed TeamingRed Teaming

I designed and executed large-scale adversarial evaluations to probe multimodal models for bias and security vulnerabilities. Leveraging prompt engineering, I systematically structured evaluation tasks and annotated model failures with human and automated scoring. My work supplied annotated RLHF datasets to improve model safety and reliability. • Created more than 500 adversarial test cases for model robustness evaluation • Developed structured prompt templates targeting specific risk domains • Enabled human reviewers to annotate model failures efficiently via custom dashboard • Automated scoring and feedback integration using GPT-4 for rapid dataset expansion

2023 - 2024

Real-Time Financial Sentiment Extraction System - Data Annotation & Model Training

TextTextClassificationClassification

I created and labeled a financial news dataset to support real-time sentiment extraction using a fine-tuned transformer model. My task involved annotating 5,000 news samples for training, achieving high performance in sentiment classification. The data annotation supported robust real-world deployment for market analysis. • Performed manual sentiment labeling on a custom financial dataset • Trained and validated FinBERT model using annotated data for real-time use • Maintained dataset organization and ensured label quality standards • Enabled high-throughput deployment through optimized feature engineering for latency control

2023 - 2023

LLM Alignment via Direct Preference Optimization (DPO) - RLHF Data Curation & Training

TextTextRLHFRLHF

I fine-tuned Mistral-7B-Instruct using preference triplets to improve model alignment with human feedback. I curated and filtered datasets by removing low-quality pairs with embedding similarity thresholds and reward model confidence, optimizing for dataset quality. My work resulted in significant metric improvements on benchmark evaluations and practical deployment of experiments. • Processed and filtered 10,000+ preference triplets for high-quality supervision • Applied preference optimization algorithms such as DPO and PPO for RLHF training • Carefully evaluated model outputs to iteratively improve dataset for LLM alignment • Used inference endpoints and experiment tracking for deployment and validation

2023 - 2023

Education

A

Amrita Vishwa Vidyapeetham

Bachelor of Technology, Computer Science and Engineering

Bachelor of Technology
2024

Work History

A

Amrita Vishwa Vidyapeetham

Student Developer — Computer Science Dept

coimabtore
2016 - Present
B

BharOS

Systems Programmer

Chennai
2025 - 2025