For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
A
Aditya S.

Aditya S.

AI Model Evaluation & LLM Output Assessment (Personal Research & Projects)

India flagDelhi, India

Key Skills

Software

Other
Data Annotation TechData Annotation Tech
Google Cloud Vertex AIGoogle Cloud Vertex AI

Top Subject Matter

LLM output evaluation and benchmark-based assessment for continual learning.
Retinal fundus image classification and uncertainty/failure-case assessment.
Continual learning research with uncertainty-based evaluation for data quality/annotation prioritization.

Top Data Types

TextText
ImageImage
AudioAudio

Top Task Types

DiagnosisDiagnosis
RLHFRLHF
ClassificationClassification

Freelancer Overview

With hands-on experience building, training, and evaluating ML systems, including a retinal classifier trained on 91,000+ images and a continual learning framework designed to catch model failure modes, I bring a level of technical depth to AI training work that goes well beyond typical annotation experience. My research focuses specifically on how and why models hallucinate, forget, and misfire. That foundation means I don't just evaluate outputs at face value, I understand what's happening underneath, which translates directly into higher-quality, more consistent annotations and a sharper eye for edge cases that matter.

Labeling Experience

The Ego Gate — Continual Learning Framework (Open Preprint)

Other

Designed a selective learning framework that requires systematic evaluation of model behavior to detect catastrophic forgetting patterns. Formalized model uncertainty (predictive entropy) and knowledge gaps as measurable signals that can inform data quality and annotation/label correction priorities. Benchmarked the framework to quantify and improve continual learning behavior. • Built evaluation-driven selective learning methods to surface forgetting. • Used measurable uncertainty/entropy signals to represent knowledge gaps. • Benchmarked model behavior to guide which data/labels to focus on. • Connected evaluation outputs to downstream data quality/annotation workflows.

2025 - Present

AI Model Evaluation & LLM Output Assessment (Personal Research & Projects)

OtherTextText

Evaluated outputs from Claude, GPT-4, and open-source LLMs by prompting and systematically assessing reasoning quality, hallucinations, and instruction-following failures. Compared generated model outputs to ground-truth benchmarks as part of continual learning research to identify discrepancies and gaps in behavior. Used these evaluations to guide selective/continual learning frameworks focused on measurable uncertainty and forgetting patterns. • Prompted and assessed LLM responses across multiple research engineering projects. • Benchmarked outputs against ground-truth continual learning setups (EWC, PNN, replay-based CL). • Analyzed uncertainty-related signals (e.g., predictive entropy) for identifying knowledge gaps. • Identified failure cases analogous to flagging incorrect outputs/labels for downstream improvement.

2024 - Present

AEGIS — AI Safety Incident Response Environment (Personal Research)

OtherRLHFRLHF

Built an OpenAI Gym-compliant reinforcement learning environment for AI safety triage that depends on careful labeling of agent actions and outcomes. Defined reward signals and evaluated policy correctness using the labeled interactions between the agent and environment. The work functioned as structured outcome labeling to support safe decision-making training logic. • Labeled agent actions and outcomes to construct reward signals. • Implemented environment dynamics in an RL setting for safety triage. • Evaluated policy correctness based on labeled interaction outcomes. • Ensured accurate mapping from observed outcomes to training/reward labels.

2026 - 2026

RetinAI — Multi-Disease Retinal Classifier (AI Craft Research Symposium)

OtherDiagnosisDiagnosis

Trained and evaluated retinal disease classification models on a large fundus image dataset spanning three disease classes. Assessed prediction confidence and uncertainty scores and examined failure cases to improve labeling-like quality control of model outputs. Validated the final pipeline on held-out test sets using standard performance metrics. • Worked with ~91,000 fundus images across Diabetic Retinopathy, Glaucoma, and Pathologic Myopia. • Evaluated confidence, uncertainty, and failure cases analogous to reviewing incorrect predictions. • Achieved 98.57% accuracy and 99.81% AUC-ROC on held-out test sets. • Focused on error analysis to flag unreliable outputs for model refinement.

2026 - 2026

Education

A

Amity University, Noida

Bachelor of Technology, Computer Science Engineering

Bachelor of Technology
2024 - 2028

Work History

I

Independent AI Research & Development

AI/ML Engineer & Researcher

Delhi
2025 - Present