For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
F
Freelancer

Freelancer

AI Engineer | LLM Evaluation & Multi-Agent Systems Specialist

India flagHyderabad, India

Key Skills

Software

Scale AIScale AI
RemotasksRemotasks
OneFormaOneForma
Micro1

Top Subject Matter

Artificial Intelligence & Machine Learning
Large Language Model (LLM) Evaluation
Retrieval-Augmented Generation (RAG)
Model Fine-Tuning
Multi-Agent Systems

Top Data Types

TextText
DocumentDocument
Computer Code ProgrammingComputer Code Programming

Top Task Types

Evaluation/RatingEvaluation/Rating
RLHFRLHF
Red TeamingRed Teaming
Computer Programming/CodingComputer Programming/Coding
Prompt + Response Writing (SFT)Prompt + Response Writing (SFT)
Fine-tuningFine-tuning
SegmentationSegmentation
ClassificationClassification

Freelancer Overview

My experience in AI training and data labeling spans several years and covers a range of task types — from basic annotation and classification to more nuanced work involving model preference ranking, instruction-following evaluation, and RLHF-style feedback. I've worked on tasks that require assessing response quality across multiple dimensions: factual accuracy, tone, helpfulness, and adherence to specific instructions. Over time I've developed a strong eye for subtle differences between model outputs, which is especially valuable when distinguishing a good response from one that merely looks correct on the surface. More recently, my work has shifted toward agentic AI evaluation — specifically assessing how well large language models plan, coordinate across tools, handle friction, and produce verifiable end-state artifacts. I'm comfortable working with trajectory-based evaluations, writing outcome-focused rubrics that hold up under scrutiny, and identifying safety failures across domains like private data handling, high-stakes actions, and ambiguous requests. I understand what it takes to design tasks that actually differentiate model capability rather than tasks where every model performs the same — which is ultimately what rigorous AI evaluation should be about.

Labeling Experience

I have experience working on AI training and data labeling tasks, including text annotation, response quality evaluation

I have experience working on AI training and data labeling tasks, including text annotation, response quality evaluation, and preference ranking between model outputs. My work has involved reviewing AI-generated content for accuracy, relevance, and instruction-following, as well as flagging responses that contain factual errors, unsafe behavior, or policy violations. I'm familiar with evaluating outputs across multiple dimensions — tone, helpfulness, factuality, and adherence to specific constraints — and providing structured feedback that helps improve model performance. More recently, I've been involved in agentic AI evaluation, where I assess how well models plan and execute multi-step tasks, coordinate across tools, and produce verifiable end-state artifacts. This includes identifying safety failures in model behavior, building evaluation rubrics, and rating model outputs based on outcome-focused criteria. I'm comfortable working with detailed guidelines and applying consistent judgment across a high volume of tasks.

Not specified

Education

D

Degree: Bachelor of Technology (B.Tech) Field of Study: Computer Science and Engineering Institution: RVR & JC College o

Degree: Bachelor of Technology (B.Tech) Field of Study: Computer Science and Engineering Institution: RVR & JC College of Engineering, Guntur, Andhra Pradesh, India Certification: Evaluating and Debugging Generative AI Issued by: DeepLearning.AI (Coursera) Relevance: Covers model evaluation, debugging generative AI outputs, and performance tracking — directly applicable to AI training and evaluation workflows.

Degree: Bachelor of Technology (B.Tech) Field of Study: Computer Science and Engineering Institution: RVR & JC College of Engineering, Guntur, Andhra Pradesh, India Certification: Evaluating and Debugging Generative AI Issued by: DeepLearning.AI (Coursera) Relevance: Covers model evaluation, debugging generative AI outputs, and performance tracking — directly applicable to AI training and evaluation workflows.
Not specified

Work History

T

Tata Consultancy Services Limited

AI Engineer

Hyderabad
2023 - Present