For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
P
Pulkit S.

Pulkit S.

Data Analysis Intern, Search Education

India flagNew Delhi, India

Key Skills

Software

Other

Top Subject Matter

CRM analytics and BI reporting
Customer behavior analytics and ROI modeling
LLM preference tuning / UGC moderation quality and trust & safety review

Top Data Types

DocumentDocument
ImageImage
TextText

Top Task Types

Data CollectionData Collection
RLHFRLHF
Emotion RecognitionEmotion Recognition

Freelancer Overview

Data Analysis Intern, Search Education. Brings 2+ years of professional experience across complex professional workflows, research, and quality-focused execution. Core strengths include Other, HuggingFace, and Internal. Education includes Bachelor of Technology, VIPS, New Delhi (2027) and High School Diploma, Dwarka International School (2023). AI-training focus includes data types such as Data Collection, Computer Code, and Programming and labeling workflows including Data Collection, RLHF, and Evaluation.

Labeling Experience

Data Analysis Intern, Search Education

OtherData CollectionData Collection

Worked on dataset preparation and analysis for churn-pattern insights using structured CRM data. The internship involved cleaning, integrating, and summarizing records to support reporting and strategy decisions. Outputs were delivered as dashboards and automated reports for stakeholders. • Analyzed ~15,000+ CRM records to surface churn patterns • Built 5+ automated MIS reports in Excel/Power BI to reduce manual effort • Cleaned and integrated data from 4 source systems into analysis-ready format • Produced BI dashboards informing quarterly strategy decisions

2025 - 2025

AI Image Caption Generator (TensorFlow, CNN-LSTM, Keras)

Trained and evaluated an image caption generation model using standard captioning metrics to identify systematic weaknesses. The workflow included scoring generated captions with BLEU and analyzing failure cases by object category. The outputs were produced and validated against dataset baselines. • Trained a CNN-LSTM encoder-decoder on MS-COCO (~83k images) • Implemented BLEU scoring and analyzed failure cases across object categories • Reported BLEU-4 of 0.29 comparable to published baselines • Generated captions in <2 seconds while surfacing model weaknesses

2024 - 2024

Real-Time Emotion Detector (DeepFace, WebSocket, FastAPI)

OtherImageImageEmotion RecognitionEmotion Recognition

Implemented a real-time facial emotion detection system and performed systematic benchmarking across operating conditions. The work included evaluating model behavior under varying frame rates and lighting and documenting failure modes. Results were suitable for ML output evaluation. • Live facial emotion detection with <100ms WebSocket latency • Achieved ~92% accuracy on FER-2013 at sustained 30fps load • Benchmarked across frame rates and lighting conditions • Documented model failure modes for output evaluation

2024 - 2024

DocuAgent System (PDF/DOCX/TXT AI pipeline)

DocumentDocument

Developed a multi-agent document processing pipeline with evaluation instrumentation for downstream output quality review. The system logs agent activity to enable post-hoc evaluation of model decisions and policy compliance auditing. It supports retrieval and concurrent querying over ingested documents. • Built an 8-agent AI pipeline for PDF/DOCX/TXT processing • Used Pinecone namespaces for per-document retrieval context • Implemented an Agent Activity Log for chain-of-thought evaluation • Enabled post-hoc output quality review and policy compliance auditing

2024 - 2024

RLHF Preference Trainer (Python, HuggingFace, TRL, LoRA/PEFT, Gradio)

RLHFRLHF

Built an RLHF preference training system to create reward signals from labeled preference pairs and then improve a base language model via PPO. The workflow included documentation of disagreement among labels, collection of preference data, and evaluation-oriented reward modeling. The project emulated trust-and-safety style review and audit practices. • Collected ~1,200 labeled preference pairs using a Gradio annotation UI • Documented inter-annotator agreement (Cohen's kappa 0.67) and disagreement categories • Trained a Bradley-Terry reward model from preference pairs • Fine-tuned GPT-2 Medium via PPO (TRL), improving mean reward by +0.31 over SFT

2024 - 2024

Education

V

VIPS, New Delhi

Bachelor of Technology, Artificial Intelligence and Data Science

Bachelor of Technology
2023 - 2027
D

Dwarka International School

High School Diploma, Secondary Education

High School Diploma
2023 - 2023

Work History

S

Search Education

Data Analysis Intern

New Delhi
2025 - 2025
C

Coratia Technologies

Data Science Intern

N/A
2024 - 2024