For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
S
Sanay S.

Sanay S.

Data Integrity Intern

USA flagAurora, Usa

Key Skills

Software

No software listed

Top Subject Matter

Legal Services & Contract Review
Regulatory Compliance & Risk Analysis
Legal Research & Document Analysis

Top Data Types

VideoVideo
TextText
DocumentDocument

Top Task Types

SegmentationSegmentation

Freelancer Overview

During my Data Integrity internship at Ubiquity, I worked directly with large-scale data quality and labeling-adjacent tasks, developing Python algorithms and SQL queries to de-duplicate and clean records for over 150,000 customers. This involved precise record matching, identifying repeat entries by address and metadata, and applying consistent labeling logic at scale — work that reduced downstream BI dashboard errors by 20% and improved customer segmentation accuracy by over 40%. Earlier, as an Internal Audit intern, I reconciled financial records with 100% accuracy across $100,000+ in monthly cash flow, reinforcing the meticulous attention to detail and consistency that high-quality data annotation demands. Beyond professional experience, I've built data-centric projects that strengthen my fit for AI training work, including a Python stock analysis tool that processed five years of historical data across 50+ companies and a financial dashboard analyzing multi-year datasets. I'm proficient in Python, SQL, and Excel, comfortable working within structured guidelines and quality-control standards, and bring strong analytical judgment from finance and accounting coursework (4.00 technical GPA). Combined with my ability to follow detailed protocols, maintain accuracy at scale, and collaborate cross-functionally, I'm well-positioned to deliver reliable, consistent labeling and contribute to building high-quality training datasets.

Labeling Experience

Data Integrity Intern

VideoVideoSegmentationSegmentation

As a Data Integrity Intern, they developed Python and SQL solutions to clean and reconcile customer data for a fiber internet service. They identified duplicates by matching records using addresses and customer metadata. They collaborated with engineering and product teams to implement data hygiene protocols that improved reporting quality and segmentation accuracy. • Developed Python-based data cleaning algorithms and SQL queries to de-duplicate over 150,000 customer records • Matched repeat entries using address and customer metadata to ensure accurate records • Implemented data hygiene protocols in collaboration with cross-functional engineering and product teams • Reduced downstream BI dashboard errors by 20% and improved customer segmentation accuracy by over 40%

2025 - 2025

Education

U

University of Illinois at Urbana-Champaign

Bachelor of Science, Finance

Bachelor of Science
2024 - 2028

Work History

U

Ubiquity

Data Integrity Intern

Dallas
2025 - 2025
S

Seba

Internal Audit Intern

Aurora
2024 - 2024