For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
F
Freelancer

Freelancer

Data Scientist – Automated Verbal Assessment (Evaluation/Rating)

India flagNoida, India

Key Skills

Software

Other
Don't disclose

Top Subject Matter

Voice AI assessment and evaluation
Voice AI pipeline operation and evaluation data generation
LLM-as-judge evaluation for coaching

Top Data Types

AudioAudio
TextText
ImageImage

Top Task Types

Fine-tuningFine-tuning
TrackingTracking

Freelancer Overview

Data Scientist – Automated Verbal Assessment (Evaluation/Rating). Brings 2+ years of professional experience across complex professional workflows, research, and quality-focused execution. Core strengths include Other and Don't disclose. Education includes Bachelor of Technology, Indian Institute of Technology, Bombay (2024). AI-training focus includes data types such as Audio, Text, and Image and labeling workflows including Evaluation, Rating, and Fine-tuning.

Labeling Experience

Data Scientist (Manager) - Astra Global

ImageImageTrackingTracking

Engineered and shipped production GenAI systems focused on voice-first retrieval, AI observability, and multi-tenant automation for enterprise users. Built end-to-end orchestration pipelines using LangChain/LangGraph and vector-store RAG with rigorous latency, cost, and safety guardrails. Managed conversational and real-time voice agent experiences including voicemail workflows, live transfer, and ongoing multimodal assistant development. • Designed multi-model RAG pipelines (OpenAI/Gemini) integrated with realtime APIs and Pinecone for sub-second voice retrieval. • Developed observability and usage tracking stacks (Chrome extension plus native host) instrumenting multiple custom GPT and NotebookLM deployments. • Automated operational workflows (dashboard crawling/scraping, voicemail) and implemented LangGraph/LangSmith/Langfuse tracing with prompt-version tracking. • Optimized virtual transfer and virtual assistant agents using eval-driven iteration (GEPA, DSPy) and self-hosted model inference on cloud infrastructure.

2025 - Present

Data Scientist – (Building) AI Supervisor / Coach (Evaluation/Rating)

OtherTextText

Designed an LLM-judge-based multi-agent coaching system that evaluates recordings and transcripts against a portfolio-specific rubric. The system cross-evaluates transcript-judge and recording-judge outputs using a meta-judge to arbitrate and produce role-specific actions. This constitutes structured AI evaluation work and the creation of labeled assessment signals for training and iteration of the grading pipeline. • Built a multi-agent LLM judge system using Google ADK. • Graded recordings and transcripts against an 11-point checklist. • Implemented cross-evaluation and meta-judge arbitration. • Generated manager/supervisor/agent actions derived from rubric-based judgments.

2025 - Present

Data Scientist – VocalDirect Voicemail Platform

Don't discloseAudioAudioFine-tuningFine-tuning

Developed and operated a voicemail and voice assistant stack that uses self-hosted speech/TTS pipelines to generate and manage voice artifacts at production scale. While not explicitly stated as model fine-tuning, the work includes configuring and running domain-specific voice processing components that can be used to prepare training/evaluation data and validate performance. This role supports continuous improvement of voice interaction quality through large-scale voice generation and throughput testing. • Built VocalDirect multi-threaded voicemail platform (5k+ voicemails/day). • Used self-hosted Qwen TTS pipeline with 300k+ weekly TTS across 20+ portfolios. • Orchestrated retrieval/processing for call/voicemail workflows. • Deployed using self-signed binaries and cloud delivery mechanisms for production readiness.

2025 - Present

Data Scientist – Automated Verbal Assessment (Evaluation/Rating)

OtherAudioAudio

Built and shipped systems that automatically score spoken responses from conversational voice interviews, converting voice signals into structured assessment outcomes. The platform focuses on evaluating pronunciation, fluency, clarity, and confidence based on a fixed question set and produces assessment records for ongoing tracking. Results are generated from speech-to-text/voice interview flows suitable for training-time evaluation and quality measurement for LLM/voice experiences. • Automated verbal assessment via a React/FastAPI/Firebase application with SpeechSuper. • Conducted 3-question conversational voice interview scoring. • Produced 100+ assessments to date for model/product evaluation. • Implemented end-to-end orchestration for repeatable assessment runs.

2025 - Present

Education

I

Indian Institute of Technology, Bombay

Bachelor of Technology, Computer Engineering

Bachelor of Technology
2020 - 2024

Work History

A

Astra Global

Data Scientist (Manager)

Noida
2025 - Present