For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
N
Niranjan S.

Niranjan S.

Virtual Teaching Assistant (VTA) — multimodal learning assistant development with RAG, OCR, and cited Q&A

USA flagN/A, Usa

Key Skills

Software

No software listed

Top Subject Matter

Educational tutoring and multimodal learning (text, vision-language, audio)

Top Data Types

DocumentDocument
ImageImage

Top Task Types

Question AnsweringQuestion Answering
Computer Programming/CodingComputer Programming/Coding
Text SummarizationText Summarization

Freelancer Overview

My experience in software development and AI engineering heavily intersects with data labeling, data curation, and AI training data preparation. Throughout my technical projects and professional roles, I have engineered critical pipelines that ingest, clean, and structure unstructured data for downstream machine learning tasks. Notably, during my Virtual Teaching Assistant role, I built a multi format document processing pipeline that integrated OCR extraction to process text, images, and tables from diverse formats like PDFs and Word documents. Additionally, I implemented a Retrieval Augmented Generation (RAG) pipeline utilizing semantic vector search and all-MiniLM-L6-v2 embeddings. This required a deep, highly analytical approach to text chunking, document parsing, and semantic data preparation to ensure high quality, ground truth aligned contexts for AI model inference. My qualification for advanced AI training data workflows is further distinguished by my hands-on experience structuring data for specialized domain tasks and evaluating model outputs. I have worked directly with medical classification datasets, such as the Wisconsin Breast Cancer Dataset, which demanded rigorous data preprocessing and validation to train a neural network to 96.5% accuracy. Furthermore, I engineered an Agile task generator that parses and translates audio transcripts into highly structured text datasets, containing distinct technical requirements and acceptance criteria, using Hugging Face's Mixtral model. Combined with my proficiency in embedding services, prompt optimization, and vector databases like ChromaDB, I possess the rigorous technical skillset, attention to detail, and analytical depth required to curate, label, and validate complex datasets for sophisticated AI training applications.

Labeling Experience

Virtual Teaching Assistant (VTA) — multimodal learning assistant development with RAG, OCR, and cited Q&A

DocumentDocumentQuestion AnsweringQuestion Answering

Built a multimodal AI-driven learning assistant that supports source-grounded student answers with citations and page references. Implemented a Retrieval-Augmented Generation pipeline using semantic vector search to route questions to appropriate AI models. Developed OCR-based processing for multi-format learning materials to enable structured retrieval and accurate responses. • Supports PDF, Word, Text, and audio (MP3, WAV, OGG, M4A, FLAC) ingestion • Performs OCR extraction for images and tables within documents • Generates cited, page-referenced answers via RAG • Implements intelligent query routing to select Gemini and Nemotron models

2026 - Present

Education

W

Washington State University

Bachelor of Science, Computer Science

Bachelor of Science
2023 - 2027

Work History

V

Virtual Teaching Assistant (VTA)

Virtual Teaching Assistant (Multimodal AI Developer)

N/A
2026 - Present
E

Employee Portal

Software Engineering Intern

Pullman
2025 - 2025