For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
D
Debmalya S.

Debmalya S.

Machine Learning Intern (bioinformatics pipeline and statistical/genomic analysis)

India flagHyderabad, India

Key Skills

Software

Other
ClickworkerClickworker
Google Cloud Vertex AIGoogle Cloud Vertex AI

Top Subject Matter

Fine-tuning LLM for Healthcare Records
LLM training / RAG
Agentic AI systems for developers

Top Data Types

TextText
DocumentDocument
ImageImage

Top Task Types

Function CallingFunction Calling
Fine-tuningFine-tuning
ClassificationClassification
Question AnsweringQuestion Answering
Text GenerationText Generation

Freelancer Overview

I have experience working with AI/ML systems, large-scale data processing, and deep learning workflows through internships and independent projects. During my Machine Learning internship at National Institute of Technology Warangal, I developed bioinformatics data-processing pipelines handling over 2 million gene expression records, focusing on data preprocessing, feature extraction, and statistical analysis for machine learning applications. I also built and trained a 45M-parameter Transformer model on a 560M-token dataset using PyTorch and Hugging Face, gaining hands-on experience with AI training pipelines, NLP, experiment tracking, and model optimization. My technical background in Python, deep learning, and large dataset handling, combined with strong analytical and problem-solving skills, makes me well-suited for AI training data and data labeling tasks. Humanity keeps inventing larger datasets to clean because apparently suffering alone wasn’t scalable enough.

Labeling Experience

High-Performance LLM Training Pipeline (Helium-Nano)

OtherFine-tuningFine-tuning

Built and trained a 45M parameter Transformer training pipeline from scratch on a 560M token dataset to support model learning. Developed mixed-precision (bfloat16) training and performance optimizations to improve throughput and stability while reducing GPU memory consumption. Integrated RAG and multi-agent orchestration components to enable retrieval-augmented workflows during experimentation. • Trained a 45M parameter Transformer on a 560M token dataset using PyTorch mixed-precision • Optimized training throughput from 25K to 409K tokens/sec using PyTorch Inductor and CUDA Graphs • Integrated Weights & Biases for experiment tracking, artifacts, and visualization • Experimented with Retrieval-Augmented Generation (LangChain) and multi-agent orchestration (LangGraph)

2024 - Present

Software Engineer Intern (agentic AI automation for CI/CD)

Function CallingFunction Calling

Designed and delivered an agentic AI system that automates CI/CD pipeline setup for GitHub and GitLab environments. Implemented autonomous tool-use workflows with task decomposition and multi-step decision-making pipelines to reduce manual configuration and human error. Engineered a scalable FastAPI backend with async orchestration to support AI-driven workflow execution and production deployment. • Built autonomous tool-use workflows with multi-step decision pipelines • Reduced manual CI/CD configuration effort by 80% and human error by 30% • Implemented an async FastAPI backend for AI workflow execution and inference management • Created composable agent tools to reduce customer support intervention by 40%

2025 - 2025

Machine Learning Intern (bioinformatics pipeline and statistical/genomic analysis)

OtherFunction CallingFunction Calling

Engineered an end-to-end bioinformatics Python package and statistical/genomic analysis workflows used by PhD researchers, supporting preparation and analysis of biological datasets. Implemented vectorized data-processing modules (e.g., Z-score matrix computation, pathway scoring, and disease module preparation) and graph-based gene networking with hypergeometric feature-selection for automated pathway analysis. The work emphasized building repeatable pipelines that reduce manual feature engineering and streamline downstream ML steps. • Built high-performance statistical modeling modules for genomic analysis • Applied graph-based gene networking and hypergeometric feature-selection pipelines • Optimized vectorized processing across datasets containing 2M+ gene expression records • Reduced manual feature engineering effort by 60% for downstream ML workflows

2024 - 2024

Education

N

National Institute of Technology, Warangal

Bachelor of Technology, Computer Science and Engineering

Bachelor of Technology
2022 - 2026

Work History

P

Publicis Sapient

Software Engineer Intern

N/A
2025 - 2025
N

National Institute of Technology, Warangal

Machine Learning Intern

Warangal
2024 - 2024