For employers

Hire this AI Trainer

Sign in or create an account to invite AI Trainers to your job.

Invite to Job
A
Aisyah R.

Aisyah R.

ML Research Engineer

Malaysia flagKuala Lumpur, Malaysia

Key Skills

Software

Other

Top Subject Matter

Speech recognition and speech synthesis (Malaysian/Arabic languages, dialectal robustness)
Natural Language Processing
GenAI chatbot

Top Data Types

AudioAudio
TextText
ImageImage

Top Task Types

Fine-tuningFine-tuning
Prompt + Response Writing (SFT)Prompt + Response Writing (SFT)
ClassificationClassification

Freelancer Overview

ML Research Engineer — Speech-to-Text / Text-to-Speech model training (Revolab AI). Brings 5+ years of professional experience across complex professional workflows, research, and quality-focused execution. Core strengths include Internal, Proprietary Tooling, and Other. Education includes Bachelor of Computer Science (Data Science), University of Malaya (2023) and Foundation in Science, Centre of Foundation Studies, University of Malaya (2019). AI-training focus includes data types such as Audio and Text and labeling workflows including Fine-tuning, Prompt + Response Writing (SFT), and Classification.

Labeling Experience

ML Research Engineer — Speech-to-Text / Text-to-Speech model training

AudioAudioFine-tuningFine-tuning

Performed end-to-end training for Speech-to-Text and Text-to-Speech systems tailored to Malaysian and Arabic dialects, improving robustness in real-world conditions. Worked with transformer architectures, self-supervised learning, and neural vocoders to build and refine ASR/TTS pipelines. Enhanced multilingual performance using data augmentation, noise-robust training, and multilingual transfer learning strategies. • Train and refine STT/TTS deep learning models • Apply transformer-based and self-supervised learning approaches • Use neural vocoders for speech generation • Optimize accuracy and latency for production use

2025 - Present

Malaysia AI Volunteer - ML Engineer — LLM/speech model training and multimodal research

TextTextFine-tuningFine-tuning

Contributed to preparing Malaysian datasets by writing crawling scripts to support pretraining and finetuning of large language models for the Malaysian context. Pretrained and fine-tuned Whisper for speech recognition and trained Mistral and Llama-based models using DeepSpeed for memory-efficient training. Conducted research on multimodal large language model training with multi-audio, multi-image, and multi-modal pretraining, integrating CLIP vision and speech components. • Crawl and prepare Malaysian dataset for LLM pretraining • Pretrain and fine-tune Whisper, Mistral, and Llama with DeepSpeed • Research multimodal multi-audio/multi-image pretraining • Implement distributed training orchestration with Ray across many GPUs

2023 - Present

Executive, Advanced Analytics — NLP fine-tuning and GenAI chatbot development (Tenaga Nasional Berhad)

OtherTextTextFine-tuningFine-tuning

Fine-tuned BERT to analyze employee feedback text for improved response efficiency and automated summarization. Built and integrated GenAI chatbot capabilities for internal staff using an LLM toolchain to support organizational workflows. Developed data-driven machine learning components for targeted analytics use cases to improve collection efficiency and reporting. • Fine-tune BERT on employee feedback for analysis and summarization • Support internal GenAI chatbot development using Claude and LangChain • Develop ML models for customer targeting and debt collection efficiency • Build analytics pipelines supporting downstream AI systems

2023 - 2025

Safe For Work Classifier For Malaysian Data — LLM-assisted dataset labeling (Malaysia AI Volunteers)

OtherTextTextClassificationClassification

Led development of a Safe-for-Work classifier for Malaysian text data with categories spanning sexist, violent, and not safe for work contexts. Constructed an annotation methodology for a large crawled dataset using active learning with LLM assistance, producing a multilingual labeled dataset for NSFW categories. Validated dataset quality through an LLM-Ops alignment-oriented workflow and communicated results via publication and conference presentation. • Define and apply NSFW labeling categories for Malaysian context • Use active learning with LLMs to label ~200k items • Produce multilingual dataset labels in English, Malay, and Indonesian • Publish and present results (Arxiv and PyCon Malaysia)

2024 - 2024

Data & AI Intern — AI-application and chatbot development (Pandai Education)

OtherTextTextPrompt + Response Writing (SFT)Prompt + Response Writing (SFT)

Developed an AI application using OpenAI models and LangChain as part of a small team. Built an AI chatbot frontend and implemented supporting APIs using FastAPI to enable conversational interactions. Contributed to analytics data preparation workflows through efficient SQL query development and dashboard creation for stakeholders. • Build AI application using OpenAI models and LangChain • Implement chatbot API services with FastAPI • Create dashboard views using Looker Studio for monitoring • Write optimized SQL for data mart creation to support analysis

2023 - 2023

Education

U

University of Malaya

Bachelor of Computer Science (Data Science), Computer Science (Data Science)

Bachelor of Computer Science (Data Science)
2019 - 2023
C

Centre of Foundation Studies, University of Malaya

Foundation in Science, Science (Mathematics and Engineering)

Foundation in Science
2018 - 2019

Work History

R

Revolab AI

ML Research Engineer

Kuala Lumpur
2025 - Present
T

Tenaga Nasional Berhad

Executive, Advanced Analytics

Bangsar
2023 - 2025