LLM Fine-Tuning Dataset Curation for Mental Health Dialogue AI
Independently curated a 564-sample JSONL fine-tuning dataset for a LLM-powered mental health dialogue system (MindJournal AI). Designed and annotated multi-turn therapeutic conversations across 3 clinical roles (CBT therapist, emotional support counselor, crisis intervention specialist), with strict quality control on tone consistency, response length, and bilingual (English/Chinese) alignment. Dataset was used to perform Supervised Fine-Tuning (SFT) on GLM-4-Plus, directly improving role consistency and multilingual output quality. Additionally handled multimodal data annotation including ASR transcription validation and handwriting OCR processing