Disfluency Restoration — Multimodal Seq2Seq (Wav2Vec2 + IndicBART)
Developed a multimodal seq2seq pipeline to restore disfluent transcripts from audio inputs. Extracted speech representations using Wav2Vec2 speech embeddings, fused them with text features, and fine-tuned IndicBART as the seq2seq generator to reproduce disfluent transcripts. Implemented a full end-to-end workflow including preprocessing, modality fusion, training, and inference. • Wav2Vec2 speech embedding extraction from audio • Modality fusion with text features • IndicBART fine-tuning for seq2seq generation • End-to-end preprocessing, training, and inference pipeline