AI Engineer Intern at Oscowl AI (Nov 2025 – Jan 2026): engineered speech recognition, voice fingerprinting, and translation pipeline components.
Built and optimized an end-to-end real-time voice cloning and translation pipeline using speech AI components. Engineered GPU-accelerated speech recognition to extract consistent voice fingerprints for few-shot cloning and integrated upstream audio preprocessing blocks. Performed evaluation-oriented tuning for high-fidelity synthesis in a low-latency setting.• Implemented GPU-accelerated transcription/recognition using a custom riva engine module.• Extracted short-duration (6-second) voice fingerprints for cloning.• Integrated Voice Activity Detection (VAD) and denoising to improve audio quality.• Wired the pipeline to NVIDIA Riva, Whisper, and Llama-3-based translation workflows.