Core Developer / Architecture Designer for Automated Multimodal AI Video Production Workflow
Designed and implemented an automated multimodal AI video production workflow based on ComfyUI. Coordinated semantic alignment and prompt engineering to enhance image understanding and text extraction for video generation. Integrated Florence and GPT for structured prompt generation and semantic rewriting, resulting in improved alignment between visual and contextual content. • Developed a modular and decoupled data pipeline for automated material and background collection. • Performed advanced semantic matching between images and extracted text descriptions to improve content consistency. • Utilized Stable Video Diffusion to automate high-quality video segment creation from image inputs. • Employed modular prompt strategies and negative prompts to minimize generation bias and enable rapid style iteration.