Open-source AI Model On-Premises Deployment & Fine-tuning Specialist
Led the on-premises deployment of large open-source AI language models, focusing on optimizing model inference, configuration, and system interoperability. Developed workflows for private AI service environments including Docker, Ollama, and web-based interfaces, enabling customizable model deployment for dialog-based applications. Established reusable process documentation and continuous optimization of model tuning and evaluation. • Oversaw the architectural design for disk allocation and resource deployment • Pulled and tested Qwen series models for hardware adaptability • Deployed and configured Open WebUI for model visualization • Integrated local APIs to ensure seamless cross-environment communications