Software Engineer (System & Knowledge Management) — AI/RAG system architecture and LLM application development (KEWASNET)
Built and deployed a production-grade Retrieval-Augmented Generation (RAG) system to enable context-aware question answering over internal documents. Implemented document chunking and vector embedding using a vector database to support retrieval for LLM responses. Optimized streaming latency and system reliability to improve end-user interactions with model outputs. • Designed RAG architecture with Django 5 and vector databases • Implemented prompt/response generation pipeline for context-aware answers • Integrated caching and rate-limiting for low-latency LLM streaming • Added monitoring via Prometheus for health and query performance