Data Science Intern, Dr. Reddy’s Laboratories (RAG-based retrieval and insight generation from documents and databases)
Built and deployed a multimodal RAG system that retrieves insights from 1000+ heterogeneous databases using embedding-based semantic retrieval. Designed an end-to-end semantic retrieval pipeline that generates high-dimensional embeddings, stores them in a vector database, and uses optimized similarity search with LLM-powered response generation. Implemented NLP processing to extract insights from documents and map them to email metrics to support content evaluation and optimization.• Processed 7000+ PDFs and JPEG inputs for downstream retrieval and insight extraction.• Used embeddings and vector database similarity search as the core retrieval mechanism.• Implemented LLM-powered response generation on top of retrieved context.• Engineered a feedback loop for content optimization using RAG-based contextual understanding.