AI/ML Engineer (LLM and RAG Systems) - Perceptyx AI
Built an AI search engine that answers user queries using real-time web search with citations. Designed and implemented a multi-layer memory system to improve context retention and response quality across sessions. Developed intelligent query routing, search aggregation, and reranking to reduce latency through parallel retrieval and caching strategies. • Linked LLM and retrieval components with LangGraph and FastAPI for end-to-end query handling. • Implemented Redis semantic caching and async Qdrant queries to improve throughput and response times. • Built scalable background workflows with monitoring and RLHF-style self-improvement loops. • Integrated PostgreSQL for supporting data storage and operational persistence.