Bio Medical Knowledge Base: Leveraging Deep learning Models on One Billion Biomedical Data
Consolidated and leveraged one-billion-scale biomedical protein records to construct a large protein-based biomedical knowledge base. Built deep learning models on the aggregated dataset to support tasks such as target identification for disease. The work required organizing heterogeneous inputs from many sources into a unified dataset for downstream AI modeling. • Consolidated data from more than 80 sources into a single protein-based database • Created a database with more than 1 billion records • Built deep learning models for target identification given disease • Designed the biomedical dataset pipeline for large-scale AI training