Data Scientist – Metadata extraction from SEM images using LLM + OCR
Developed an optimized workflow for extracting metadata from SEM images using an internally built LLM and OCR. Used PyTesseract in the pipeline to convert visual information into machine-readable text for subsequent processing. Achieved a reported 40% increase in efficiency of extraction and transformation activities supporting dataset preparation.• Extracted SEM image metadata via an internal LLM workflow.• Applied OCR using PyTesseract to structure extracted information.• Improved efficiency of extraction and transformation by 40% for downstream use.