AI Data Analyst / Technical Fellow (Remote)
As an AI Data Analyst/Technical Fellow, I authored multi-agent benchmark tasks and evaluated LLM-generated outputs for accuracy and logic. I curated and cleaned diverse real-world datasets for use in AI model training and validation processes. My work included identifying hallucinations, designing rigorous evaluation metrics, and improving LLM grounding. • Designed and scored multi-agent benchmark tasks for language models • Cleaned and prepared datasets in formats such as CSV, JSON, and logs • Evaluated outputs for logical reasoning, accuracy, and hallucination detection • Used Python and proprietary/internal tools for automation and evaluation