I have hands-on experience in AI training and data annotation through freelance roles as an AI Trainer at Outlier AI and
I have hands-on experience in AI training and data annotation through freelance roles as an AI Trainer at Outlier AI and an AI Data Contributor at Remotasks. In these positions, I actively evaluate and rank model outputs across dimensions of accuracy, helpfulness, and safety to support reinforcement learning from human feedback (RLHF) workflows. Leveraging my academic background in chemical sciences, I apply specialized domain knowledge to assess the factual accuracy, reasoning quality, and response completeness of complex LLM outputs, alongside crafting highly targeted prompts for model refinement and capability improvement. Complementing my manual data annotation experience is a technical background in designing scalable LLM evaluation environments. As an AI SDE Intern at Xelron AI, I engineered containerized benchmark tasks and authored complex, HLE-style questions spanning math, physics, chemistry, and software engineering to support frontier-level capability assessments. I also validated these benchmarks through an 11-gate CI/LLM-as-Judge pipeline to enforce reproducibility and rigorous anti-cheat integrity. This dual perspective—combining meticulous data labeling with the technical architecture of AI benchmarking—equips me to deliver the high-fidelity, domain-accurate training data required at OpenTrain AI.