Technical Research Fellow (AI Evaluation) — Handshake AI (November 2025 – Present)
Role involves evaluating and ranking outputs from Large Language Models (LLMs) to refine model behavior and ensure training-quality responses. You assess complex multimedia content and assign ratings/ordering to support high-fidelity inputs for neural network training. The work includes collaborative testing with frontier AI research labs to improve evaluation rigor and output reliability. • Evaluate and rank LLM outputs via rigorous output assessment. • Rate and order complex multimedia content (image, audio, video) for training suitability. • Contribute to refining LLMs through structured comparisons. • Ensure high-fidelity training inputs for neural networks.