Snorkel AI - AI Engineer / Task Author
Designed and implemented multi-step coding and reasoning tasks for the Terminal-Bench benchmark under the Terminus framework. Curated datasets and built automated evaluation pipelines to test complex problem-solving and reasoning capabilities in frontier GPT-level systems. Focused on dataset preparation and benchmark task authoring to enable consistent model scoring and comparison.• Authored multi-step coding/reasoning tasks for CLI/agentic evaluation.• Built automated evaluation pipelines for benchmark runs.• Curated high-difficulty datasets for LLM benchmarking.• Supported assessment of complex reasoning behaviors.