AI Data — SWE-bench Task Authoring & Annotation (Outlier.ai / Bespoke Labs)
Authored and validated SWE-bench software-engineering tasks and annotated training data to support LLM evaluation and fine-tuning. The work focused on producing high-quality labeled datasets for model assessment and improvement in coding-related tasks. Output labeling included task definitions and corresponding annotations suitable for benchmark and training pipelines. • SWE-bench task authoring • Training-data annotation for LLM evaluation • Support for fine-tuning workflows • Ensuring validation/quality of labeled tasks