You will test how large language models handle single-turn image-edit requests. You will classify prompts and outputs, find safety failures, and write clear explanations that can be used to improve model behavior.
- Classify image-edit prompts and outputs using project guidelines and a defined safety taxonomy.
- Review ambiguous, borderline, and benign requests consistently.
- Write precise rationales that support each classification decision.
- Document bypassed safeguards and subtle policy violations.
- Create adversarial prompts and compare multiple model outputs.
- Identify unclear or conflicting guidelines and suggest clarifications.
What It Pays And Takes
This is a project-based independent contractor engagement. The listing does not provide a pay rate. The role is marked entry level, but it requires strong analytical judgment and familiarity with AI safety concepts.
- Hours: 20 or more hours per week.
- Term: 1–2 weeks per statement of work.
- Work arrangement: Part-time, independent contractor, with a self-set schedule.
- Location: Open worldwide.
- Language: Fluent English required.
- Equipment: Your own desktop or laptop and a reliable internet connection.
- Core skills: Policy-based analysis, careful classification, precise writing, and the ability to assess complex or ambiguous information.
- Safety experience: Red teaming, prompt engineering, or designing challenge prompts to test AI safety filters.
- Helpful background: Content moderation, policy analysis, AI safety evaluation, RLHF, or data annotation.
- Relevant education or experience may include policy, law, ethics, linguistics, journalism, or computer science.
How It Works
Apply on OpenTrain with your resume, then complete the application on the hiring site.
About AI Training Work
AI training work uses human judgments, examples, and written feedback to improve how artificial intelligence systems behave. Safety evaluators are paid for careful policy analysis because their decisions help identify harmful outputs and improve model responses.