Skip to content
OpenTrain AIFor AI Companies

GenAI Security Evaluation Engineer

Build and test vulnerable GenAI agent and RAG codebases, annotate security risks, and evaluate detection tools at $150 per hour. This contract role requires 20+ hours per week and strong application security experience.

Apply now
OpenTrain AI

Coding & Software

100% Remote Hourly · $150/hr

$150/hr

Compensation

Contract, part-time

Engagement

Remote

Location

Oct 7, 2026

Posted

Open worldwide

The work

You will create small, runnable agent and retrieval-augmented generation (RAG) codebases with realistic security weaknesses. You will test static analysis tools and document how well they detect vulnerabilities in tool calling, memory, and Model Context Protocol (MCP) systems.

  • Build agent and RAG repositories with code-reachable Sensitive Information Disclosure and Excessive Agency vulnerabilities.
  • Create vulnerable, fixed, and hard-negative versions with only small security-relevant differences.
  • Trace and annotate assets, data and action paths, controls, root causes, severity, and remaining risk.
  • Define authorization contexts and write deterministic tests for vulnerable, fixed, and negative behavior.
  • Recommend security controls and take part in calibration and peer review.

What it pays and takes

This is contract, part-time work for someone with deep experience in application security and LLM agent frameworks. The project requires at least 20 hours per week and professional fluency in English.

  • Pay: $150 per hour.
  • Time: 20+ hours per week.
  • Location: Open worldwide.
  • Work type: Contractor and part time.
  • Experience: 5+ years in application or product security, or security-focused software engineering, including secure code review.
  • Security analysis: Experience with source-to-sink analysis, taint analysis, SAST, CodeQL, or Semgrep.
  • GenAI systems: Hands-on experience building LLM agents or RAG systems with tools such as LangChain, LlamaIndex, OpenAI or Anthropic SDKs, or MCP.
  • Authorization: Strong knowledge of actors, trust boundaries, tenants, OAuth, IAM, identity, permitted data and actions, purposes, recipients, and document-level access control.
  • Programming: Production experience with Python and/or TypeScript.
  • Helpful background: OWASP LLM security risks, MCP, threat modeling, or security evaluation.

How it works

Apply on OpenTrain with your resume, then complete the application on the hiring site.

About AI training work

AI training work is the human work behind modern artificial intelligence, including testing systems, reviewing model behavior, and preparing examples that help software improve. Experienced engineers are needed for specialized projects such as security evaluation because they can identify realistic risks and judge whether safeguards work.

Requirements

  • Experience: Entry level
  • Languages: English

How to apply

  1. Apply here on OpenTrain. You create a free account, and we send you to the hiring platform.
  2. Complete your application on the hiring platform.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar open roles

View all AI training jobs

Coding-Agent Benchmark Engineer

Design and evaluate realistic coding-agent tasks using production-like repositories, tests, evaluators, and scoring rubrics. This contract role is open worldwide, requires English fluency and 20+ hours per week, and calls for deep software engineering experience.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Oct 7, 2026

Bioinformatics AI Evaluation Task Designer

Design rigorous bioinformatics and computational genomics tasks that test whether AI models can analyze data, write code, and produce verifiable scientific results. This remote five-week contractor assignment pays $150 per approved task.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level
Per task · $150/label

Posted Oct 7, 2026

Backend AI Coding Task Creator

Create realistic backend coding tasks and reliable verifiers for AI systems. This contractor role offers a flexible schedule of about 15 hours per week and pays $30 to $100 per hour equivalent, paid per task that meets project specifications.

Coding & Software
Computer Code Programming
Remote · United Arab Emirates, Argentina, Austria +50 more
English
Part-time · Flexible
Mid-Senior level
Per task · $30–$100/hr equivalent

Posted Oct 6, 2026

Python Backend Developer for AI Agent Data

Build and test Python backend connectors or create realistic AI agent tasks and evaluation rubrics. This remote, four-week contract pays $300 per approved task and requires at least 20 hours per week.

Coding & Software
Computer Code Programming
Remote · Worldwide
English
Part-time · Flexible
Entry level
Per task · $300/label

Posted Oct 4, 2026

Backend Coding Task Creator and AI Evaluator

Create realistic backend coding tasks, build verifiers, and evaluate AI-generated solutions for correctness, performance, testing, and maintainability. Pay is $30 to $100 per hour equivalent, paid per task that meets the project specifications.

Coding & Software
Computer Code Programming
Remote · United Arab Emirates, Argentina, Austria +50 more
English
Part-time · Flexible
Mid-Senior level
Per task · $30–$100/hr equivalent

Posted Oct 3, 2026