Skip to content
OpenTrain AIFor AI Companies

LLM Conversation Evaluation Engineering Manager

Evaluate multi-turn LLM conversations, tool-use scenarios, and model responses while applying engineering judgment and clear feedback. This remote contractor role offers flexible AI training work through OpenTrain.

OpenTrain AI

Generative AI & RLHF

100% Remote

Worldwide

Eligibility

Entry

Experience

Sep 16, 2026

Posted

Open worldwide

About OpenTrain

OpenTrain AI is the hiring and contracting organization for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping people discover projects, build a professional AI training profile, and apply in minutes.

Creating an OpenTrain account is free, and this role gives technical professionals an opportunity to contribute to the development of modern AI systems through structured evaluation work.

About AI Training Work

AI training is the human side of building artificial intelligence. Contributors review model outputs, write or improve responses, evaluate conversations, and provide clear feedback that helps AI systems become more accurate, useful, and consistent.

This fast-growing field supports remote, flexible work for people with technical expertise, strong communication skills, and careful judgment. Your evaluations can directly influence how advanced AI systems behave.

The Role

OpenTrain is recruiting an LLM Conversation Evaluation Engineering Manager to support AI training focused on LLM Conversation Data Evaluation. You will work with multi-turn LLM conversations and tool-use interaction scenarios, applying technical expertise and structured human judgment to assess model behavior.

The role calls for strong practical knowledge of APIs, JSON, structured workflows, assistant behavior, and function-call or tool-use design. Prior AI training experience is not required unless stated elsewhere in this posting, but relevant professional, academic, or language expertise in LLM Conversation Data Evaluation is required.

  • Employment type: Contractor and part-time
  • Experience level: Entry level
  • Work arrangement: Remote and worldwide
  • Language: English fluency required
  • Data type: Text
  • Schedule: At least 4 hours per day and a minimum of 40 hours per week
  • Time-zone overlap: At least 4 hours with PST
  • Subject area: LLM Conversation Data Evaluation

What You'll Do

You will evaluate realistic, multi-turn interactions and apply project guidelines consistently. Your written judgments and actionable feedback will help improve model behavior and maintain quality across repeated review tasks.

  • Write or improve expert responses and explain your reasoning clearly.
  • Review source materials and model outputs against project guidelines.
  • Evaluate multi-turn conversations and tool-use interaction scenarios.
  • Assess assistant behavior, function calls, structured workflows, and response quality.
  • Explain judgments clearly so feedback can improve model behavior.
  • Enforce quality standards and provide actionable feedback.
  • Maintain consistency across repeated review tasks.
  • Communicate effectively in English while working independently.

Requirements

This role is suited to someone with technical team leadership or mentoring experience and at least five years in software engineering, technical operations, data, AI, or a related technical field. You should be comfortable applying detailed rubrics, reviewing complex interactions, and communicating decisions in writing.

  • Relevant professional, academic, or language expertise in LLM Conversation Data Evaluation.
  • Technical team leadership or mentoring experience.
  • At least five years of experience in software engineering, technical operations, data, AI, or a related technical field.
  • Strong practical knowledge of APIs, JSON, structured workflows, and assistant behavior.
  • Knowledge of function-call or tool-use design.
  • Strong written communication and careful attention to detail.
  • Ability to follow detailed project instructions and apply rubrics consistently.
  • Ability to evaluate realistic multi-turn conversations and enforce quality standards.
  • Fluency in English.
  • Comfort working independently on remote contractor tasks.

Schedule and Compensation

The expected schedule is at least four hours per day and a minimum of 40 hours per week, including four hours of overlap with PST. Compensation is handled through the OpenTrain project budget fields shown on this job page.

  • Minimum daily commitment: 4 hours.
  • Minimum weekly commitment: 40 hours.
  • Required overlap: 4 hours with PST.
  • Compensation information: See the OpenTrain project budget fields.

Build Your AI Training Career

OpenTrain gives contractors one place to find AI training work, build a unified AI training portfolio, and use their work history to qualify for more opportunities over time. For engineers, technical operators, data specialists, and reviewers, AI training creates a way to apply specialized judgment to the rapidly growing field of artificial intelligence.

Apply through OpenTrain to begin contributing to LLM evaluation work that helps shape how AI assistants understand instructions, use tools, and respond in realistic conversations.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar jobs

View all AI training jobs

AI Workflow Engineer Prompt Engineering

Build and refine LLM-powered automation workflows as an intermediate AI Workflow Engineer. Evaluate outputs, design prompts, and improve communication and content automation for $15 to $45 per hour.

Generative AI & RLHF
Text
Remote · Worldwide
Part-time · Flexible
Intermediate level
Hourly · $15–$45/hr

Posted Mar 29, 2026

Physics Expert for LLM Evaluation

Use advanced physics knowledge to design challenging problems, solve them step by step, and help evaluate how large language models reason. This flexible remote contractor role is open worldwide.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Entry level

Posted Jul 17, 2026

Physics LLM Evaluation Expert

Help advance large language models by designing challenging physics problems, writing rigorous solutions, and shaping evaluation benchmarks. This expert-level remote contract offers 20+ hours per week for graduate-level STEM specialists.

Generative AI & RLHF
Text
Remote · Worldwide
English
Part-time · Flexible
Expert level

Posted Jul 17, 2026