Use your video game industry expertise to evaluate and improve large language models through advanced prompts, benchmark datasets, and evidence-based feedback. This US-based contractor project runs for 8 weeks with at least 4 hours of PST overlap.
Generative AI & RLHF
Remote
1 country
Eligibility
Entry
Experience
Aug 19, 2026
Posted
Open to applicants in
United States
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI hires and contracts contributors for specialized projects where human expertise helps improve modern artificial intelligence systems.
Work on cutting-edge AI projects that match your expertise.
Build experience in a fast-growing field shaping how AI systems behave.
Create an OpenTrain account for free and apply in minutes.
About AI Training Work
Large language models improve when experienced people create challenging examples, review model-generated responses, and explain what makes an answer accurate, complete, and useful. This role applies that human evaluation process to the video games domain, including development, design, esports, platforms, publishers, and gaming communities.
Design prompts that test knowledge and reasoning.
Rate model outputs for accuracy, completeness, nuance, and consistency.
Identify hallucinations, outdated information, and difficult edge cases.
The Role
OpenTrain is recruiting a Video Game AI Evaluation Expert to evaluate and improve large language models. You will use deep knowledge of the video games industry to create advanced prompts, assess AI-generated responses, develop evaluation datasets, and provide evidence-based feedback to AI researchers.
This is a contractor, part-time project for candidates located in the United States. The project has a planned commitment of 40 hours per week for 8 weeks, requires at least 4 hours of overlap with Pacific Standard Time, and lists a minimum time requirement of 20+ hours per week.
Location: United States
Language: English
Project length: 8 weeks
Schedule: 20+ hours per week listed; 40 hours per week planned
Availability: At least 4 hours of PST overlap
Engagement: Contractor and part-time
What You'll Do
You will create and review evaluation material designed to reveal how well AI systems understand the video game industry. Your work will combine subject-matter expertise, analytical judgment, careful documentation, and clear written feedback.
Create advanced prompts covering game development, esports, platforms, game design, publishers, and communities.
Evaluate AI-generated responses for factual accuracy, reasoning quality, completeness, and nuance.
Identify hallucinations, logical inconsistencies, outdated information, and edge cases.
Develop benchmark datasets and adversarial test cases.
Provide evidence-based feedback supported by reliable references.
Collaborate with AI researchers to improve model performance.
Maintain high annotation quality and clear documentation.
Requirements
This role requires substantial video game industry knowledge and professional experience. A master's degree or higher is required, preferably in a gaming-related discipline, along with at least three years of experience in a relevant field.
Master's degree or higher, preferably in a gaming-related discipline.
Strong knowledge of game development, game design, platforms, genres, publishers, studios, esports, and gaming communities.
At least 3 years of professional experience in game development, design, quality assurance, esports, gaming journalism, content creation, research, community management, or a related field.
Excellent written English and analytical skills.
Ability to evaluate gaming-related information accurately and identify factual inconsistencies.
Ability to evaluate AI-generated responses for factual accuracy and reasoning quality.
Helpful Background
Experience with AI systems can help you contribute to model evaluation work, although the core requirements center on video game expertise, analytical ability, and professional experience.
Experience with large language models or generative AI.
Prompt engineering experience.
Previous AI evaluation experience.
Published research, industry recognition, or teaching experience.
Why This Work Matters
AI training is the human side of building artificial intelligence. By writing prompts, checking model responses, and documenting errors, contributors help shape how advanced systems understand specialized subjects such as video games.
The work is part of a rapidly growing technology field and can provide a way to apply professional knowledge directly to the development of state-of-the-art AI.
Help improve AI accuracy and reasoning in a specialized domain.
Apply your video game expertise to emerging technology.
Work remotely with a flexible contractor model, subject to this project's schedule and overlap requirements.
How to Apply Through OpenTrain
Create a free OpenTrain account, build your profile around your video game expertise and professional background, and apply in minutes. Highlight relevant experience in development, design, QA, esports, journalism, research, content creation, or community management, along with any AI evaluation or prompt engineering work.
Confirm that you are located in the United States.
Confirm your English writing ability and required PST overlap.
Showcase your video game industry knowledge and analytical experience.
Apply for the Video Game AI Evaluation Expert contractor project through OpenTrain.
Help improve next-generation AI by creating and evaluating board game reasoning tasks, analyzing strategic scenarios, and identifying logical errors. This worldwide, two-month contractor assignment requires 20+ hours weekly with PST overlap.
Use PhD-level physics expertise to adjudicate competing AI model solutions, assess assumptions and approximation limits, and write rigorous evaluations. This worldwide, part-time contract pays $80 to $160 per hour.
Use your professional expertise to evaluate AI-generated responses, apply detailed rubrics, and provide feedback that improves model behavior. This part-time remote contract offers 20+ hours per week and pays $140-$200 per hour.