Use broad music expertise to create challenging prompts, evaluate large language model responses, and build evidence-backed benchmarks. This worldwide contractor assignment runs 40 hours per week for eight weeks.
Generative AI & RLHF
100% Remote
Worldwide
Eligibility
Entry
Experience
Aug 10, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. OpenTrain AI is hiring and contracting for this specialized project, helping contributors apply their expertise to cutting-edge AI development and build a durable portfolio through project work.
Create an OpenTrain account for free.
Apply your subject-matter expertise to advanced AI training and evaluation.
Build experience in a rapidly growing field where people help shape how AI systems behave.
About AI Training and LLM Evaluation
Large language models improve through carefully designed examples and structured human feedback. AI training contributors write prompts, assess model responses, identify errors, and document the reasoning behind their judgments.
This work focuses on the human side of artificial intelligence. Your music knowledge will help test whether models can produce accurate, complete, well-reasoned, and nuanced answers across a demanding specialist domain.
Work on evaluation methods used to improve generative AI.
Help identify hallucinations, outdated information, and difficult edge cases.
Contribute to benchmarks that measure model quality in meaningful ways.
The Role
OpenTrain is seeking a Music Domain Expert for LLM Evaluation. You will combine broad knowledge of music with research, prompt development, adversarial testing, and objective review to evaluate and improve large language models.
This is a worldwide contractor assignment scheduled for 40 hours per week, including at least four hours of overlap with Pacific Time. The contract lasts eight weeks.
Contractor engagement
Eight-week contract
40 hours per week
At least four hours of Pacific Time overlap each week
Worldwide eligibility
English-language work
What You'll Do
You will create advanced music prompts and assess generated answers against clear standards for factual accuracy, reasoning quality, completeness, and musical nuance. You will also help develop reliable evaluation resources and communicate findings clearly to AI researchers.
Create prompts covering music theory, composition, genres, composers, artists, instruments, music history, notation, production, and contemporary music trends.
Evaluate AI-generated responses for factual accuracy, reasoning quality, completeness, and musical nuance.
Identify hallucinations, logical inconsistencies, outdated information, and challenging edge cases.
Develop benchmark datasets and adversarial test cases for language-model evaluation.
Provide objective feedback supported by reliable references and evidence-based research.
Collaborate with AI researchers to improve model performance.
Maintain high annotation quality and clear documentation.
Requirements
This role requires advanced, broad music knowledge and the ability to review specialist content carefully. You should be comfortable researching claims, checking references, and distinguishing subtle problems in factual accuracy or reasoning.
A master's degree or higher in any field is required, along with at least three years of relevant professional experience. The structured listing identifies this opportunity as entry level, but the stated education and experience requirements still apply.
Master's degree or higher in any field.
At least three years of relevant professional experience in music education, performance, composition, production, journalism, research, content creation, audio engineering, or a related area.
Strong knowledge of music theory, composition, genres, composers, artists, instruments, production, notation, music history, and contemporary music.
Excellent written English.
Strong research, analytical, and detail-oriented review skills.
Ability to identify factual inconsistencies, reasoning problems, hallucinations, outdated information, and edge cases.
Ability to provide objective, evidence-backed judgments and reliable citations.
Helpful Background
Experience with large language models, generative AI, prompt engineering, or AI evaluation is helpful. Published research, industry recognition, or teaching experience may also support your application.
Independent work and evidence-backed reviewing are important because evaluation results and training data must be consistent, well documented, and reliable.
Large language model experience
Generative AI experience
Prompt engineering experience
AI evaluation experience
Published research
Industry recognition
Teaching experience
Why This Work Matters
Every major AI system depends on people who prepare examples and review model behavior. By applying deep music expertise to prompt creation and evaluation, you will help make generative AI more accurate, useful, and capable of handling specialist knowledge.
Work remotely on a worldwide assignment.
Use professional music knowledge in a growing technology field.
Help shape how state-of-the-art AI systems understand and communicate about music.
Gain specialized AI training experience through project-based contractor work.
Use your expertise in art history, visual arts, architecture, and design to create challenging prompts and evaluate large language model responses. This remote contract assignment combines advanced research, cultural knowledge, and evidence-based AI feedback.
Use political science and governance expertise to evaluate, challenge, and improve large language models. This remote contractor assignment runs for eight weeks and requires 40 hours weekly with Pacific Time overlap.
Use deep sports knowledge to write challenging prompts, evaluate large language model responses, and identify factual or reasoning issues. This remote contractor role offers 20+ hours per week for experts with a master's degree and three years of relevant experience.