Join OpenTrain as a remote Clinical Medicine AI Evaluation Physician to design and run clinical evaluations that test AI medical reasoning. Flexible contract work (20–30 hrs/week), 1-month engagement with possible extensions; licensed physicians only.
Generative AI & RLHF
100% Remote
Worldwide
Eligibility
Entry
Experience
Jul 25, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for people building careers in AI training and data labeling. We connect clinicians and other specialists with short-term, flexible projects where your expertise directly shapes how AI systems behave.
For this role, OpenTrain is the hiring and contracting organization and will manage the engagement, payments, and project onboarding.
Why AI Training Matters in Medicine
AI systems learn from examples and human feedback. Medical AI must be evaluated by clinicians to ensure it reasons safely and accurately in real-world clinical situations.
Participating in AI evaluation is a way to influence how these tools perform in practice, while fitting work around clinical responsibilities.
The Role
As a Clinical Medicine AI Evaluation Physician you will design evaluation frameworks and clinical scenarios, assess model reasoning, and report gaps that affect patient care. This is remote, contract work intended to fit alongside clinical duties.
OpenTrain is looking for physicians who can bring real-world clinical judgment and evidence-based reasoning to test and improve medical AI behavior.
Position type: Independent contractor, part-time
Location: Remote / work from anywhere (English required)
Duration: 1 month, with potential extensions based on performance and fit
What You'll Do
Your core work is to evaluate how medical AI systems handle clinical reasoning and decision-making. You will create realistic clinical test cases, apply assessment methods, and collaborate with researchers to refine evaluations.
Design systematic evaluation frameworks for medical AI systems
Create clinical scenarios that test AI reasoning and decision-making capabilities
Build assessment methods that capture the nuance of clinical practice
Identify gaps in AI medical knowledge and reasoning and document findings
Collaborate with AI researchers to improve model performance and iterate on tests
Requirements
This role requires an active, licensed physician who regularly practices clinical medicine and can apply evidence-based reasoning to evaluate AI outputs.
Licensed physician (MD/DO) with active clinical practice in any specialty
Experience with evidence-based clinical decision-making
Strong analytical skills for evaluating AI reasoning output
Clear written and verbal communication to document findings and recommendations
Ability to design clinical scenarios that effectively test AI capabilities
Helpful Background
Candidates with interest or prior exposure to medical AI, digital health, or clinical research will find this work especially rewarding. A background in biology is a plus but not required.
Helpful: prior interest in how AI can support clinical practice
Helpful: background in Biology
Engagement Details & Data Work
Time commitment is flexible but substantial: expect roughly 20–30 hours per week, remote, with a one-month initial engagement and the possibility of extension. This is a contract, part-time role paid through OpenTrain (rates are set in the platform).
The work is text-focused evaluation: you will review model outputs and assign evaluation ratings according to predefined rubrics.
Typical commitment: 20+ hours/week (up to 30 hrs/week), flexible scheduling
Data type: TEXT; Labeling tasks: EVALUATION_RATING
Languages: English required
Employment types: CONTRACTOR, PART_TIME
Who Should Apply & How To Get Started
Apply if you are a practicing physician interested in shaping safe, clinically reliable AI. This project welcomes clinicians who are new to AI evaluation as well as those with prior experience.
To apply, create or update your OpenTrain profile so it reflects your medical license and clinical experience, then submit your application through the OpenTrain platform. Successful applicants will complete onboarding and a short evaluation task before starting.
Experience level: Entry level (project suitable for clinicians new to AI evaluation)
Worldwide applicants accepted; ensure your OpenTrain profile lists your medical license and specialty
Evaluate AI-generated clinical responses and write gold-standard medical answers as a remote Nurse Practitioner Clinical AI Reviewer. Contract, part-time work (20+ hrs/week) paying $40–90/hr; requires NP/APRN qualification, active licensure or eligibility, and advanced clinical reasoning.
Evaluate and improve AI-generated biological answers as a remote contract specialist — MS/PhD in biology required. Ongoing part-time work (~20 hrs/week) at $70/hr (USD); English C1+ and applicants from specified countries are welcome.
Use your computational biology expertise to evaluate and annotate AI-generated outputs across genomics, structural biology, and systems biology. Remote contractor role, 20+ hrs/week, $40–$60/hr — help shape safer, more accurate scientific AI.