Compare Model and Expert Answers for Dutch Corporate Law
Compare 80 model responses to expert (gold) answers for Dutch corporate financial law documents, identify mismatches, and deliver clear top-level insights; remote, contractor role at $40/hr, ~20+ hours/week.
Legal & Finance
100% Remote Hourly · $40/hr
$40/hr
Compensation
Worldwide
Eligibility
Intermediate
Experience
Jan 28, 2025
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We connect people with hands-on projects that teach and improve artificial intelligence by giving models curated human examples and feedback.
This role is hired and contracted through OpenTrain. Contributors work remotely, build experience in a fast-growing field, and help shape how state-of-the-art AI systems behave.
About AI training and this work
AI training (also called data labeling or human feedback work) is the human side of building AI systems: people prepare, review, and score examples that models learn from. Tasks range from annotating text to evaluating model outputs against expert answers.
This project focuses on comparing model-generated responses to expert gold answers for legal text. Your comparisons will directly inform model improvements in a high-value domain.
100% remote, flexible scheduling available.
Work directly influences model quality in a regulated, specialist area.
The role
You will compare 80 pairs of responses: each pair contains a model response and a gold (subject-matter expert) response tied to Dutch corporate financial law documents. After evaluating each pair, you will produce concise, top-level insights summarizing common error types and overall performance.
This is a contractor, part-time role requiring about 20+ hours per week and an intermediate level of legal familiarity. The project uses text data and a PROMPT_RESPONSE_WRITING_SFT labeling task type.
Project size: 80 query/response pairs.
Data type: TEXT; label task: PROMPT_RESPONSE_WRITING_SFT.
Carefully read each query, the supporting documents, the gold expert answer, and the model's answer. Identify where the model matches the gold, where it deviates, and why.
Document and summarize patterns across the 80 comparisons, highlighting recurring mistakes, missing facts, incorrect legal conclusions, clarity issues, or any systematic biases.
Evaluate each pair against the gold standard and note agreement or divergence.
Annotate specific error types (factual, legal interpretation, omission, clarity, style) for each pair.
Produce a concise summary report with top-level findings and examples.
Requirements
Experience level: Intermediate. You should be comfortable reading and interpreting legal text and comparing written answers for accuracy and completeness.
Time commitment: Approximately 20+ hours per week for the duration of the task. This role is paid hourly at the rate below and hired as a contractor.
Domain: Dutch corporate financial law (familiarity helpful but not strictly mandatory).
Language: English fluency required; understanding of Dutch legal terms is a plus.
Attention to detail and the ability to write clear, evidence-based feedback.
Who should apply
Ideal applicants are English-speaking Netherlands-based or Dutch-knowledgeable master’s students in law, paralegals, or early-career legal researchers with familiarity in corporate/financial law. Candidates with prior evaluation or annotation experience are a plus.
If you are worldwide but can reliably read Dutch legal documents or work comfortably with the English text and supporting materials, you are encouraged to apply.
Preferred: Dutch master’s student in law or equivalent legal training.
Open to contractors worldwide who meet the language and time-commitment needs.
How the project is delivered
You will receive the 80 query packages (query, supporting documents, gold answer, model response) and a template for recording per-pair judgments. Deliverables: per-pair annotations and a top-level insights report.
Labeling software is not specified; OpenTrain will share instructions and the submission workflow after hiring. Maintain clear, reproducible notes for every comparison so findings can be audited.
Deliverables: annotated comparisons for all 80 pairs and a concise summary of top-level insights.
Follow provided templates and submission procedures from OpenTrain.
Compensation and application
Pay: USD 40 per hour, paid per hour as a contractor. Employment types: CONTRACTOR and PART_TIME. The role is remote and open worldwide.
To apply, highlight relevant legal education or experience, availability for 20+ hours/week, and any experience comparing model outputs to expert references. Specify comfort with English and any Dutch-language skills.
OpenTrain is hiring a US Corporate Tax Review Specialist to evaluate ~120 AI-generated tax research answers and rubrics; part-time contractor work (<20 hrs/week), US-only, $120–$140/hr. Must have 5+ years of US tax experience and CPA, JD, or equivalent.
Remote contract role for a US‑licensed attorney to evaluate corporate, M&A, and commercial documents for AI training; 20+ hrs/week at $90–$150/hr. JD and active U.S. bar membership required.
Create 150 private-wealth Q&A pairs and evaluate model responses against certified-advisor standards, assigning 1–10 ratings. Remote, contract, part-time project (under 20 hrs/week) with a fixed fee of $500, open worldwide to finance professionals.