AI Personalization Evaluation Analyst (Polish)
Evaluate a personalization feature in Polish by designing short multi-turn prompts, comparing paired model responses, and writing concise, defensible quality rationales. Contractor role, $20/hr, remote, ~4 hours/day with 4-hour overlap with PST for a 1-month engagement.
Generative AI & RLHF
$20/hr
Compensation
Worldwide
Eligibility
Entry
Experience
Jul 20, 2026
Posted
Open worldwide
About OpenTrain
OpenTrain is the #1 platform for people who build careers in AI training and data labeling. We connect contributors with specialized AI training work, help them build a unified portfolio, and support growth into durable freelance careers.
Working with OpenTrain means contributing directly to how modern AI systems learn. We hire contractors to perform evaluation and annotation tasks that shape model behavior, with flexible remote schedules and clear project guidance.
About AI training and this project
AI training (also called data labeling or human feedback work) is the human side of building intelligent systems. Contributors create and review examples—writing prompts, judging responses, and documenting why one output is better—to teach models to respond more helpfully and accurately.
This project focuses on evaluating a personalization feature: you will test whether a model uses personal context correctly, avoids flawed inferences, and produces grounded, useful responses. The work is hands-on and directly influences model quality.
- 100% remote work you can do from anywhere with a computer and internet.
- Flexible part-time or fuller schedules; great for freelancers or those balancing other commitments.
The role
As an AI Personalization Evaluation Analyst you will design short multi-turn prompts drawn from your own experience, review how the model uses personal signals, compare paired responses side-by-side, and write concise, defensible rationales that reference specific turns.
You will also verify debug information and keep evaluation chats organized and clean so future assessments remain reliable and reproducible.
What you'll do
- Create short multi-turn prompts (typically 1–5 turns) that test personalization behavior using realistic personal-context scenarios.
- Judge whether the model uses personal context appropriately and avoids flawed inferences or hallucinations.
- Compare two model responses side-by-side and select the one that is more helpful, natural, and usable.
- Write concise explanations that reference specific turns and explain quality differences with defensible reasoning.
- Check supporting debug information and maintain strict data hygiene across evaluation conversations.
Requirements
You must be fluent in Polish for reading and writing evaluation content, and you should be comfortable working with personal-context tasks and structured rationales.
- Fluency in Polish for reading and writing evaluation notes (required).
- Strong analytical judgment for nuanced model assessment.
- Experience with prompt design, personalization review, or side-by-side response evaluation.
- Ability to write clear, concise rationales that reference specific conversation turns.
- Independent, reliable remote work habits and clear communication.
Helpful background
- Experience in data annotation, AI quality evaluation, or content moderation.
- BS/BA or equivalent experience in an analytical field such as policy, law, ethics, linguistics, journalism, or computer science.
Pay & schedule
This is a contractor engagement at $20.00 USD per hour. Expect about 4 hours per day (roughly 20+ hours/week) and up to 40 hours per week depending on project needs.
The engagement is approximately 1 month. You must be available full-time in your local time zone with at least 4 hours of overlap with Pacific Standard Time (PST).
- Payment type: hourly contractor work at $20.00 USD/hour.
- Typical cadence: ~4 hours/day; roughly 20+ hours/week and up to 40 hours/week.
- Engagement length: 1 month; part-time or fuller hours within that window.
How it works
OpenTrain AI will contract you for this evaluation work. You will receive project instructions, example prompts, and evaluation rubrics during onboarding.
Work is done remotely. After onboarding you will perform evaluations, submit ratings and rationales through the project interface, and follow data hygiene practices to keep assessment chats organized and reproducible.
- Apply with your OpenTrain profile and language competence; onboarding and a brief calibration task will verify alignment with project guidelines.
- You will complete evaluations, record rationales, and verify debug metadata as part of each assignment.
- Maintain confidentiality and follow provided data handling instructions.