Skip to content
OpenTrain AIFor AI Companies

AI Personalization Evaluation Analyst (Polish)

Evaluate a personalization feature in Polish by designing short multi-turn prompts, comparing paired model responses, and writing concise, defensible quality rationales. Contractor role, $20/hr, remote, ~4 hours/day with 4-hour overlap with PST for a 1-month engagement.

OpenTrain AI

Generative AI & RLHF

100% Remote Hourly · $20/hr

$20/hr

Compensation

Worldwide

Eligibility

Entry

Experience

Jul 20, 2026

Posted

Open worldwide

Interested in this role?

Create a free OpenTrain account and apply in minutes.

About OpenTrain

OpenTrain is the #1 platform for people who build careers in AI training and data labeling. We connect contributors with specialized AI training work, help them build a unified portfolio, and support growth into durable freelance careers.

Working with OpenTrain means contributing directly to how modern AI systems learn. We hire contractors to perform evaluation and annotation tasks that shape model behavior, with flexible remote schedules and clear project guidance.

About AI training and this project

AI training (also called data labeling or human feedback work) is the human side of building intelligent systems. Contributors create and review examples—writing prompts, judging responses, and documenting why one output is better—to teach models to respond more helpfully and accurately.

This project focuses on evaluating a personalization feature: you will test whether a model uses personal context correctly, avoids flawed inferences, and produces grounded, useful responses. The work is hands-on and directly influences model quality.

  • 100% remote work you can do from anywhere with a computer and internet.
  • Flexible part-time or fuller schedules; great for freelancers or those balancing other commitments.

The role

As an AI Personalization Evaluation Analyst you will design short multi-turn prompts drawn from your own experience, review how the model uses personal signals, compare paired responses side-by-side, and write concise, defensible rationales that reference specific turns.

You will also verify debug information and keep evaluation chats organized and clean so future assessments remain reliable and reproducible.

What you'll do

  • Create short multi-turn prompts (typically 1–5 turns) that test personalization behavior using realistic personal-context scenarios.
  • Judge whether the model uses personal context appropriately and avoids flawed inferences or hallucinations.
  • Compare two model responses side-by-side and select the one that is more helpful, natural, and usable.
  • Write concise explanations that reference specific turns and explain quality differences with defensible reasoning.
  • Check supporting debug information and maintain strict data hygiene across evaluation conversations.

Requirements

You must be fluent in Polish for reading and writing evaluation content, and you should be comfortable working with personal-context tasks and structured rationales.

  • Fluency in Polish for reading and writing evaluation notes (required).
  • Strong analytical judgment for nuanced model assessment.
  • Experience with prompt design, personalization review, or side-by-side response evaluation.
  • Ability to write clear, concise rationales that reference specific conversation turns.
  • Independent, reliable remote work habits and clear communication.

Helpful background

  • Experience in data annotation, AI quality evaluation, or content moderation.
  • BS/BA or equivalent experience in an analytical field such as policy, law, ethics, linguistics, journalism, or computer science.

Pay & schedule

This is a contractor engagement at $20.00 USD per hour. Expect about 4 hours per day (roughly 20+ hours/week) and up to 40 hours per week depending on project needs.

The engagement is approximately 1 month. You must be available full-time in your local time zone with at least 4 hours of overlap with Pacific Standard Time (PST).

  • Payment type: hourly contractor work at $20.00 USD/hour.
  • Typical cadence: ~4 hours/day; roughly 20+ hours/week and up to 40 hours/week.
  • Engagement length: 1 month; part-time or fuller hours within that window.

How it works

OpenTrain AI will contract you for this evaluation work. You will receive project instructions, example prompts, and evaluation rubrics during onboarding.

Work is done remotely. After onboarding you will perform evaluations, submit ratings and rationales through the project interface, and follow data hygiene practices to keep assessment chats organized and reproducible.

  • Apply with your OpenTrain profile and language competence; onboarding and a brief calibration task will verify alignment with project guidelines.
  • You will complete evaluations, record rationales, and verify debug metadata as part of each assignment.
  • Maintain confidentiality and follow provided data handling instructions.

Ready to apply?

Create a free OpenTrain account and apply for this role in minutes.

Keep exploring

Similar Jobs

View all jobs

AI Quality Analyst — Personalization (German)

Evaluate a new personalization feature for a conversational AI in German: design multi-turn prompts, compare side-by-side responses, and write clear rationales. Contractor role, $15/hr, 3-month engagement, remote with required PST overlap.

Generative AI & RLHF
Text
Remote · Worldwide
German
Part-time · Flexible
Entry level
Hourly · $15/hr

Posted Jul 15, 2026

AI Quality Analyst, Personalization (Indonesian)

Evaluate and shape AI personalization in Indonesian — $15/hr, contractor work with a minimum of 20 hours/week; ideal candidates can commit 30+ hours with PST overlap. Design multi-turn prompts, rate model outputs, write clear rationales, and verify grounding.

Generative AI & RLHF
Text
Remote · Worldwide
Indonesian
Part-time · Flexible
Intermediate level
Hourly · $15/hr

Posted Jul 16, 2026

Polish Quality Assurance Lead

Lead quality for Polish AI training projects by reviewing AI-generated Polish content, coaching contributors, and maintaining style guides. Remote (Poland), up to $35/hr, ~20+ hours/week — ideal for experienced Polish-language reviewers.

Generative AI & RLHF
Text
Remote · Poland
Polish, English
Part-time · Flexible
Expert level
Hourly · $35/hr

Posted Jul 9, 2026