You will test two AI agents side by side on a smartphone while completing realistic everyday errands. Your feedback will show where agents complete tasks well, lose control, or create a poor user experience.
- Complete tasks such as paying bills, processing returns, sending follow-up emails, making bookings, shopping, managing groceries, planning travel, coordinating family logistics, attending outings, and completing forms.
- Use personal accounts to run realistic scenarios with both AI agents.
- Approve or stop an agent when an important action requires your control.
- Screen record every attempt from a smartphone.
- Score each run for completion, quality, control issues, satisfaction, and time.
- Write clear notes about what worked, what failed, and how the experience felt.
- Follow detailed instructions and document every step accurately.
What It Pays and Takes
This is a remote, part-time contractor role completed from a smartphone. Each run takes about 1.5 hours, including evidence collection and scoring, and the schedule is flexible.
- Pay: $25 to $45 per hour.
- Location: Open to candidates in the United States.
- Language: Strong written English communication.
- Education: Bachelor's degree required.
- Experience: Two to three years of relevant experience.
- Regular use of AI tools in daily life.
- Strong attention to detail and reliable instruction-following.
- Clear written communication and organized documentation.
- Ability to record a smartphone screen.
- Active personal AI accounts that can be used on a phone.
- Generalists are welcome. Experience in engineering, business management, office operations, digital workflows, or another area related to everyday errands is helpful.
- No coding background or previous AI industry experience is required.
How It Works
Apply on OpenTrain with your resume, then complete the application on the hiring site.
About AI Training Work
AI training is the human work behind artificial intelligence, including testing systems, rating their responses, and recording clear feedback. People who carefully evaluate real-world behavior help make AI tools more accurate, useful, and easier to control.