Annotate first-person videos with three levels of captions, action labels, and precise body-part trajectories. This part-time contractor role pays $8 per hour and requires C1 English, video annotation experience, and 20+ hours weekly.
Image & Video Annotation
Remote Hourly · $8/hr
$8/hr
Compensation
2 countries
Eligibility
Intermediate
Experience
Feb 2, 2026
Posted
Open to applicants in
India Philippines
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain AI is the organization hiring and contracting for this role. OpenTrain is the #1 platform for finding and building careers in AI training and data labeling, helping contributors discover projects, build profiles, and apply to cutting-edge work. Creating an OpenTrain account is free.
About AI Training Work
AI systems learn from examples prepared and reviewed by people. In video annotation, contributors identify actions, describe what happens, and align labels precisely with the underlying footage so models can better understand human activity and motion.
Remote work completed with a computer, internet connection, and the OpenTrain platform
Flexible, part-time contract work in a fast-growing technology field
Your annotations help evaluate and improve how AI interprets real-world video
The Role
As an Egocentric Video Annotation Specialist, you will label first-person videos showing human actions and motion trajectories. The project uses a three-tier captioning scheme and requires careful adherence to detailed guidelines, strong written English, and high-quality temporal annotation.
The work is intermediate-level and is available to contractors in India and the Philippines. The rate is $8 USD per hour, with an expected commitment of at least 20 hours per week. A pilot is planned for February 3, with potential to scale from 1,000 to as many as 10,000 video hours.
Pay: $8 USD per hour
Engagement: Part-time contractor
Location: India or the Philippines
Time requirement: 20+ hours per week
Data type: Video
Labels: Text generation, action recognition, and tracking
What You’ll Do
You will work with pre-labeled first-person video and improve or create detailed annotations. Each tier must remain grounded in what is visibly shown, without guessing about a person’s intent.
Your work will be assessed against description accuracy, description completeness, and timestamp precision, with a target quality level of 95% or higher.
Write one high-level video summary of one to two sentences without timestamps
Improve pre-annotated action-level segments with clear start and end timestamps
Write atomic action labels in verb plus object format, such as “grasp cup” or “place lid”
Create trajectory-level descriptions from scratch using sub-second segments
Describe visible motion at the body-part level, including the left hand, right hand, or torso
Use overlapping segments when needed to represent simultaneous limb movements
Maintain precise temporal alignment and follow all project guidelines
Requirements
This role requires C1 English proficiency or higher and the ability to write precise, natural descriptions. You should have prior experience with video annotation, especially action segmentation or temporal labeling, and be comfortable working with detailed customer-provided instructions.
Applicants should be able to complete a short qualification check before starting and support the February pilot. Familiarity with OpenTrain or an equivalent annotation tool is expected.
C1 English proficiency or higher
Prior experience with video annotation, action segmentation, or temporal labels
Experience timestamping actions with tight start and end alignment
Ability to write observable verb plus object action labels
Ability to create body-part-level trajectory descriptions using segments shorter than one second
Familiarity with OpenTrain or equivalent annotation tools
A demonstrated record of high-quality work, targeting 95% or higher accuracy and low rework
Availability for the pilot beginning February 3 and potential subsequent scale-up
Availability for at least 20 hours per week
Who Should Apply
This opportunity suits experienced video annotators and quality-focused review specialists who notice fine-grained movement, distinguish observable actions from assumptions, and can consistently apply a structured captioning system. It may be a strong fit if you enjoy detailed visual work and want to contribute directly to the development of AI video understanding.
Video annotators with action-level labeling experience
Reviewers accustomed to precision, completeness, and low rework
Contributors comfortable describing hand, torso, and other body-part movements
Writers with strong English comprehension and careful attention to detail
How to Apply Through OpenTrain
Create or use your free OpenTrain account to apply. During screening, be prepared to confirm your C1 English level, describe relevant video annotation experience, and complete the required qualification check before beginning project work.
Review the project requirements carefully
Confirm your English proficiency and relevant experience
Indicate that you can support at least 20 hours per week
Be ready to work in the OpenTrain annotation tool and follow detailed guidelines
Annotate videos of robotic arms performing assigned tasks and help train next-generation AI systems. This remote contractor project pays $50-$90 per hour and requires under 20 hours weekly for 1-3 months.
Join a 4–6 week contract project labeling exact scene boundaries across approximately 1,000 YouTube videos. Experienced video annotators and teams will use provided guidelines and tools to create high-quality temporal segmentation data.
Review robotic-arm footage, identify key actions, and create accurate video annotations that support machine learning and AI training. This remote contractor role pays $50-$90 per hour and runs an expected 3 to 6 months.