Join OpenTrain AI to label humanoid robotics video for VLA models as a Level B annotator — 20+ hrs/week, contractor, paid USD $6/hr (range $5–$8); Encord experience preferred but other platform experience is accepted.
Image & Video Annotation
100% Remote Hourly · $5–$8/hr
$5–$8/hr
Compensation
Worldwide
Eligibility
Intermediate
Experience
Jun 27, 2026
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for building careers in AI training and data labeling. We connect people to flexible, remote annotation work and manage hiring and contracts for every role.
This role is offered by OpenTrain AI. Create a free OpenTrain account to apply, track work, and grow your annotation portfolio in a fast-growing industry.
About AI Training Work
AI training (data labeling/annotation) is the human side of building intelligent systems. Annotators create the examples and ground truth that modern models learn from — here, you’ll help shape vision-language-action (VLA) models by labeling video data of humanoid robots.
Many projects are 100% remote and flexible, making this a good fit for part-time work. Specialist projects that require domain knowledge (robotics, biomechanics) often pay more or ask for higher annotation skill.
The Role
We are hiring an intermediate (Level B) video annotator to label humanoid robotics videos for VLA model training. The work focuses on precise, frame-by-frame video annotation using Encord.
This is a contractor, part-time role with a minimum expectation of 20+ hours per week. You will be paid per hour (USD $6/hr listed; typical range $5–$8/hr).
What You’ll Do
Annotate video of humanoid robots using Encord according to project guidelines. Tasks will include multiple label types and may require careful temporal consistency across frames.
Create and edit bounding boxes for moving objects and body parts.
Apply segmentation masks where required to separate subjects from background.
Label point keypoints for humanoid joints and track them through frames.
Perform object detection and classification tasks on video sequences.
Follow detailed annotation guidelines and maintain high accuracy and consistency.
Requirements
Candidates must have prior experience annotating video for VLA or related vision-language-action models and meet Level B skill expectations. This is listed as an intermediate role.
Level B (intermediate) annotation experience with video data required.
Hands-on experience with Encord is preferred; if you have not used Encord, state which labeling platform(s) you have worked with previously.
Ability to work 20+ hours per week and meet contractor timelines.
Comfort working remotely and following detailed annotation instructions.
Who Should Apply
Apply if you have prior video annotation experience, familiarity with humanoid or robotics recordings, and the discipline to deliver consistent, high-quality labels. This role is a good fit for annotators who want deeper, more technical annotation work in robotics-related projects.
Intermediate annotators with experience on Encord or similar tools.
People who understand human/robot kinematics or have labeled joint/keypoint data.
Annotators seeking steady part-time contractor work with clear, technical guidelines.
How It Works & Pay
This is a contractor, part-time position hired and managed through OpenTrain AI. You will receive project materials (humanoid videos) and detailed annotation instructions after onboarding.
Pay is hourly in USD. The listed hourly rate is $6/hr with an expected range of $5–$8/hr. Work is remote and open worldwide.
Apply via your OpenTrain account and mention your Encord experience (or list other platforms you’ve used).
Onboarding will include sample tasks, quality checks, and access to the Encord workspace for the project.
Label short video clips for event classification and precise timestamping in a flexible, remote contractor role. Work 20+ hours/week with pay from $40–$80/hr, using strong attention to detail and English to produce high-quality training data.
Join a long-term, remote video-annotation contract to help train vision-language action models by classifying activities and marking objects in everyday household videos using Encord. Fixed-price contract ($58,000 USD) with up to 180 hours/month per annotator and ongoing work for top performers.