Join OpenTrain AI to collect 30-minute egocentric videos of T-shirt folding for an AI training project — earn a fixed $80 per completed video. This entry-level, remote contract task requires a smartphone, a helper to hold the camera, and following step-by-step audio and folding cues exactly.
Data Collection
100% Remote Fixed price · $80
$80 fixed price
Compensation
Worldwide
Eligibility
Entry
Experience
Mar 18, 2025
Posted
Open worldwide
Interested in this role?
Create a free OpenTrain account and apply in minutes.
OpenTrain is the #1 platform for building careers in AI training and data labeling. We connect contributors with short-term and part-time projects that help teach cutting-edge models how to understand the world.
This role is contracted through OpenTrain AI. You'll work remotely and contribute directly to how future AI systems learn from human demonstrations.
About AI Training and Data Collection
AI training (also called data collection or annotation) is the human work that teaches models by providing real examples. Collecting high-quality demonstration videos helps robots and systems learn physical tasks like folding clothing.
These projects are often flexible and accessible: many need no prior experience beyond following clear instructions, making them ideal part‑time tasks you can do from home.
The Role — What We Need
Record a single, compiled 30-minute egocentric (first-person) video of yourself folding short-sleeved T-shirts at least 120 times following the exact step sequence and audio cues described below. You will be paid a fixed $80 for each completed and accepted video.
Task type: Video data collection (demonstration of T-shirt folding).
Employment: Contractor, part-time, entry level.
Time: Each submitted video is ~30 minutes; project fits under Less than 20 hours/week.
Pay: Fixed price of USD 80 per accepted video.
How To Film — Camera, Setup, and Quality
Camera view must be egocentric (first-person). Have someone hold the phone at your eye level behind you with the bottom part of the phone touching your forehead so the view matches the sample.
Any smartphone or camera is acceptable. Set video quality to HD (preferably 30 FPS) and use 0.5x zoom if available. The video must be well-lit with clear video and sound, and your hands must remain visible throughout.
Record from first-person perspective with a helper holding the camera behind you.
Use HD video (preferably 30 FPS) and 0.5x zoom when possible.
Only use a flat surface for folding and only short-sleeved T-shirts.
Exact Folding Sequence and Audio Cues
Follow this sequence precisely for every T-shirt. Speak the numbered cue word out loud at each step, and at the end make a loud table hit as described.
You must perform a continuous series of foldings totaling at least 120 foldings in the compiled video (multiple raw cuts may be used during filming but upload only the compiled video). Do not take breaks between foldings; fold continuously for the entire duration.
1) Start with a crumpled T-shirt and say only "ONE".
2) Flatten it on its back and say only "TWO".
3) Fold from the left and say only "THREE".
4) Fold from the right and say only "FOUR".
5) Fold from the middle and say only "FIVE".
6) After finishing folding say only "SIX," then hit the table to make a loud sound.
Additional Requirements and Good Practices
Use a variety of short-sleeved T-shirts in different colors and materials so the dataset is diverse. Keep your hands clearly visible and avoid obstructions in the frame.
Multiple raw takes are allowed during filming, but submit a single compiled video file or link. The sample video must be followed exactly as shown.
Use many different short-sleeved T-shirts (colors & materials).
Keep consistent lighting and clear audio for the spoken cues and table hit.
Do not include other people in frame except the helper holding the camera behind you.
Requirements, Eligibility, and Submission
Basic requirements: a smartphone with a camera, a helper to hold the phone at eye level behind you, and the ability to record a continuous 30-minute compiled video with at least 120 foldings while following the script exactly.
This project is open worldwide. You will upload the compiled video (the instructions reference uploading to YouTube) and provide the link or submission according to OpenTrain AI's submission workflow using our internal proprietary tooling.
Experience level: Entry level — no prior annotation experience required.
Equipment: Smartphone or camera capable of HD video.
Location: Worldwide — you may participate from any country.
How To Apply
If you meet the requirements and are ready to record, apply through OpenTrain AI and follow the project onboarding steps. Review the sample carefully and practice the sequence before recording the full 30-minute video.
Sample reference video to follow precisely: https://youtube.com/shorts/D6MEMEjIKAo
Prepare a helper to hold the camera and practice the spoken cues.
Record, compile into a single 30-minute video with at least 120 foldings, and upload per the submission instructions.
Submit your video link and any required metadata through OpenTrain's submission tool to receive payment upon acceptance.
Record first-person smartphone video of household tasks to help train AI systems; part-time contractor work at $13/hr for US-based contributors. No prior AI experience required — strong attention to detail and a compatible phone/head strap are essential.
Record specification-compliant household task videos using a UMI gripper to train personal robotics systems; contractor, part-time, U.S.-only, $30/hr, 20+ hours/week with equipment shipped to you and onboarding starting soon.
Record first-person (POV) video of daily activities to help train AI and humanoid robots. This entry-level, part-time contractor role pays $6/hour (USD) and requires a smartphone, access to a real home or workspace, and a 20+ hour/week commitment.