Data Labeler at Atlas Capture
Videos are passed through an AI model to say what action an ego (human) is performing only with the hands by segments of 1-20seconds. My job is to review the results and correct where needed. An example would be an ego washing dishes. Segments would look like "pick up cup, scrub cup with sponge". Another example would be an ego doing laundry. Segments would look like "pick up cloth from pile", "place cloth on ironing table", "fold cloth". Most videos are from 1-2min long. Some quality measures include: 1. No missed action 2. The Hold Rule: It must be stated when an ego is holding an object even if they are not using it 3. The segments should not exceed 0 seconds and two actions per segment 4. Consistency Rule. You shouldn't say, "pick up yellow cup, place cup on table" when it's same object. Keep the naming consistent throughout the episode. 5. No articles rule: You know those articles right? Don't use them.