OpenTrain AI-style response evaluation, dataset labeling, prompt-output comparison, and structured quality assessment (candidate)
Performed rubric-based evaluation of AI responses, including checking instruction-following and whether answers directly addressed the prompt. Compared prompt and output pairs to identify inconsistencies, unsupported claims, and weak reasoning, then provided structured assessment notes. Practiced quality assurance through consistency checks and categorization to support repeatable scoring across technical and general content. • Evaluated accuracy, completeness, clarity, and safety criteria during response rating • Identified errors and hallucination-like issues by cross-checking outputs against requirements • Wrote concise rationales and structured feedback aligned to written guidelines • Supported dataset/task improvement via systematic prompt-output comparisons