AI training / data annotation & evaluation preparation (LLM output evaluation and rubrics)
1. Engineering Data Validation & Error Annotation In my role as a mechanical engineer, I conduct simulation analyses that generate large datasets across hundreds of operating conditions. My task is to systematically review these outputs — identifying anomalies, flagging data points that deviate from expected ranges, and categorizing results by severity and risk level. This process mirrors large-scale data labeling workflows: applying a consistent rubric to evaluate each data point, maintaining traceability of decisions, and ensuring that downstream decisions (design approvals, safety sign-offs) rest on accurately annotated data. Over two years, I have reviewed thousands of such data records with zero compliance incidents traced to misjudgment. 2. Technical Document Review & Content Annotation I regularly author and peer-review technical documentation including design specifications, calculation reports, and compliance submissions. Part of this review involves verifying that each statement is factually supported, cross-referencing claims against source data, and flagging sections that require revision — functionally equivalent to factuality annotation and content quality scoring in AI training. This has trained me to apply detailed annotation guidelines with high inter-rater consistency and to produce clear, actionable feedback on content quality. 3. LLM Output Evaluation (Self-Directed) As an active user of Claude and DeepSeek for over two years, I have developed a systematic practice of evaluating AI-generated responses for accuracy, helpfulness, safety, and instruction adherence. I regularly assess whether outputs contain hallucinations, logical gaps, or tone mismatches, and I iteratively refine prompts based on these evaluations. While this is self-directed rather than platform-based, the evaluation methodology — comparing outputs against intent, applying multi-dimensional rubrics, and maintaining consistency across rounds — is directly applicable to professional RLHF and data annotation projects.• Reviewed LLM outputs (e.g., Claude and DeepSeek) for quality and correctness. • Used structured guidelines, rubrics, and scoring criteria to judge responses. • Assessed output consistency and sensitivity to data accuracy. • Planned to deliver consistently high-quality annotations and evaluations for AI training platforms.