AI Training Self-Practice Project: RLHF Simulation & Dataset Annotation Practice
Self-directedly simulated an RLHF pipeline by generating prompt-response pairs, scoring and ranking AI-generated texts, and evaluating outputs against helpfulness, accuracy, and coherence criteria. Built a small-scale annotated dataset of 200+ Chinese-language prompt-response pairs with quality scoring and documentation of correction rationale to mimic real annotation and prompt tuning work. Practiced additional annotation tasks including sentiment labeling and multimodal guideline-based annotation for image-description alignment and video content summaries. • Generated prompts and compared model responses • Assigned quality scores and flagged factual inaccuracies • Corrected logical inconsistencies and recorded annotation rationale • Performed sentiment and multimodal alignment checks