RLHF Feedback Interface, Human Feedback Dataset & Model Validation — Master’s Thesis
Conducted Chinese-English LLM evaluation and AI output review tasks as part of interaction and interface research. Applied bilingual prompt-response evaluation to assess output issues and measure alignment-related quality signals. Produced structured evaluations covering hallucination, style mismatch, logic errors, redundancy, missing information, and constraint adherence. • Performed response ranking, quality scoring, and preference comparison for evaluated LLM outputs. • Wrote short rationales to justify evaluation decisions and label quality. • Reviewed Chinese and English academic or study materials relevant to evaluation protocols. • Applied evaluation categories to support model comparison and cognitive-load study conclusions.