Disambiguating “Feel” in MBTI Thinking vs. Feeling Types
In my CBS5502 final project, I worked on an NLP data preparation and manual annotation workflow using the Kaggle MBTI English corpus. The project processed 8,675 raw rows into over 421,000 post segments, extracted target-word sentences for “think” and “feel,” and produced cleaned datasets with traceable audit logs. For the “feel” subset, we identified 22,587 candidate sentences, kept 18,636 usable items, and created a balanced 400-row gold set for manual review. My annotation work focused on distinguishing A: experiential/state uses from B: propositional/judgment uses of “feel.” I followed detailed review guidelines, checked whether each sentence should remain in the gold set, confirmed or corrected labels, marked exclusions such as truncated fragments or non-target usage, and supported final quality control. This experience strengthened my ability to follow annotation standards, resolve linguistic ambiguity, document decisions, and produce reliable training data for AI/NLP systems.