In my undergraduate thesis, I conducted a data labeling process as part of a multimodal cyberbullying detection research
In my undergraduate thesis, I conducted a data labeling process as part of a multimodal cyberbullying detection research project on the X (Twitter) platform. I was involved in manually annotating over 34,000 text and image samples collected via the X API, classifying each entry into two categories: cyberbullying (label 1) and non-cyberbullying (label 0). The labeling was carried out using a majority-voting principle among three annotators to ensure objectivity, consistency, and reduced bias, resulting in a nearly balanced dataset of 17,195 cyberbullying and 17,447 non-cyberbullying samples.