My work involved reviewing and comparing pairs of AI-generated responses (including Polish-language content, such as poe
My work involved reviewing and comparing pairs of AI-generated responses (including Polish-language content, such as poetry), rating their quality based on criteria like coherence, fluency, and relevance, and writing detailed justifications in both English and Polish explaining which response was better and why. This required careful attention to language nuance, accuracy, and consistency in applying evaluation guidelines.