i do core data labeling task including reinforcement learning where i compare two AI-generated responses to see which on
i do core data labeling task including reinforcement learning where i compare two AI-generated responses to see which one is more helpful, accurate, and safe. I sometimes supervise fine-tuning by writing the "perfect" ideal response to a complex user prompt so the machine can learn how to write properly. i do fact-checking and verification to ensure claims are true and reduce output hallucination. i sometimes work on identification and flagging of toxic, biased, illegal , or harmful context (from guidelines provided by the AI company)