AI Data Annotator
I evaluated two model responses regarding the given user prompt, selected the better response, and rated both responses on 5 different dimensions: instruction following, harmfulness, helpfulness, factuality, and conciseness and coherence to the prompt.