Mercor
main project involved evaluating and ranking AI-generated outputs for a specific use case, basically assessing how well the model understood a brief and whether the output was actually useful or just looked good on the surface. Scored responses from 1-7 on a scale and flagged issues with reasoning, structure and relevance.