AI Evaluator & Writer | Mercor (Remote)
Audited long-form analytical responses and synthetic text generated by large language models for contextual truthfulness and readability. Designed detailed benchmarking prompts to test domain-specific limitations under complex creative and professional constraints. Supplied granular structured feedback and rewrite variants to optimize model execution parameters and tone tracking. • Evaluated LLM outputs for truthfulness, formatting, and target readability • Built benchmarking prompt suites to measure model limitations across constraints • Produced structured feedback and alternative rewrites to guide improvements • Tuned guidance toward better tone tracking and execution parameters