LLM/AI Trainer at Turing (Remote)
Trained and evaluated LLM outputs to ensure alignment with task objectives, safety standards, and business requirements. Tested Tencent AI models for adherence to constraints by probing edge cases and validating policy-compliant behavior. Reviewed and graded AI agent responses on specialized workflows to provide detailed performance analysis and improve reliability. • Evaluated and refined enterprise-scale AI model behavior through large-scale testing. • Assessed constraint-following and detection of improper bypass attempts. • Performed systematic review and scoring of AI responses for real-world task execution. • Contributed to benchmarking and readiness improvements for deployment.