Software Engineer / AI Trainer (AI evaluation, RLHF/RL preference training, red teaming, and rubric-based verification) — Softtek, Remote
Led RLHF preference ranking and DPO pair generation for an LLM alignment project, producing 12K+ ranked response pairs to improve reward model accuracy. Performed code generation review and reasoning step verification using custom rubrics to audit chain-of-thought outputs across JavaScript, TypeScript, Python, and SQL. Developed adversarial prompts for red teaming to uncover safety failures across prompt injection, bias, and harmful-content categories. • RLHF preference ranking and DPO pair generation from model responses • Reasoning step verification and rubric-based auditing • Red teaming prompt design and safety failure identification • Function/tool-use evaluation validation against schema and grounding constraints