Prompt Engineer – Soul AI (Model Alignment, CoT & SFT)
Performed reinforcement learning from human feedback (RLHF) to rank model outputs for technical accuracy, logical coherence, and safety, improving alignment with human intent. Engineered supervised fine-tuning (SFT) training datasets using LaTeX-audited chain-of-thought solutions to support multi-step reasoning. Conducted evaluation and quality strengthening through safety-centric workflows. • Ranked generated outputs using human-feedback criteria (accuracy, coherence, safety). • Built SFT datasets with LaTeX-audited chain-of-thought reasoning. • Ran red-teaming protocols to surface edge-case vulnerabilities and hallucinations. • Managed high-velocity training task execution with review/acceptance loops (1,000+ tasks).