Senior LLM Trainer (Supervised Fine-Tuning and RLHF)
Conducted supervised fine-tuning of large language models (LLMs) and applied RLHF techniques to align model behavior with specific tasks and domains. Performed detailed evaluation of generated outputs using metrics such as BLEU scores and perplexity, followed by iterative optimization to improve robustness and generalization. Developed and maintained reusable training notebooks with documentation to enable reproducible training runs across experiments. • Supervised fine-tuning for task/domain alignment • RLHF-based refinement using user-feedback loops • Evaluation using BLEU and perplexity • Reproducible training notebooks and documentation