LLM Prompt Evaluation Tool (internal data collection for RLHF)
Developed an internal LLM prompt evaluation workflow to enable non-technical team members to assess and flag model responses. Collected evaluation outputs and fed them back into an RLHF loop to support iterative model improvement. Enabled structured rating/flagging during prompt-response testing through a web-based interface. • Created an internal web tool for model response evaluation by team members. • Structured evaluation/flagging outputs for downstream RLHF feedback. • Supported non-technical evaluators with usable UI and guidance. • Integrated evaluation results into the iterative training workflow.