NLP research papers (Inference-Time Pareto-Guided Alignment) — Person in charge
Worked as the person in charge for an NLP research project on empathy dialogue alignment. Focused on reducing distribution collapse in parameter-based alignment approaches such as RLHF-style methods. Implemented an inference-time strategy that generates multiple candidate responses and filters Pareto-optimal outputs using a multi-target discriminator. • Generated multiple candidate replies for empathy dialogue inference. • Applied Pareto-guided selection among candidates. • Used a multi-target discriminator for response filtering. • Aimed to improve alignment behavior in dialogue tasks.