Seminar
Special Talk Series — RL for LLM Alignment, Risk-sensitive RL, and more

A special talk series by Prof. Pascal Poupart (University of Waterloo), visiting 6/9–6/20, 2025. Joint-research theme: Reinforcement Learning Foundation Models.
Reinforcement Learning for LLM Alignment and Reasoning. How RL from human feedback improves the alignment of LLMs, recent advances in reward-guided text generation, and leveraging reward process models to improve inference-time reasoning.
Risk-Sensitive Reinforcement Learning. Most RL techniques maximize expected returns, risking catastrophic outcomes. This talk discusses recent advances in risk-sensitive RL where the search balances expected returns with various notions of risk and variability.
Training Machines to Know What They Don’t Know. Why neural predictors with a softmax output layer exhibit arbitrarily high confidence away from the training data, a simple modification to prevent it, and interpolation techniques to enhance calibration in distributed/federated learning.
When Should Reinforcement Learning Use Causal Reasoning? RL and causal reasoning complement each other; we examine which RL settings can benefit from causal reasoning, and how.
Speaker
Pascal Poupart is a Professor in the David R. Cheriton School of Computer Science at the University of Waterloo and a Canada CIFAR AI Chair at the Vector Institute. He received his Ph.D. from the University of Toronto (2005) and is best known for his contributions to reinforcement learning. He serves on the editorial boards of JMLR and TMLR and routinely as (senior) area chair for NeurIPS, ICML, AISTATS, ICLR, IJCAI, AAAI and UAI.
