Seminar

Dr. Constance van Eeden Seminar

The van Eeden seminar is a yearly event in which graduate students vote for their favorite statisticians. The winner is contacted by the organizing committee and invited to give a talk in the department’s seminar. The speaker spends one or two days on-campus, and graduate students have the opportunity to have lunch and dinner with them.

THIS YEAR'S SPEAKER

The Constance van Eeden Speaker for 2026 is Dr. Ryan Tibshirani, Professor in the Department of Statistics at the University of California, Berkeley, and Principal Investigator in the Delphi Research Group. Before joining Berkeley, Dr. Tibshirani served as a faculty member in the Departments of Statistics and Machine Learning at Carnegie Mellon University from 2011 to 2022. He earned his Ph.D. in Statistics from Stanford University in 2011 under the supervision of Professor Jonathan Taylor, and his B.S. in Mathematics from Stanford University in 2007. 

Seminar Title: Online Conformal Prediction, Multi-Level Quantile Tracking, and Gradient Equilibrium

Event Date: Thursday, April 2nd, 2026. 10:30-12:00.
Location: ESB 5104, University of British Columbia*

Event registration: https://ubc.zoom.us/meeting/register/htGYWnrFSCWhHIs8-d4wqw

Abstract:

This talk is about uncertainty quantification for time series prediction.

The overarching goal is to provide easy-to-use algorithms with formal guarantees. The algorithms we present build upon ideas from conformal prediction and control theory, are able to prospectively model conformal scores in an online setting, and adapt to the presence of systematic errors due to seasonality, trends, and general distribution shifts. We will then discuss an extension of these ideas to the setting of probabilistic forecasting, which is essentially a generalization of the framework to handle vector-valued predictions, i.e., predictions which take the form of a set of ordered quantile forecasts at different probability levels. Finally, we will generalize this even further to discuss an abstract property in online learning called gradient equilibrium, which encapsulates these settings, and more.

Dr. Ryan Tibshirani has been invited to be this year’s van Eeden speaker by the graduate students in the Department of Statistics at the University of British Columbia. A van Eeden speaker is a prominent statistician who is chosen each year to give a lecture, supported by the UBC Constance van Eeden Fund. The 2024 seminar is additionally sponsored by the Canadian Statistical Sciences Institute (CANSSI), the Pacific Institute for the Mathematical Sciences (PIMS), and the Walter H. Gage Memorial Fund.

*The room location may change.

Event Photo
Dr. Ryan Tibshirani

Distributional Balancing for Causal Inference: A Unified Framework via Characteristic Function Distance

Weighting methods are essential tools for estimating causal effects in observational studies, with the goal of balancing pre-treatment covariates across treatment groups. Traditional approaches pursue this objective indirectly, for example, via inverse propensity score weighting or by matching a finite number of covariate moments, and therefore do not guarantee balance of the full joint covariate distributions. Recently, distributional balancing methods have emerged as robust, nonparametric alternatives that directly target alignment of entire covariate distributions, but they lack a unified framework, formal theoretical guarantees, and valid inferential procedures. We introduce a unified framework for nonparametric distributional balancing based on the characteristic function distance (CFD) and show that widely used discrepancy measures, including the maximum mean discrepancy and energy distance, arise as special cases. Our theoretical analysis establishes conditions under which the resulting CFD-based weighting estimator achieves root-N consistency. Since the standard bootstrap may fail for this estimator, we propose subsampling as a valid alternative for inference. We further extend our approach to an instrumental variable setting to address potential unmeasured confounding. Finally, we evaluate the performance of our method through simulation studies and a real-world application, where the proposed estimator performs well and exhibits results consistent with our theoretical predictions.

The paper is available at https://arxiv.org/abs/2601.15449

Bio:

Dr. Chan Park is an assistant professor at the University of Illinois Urbana-Champaign. His research focuses on causal inference in complex settings, including dependence among units and omitted variables. He specializes in applying nonparametric methods and semiparametric theory to address these challenges.

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca.

Variational Inference for Variable Selection in Scalar-on-Function Regression

In practical regression applications, multiple covariates are often measured, but not all may be associated with the response variable. Identifying and including only the relevant covariates in the model is crucial for improving prediction accuracy. In this work, we develop a variational inference approach for estimation and variable selection in scalar-on-function regression, involving only functional covariates, and in partially functional regression models that also include scalar covariates. Specifically, we develop a variational expectation–maximization algorithm, with a variational Bayes procedure implemented in the E-step to obtain approximate marginal posterior distributions for most model parameters, except for the regularization parameters, which are updated in the M-step. Our method accurately identifies relevant covariates while maintaining strong predictive performance, as demonstrated through extensive simulation studies across diverse scenarios. Compared with alternative approaches, including BGLSS (Bayesian Group Lasso with Spike-and-Slab priors), grLASSO (group Least Absolute Shrinkage and Selection Operator), grMCP (group Minimax Concave Penalty), and grSCAD (group Smoothly Clipped Absolute Deviation), our approach achieves a superior balance between goodness-of-fit and sparsity in most scenarios. We further illustrate its practical utility through real-data applications involving spectral analysis of sugar samples and weather measurements from Japan.
 

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca. 

UBC Statistics Department Colloquium Series: A Debiased Machine Learning Single-Imputation Framework for Item Nonresponse in Surveys

Machine learning methods are now increasingly studied and used in National Statistical Offices, in particular to handle item nonresponse, where some survey respondents answer certain questions but leave others missing. In most surveys, item nonresponse affects key study variables, and imputation is routinely used to handle the resulting missing data. Standard parametric imputation methods can support rigorous inference when their modeling assumptions are approximately correct. However, when the imputation model is misspecified, the resulting inferences may be potentially misleading. Machine learning offers a flexible alternative by learning complex relationships between variables from the data, which can reduce the risk of misspecification. At the same time, this flexibility introduces new challenges for survey inference, since modern learning algorithms may converge more slowly than classical parametric models and may not automatically deliver valid uncertainty quantification. In this talk, I will present a survey sampling extension of the double/debiased machine learning framework of Chernozhukov et al. (2018). The proposed approach combines machine learning-based imputation with design-based survey weighting and an orthogonalized estimating strategy, leading to root-$n$ consistent and asymptotically normal estimation of population means under realistic conditions. We also develop a consistent variance estimator, yielding asymptotically valid confidence intervals while allowing the use of a wide range of machine learning algorithms. I will briefly discuss aggregation procedures and conclude with simulation results illustrating the performance of the proposed methodology.
 

This talk is part of the UBC Statistics Colloquium Series, which features broad and accessible seminars throughout the term and is sponsored in part by the Constance van Eeden Endowment.

Functional State Space Models and the Kalman Filter

In this talk, we propose a state space model for functional time series data, which extends many time series models to the realm of functional data. Most notably, we introduce the Functional ARMAX process (FARMAX), which is developed in the fully functional setting, i.e. without relying on projection onto a finite number of basis functions. These models are fit via our fully functional variant of the Kalman filter and smoother methods. The theoretical soundness of this approach is proven using tools from the theory of Gaussian measures in locally convex spaces. As an application, we consider signals data collected from small wearable medical tri-axial accelerometers affixed to a patient's wrists or ankles. Each device collects three time series (x, y, z directions) at 100Hz and can continuously collect data for 14 days.

//--Note updated time is 11 AM, March 10th--//

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca. 

 

An Economical Approach to Design Posterior Analyses

To design Bayesian studies, criteria for the operating characteristics of posterior analyses—such as power and the Type I error rate—are often assessed by estimating sampling distributions of posterior probabilities via simulation. In this work, we propose an economical method to determine optimal sample sizes and decision for such studies. Using our theoretical results that model posterior probabilities as a function of the sample size, we assess operating characteristics throughout the sample size space given simulations conducted at only two sample sizes. These theoretical results are used to construct bootstrap confidence intervals for the sample sizes and decision criteria that reflect the stochastic nature of simulation-based design. The broad applicability and wide impact of our methodology is illustrated using two clinical examples.

To join this seminar virtually, please request Zoom connection details from ea@stat.ubc.ca. 

Randomization Tests for Distributional Group Symmetry

Symmetry plays a central role in the sciences and in statistics. Yet, identifying distributional symmetry from a single sample of data can be challenging. Inferential tools for group symmetry of a probability measure exist in the form of hypothesis tests, but analogous tools for the symmetry of a conditional distribution are absent from the literature. This thesis initiates the study of nonparametric tests for equivariance and invariance of a conditional distribution under the action of a locally compact group. By characterizing conditional symmetry in terms of a conditional independence statement, we leverage the existing conditional randomization testing framework to construct consistent randomization tests for conditional symmetry. We instantiate such tests using kernel methods and derive finite-sample power lower bounds. Furthermore, we show that kernel-based tests for invariance of a probability measure can be unified with our tests under the conditional randomization framework, extending our theoretical results to those tests. We evaluate our tests for conditional symmetry on synthetic examples and demonstrate their use in particle physics applications.

To join this seminar virtually, please request Zoom connection details from ea@stat.ubc.ca. 

Automated Tuning and Analysis for Non-Reversible Parallel Tempering

Non-reversible parallel tempering (NRPT) is an effective algorithm for sampling from distributions with complex geometry, such as those arising from posterior distributions of weakly identifiable and high-dimensional Bayesian models or Gibbs distributions in statistical mechanics. In this work we introduce methods for the automated tuning of NRPT and establish convergence results that explain its observed empirical success. A central feature of all methods that we consider is that they can be fully automated and are robust, enabling them to be used in software with minimal hassle for the user, as evidenced by their application to open problems in astrophysics by our collaborators. Furthermore, the methods are all parallelizable and scale well with modern computational resources.

We begin with a study of how to bridge NRPT and variational inference in order to obtain more effective samplers. To do so, we introduce a generalized annealing path connecting the posterior to an adaptively tuned variational reference, where the reference is tuned to minimize the forward (inclusive) KL divergence to the posterior. To easily tune a general class of such variational families, we introduce AutoGD: a gradient descent method that automatically adapts its learning rate at each iteration. Our theory establishes the convergence of AutoGD, recovering the optimal rate of gradient descent (up to a constant) for a broad class of functions. Finally, to shed light on the empirical success of NRPT, we establish its uniform (geometric) ergodicity under a model of efficient local exploration. We obtain analogous ergodicity results for classical reversible parallel tempering, providing new evidence that NRPT dominates its reversible counterpart. 

To join this seminar virtually, please request Zoom connection details from ea@stat.ubc.ca. 

Nature-inspired Metaheuristics as a General-Purpose Optimization Tool in Statistical Research

Nature-metaheuristics have been widely used in engineering, computer science and artificial intelligence to tackle various types of challenging optimization problems for decades and are increasingly used across disciplines. Interestingly, metaheuristics seem to be still relatively underused in the statistical research community.

I present an overview of nature-inspired metaheuristics and describe their main appealing features, which are their speed, flexibility, availability of codes in different platforms, and ease of implementation and usage. Above all, they are virtually assumptions free, enabling us to apply these general-purpose algorithms to tackle all kinds of high-dimensional optimization tasks. I will highlight some recent applications of these algorithms to construct theory-based early phase clinical trials, that are more realistic and flexible for dose response studies. If time permits, I will provide demonstrations to show how the codes work to find user-tailored optimal experimental designs.

To join this seminar virtually, please request Zoom connection details from ea@stat.ubc.ca. 

Bio: Professor Wong is a Professor at UCLA since 1990 and over the years, he has done collaborative work in dentistry, environment health science, rheumatology, and various domains in oncology, including in the design and analysis of cancer control and prevention trials for controlling Hepatitis B among Asians, colorectal cancer for Hispanics, and fighting obesity and promoting health of minorities at workplace. His main methodology research is in the construction of model-based optimal experimental designs for various biostatistical applications. His recent interests are in the applications of nature-inspired metaheuristics to tackle challenging estimation and design problems in toxicology, clinical trials and other areas of statistics. He has delivered more than 250 presentations globally, including recent short courses in design at Seoul National University and at the Toxicology Center in TU Dortmund University in Germany. Professor Wong has received grant awards from NSF and private foundations, along with several R01 grant awards from NIH as a principal investigator. He is fellow of the American Statistical Association, the Institute of Mathematical Statistics, the American Association for the Advancement of Science, an elected member of the International Statistical Institute and a full member of the Sigma Xi - The Scientific Research Honor Society. He has also just completed a 3-year Yushan Scholarship Award from the Ministry of Education in Taiwan.

Event Photo
Weng Kee Wong

Category tree Gaussian process for computer experiments with many-category qualitative factors and application to cooling system design

In computer experiments, Gaussian process (GP) models are widely employed for emulation. However, when both qualitative and quantitative factors are involved, especially when qualitative factors have many categories, GP-based emulation becomes challenging, and existing methods can become unwieldy due to the curse of dimensionality. Motivated by computer experiments for the design of a cooling system, we introduce a new tree-based GP model for emulating computer codes with high-cardinality qualitative factors, referred to as the category tree GP (ctGP). The proposed approach incorporates a tree structure to partition the categories of the qualitative factors, after which GP or mixed-input GP models are fitted to the simulation outputs within the leaf nodes. The splitting rule is designed to reflect the cross-correlations among the categories of the qualitative factors, which a recent theoretical study has identified as a key component for improving prediction accuracy, and a pruning procedure based on cross-validation error is introduced to further ensure strong predictive performance. An application to the design of a cooling system demonstrates that the proposed method not only yields substantial computational gains and accurate predictions, but also offers meaningful insights into the system by uncovering an interpretable tree structure. Furthermore, in this cooling system design problem, the computer code is capable of generating multiple responses in addition to a single objective response; to accommodate this, we extend the ctGP framework to handle multiple responses by introducing an additional categorical variable that indicates which response is associated with each experimental point. Finally, we complete the cooling system design study by addressing the corresponding global optimization problem using Bayesian optimization with ctGP and an expected-improvement-type criterion.

To join this seminar virtually, please request Zoom connection details from ea@stat.ubc.ca. 

Bio: Ray-Bing Chen is a Professor in the Institute of Statistics and Data Science at National Tsing Hua University. He received his Ph.D. in Statistics from the University of California, Los Angeles in 2003. Prof. Chen’s research interests include statistical and machine learning, statistical modeling, computer experiments, and optimal design. His work has been published in leading journals such as the Annals of Applied Statistics, Journal of Computational and Graphical Statistics, Statistics and Computing, Technometrics, Journal of Quality Technology and Computational Statistics and Data Science. In recognition of his contributions to the field, he was elected as an Elected Member of the International Statistical Institute in 2020.

Event Photo
Ray-Bing Chen