Seminar

Two MSc student presentations: Christine Chuong & Sarah Masri

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca

Presentation 1

Time: 11:00am – 11:30am

Speaker: Sarah Masri, UBC Statistics MSc student

Title: Compartmental models and Hawkes processes: equivalence and computational advantages in epidemiological modelling

Abstract: Epidemiological modelling is crucial for understanding and responding to the spread of infectious diseases, helping public health officials assess the impact of interventions and inform policy decisions. Some prominent models within this field, such as the SIR and SEIR compartmental models, rely on unobserved measurements and can be computationally intensive. This thesis investigates the equivalence between the stochastic SIR and SEIR compartmental models and the Hawkes process, a self-exciting point process, in the epidemiological setting. The research demonstrates that, under specified conditions, the SIR and SEIR models can be interpreted as special cases of the finite population Hawkes process, offering a unified framework for disease modelling that does not rely on latent measurements. This thesis contributes to the growing body of literature on stochastic epidemic models by providing an alternative approach to complement compartmental models, highlighting how inference under the Hawkes process, when fitting the process to data, is consistent with the parameters associated with the SIR and SEIR models. The findings suggest that the Hawkes process can approximate some compartmental models, offering a promising tool for epidemiological modelling.

Presentation 2

Time: 11:30am – 12:00pm

Speaker: Christine Chuong, UBC Statistics MSc student

Title: Forecasting Influenza, COVID-19 and Respiratory Syncytial Virus Detections in Canada

Abstract: Respiratory illnesses such as influenza, respiratory syncytial virus, and COVID-19 result in many hospitalizations and deaths per year in Canada. Anticipating the behaviour of viruses can be challenging as behaviour can differ between seasons, so accurate forecasts of future behaviour can help reduce uncertainty. Modelling hubs can be a useful tool in collecting and evaluating multiple forecasts in one place. Hubs run forecasting challenges that invite teams to make weekly probabilistic short-term forecasts for illnesses of interest, but no national challenge predicting respiratory viruses exists in Canada. Using data on respiratory virus detections taken from historic reports and an interactive dashboard maintained by the Public Health Agency of Canada, we established a forecasting hub for the 2024-2025 respiratory illness season. We submitted forecasts for a climatological model using historical data and an ensemble of this climatological model and an autoregressive with exogenous covariate (ARX) model with the predicted climatological medians. Forecasts were evaluated using the absolute error of the predicted median, weighted interval score (WIS) and empirical coverage of prediction intervals.

Multivariate extreme inference with application to systemic risk

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca

Abstract: In complex systems such as financial networks, the failure of a single entity can trigger cascading effects that threaten the stability of the entire system, and this is known as the systemic risk. A common measure for quantifying systemic risk is CoVaR, the Value-at-Risk of the system conditional on the distress of a single component.

In this talk, I will present a new approach to CoVaR based on tail expansions of copulas. Tail expansions of copulas provide a systematic way to characterize the joint tail behavior of multiple dependent random variables. This characterization naturally integrates and extends classical extreme value theory, offering a more flexible and interpretable representation of extremal dependence.

I will highlight the theoretical value of tail expansions in understanding the asymptotic behavior of CoVaR and demonstrate their practical use in developing new extreme value estimation methods. The talk also includes an empirical study that illustrates how the proposed approach can be used to assess the systemic risk in the U.S. financial industry.

Multilayer random dot product graphs: Estimation and online change point detection

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca

Abstract: We study the multilayer random dot product graph (MRDPG) model, an extension of the random dot product graph to multilayer networks. To estimate the edge probabilities, we deploy a tensor-based methodology and demonstrate its superiority over existing approaches. Moving to dynamic MRDPGs, we formulate and analyse an online change point detection framework. At every time point, we observe a realization from an MRDPG. Across layers, we assume fixed shared common node sets and latent positions but allow for different connectivity matrices. We propose efficient tensor algorithms under both fixed and random latent position cases to minimize the detection delay while controlling false alarms. Notably, in the random latent position case, we devise a novel nonparametric change point detection algorithm based on density kernel estimation that is applicable to a wide range of scenarios, including stochastic block models as special cases. Our theoretical findings are supported by extensive numerical experiments, with the code available online.

Two MSc student presentations (Zefan Liu & Tom Tang)

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca

Presentation 1

Time: 11:00 am - 11:30 am

Speaker: Zefan (Steve) Liu, UBC Statistics MSc student

Title: Modelling peaks over thresholds in panel data: a grouped panel generalized Pareto regression model

Abstract: Extreme Value Theory (EVT) provides probabilistic tools to understand the behaviour of extreme events, making it widely applicable across various fields. When modelling the marginal distributions of the extremes in panel data, one may wish to balance the flexibility to capture the heterogeneity among margins and the efficiency of estimation through a combination of regression technique and assuming a latent group structure among subjects. This group structure facilitating information pooling may not be known a priori and needs to be estimated from data, which may then lead to potential physical interpretations. One existing approach addressing this modelling idea builds on the Block Maxima (BM) method in EVT, which can result in a loss of valuable information. Moreover, similar to the classic k-means clustering method, the current algorithm for estimating group structure is prone to converging to locally optimal solutions. We extend the current approach to a new framework called the grouped panel generalized Pareto regression model, which utilizes the Peaks Over Threshold (POT) method to model excesses over high thresholds, thereby leveraging extreme event information more exhaustively. To account for the conditional dependence structure within clusters of excesses, we introduce a dependence-window-based sandwich estimator for standard error estimation. Taking advantage of the POT method, we develop a new grouping algorithm inspired by hierarchical clustering, which relies on a pre-determined linkage and stopping rule. This algorithm estimates the latent number of groups, the group structure and associated parameters simultaneously, and it demonstrates improved performance in identifying the globally optimal structure and balancing the goodness of fit across subjects under reasonable conditions. The finite-sample performance of our methodology is carefully evaluated through simulation studies, and an application to the river flow data from 31 hydrological stations in Upper Danube river basin is used to illustrate the real-world applicability of our modelling strategy, where the estimation efficiency is notably improved and physically interpretable group structures are identified.

Presentation 2

Time: 11:30 am – 12:00 pm

Speaker: Tom Tang, UBC Statistics MSc student

Title: The challenges of non-identifiability and a penalized maximum likelihood estimator for the beta mixture model

Abstract: This thesis explores statistical inference for the finite mixture models, with a particular focus on beta mixture models, which are widely used in biostatistics, bioinformatics, and computer science. It addresses significant issues such as unbounded likelihood and non-identifiability, which can complicate parameter estimation. To overcome the obstacle caused by the unbounded likelihood, we propose a penalized maximum likelihood estimation approach by adding a penalty term to the log-likelihood function, leading to stable parameter estimation. Additionally, we derive a closed-form expression for testing non-identifiability in beta mixture models. The effectiveness of our penalized approach is evaluated through simulation studies and compared with alternative approaches, such as the method of moments. Practical applicability is demonstrated through applications to DNA methylation analysis and local false discovery rate estimation. Finally, we suggest several directions for future research.

Statistical models that are known or suspected to be partially identified: Issues of parameterization, computation, and software development

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca

Abstract: With skyrocketing improvements in the computational performance of modern computing machines, the area of Bayesian inference applications is outstandingly improved, and Bayesian statistical analysis is used more frequently. In Bayesian inference, most computational resources are applied to running Markov chain Monte Carlo (MCMC) algorithms to obtain samples from posterior distributions. The MCMC algorithm is the main route to implement Bayesian inference. It allows for high-dimensional and flexible sampling. However, at the same time, researchers can undergo poorer computational performance when Bayesian statistical inference is performed using some specific families of models, namely partially identified models. This is because the good computational performance of the MCMC algorithm is not guaranteed. The parameters of the partially identified model are not uniquely identified, which makes the off-she-shelf MCMC algorithm hard to sample from posterior distributions. Importance sampling with transparent reparameterization (ISTP) is a good computational remedy for posterior inference with partially identified models. With the ISTP algorithm, researchers could obtain better and more stable computational performance while having samples in their original parameterization. In this talk, we first traverse scenarios of worsening computational performance with partially identified models and compare the results of ISTP with an off-the-shelf MCMC algorithm. Then, we discuss the general usability of ISTP and develop the diagnostic method for models suspected to have partial or weak identification. Along with ISTP, we introduce an R package for the Bayesian inference with the partially identified model.  Lastly, we discuss what was completed, its limitations, and possible future improvements.

Statistics for Satellite Conjunctions

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca

Abstract: The current and projected growth of the space industry has brought risk assessment for near-space encounters into sharper focus.  The dominant paradigm is the computation of a so-called collision probability, but this has a number of drawbacks, including a `dilution paradox'. I shall describe an alternative statistically-based approach to the problem that seems to have numerous advantages over existing procedures, though it brings some issues of its own.  The work is joint with Soumaya Elkantassi, Russell Carpenter and Matt Hejduk.

Efficient smoothness selection for Markov-switching models

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca

Abstract: Markov-switching models are powerful tools that allow capturing complex patterns from time series data driven by latent states. Recent work has highlighted the benefits of estimating components of these models nonparametrically, enhancing their flexibility and reducing biases, which in turn can improve state decoding, forecasting, and overall inference. Formulating such models using penalised splines is straightforward, but practically feasible methods for a data-driven smoothness selection in these models are still lacking. Traditional techniques, such as cross-validation and information criteria-based selection suffer from major drawbacks, most importantly their reliance on computationally expensive grid search methods, hampering practical usability for Markov-switching models. Michelot (2022) suggested treating spline coefficients as random effects with a multivariate normal distribution and using the R package TMB (Kristensen et al., 2015) for marginal likelihood maximisation. While this method avoids grid search and typically results in adequate smoothness selection, it entails a nested optimisation problem, thus being computationally demanding. We propose to exploit the simple structure of penalised splines treated as random effects, thereby greatly reducing the computational burden while potentially improving fixed effects parameter estimation accuracy. The proposed method offers a reliable and efficient mechanism for smoothness selection, rendering the estimation of Markov-switching models involving penalised splines feasible for complex data structures.

Approximate posterior inference for Bayesian nonparametrics with guarantees

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca

Abstract: Bayesian nonparametric (BNP) models provide a flexible and powerful framework for statistical modeling by allowing the number of features or subgroups within a population to grow with the data volume. However, posterior inference in BNP models is challenging due to the infinite parameters involved, and the lack of general and efficient inference procedures impedes their practical application. Exact posterior inference methods either analytically marginalize out the infinitely many parameters, or introduce auxiliary variables to adaptively adjust the model size during inference. The former approach relies on conjugacy relationships between priors and likelihoods and suffers from high computational costs. Similarly, the latter approach is also computationally demanding, as it requires numerical integration during sampling for nonconjugate models.

An alternative common practice in fitting BNP models involves approximating the nonparametric model with a parametric one, and subsequently applying a standard inference algorithm. While this is practical, parametric truncation can lead to significant unknown posterior approximation errors, particularly for BNP models with heavy tails that support the power-law behavior of the population. Previous work on truncated inference in BNP models has determined the truncation level via analysis of the forward generative model, which does not accurately reflect the error of approximation of the target posterior distribution. This thesis aims to develop approximate inference algorithms that can be directly used for posterior inference for general BNP models. We propose truncated inference methods and provide estimates of the posterior truncation error. Rather than setting the truncation level based on prior approximation error, we establish a desired posterior truncation error level, allowing the algorithm to adapt the truncation level until the desired truncation error is reached. The proposed algorithms are general in that they can be applied to a wide range of BNP models with completely random measure (CRM) priors. We have applied these algorithms to edge-exchangeable network models, where feature assignment variables are observed, and to latent feature models with latent feature assignment variables.

van Eeden seminar: From Diffusion Models to Schrödinger Bridges - When Generative Modeling meets Optimal Transport

Zoom Registration

https://ubc.zoom.us/meeting/register/Z_eCE0H9QqGknxiuC66eBg  

Title

From Diffusion Models to Schrödinger Bridges - When Generative Modeling meets Optimal Transport

Abstract

Denoising Diffusion models have revolutionized generative modeling. Conceptually, these methods define a transport mechanism from a noise distribution to a data distribution. Recent advancements have extended this framework to define transport maps between arbitrary distributions, significantly expanding the potential for unpaired data translation. However, existing methods often fail to approximate optimal transport maps, which are theoretically known to possess advantageous properties. In this talk, we will show how one can modify current methodologies to compute Schrödinger bridges—an entropy-regularized variant of dynamic optimal transport. We will demonstrate this methodology on a variety of unpaired data translation tasks.

van Eeden speakers

Dr. Arnaud Doucet has been invited to be this year's van Eeden speaker by the graduate students in the Department of Statistics at the University of British Columbia. A van Eeden speaker is a prominent statistician who is chosen each year to give a lecture, supported by the UBC Constance van Eeden Fund (https://www.stat.ubc.ca/constance-van-eeden-fund). The 2025 seminar is additionally sponsored by the Canadian Statistical Sciences Institute (CANSSI), the Pacific Institute for the Mathematical Sciences (PIMS), and the Walter H. Gage Memorial Fund.

 

 

CANCELLED: Policy Evaluation in Dynamic Experiments

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca

Abstract: Experiments where treatment assignment varies over time, such as micro-randomized trials and switchback experiments, are essential for guiding dynamic decisions. These experiments often exhibit nonstationarity due to factors like hidden states or unstable environments, posing substantial challenges for accurate policy evaluation.

In this talk, I will discuss how Partially Observed Markov Decision Processes (POMDPs) with explicit mixing assumptions provide a natural framework for modeling dynamic experiments and can guide both the design and analysis of these experiments. In the first part of the talk, I will discuss properties of switchback experiments in finite-population, nonstationary dynamic systems. We find that, in this setting, standard switchback designs suffer considerably from carryover bias, but judicious use of burn-in periods can considerably improve the situation and enable errors that decay nearly at the parametric rate. In the second part of the talk, I will discuss policy evaluation in micro-randomized experiments and provide further theoretical grounding on mixing-based policy evaluation methodologies. Under a sequential ignorability assumption, we provide rate-matching upper and lower bounds that sharply characterize the hardness of off-policy evaluation in POMDPs. These findings demonstrate the promise of using stochastic modeling techniques to enhance tools for causal inference. Our formal results are mirrored in empirical evaluations using ride-sharing and mobile health simulators.

***

Dear STAT news subscribers:

Please be advised that this seminar has been cancelled. We sincerely apologize for any inconvenience.

Best wishes,

UBC Statistics Department