Seminar

Modelling Complex Biologging Data with Hidden Markov Models

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Abstract:  Hidden Markov models (HMMs) are commonly used to identify latent processes from observed time series, but it is challenging to fit them to large and complex time series collected by modern sensors. Using data from threatened resident killer whales (Orcinus orca) off the western coast of Canada as a case study, we provide solutions to three common challenges faced when identifying latent behaviour from complicated biologging data. First, biologging time series often violate common assumptions of HMMs when collected at high frequencies. We thus propose a hierarchical approach which utilizes moving-window Fourier analysis to capture fine-scale dependence structures. Second, modern technology allows researchers to directly label the latent process of interest, but rare labels can have a negligible influence on parameter estimates. We introduce a weighted likelihood approach that increases the relative influence of labelled observations. Third, applying HMMs to large time series is computationally demanding, so we propose a novel EM algorithm that combines a partial E step with variance-reduced stochastic optimization within the M step. These solutions allow researchers to model biologging data with HMMs that are more interpretable, accurate, and efficient to fit than existing methods.

Statistical Analysis with Non-Probability Survey Samples

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Abstract: We discuss issues arising from methodological developments related to inverse probability weighting and model-based prediction with non-probability survey samples, with focuses on the validity of estimation procedures for participation probabilities (propensity scores) and the impact of key assumptions on inferential approaches. We also explore strategies for dealing with undercoverage problems due to violations of the positivity assumption.   

Two MSc student presentations (Charlotte Edgar & Graeme Kempf)

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Presentation 1

Time: 11:00am – 11:30am

Speaker: Charlotte Edgar, UBC Statistics MSc student

Title: Cellwise Robust Covariance-Regularized Regression for High-Dimensional Data

Abstract: It is common to use regularization methods when dealing with high-dimensional regression problems. The scout family, developed by Witten and Tibshirani in 2009, is a class of covariance-regularized regression procedures suitable for prediction in high-dimensional settings. The scout procedure estimates the inverse covariance matrix through two log-likelihood maximization steps that each allow for regularization and then uses the estimated inverse covariance matrix to obtain estimates of the regression coefficients. The aim of this project was to make the scout procedure robust to cellwise outliers. Cellwise outliers are common in high-dimensional datasets and recent work has led to cellwise robust covariance estimates that could be used in the scout procedure. We assess the predictive performance of robust plug-in estimators and outlier detection methods. The development of a regression method that is robust to cellwise outliers, encourages sparsity, and can be applied in high-dimensional settings would be valuable to many fields, such as genomics, and is an area undergoing current research.

Presentation 2

Time: 11:30am – 12:00pm

Speaker: Graeme Kempf, UBC Statistics MSc student

Title: The impact of disease-modifying drugs for multiple sclerosis on hospitalizations and mortality in British Columbia: A retrospective study using an illness-death multi-state model

Abstract: The efficacy of disease-modifying drugs (DMDs) for multiple sclerosis was established in clinical trials that were short and excluded older individuals and individuals living with comorbidities. This has led to a lack of knowledge of the effects of chronic DMD use and the effects of DMDs on individuals that do not meet the traditional eligibility criteria for clinical trials. Multi-state models are a technique which can advance the understanding of a disease beyond that offered by time-to-event models alone. The long-term, real-world efficacy of DMDs was explored by applying a multi-state model to administrative healthcare data. Whether exposure to any DMD is associated with fewer hospitalizations, shorter hospitalizations, and/or a reduction in the chance of dying inside or outside the hospital was investigated using multi-state techniques such as intensity-based analysis and pseudo-value regression.

Tags

On Bayesian quadrature estimators

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Abstract: Computationally expensive integration problems are ubiquitous across statistics and machine learning. This creates a need for methods that approximate integrals well with as few samples as possible. Bayesian quadrature is a probabilistic integration method in which a Gaussian process prior is placed on the integrand, allowing information about properties of the integrand – such as smoothness –  to be used for improved sample efficiency. I will discuss two projects where we used Bayesian quadrature to create better estimators: (1) an improved estimator for maximum mean discrepancy when the measure is a pushforward, and (2) estimators for conditional expectation. In addition, I'll discuss how the choice of prior kernel affects the quality of uncertainty quantification in Gaussian process interpolation (and consequently, Bayesian quadrature), and present a comparison of maximum likelihood and cross-validation estimators that shows that cross-validation is more robust to smoothness misspecification.

Automatic Massively Parallel MCMC with Quantifiable Error

To Join via Zoom: To join this seminar virtually, please register here.

Abstract: Simulated Tempering (ST) is an MCMC algorithm for complex target distributions that operates on a path between the target and an amenable reference distribution. Crucially, if the reference enables i.i.d. sampling, ST is regenerative and therefore embarrassingly parallel. However, the difficulty of tuning ST has hindered its widespread adoption. In this work, we develop a simple nonreversible ST (NRST) algorithm, a general theoretical analysis of ST, and an automated tuning procedure for ST. This procedure enables straightforward integration of NRST into existing probabilistic programming languages. We provide extensive experimental evidence that our tuning scheme improves the performance and robustness of NRST algorithms on a diverse set of probabilistic models.

NRST can be seen as a meta-MCMC algorithm, in that an explorer Markov chain is required to make local moves within distributions in the path, while NRST orchestrates movement along the path. Gradient-based methods like Metropolis-adjusted Langevin algorithm (MALA) produce Markov chains that scale favorably with dimension. However, MALA depends critically on a step size parameter, and tuning it requires too much work to be useful for NRST. To resolve this issue we introduce autoMALA, an improved version of MALA that automatically sets its step size at each iteration based on the local geometry of the target distribution. We prove that autoMALA preserves the target measure despite continual adjustments of the step size. Our experiments demonstrate that autoMALA is competitive with related state-of-the-art MCMC methods, in terms of the number of density evaluations per effective sample, and it outperforms state-of-the-art samplers on targets with varying geometries.

CANCELLED: On False Positive Error

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Abstract: Controlling the false positive error in model selection is a prominent paradigm for gathering evidence in data-driven science.  In model selection problems such as variable selection and graph estimation, models are characterized by an underlying Boolean structure such as presence or absence of a variable or an edge.  Therefore, false positive error or false negative error can be conveniently specified as the number of variables/edges that are incorrectly included or excluded in an estimated model.  However, the increasing complexity of modern datasets has been accompanied by the use of sophisticated modeling paradigms in which defining false positive error is a significant challenge.  For example, models specified by structures such as partitions (for clustering), permutations (for ranking), directed acyclic graphs (for causal inference), or subspaces (for principal components analysis) are not characterized by a simple Boolean logical structure, which leads to difficulties with formalizing and controlling false positive error.  We present a generic approach to endow a collection of models with partial order structure, which leads to systematic approaches for defining natural generalizations of false positive error and methodology for controlling this error. (Joint work with Peter Bühlmann, Venkat Chandrasekaran, and Parikshit Shah)

***

Dear STAT seminars and events subscribers:

Please be advised that this seminar has been cancelled. We will reschedule the seminar in the near future. We sincerely apologize for any inconvenience.

Best wishes,

UBC Statistics Department

Modular and Efficient Compilation of Probabilistic Programs

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Abstract: Probabilistic programming languages (PPLs) enable a clean separation between probabilistic models and Bayesian inference algorithms. Ideally, such separation makes it possible for the modeler to focus on the probabilistic modeling task without knowing the details of how to implement the inference method. However, making such general inference machinery efficient and automatic is challenging. In this talk, I will discuss our ongoing work on developing a modular framework for efficient compilation of domain-specific languages (DSLs) targeting PPLs. Specifically, I will discuss techniques that enable modular design and compilation algorithms for efficient inference. Moreover, I will also briefly discuss two application areas, including probabilistic programming of real-time systems and statistical phylogenetics.

Bio: David Broman is a Professor at the Department of Computer Science, KTH Royal Institute of Technology, a Visiting Professor at the Computer Science Department, Stanford University, and an Associate Director Faculty for Digital Futures. He received his Ph.D. in Computer Science in 2010 from Linköping University, Sweden. Between 2012 and 2014, he was a visiting scholar at the University of California, Berkeley, where he also was employed as a part-time researcher until 2016. His research focuses on the intersection of (i) programming languages and compilers, (ii) real-time and cyber-physical systems, and (iii) probabilistic machine learning. David has received the Best ETAPS paper award on on programming languages and systems (the EAPLS Award, co-authored 2023), a Distinguished Artifact Award at ESOP (co-authored 2022), an outstanding paper award at RTAS (co-authored 2018), a best paper award in the journal Software & Systems Modeling (SoSyM award 2018), the award as teacher of the year, selected by the student union at KTH (2017), the best paper award at IoTDI (co-authored 2017), and awarded the Swedish Foundation for Strategic Research's individual grant for future research leaders (2016). He has worked several years within the software industry, co-founded companies, co-founded the EOOLT workshop series, and is a member of IFIP WG 2.4, Modelica Association, a senior member of IEEE, and a former board member of Forskning och Framsteg.

Generalized Data Thinning Using Sufficient Statistics

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Abstract: Sample splitting is one of the most tried-and-true tools in the data scientist toolbox. It breaks a data set into two independent parts, allowing one to perform valid inference after an exploratory analysis or after training a model. A recent paper (Neufeld, et al. 2023) provided a remarkable alternative to sample splitting, which the authors showed to be attractive in situations where sample splitting is not possible. Their method, called convolution-closed data thinning, proceeds very differently from sample splitting, and yet it also produces two statistically independent data sets from the original. In this talk, we will show that sufficiency is the key underlying principle that makes their approach possible. This insight leads naturally to a new framework, which we call generalized data thinning. This generalization unifies both sample splitting and convolution-closed data thinning as different applications of the same procedure. Furthermore, we show that this generalization greatly widens the scope of distributions where thinning is possible.

Testing for Change-points in Heavy-tailed Time Series – A Winsorized CUSUM Approach

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Abstract: It is well-known how to detect the change-point in heavy-tailed time series is an open problem since the traditional tests may not have a power. This article proposes a winsorized CUSUM approach to solve this problem. We begin by investigating the winsorized CUSUM process and deriving the limiting distributions of the Kolmogorov-Smirnov test and the Self-normalized test under the null hypothesis. Under the alternative hypothesis, we firstly uncover the behavior of change-point magnitude after the winsorized data and show that our tests have a power approaching to 1 as the sample size $n \to\infty$. We then extend the winsorizing technique to tests for multiple change-points without the prior information on the number of actual change points. Our framework is quite general and its assumption is very weak. This enables the application of our tests to both linear time series and nonlinear time series, such as TAR and G-GARCH processes. The empirical results illustrate the effectiveness of our proposed procedures for change-point detection. (This is a joint work with Rui She and Linlin Dai)

Two MSc student presentations (Hannah Bobst & Nathaniel Dyrkton)

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca.

Presentation 1

Time: 3:00pm – 3:30pm

Speaker: Hannah Bobst, UBC Statistics MSc student

Title: Analysis of Density Ratio Model Performance for Quantile Estimation

Abstract: Quantiles are important descriptive statistics in many applications. In practice, quantiles must be estimated using a representative sample from the population of interest. While parametric estimators are preferable in general, they become inaccurate in the case of even minor model misspecification. Nonparametric estimators, which requires no specific model, are common alternative to parametric estimators. However, empirical-based estimators often require large sample sizes to attain satisfactory precision. The density ratio model (DRM) is a semi-parametric model which balances the trade-off between model misspecification risks and statistical efficiency in the presence of multiple related populations. This model allows for related populations to be used alongside the population of interest, which is particularly beneficial when the sample of interest is not large enough. This paper examines the perceived benefit of the DRM quantile estimator. We compare the DRM estimator with the parametric and empirical based estimators based on their bias and variance via simulation. We created five scenarios in terms of sample sizes. We generated data from the normal and gamma distributions, but the distribution information is only used for parametric estimation. We also generated data the same distributions but added some noise. Namely, the model is now mildly misspecified. We confirm that the parametric estimator is preferred under the true model, while DRM estimator is not far behind and it improves substantially over the nonparametric estimator. When the parametric model is wrong, DRM estimator is the overall winner. We also repeated the simulation using data of daily price changes of some technology stocks and found the DRM estimator has the best overall performance.

Presentation 2

Time: 3:30pm – 4:00pm

Speaker: Nathaniel Dyrkton, UBC Statistics MSc student

Title: Integrating representative and non-representative survey data for efficient inference

Abstract: Non-representative surveys are commonly used and widely available but suffer from selection bias that generally cannot be entirely eliminated using weighting techniques. Instead, we propose a Bayesian method to synthesize longitudinal representative and unbiased surveys with non-representative biased surveys by estimating the degree of selection bias over time. We show using a simulation study that synthesizing biased and unbiased surveys together out-performs using the unbiased surveys alone, even if the selection bias may evolve in a complex manner over time. Using COVID-19 vaccination data, we are able to synthesize two large sample biased surveys with an unbiased survey to reduce uncertainty in now-casting and inference estimates while simultaneously retaining the empirical credible interval coverage. Ultimately, we are able to conceptually obtain the properties of a large sample unbiased survey if the assumed unbiased survey, used to anchor the estimates, is unbiased for all time-points.