Seminar

A simple measure of conditional dependence

To join via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca

Title: A simple measure of conditional dependence

Abstract: We propose a coefficient of conditional dependence between two random variables Y and Z given a set of other variables X1,…,Xp, based on an i.i.d. sample. The coefficient has a long list of desirable properties, the most important of which is that under absolutely no distributional assumptions, it converges to a limit in [0,1], where the limit is 0 if and only if Y and Z are conditionally independent given X1,…,Xp, and is 1 if and only if Y is equal to a measurable function of Z given X1,…,Xp. Moreover, it has a natural interpretation as a nonlinear generalization of the familiar partial R2 statistic for measuring conditional dependence by regression. Using this statistic, we devise a new variable selection algorithm, called Feature Ordering by Conditional Independence (FOCI), which is model-free, has no tuning parameters, and is provably consistent under sparsity assumptions. A number of applications to synthetic and real datasets are worked out.

Some probability “paradoxes” to warm up your winter

To join via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca

To join in-person: To join this seminar in-person, online registration is required (limited seating)

Title: Some probability “paradoxes” to warm up your winter

Abstract: I will present 2-3 counterintuitive examples from elementary probability and Markov chains that I have learned about in the last few years. This includes joint work with Omer Angel (UBC Mathematics Department), Norbert Henze (Karlsruhe Institute of Technology), and Peter Taylor (University of Melbourne).

Semiparametric inference under a density ratio model

To join via Zoom: To join this seminar via Zoom, please request Zoom connection details from headsec@stat.ubc.ca.

To join in-person: To join this seminar in-person, online registration is required (limited seating)

Title: Semiparametric inference under a density ratio model

Abstract: In many applications, we collect independent samples from interconnected populations. These population distributions share some latent structure, so it is advantageous to jointly analyze the samples. Recently, many researchers have advocated the use of the semiparametric density ratio model (DRM) to account for the latent structure these distributions share and have developed more efficient data analysis procedures based on pooled data. Advantages and several asymptotic properties of the DRM-based inferences have been demonstrated in many fields and studies, and they show that the DRM helps to improve statistical efficiency. In this thesis, we investigate several inference problems related to the DRM.

The first research problem we study is on the efficiency of the inference under a two-sample DRM. We consider a scenario where we have two samples whose sizes grow to infinity at different rates. The DRM-based inferences for the smaller-sized sample are studied. We find that some DRM-based estimates achieve the same asymptotic efficiency as the parametric estimates under some parametric model. Our simulation studies confirm our theoretical results.

Our second work studies hypothesis test problems on population quantiles when we have multiple samples whose population distributions are connected via a DRM. We explore the use of the empirical likelihood ratio test for these hypotheses, which fills a gap in the literature in this context. Our major contribution is the derivation of the limiting chi-square distribution of the test statistic. Simulation experiments and a real-data example illustrate the efficacy of the proposed method.

Finally, we solve an important open problem in the literature of DRM. The DRM postulates that the log density ratios are linear combinations of prespecified basis functions. The benefit of DRM relies on correctly specifying the basis functions. However, in applications, we do not have complete knowledge to enable a perfect choice of the basis functions. A data-adaptive choice can alleviate the risk of model misspecification, and it remains an open problem. We propose a data-adaptive approach to the choice of basis functions based on functional principal component analysis. Our simulations and real-data analyses demonstrate that our proposed method leads to an efficiency gain.

Model Projections in Model Space (MPMS): A geometric interpretation of the AIC and estimating the distance between the generating process and the best approximating model

To join via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca

Title: Model Projections in Model Space (MPMS): A geometric interpretation of the AIC and estimating the distance between the generating process and the best approximating model

Abstract: Information criteria have had a profound impact on modern ecological and I dare to say, biological science. They allow researchers to estimate which probabilistic approximating models are closest to the generating process. Unfortunately, information criterion comparison does not tell how good the best model is. This caveat is all the more important when it is considered that in science, models are commonly misspecified. In this work, my co-author Mark Taper and I show that this shortcoming can be resolved by extending the geometric interpretation of Hirotugu Akaike’s original work. Standard information criterion analysis considers only the divergences of each model from the generating process. It is ignored that there are also estimable divergence relationships amongst all of the approximating models. We then show that using both sets of divergences and an estimator of the negative self entropy, a model space can be constructed that includes an estimated location for the generating process. Thus, not only can an analyst determine which model is closest to the generating process, she/he can also determine how close to the generating process the best approximating model is. Properties of the generating process estimated from these projections are more accurate than those estimated by model averaging. We illustrate in detail our findings and our methods with two ecological examples for which we use and test two different neg-selfentropy estimators. The applications of our proposed model projection in model space extend to all areas of science where model selection through information criteria is done. Finally, the implications of our results are discussed in the context of the evidential approach to statistical and scientific inference.

Kriging Performance Without Kriging Complexity

To join via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca

To join in-person: To join this seminar in-person, online registration is required (limited seating)

Title: Kriging Performance Without Kriging Complexity

Abstract: Among methods for interpolating scattered data, the method of kriging is often considered special because it not only makes highly accurate predictions, but also provides credible confidence intervals around those predictions. Alternative interpolation methods, such as radial basis functions or inverse-distance-weighted interpolation, make predictions that are typically less accurate and come without confidence intervals. However, kriging's advantages come at the cost of high conceptual complexity. While it is possible to explain radial-basis-function interpolation in one phrase – find a weighted combination of smooth basis functions that interpolates the data – explaining kriging in an intuitive manner is difficult. Simple short explanations such as "kriging models the unknown function as a realization of a stochastic process," while accurate, convey very little intuition to the typical engineer.
In this seminar, I show how to extend radial-basis-function methodology to improve its accuracy and provide confidence intervals without sacrificing the method's conceptual simplicity. Because the extended radial-basis-function methodology gives formulas for the predictor and confidence intervals that are essentially the same as those from kriging, it provides another way to understand these formulas, one that is more accessible to the typical engineer and may provide additional insights for the professional statistician.

CANSSI Ontario / U of T Statistical Sciences ARES Seminar: Tiffany Timbers

Registration & talk details

This talk has been organized by the Canadian Statistical Sciences Institute (CANSSI) Ontario and U of T Statistical Sciences. Learn more and register for this talk here.

Talk Title: Opinionated practices for teaching reproducibility: motivation, guided instruction and practice

Abstract: In the data science courses at the University of British Columbia, we define data science as the study, development and practice of reproducible and auditable processes to obtain insight from data. While reproducibility is core to our definition, most data science learners enter the field with other aspects of data science in mind, for example predictive modelling, which is often one of the most interesting topic to novices. This fact, along with the highly technical nature of the industry standard reproducibility tools currently employed in data science, present out-of-the gate challenges in teaching reproducibility in the data science classroom. Put simply, students are not as intrinsically motivated to learn this topic, and it is not an easy one for them to learn. What can a data science educator do? Over several iterations of teaching courses focused on reproducible data science tools and workflows, we have found that providing extra motivation, guided instruction and lots of practice are key to effectively teaching this challenging, yet important subject. Here we present examples of how we deeply motivate, effectively guide and provide ample practice opportunities to data science students to effectively engage them in learning about this topic.

Optimal Transportation in the Inference of Finite Mixture Models

To Join via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca.

Title: Optimal Transportation in the Inference of Finite Mixture Models

Abstract: Finite mixture models are widely used to model data that exhibit heterogeneity. They are also used to approximate density functions of various shapes. The learning of the model is the most fundamental task. In this thesis, we 1) investigate the learning of the finite location-scale mixtures where the maximum likelihood estimate (MLE) is not well defined; 2) develop novel procedures for distributed learning of finite mixtures; and 3) apply mixture models to approximate inference in graphical models and develop algorithms to make the inference tractable. We find the transportation divergence, which is a byproduct of the optimal transportation theory, is useful for these problems.

We study the minimum Wasserstein distance estimator (MWDE) to learn the finite location-scale mixtures. We show that the MWDE is well defined and consistent. Simulation study shows it suffers some efficiency loss against a penalized version of MLE in general without noticeable gain in robustness. The MWDE is also computationally more expensive than the penalized MLE. 

Under finite mixture models, the general split-and-conquer approach for the distributed learning cannot be directly used, since the parameter space is non-Euclidean. We develop a novel split-and-conquer approach. We show that the estimator is root-n-consistent under some general conditions. Experiments show that the proposed approach has comparable statistical performance to the global estimator based on the full dataset, if the latter is feasible. It can even outperform the global estimator if the model assumption does not match the real-world data. It also has better statistical and computational performance than existing methods.

When mixtures are used in graphical models to approximate density functions, the order of the mixture increases exponentially due to recursive procedures and the inference becomes intractable. One way to make the inference tractable is to approximate the mixture by one with lower order. We propose to approximate by minimizing the composite transportation divergence (CTD) between two mixtures. The optimization problem can be solved by a majorization maximization algorithm. We show that many existing approaches are special cases of our approach. We further show that the performance of some existing algorithms can be further improved by choosing various cost functions in the CTD.

Data Journalism and the Pandemic: Observations from 19 months of Charts

To Join Via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca

To join in-person: To join this seminar in-person, online registration is required (limited seating)

Title: Data Journalism and the Pandemic: Observations from 19 months of Charts

Abstract: Why has data transparency mattered so much to people’s understanding of this pandemic? Why has British Columbia consistently been criticized for its handling of this issue? And what lessons did somebody who made daily charts about the pandemic in this province learn? Justin McElroy will provide observations on covering the pandemic over the last 19 months, what he’s learned in communicating quantitative information, and best practices for data visualization for large audiences whether in government, academia or journalism.

Computational methods for genomics data analysis: Cancer detection using circulating tumor DNA and miRNA expression inference in single cells

To join via Zoom: To join this seminar via Zoom, please request connection details from headsec@stat.ubc.ca

To join in-person: To join this seminar in-person, online registration is required (limited seating)

Title: Computational methods for genomics data analysis: Cancer detection using circulating tumor DNA and miRNA expression inference in single cells

Abstract: I will present examples of projects involving statistical modelling within two areas of ongoing interest to my group.
Tumors continuously shed DNA into the bloodstream, though typically in minute amounts. Sequencing technology can capture low-frequency circulating tumor DNA (ctDNA) fragments. In principle, we may therefore be able to detect cancer based on a blood sample, which can be collected with ease, low risk, and low costs. I will give an introduction to the clinical opportunities offered by ctDNA, the technological advances, the current data types, the main statistical challenges, and the currently applied statistical approaches. I will further present the clinical translational setting at my home department (Aarhus, Denmark) and some of our efforts to improve ctDNA detection, which include development of improved null models for DNA sequencing errors and models for mutational signal aggregation across select cancer genes or genome-wide.
MicroRNAs (miRNAs) are short RNA molecules (~22 nucleotides long). They show highly tissue- and cell-type specific expression patterns. Each miRNA regulates a specific set of genes by destabilising their mRNAs. Recent progress in single-cell sequencing allows mRNA expression profiles to be routinely obtained. However, single-cell miRNA expression cannot be quantified in high throughput settings. We have developed a method for inferring miRNA expression from mRNA expression profiles, by modelling the regulatory effect of miRNAs on target mRNAs [1]. This approach has allowed us to infer miRNA expression at the single cell level [2]. I will briefly outline the core ideas of the approach and show some results from its application on large single cells data sets.
I’m visiting UBC and the Department of Statistics until August 2022 and hope to interact with many of you during this time.

[1]: Nielsen MM, Tataru P, Madsen T, Hobolth A & Pedersen JS. Regmex: a statistical tool for exploring motifs in ranked sequence lists from genomics experiments. Algorithms Mol. Biol. 13, 17 (2018). 
[2]: Nielsen, M. M. & Pedersen, J. S. miRNA activity inferred from single cell mRNA expression. Sci. Rep. 11, 9170 (2021).

COVID-19 Modelling and Forecasting in the US and Canada: A statistician’s pro(retro)spective

To join via Zoom: To join this seminar via Zoom, please request connection details from headsec@stat.ubc.ca

To Join in-person: To join this seminar in-person, online registration is required here (limited seating)

Title: COVID-19 Modelling and Forecasting in the US and Canada: A statistician’s pro(retro)spective

Abstract: Since the COVID pandemic came to North America in March 2020, researchers and research groups have pivoted or sprung up to aid and nudge policy decision making. For the last year-plus, I’ve worked regularly with two of these – the BC COVID Modelling Group and Carnegie Mellon University’s Delphi Research Group – groups that share a number of similarities but which could not be more different. In this talk, I’ll give an overview of the structure, goals, and accomplishments of these groups and highlight many of the difficulties faced, especially when it comes to data. I’ll also describe the statistical tools used for forecasting and modeling, efforts made to evaluate models and develop better forecasters, and some major lessons learned. I’ll conclude by describing some of the many open projects that need statistical input going forward.