Seminar

van Eeden seminar: How to represent part-whole hierarchies in a neural network

Registration

To join this seminar, please register via Zoom. Once your registration is approved, you'll receive an email with details on how to join the meeting.

If you have any questions about your registration or the seminar, please contact headsec@stat.ubc.ca.

Abstract

I will present a single idea about representation which allows advances made by several different groups to be combined into an imaginary system called GLOM. The advances include transformers, implicit functions, contrastive representation learning, distillation and capsules. GLOM answers the question: How can a neural network with a fixed architecture parse an image into a part-whole hierarchy which has a different structure for each image? The idea is simply to use islands of identical vectors to represent the nodes in the parse tree. The talk will discuss the many ramifications of this idea. If GLOM can be made to work, it should significantly improve the interpretability of the representations produced by transformer-like systems when applied to vision or language.

van Eeden speakers

Professor Geoffrey Hinton has been invited by our department's graduate students to be this year's van Eeden speaker. A van Eeden speaker is a prominent statistician who is chosen by our graduate students each year to give a lecture, supported by the Constance van Eeden Fund.

Vintage Factor Analysis with Varimax Performs Statistical Inference

To Join Via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca.

Post-Seminar Q&A: Graduate students are invited to stay after the seminar for a Q&A session with the speaker (~12pm–12:30pm PST).

Abstract: Vintage Factor Analysis is nearly a century old and remains popular today with practitioners. A key step, the factor rotation, is historically controversial because it appears to be unidentifiable. This controversy goes back as far as Charles Spearman. The unidentifiability is still reported in all modern multivariate textbooks. This talk will overturn this controversy and provide a positive theory for PCA with a varimax rotation. Just as sparsity helps to find a solution in p>n regression, we show that sparsity resolves the rotational invariance of factor analysis. PCA + varimax is fast to compute and provides a unified spectral estimation strategy for Stochastic Blockmodels, topic models (LDA), and nonnegative matrix factorization. Moreover, the estimator is consistent for an even broader class of models and the old factor analysis diagnostics (which have been used for nearly a century) assess the identifiability.  https://arxiv.org/abs/2004.05387

CANSSI National Seminar Series: April speaker, Grace Yi

Talk Title: Learning Noisy Data

Time

Seminar: 10:00am–11:30am PST
Student Session: 12:00pm–1:00pm PST

Registration & talk details

This talk has been organized by the Canadian Statistical Sciences Institute (CANSSI).

Learn about and register for this talk and the student session on the Canadian Statistical Sciences Institute (CANSSI) website. (Once registration opens, you'll see a green Seminar Registration button at the top of the CANSSI page.)

The CANSSI National Seminar Series

Grace Yi is the final speaker and Distinguished Lecturer of the winter/spring season of the CANSSI National Seminar Series.

All speakers in the series will provide suggested read-ahead journal articles, as part of a journal club. A few weeks before Grace’s seminar, check the CANSSI seminar page for Grace's seminar to see her suggested article.

CANSSI National Seminar Series: March speaker, Fabrizia Mealli

Talk Title: Bipartite Interference and Air Pollution Transport: Estimating Health Effects of Power Plant Interventions

Time

Seminar: 10:00am–11:30am PST
Student Session: 12:00pm–1:00pm PST

Registration & talk details

This talk has been organized by the Canadian Statistical Sciences Institute (CANSSI).

Learn about and register for this talk and the student session on the Canadian Statistical Sciences Institute (CANSSI) website.

The CANSSI National Seminar Series

Fabrizia Mealli is the third speaker of the winter/spring season of the CANSSI National Seminar Series. This winter/spring, the theme of the series is causal inference.

All speakers in the series will provide suggested read-ahead journal articles, as part of a journal club. Before this talk, have a look at Fabrizia’s suggested article.

CANSSI National Seminar Series: February speaker, Dylan Small

Talk Title: Testing an Elaborate Theory of a Causal Hypothesis

Time

Seminar: 10:00am–11:30am PST
Student Session: 12:00pm–1:00pm PST

Registration & talk details

This talk has been organized by the Canadian Statistical Sciences Institute (CANSSI).

Learn about and register for this talk and the student session on the Canadian Statistical Sciences Institute (CANSSI) website.

The CANSSI National Seminar Series

Dylan Small is the second speaker of the winter/spring season of the CANSSI National Seminar Series. This winter/spring, the theme of the series is causal inference.

All speakers in the series will provide suggested read-ahead journal articles, as part of a journal club. Before this talk, have a look at Dylan’s suggested article.

CANSSI National Seminar Series—January speaker, Linbo Wang

Talk Title: The Promises of Multiple Outcomes

Time

Seminar: 10:00am–11:30am PST
Student Session: 12:00pm–1:00pm PST

Registration & talk details

This talk has been organized by the Canadian Statistical Sciences Institute (CANSSI), as part of their National Seminar Series.

Learn about and register for this talk and the student session on the Canadian Statistical Sciences Institute (CANSSI) website.

The CANSSI National Seminar Series

Linbo Wang is the first speaker of the winter/spring season of the CANSSI National Seminar Series. This winter/spring, the theme of the series is causal inference.

All speakers in the series will provide suggested read-ahead journal articles, as part of a journal club. Details on Linbo’s read-ahead article and other suggested reading can be found here.

Addressing Open Challenges in Data Science Education

*To join this seminar, attendees will need to request Zoom connection details from headsec@stat.ubc.ca.

Post-seminar Q&A: Graduate students are invited to stay after the seminar for a Q&A session with the speaker (~12pm–12:30pm).

Abstract: In this talk, I will give an overview of the Johns Hopkins Data Science Lab: who we are, what are our goals, and the types of projects we are working on to make data science accessible world-wide. Then, I will discuss projects that I have focused on related to data science education, including developing the Open Case Studies educational resource that educators can use in the classroom to teach students how to effectively derive knowledge from data derived from real-world challenges.

The CHIME Telescope and the Search for Fast Radio Bursts

*To join this seminar, attendees will need to request Zoom connection details from headsec@stat.ubc.ca.

Post-seminar Q&A: Graduate students are invited to stay after the seminar for a Q&A session with the speaker (~12pm–12:30pm).

Abstract: The Canadian Hydrogen Intensity Mapping Experiment (CHIME) is a radio-telescope array located at the Dominion Radio Astrophysical Observatory (DRAO) near Penticton, BC. CHIME’s novel, digitally-driven design, combined with a massive processing capacity built from consumer computer hardware, make it a versatile instrument that can be used in a variety of astrophysical experiments.

In this talk, I will give an overview of CHIME, with a focus on the CHIME/FRB experiment. Fast Radio Bursts (FRBs) are millisecond-duration flashes of radio-frequency signal with unknown extra-galactic origins. CHIME/FRB’s realtime processing pipeline enabled us to expand the sample of known FRBs by order of magnitude within a couple of years of its construction, gathering important insights into their characteristics and likely origins.

Bayesian sparse regression for large-scale observational healthcare analytics

*To join this seminar via Zoom, attendees will need to request connection details from headsec@stat.ubc.ca.

Post-seminar Q&A: Graduate students are invited to stay after the seminar for a Q&A with the speaker (~12pm12:30pm).

Abstract: Growing availability of large healthcare databases presents opportunities to investigate how patients' response to treatments vary across subgroups. Even with a large cohort size found in these databases, however, low incidence rates make it difficult to identify causes of treatment effect heterogeneity among a large number of clinical covariates. Sparse regression provides a potential solution. The Bayesian approach is particularly attractive in our setting, where the signals are weak and heterogeneity across databases are substantial. Applications of Bayesian sparse regression to large-scale data sets, however, have been hampered by the lack of scalable computational techniques. We adapt ideas from numerical linear algebra and computational physics to tackle the critical bottleneck in computing posteriors under Bayesian sparse regression. For linear and logistic models, we develop the conjugate gradient sampler for high-dimensional Gaussians along with the theory of prior-preconditioning. For more general regression and survival models, we develop the curvature-adaptive Hamiltonian Monte Carlo to efficiently sample from high-dimensional log-concave distributions. We demonstrate the scalability of our method on an observational study involving n = 1,065,745 patients and p = 15,779 clinical covariates, designed to compare effectiveness of the most common first-line hypertension treatments. The large cohort size allows us to detect an evidence of treatment effect heterogeneity previously unreported by clinical trials.

Accounting for Preferential Sampling in the Statistical Analysis of Spatio-temporal Data

*To join this seminar via Zoom, attendees will need to request connection details from headsec@stat.ubc.ca.

Abstract: Spatio-temporal statistical methods are widely used to model natural phenomena across both space and time. Example phenomena include the concentrations of airborne pollutants and the distributions of endangered species. A spatio-temporal process is said to have been preferentially sampled when the locations and/or times chosen to observe it depend stochastically on the values of the process at the chosen locations and/or times. When standard statistical methodologies are used, predictions of a preferentially sampled spatio-temporal process into unsampled regions and times may be severely biased.

In this talk, we begin by providing a visual demonstration of preferential sampling in continuous-space, discrete-space, and point-pattern data. Next, we argue that preferential sampling is highly prevalent in real-world data, and in some cases, national laws may even dictate that data be preferentially sampled. Following this, we introduce two case studies: estimating historical UK black smoke pollution levels using data collected from a preferentially sampled air quality network and estimating the space-use of an endangered ecotype of killer whales using sightings data collected from the commercial whale-watching industry. For the first dataset, to confirm the presence of preferential sampling, we develop a fast, intuitive, powerful, and general test for preferential sampling. Then, to estimate bias-corrected air pollution levels across the UK, we develop the first general framework for modelling preferentially sampled spatio-temporal data. We demonstrate that existing estimates of population-level black smoke exposures may be highly inaccurate due to preferential sampling. Finally, for modelling the killer whale space-use, we develop a point process framework for modelling preferentially sampled spatio-temporal point-pattern data. We successfully develop maps that identify core areas of high activity that will hopefully prove useful for conservation purposes.

Statisticians from almost every domain routinely scrutinise the data collection protocols before analysing data. Yet within the domain of spatio-temporal modelling, few questions are typically asked about how and why the sampled locations and times were chosen. This needs to change. Ultimately, we hope that investigations into preferential sampling will become an essential component within spatio-temporal analyses, akin to model diagnostics. The methods presented in this talk are widely applicable, allowing researchers to routinely perform such investigations.