Seminar

Statistics and Big Data at Google

Speaker Page

Abstract:  Google lives on data. Search, Ads, YouTube, Maps, ... - they all live on data. I'll tell some stories about how Google uses data and statistics, how Google is always experimenting to make improvements (yes, this includes your searches), and how Google adapts statistical ideas to do things that have never been done before.

No statistical background is required for this talk.

Afterwards, there will be opportunity for questions and for hearing about the formation of a UBC Data Science Club.

Intergenerational equity and sustainability in a collective defined contribution plan

A collective defined contribution (CDC) plan is a group pension scheme where fixed annual contributions are made each year to a collective fund and where the benefits payable to each cohort depend on the plan’s investment experience. It differs from an individual defined contribution plan in that risks are shared among members, both within the same cohort and across different cohorts, rather than being borne separately by individual members.  The design of the mechanism used to adjust benefit levels in response to emerging experience has a direct impact on the relative benefit levels of successive generations. When the distribution of future benefit levels can be shown to be significantly different ex ante from one generation to another, the equity and sustainability of the entire scheme is called into question.

 

We study the operation of a number of different model CDC plans with simple benefit adjustment mechanisms and explore intergenerational equity and sustainability over a horizon of 100 years. In this context, adjustments based on future service are found to be inferior to adjustments affecting past service benefits. 

Interactive Engagement in the Classroom: Our Experiences with Teaching Upper-Level Statistics


Paul Gustafson co-taught STAT 300 in 2012/13 Term 1 using in-class activities in all classes. Students worked in pre-selected groups during lectures. Clicker questions were used to receive and give feedback to students during the worksheet activities. In addition, there was a two-stage midterm exam in this class, where students wrote the first part individually, and then wrote part 2 in groups.  Paul has also used worksheet activities in STAT 536 the past two years.

In 2012/13 Term 2, Will Welch taught STAT 305 using a similar model of in-class activities. He also used weekly Lab TA surveys to gather additional information to address student difficulties and created a Post Course Knowledge Retention Survey to interview past STAT 305 students. The latter gave surprising results.

Mainly via examples, Paul and Will will share their findings.

(This is joint work with Bruce Dunham and Gaitri Yapa.)

Dimensional Analysis and Its Applications in Statistics

Abstract:  Dimensional Analysis (DA) is a fundamental method in the engineering and physical sciences for analytically reducing the number of experimental variables prior to the experimentation.  The principle use of dimensional analysis is to reduce from a study of the dimensions of the variables on the form of any possible relationship between those variables.  The method is of great generality.  In this talk, an overview/introduction of DA will be first given.  A basic guideline for applying DA will be proposed, using examples for illustration.  Some initial ideas on using DA for Data Analysis and Data Collection will be discussed.  Future research issues will be proposed.
 
 
Bio:  Dr. Dennis Lin is a University Distinguished Professor of Supply Chain and Statistics at Penn State University.  His research interests are quality engineering, industrial statistics, data mining and Statistical Inference.  He has published near 200 professional (SCI/SSCI) papers in a wide variety of prestigious journals (such as, Technometrics, Annals of Statistics, Biometrika, Statistica Sinica, etc).  He has served as a co-editor for ASMBI as well as an associate editor for various (about 10) top journals.  Dr. Lin is an elected fellow of ASA, IMS and ASQ, an elected member of ISI, a fellow of RSS, and a lifetime member of ICSA.  He is the recipient of the 2004 Faculty Scholar Medal Award at Penn State University.  He is also an honorary chair professor for various universities, including a Chang-Jiang Scholar at Renmin University of China.  His recent awards include Don Owen Award (ASA), Youden Address (ASQ) and Loutit Lecturer (SSC).

Change-point Estimation and Inference in Nonparametric Regression Using Different Regularization Concepts

Classical regression techniques require a smoothness assumption to be satisfied. It makes the theoretical justification easier and the model more straightforward to interpret. In many situations however, statisticians need to deal with more complex dependence structures where the underlying functional form is non-smooth, or even discontinuous. Such models are in statistics referred to as change-point models as locations where the smoothness (continuity) assumption is not satisfied are commonly said to be change-points. Unfortunately, many existing methods require a
prior knowledge for the location of change-points in a model, which can be quite limiting in practical situations.


We will propose a new approach to change-point estimation in regression: the main advantage of our method is that it introduces a fully data-driven approach with no requirement on prior knowledge for change-point locations. It combines nonparametric regression estimation with different concepts of an L1-norm regularization.

Different alternatives are proposed, a proper statistical inference is discussed and theoretical results are derived. Finite sample performance is investigated using simulated data and real examples as well.


Edmonton, AB – 08.10.2013

S-estimators for functional principal component analysis

Principal components analysis is a widely used technique that provides an optimal lower-dimensional approximation to multivariate observations. In the  functional case, a new and simple characterization of elliptical distributions on separable Hilbert spaces allows us to obtain an equivalent stochastic optimality property for the principal component subspaces associated with elliptically distributed random elements. This property holds even when second moments do not exist.

These lower-dimensional approximations can be very useful in identifying potential outliers among high-dimensional or functional observations. In this talk we propose a new class of robust estimators for principal components. For a fixed dimension q, we robustly estimate the q-dimensional linear space that best fits the data, in the sense of minimizing the sum of coordinate-wise robust residual scale estimators. The extension to the infinite-dimensional case is also studied. In analogy to the linear regression case, we call this proposal S-estimators. Our method is consistent for elliptical random vectors, and is Fisher-consistent for elliptically distributed random elements on arbitrary Hilbert spaces. Numerical experiments show that our proposal is highly competitive when compared with other existing methods when the data are generated both by finite- or infinite-rank stochastic processes. We also illustrate our approach using two real functional data sets, where the robust estimator is able to discover atypical observations in the data that would have been missed otherwise.

This talk is the result of recent collaborations with Graciela Boente and David Tyler.

Interpretable Functional Principal Component Analysis

Functional principal component analysis (FPCA) is popularly used to explore major sources of variation in a sample of random curves. The major sources of variation in these curves are represented by the estimated functional principal components (FPCs). The intervals with high values of FPCs are interpreted as where sample curves have major variations. However, these intervals are often hard to be identified by naive users, because of the vague definition of "high value". We develop a novel penalty-based method to derive FPCs that are only non-zero on short intervals, and strictly zero on other intervals. Our derived FPCs are easier to be interpreted: those non-zero intervals are where sample curves have major variations. We propose an efficient algorithm to estimate interpretable FPCs using projection deflation. The estimated interpretable FPCs are shown to be strongly consistent and asymptotically normal under mild conditions. Our simulation studies show that our method can obtain more interpretable FPCs than other FPCA methods, while these FPCs explain similar variations of sample curves as FPCs estimated from other methods. Our method is also demonstrated by analyzing real data in two applications.

Marine mammal research in the 21st century: big data, big animals, and big issues

MMRU Page

Abstract:  Historically marine mammals were hunted for their valuable meat, oil and fur; or were culled to reduce perceived competition with fisheries.  Today, marine mammals are protected in most parts of the world from the effects of fishing, hunting, shipping, oil exploration, and other disturbances. Marine mammal research has therefore been placing greater emphasis in recent years on assessing the needs of marine mammals and the ability of marine mammals to meet them.  This includes identifying critical habitat needed by marine mammals to survive; documenting how and where they find their food; and estimating their energy requirements.  Much of this information is being gathered using tracking devices attached to individual animals.  However, the amount of data that an individual animal gathers has increased exponentially from a few at-sea locations per day to millions of lines of information collected as frequently as 16 times per second.  Marine mammals are now carrying cell-phone size data loggers that can record acceleration, pitch, roll, speed, time, depth, light levels, temperature, sound, and even video and still images. Unfortunately, most biologists are ill equipped to process the amount of data that is streaming onto their computers.  This seminar will focus on some of the bio-logging studies being carried out at UBC on killer whales, Steller sea lions, northern fur seals, and  Australian fur seals—and will discuss the statistical challenges of big data and the need for inter-department collaborations to resolve these and other big issues facing marine mammal conservation in the 21st century.

Data and Decision Making at Yammer

Abstract:
The Harvard Business Review recently called data scientist the 'sexiest job of the 21st century'.  As a data scientist at Yammer, I'll give a brief overview of what data science is and what the day-to-day life of a data scientist looks like.  I'll talk about some of the challenges present in our data, and try to give some insight into how we use data to better understand our users and make decisions at Yammer.
About the speaker:
Matthew is currently a data scientist at Yammer.  Previously, he completed a Ph.D in mathematics at UBC, studying probability theory, and was a Fellow in the Insight Data Science Fellows Program.
About Yammer:
Yammer is an Enterprise Social Network that brings together people, conversations, content, and business data in a single location.  Founded in 2008, Yammer was acquired by Microsoft Corporation in 2012 and is now part of the Microsoft Office Division.
There won't be any statistical background required, and Matthew will stay after the talk to answer any questions that people might have about data science/Yammer.  

The Assessment of Mediation in Epidemiology: What is the contribution of Bayesian Thinking?

Speaker's Page

Abstract:  Causal mediation analysis is an exciting new area innovation in biostatistics. It concerns the analysis of causal pathways that link the exposure variable to the outcome.  For example, when investigating the effect of obesity on mortality, we may be interested in the mediating role of high blood pressure.   Does obesity affect mortality through pathways other than hypertension? In recent years there have been new developments in mediation analysis techniques using the framework of counterfactuals (potential outcomes). In this talk, I will give an overview of causal mediation analysis, and opportunities for Bayesian contributions to new methodology. I will illustrate mediation analysis in a study that examines the determinants of mortality among offenders with mental illness in British Columbia. This is joint work with Julian Somers and Scott Venners from the Faculty of Health Sciences @SFU.