Seminar

Fast robust correlation for high dimensional data

The product moment covariance is a cornerstone of multivariate data analysis, from which one can derive correlations, principal components, Mahalanobis distances and many other results. Unfortunately the product moment covariance and the corresponding Pearson correlation are very susceptible to outliers (anomalies) in the data. Several robust measures of covariance have been developed, but few are suitable for the ultrahigh dimensional data that are becoming more prevalent nowadays. For that one needs methods whose computation scales well with the dimension, are guaranteed to yield a positive semidefinite covariance matrix, and are sufficiently robust to outliers as well as sufficiently accurate in the statistical sense of low variability. We construct such methods using data transformation. The resulting approach is simple, fast and widely applicable. We study its robustness by deriving influnce functions and breakdown values, and computing the mean squared error on contaminated data. Using these results we select a method that performs well overall, which we call wrapping [1] and which is available in the R package cellWise [2]. Wrapping allows a very substantial speedup of the DetectDeviatingCells [3] technique for flagging cellwise outliers, which is applied to genomic data with 12,000 variables. Wrapping is able to deal with even higher dimensional data, which is illustrated on color video data with 920,000 dimensions.

Keywords: Anomaly detection, Covariance matrix, Data transformation.

References

[1] J. Raymaekers and P.J. Rousseeuw (2018). Fast robust correlation for high dimensional data, arXiv:1712.05151.

[2] J. Raymaekers, P.J. Rousseeuw and W. Van den Bossche (2018), Package cellWise version 2.0.8, CRAN.

[3] P.J. Rousseeuw and W. Van den Bossche (2018), Detecting deviating data cells, Technometics, 60:2, 135-145.

Analyzing developmental processes with optimal transport

In this talk we introduce a mathematical model to describe temporal processes like embryonic development and cellular reprogramming. We consider stochastic processes in gene expression space to represent developing populations of cells, and we use optimal transport to recover the temporal couplings of the process. We apply these ideas to study 315,000 single-cell RNA-sequencing profiles collected at 40 time points over 18 days of reprogramming fibroblasts into induced pluripotent stem cells. To validate the optimal transport model, we demonstrate that it can accurately predict developmental states at held-out time points. We construct a high-resolution map of reprogramming that rediscovers known features; uncovers new alternative cell fates including neural- and placental-like cells; predicts the origin and fate of any cell class; and implicates regulatory models in particular trajectories. Of these findings, we highlight the transcription factor Obox6 and the paracrine signaling factor GDF9, which we experimentally show enhance reprogramming efficiency. Our approach provides a general framework for investigating cellular differentiation, and poses some interesting theoretical questions.

Two UBC Statistics M.Sc. Student Presentations

4pm4:30pm

Speaker: Eric Sanders, UBC Statistics M.Sc. student

Title: Incorporating Partial Adherence Into the Principal Stratification Analysis Framework

Abstract: Participants in pragmatic clinical trials often partially adhere to treatment. Simple statistical analyses of binary adherence (receiving either full or no treatment) introduce biases in the presence of partial adherence. We developed a framework which expands the principal stratification approach to allow partial adherers to have their own principal stratum and treatment level. We derived consistent estimates for bounds on population values of interest. A Monte Carlo posterior sampling method was derived that is computationally faster than Markov Chain Monte Carlo sampling, with confirmed equivalent results. Simulations indicate that the two methods agree with each other and are superior in most cases to the biased estimators created through standard principal stratification. The results suggest that these new methods may lead to increased accuracy of inference in settings where study participants only partially adhere to assigned treatment.

***

4:30pm–5pm

Speaker: Nikolas Krstic, UBC Statistics M.Sc. student

Title: Prediction of renal transplant rejection in pediatric patients using urinary metabolite data

Abstract: T-cell mediated rejection (TCMR) is a form of organ transplant rejection that can develop in renal transplant pediatric patients. Due to difficulties of correctly diagnosing TCMR, we use urinary metabolite data to predict the presence of TCMR in these patients. We use multiple different estimation methods to handle the high dimensionality and correlation present within the metabolite data, such as regularized regression and partial least squares. We also investigate how eliminating low quality samples (using sample quality metrics) or normalizing the metabolites by creatinine can affect predictive performance. Of the estimation methods used, PLS seems to be the best method for predicting TCMR when only using metabolite data. However, the composite LASSO model that incorporates both metabolite data and medical history data achieves similar predictive performance. We also observe that not normalizing metabolites by creatinine yields an equivalent or slightly improved predictive performance when compared to results obtained from metabolite normalization. Removal of low quality samples in the data (even at differing quality thresholds) does not seem to improve predictive performance, likely due to information loss.

$L_q$-type Penalty Estimation Under the Linear Restriction

This study is concerned with the Bridge Regression, which is a special family of penalized regressions of a penalty function $\sum_{j=1}^{p}|\beta_j|^q$ with $q>0$, in a linear model with linear restrictions. This estimator helps to estimate when it is known prior information regarding data that may come from both low dimensional case or high dimensional case. Using local quadratic approximation, the penalty term can be approximated around a local initial values vector and the restricted bridge estimation has written a closed-form which can be solved when $q>0$. Special cases of our suggested estimation are restricted LASSO ($q=1$) and restricted RIDGE ($q=2$) and restricted Elastic Net ($1< q < 2$) estimators. It is given some theoretical property of the restricted bridge estimator as well as computational details. A Monte Carlo simulation study is conducted based on different prior pieces of information and compared its performance with some competitive penalty estimators as well as ORACLE. Also, it is considered four real-world data examples. The numerical results show that the suggested estimation outstandingly performs when the preliminary information is correct or close to accuracy.

Detection of outbreaks in notifiable disease data

An essential function of the public health system is to detect and control disease outbreaks. The British Columbia Centre for Disease Control (BC CDC) monitors approximately 60 notifiable disease counts from 8 branch offices in 16 health service delivery areas. These disease counts exhibit a variety of characteristics, such as seasonality in meningococcal and a long-term trend in acute hepatitis A. As staff need to determine whether the reported counts are higher than expected, the detection process is both costly and fallible. To alleviate this problem, in the early 2000’s the BC CDC commissioned an automated statistical method to detect disease outbreaks. The method is based on a generalized additive partially linear model and appears to capture the characteristics of disease counts. However, it relies on certain ad-hoc criteria to flag counts for an outbreak. The BC CDC is interested in considering other alternatives. In this talk, we discuss an outbreak detection method based on robust estimators. It builds on recently proposed robust estimators for additive, generalized additive, and generalized linear models. Using real and simulated data we compare our method with that of the BC CDC and other natural competitors.

Spatial Cauchy Processes with Local Tail Dependence

We study a class of models for spatial data obtained using Cauchy convolution processes with random indicator kernel functions. We show that the resulting spatial processes have some appealing dependence properties including tail dependence at smaller distances and asymptotic independence at larger distances. We derive extreme-value limits of these processes and consider some interesting special cases. We show that estimation is feasible in high dimensions and the proposed class of models allows for a wide range of dependence structures.

A Behavioral Risk Model for Deposit Only Customer


Behavioural risk models are broadly used by financial institutions to manage risk and provide additional credits to grow profitability. The objective of the project is to build an internal behavioural risk model for customers of the BNS who only hold deposit products to calculate customer level credit scores, which will subsequently be used to predict credit risk for future customers. It is the first time that the BNS builds an internal behavioural risk model for deposit only customers. The objectives of this project are to explore if the available behavioural information on deposit only customers can be used to build an accurate model for credit risk prediction, and to compare the performance of alternative modelling approaches, including the logistic regression and alternative machine learning methods such as adaptive boosting.

Characterizing knots in lumber for strength prediction

For new structural applications and greater efficiency in the use of lumber in traditional applications, we’re required to predict the strength of lumber to assign grades more accurately. The presence of growth characteristics such as knots, shakes and the slope of gain, influence that strength. Sawmills can with modern scanning technology assess lumber online at the rate of one every few seconds and so in principle can predict the strength of each piece. Despite that technology, potential modern machine grading is still done based on models developed long ago for human graders. An important advance was made recently by Seong-Hwan Jun, a former UBC Statistics PhD student working with engineers and wood scientists at FPInnovations, a local industrial laboratory. Until his graduation, Seong was a member of the Forest Products Stochastic Modelling Group (FPSMG) jointly funded by NSERC and FPInnovations. The outcome of that work was a methodology for detecting knots from lumber scans. In particular, Seong and his collaborators at UBC and FPInnovations developed an elaborate library of computational tools for implementing that methodology. However, due to time constraints, that library of programs could not be documented for the community of potential users; so, the goal of the project to be described by the speaker, is the first steps toward the documentation of that library, a project which is being done in consultation with Dr. Jun. 

The talk will provide a background for the work being done by the Grading Enhancement Research Subgroup of the FPSMG along with illustrative examples. Work on the project started with a lengthy review of Canada’s forest products industry, lumber manufacturing, and the methods used to classify lumber into grades. Next, we had to learn about image processing and the associated software for doing so. The ultimate goal involved processing a laser light scattered image, consisting of a succession of spots running across a board at high resolution. We first explored the images of those tracheid laser dots by using java+imageJ software. Then we fitted an ellipse to each of these dots and extracted out their shape (namely, minor and major axes) and the rotation angle. These dots were in turn clustered to form larger ellipses that, roughly speaking, provides on each surface an outline of the knot face as it expresses itself on that surface of the board. After fitting the ellipses, we get a .csv file labelled tracheids.csv that can be used for knot identification and matching. By using the eccentricities and sines of rotation of the angle plots of the individual spots of laser dot images, it is possible to give a conclusion about whether a knot is present or absent of the board. A 3-dimensional plot then stitches the four surfaces of board together so that their position in the original (real) coordinate is preserved. Current work is underway by another research subgroup to use these images to predict lumber strength.

Modeling residential electricity expenditure in Mexico as a function of income and use of air conditioning

According to the International Energy Agency (IEA), "the growth in global demand for space cooling is one of the most critical yet often overlooked energy issues of our time". This growth in demand is mainly driven by economic and population growth in hot countries.

This project focuses on the economic growth aspect and its impact on electricity consumption in Mexico. A sensitivity analysis to study the relationship between a residence’s electricity expenditure and their income and presence of air conditioning was carried out. Depending on the state in Mexico, a linear model or a combination of a linear model and classification model were used.
The analysis shows that income increases of 5%, 10%, 15%, and to the median would result in mean electricity expenditure increases of 2%, 4%, 6%, and 9% respectively.

If Journals Embraced Conditional Equivalence Testing, Would Research be Better?

Motivated by recent concerns with the reproducibility and reliability of scientific research, we introduce a publication policy that incorporates “conditional equivalence testing” (CET), a two-stage testing scheme that combines standard null hypothesis significance testing and equivalence testing. We explain how such a policy could address issues of publication bias. We then develop a model that, given current incentives to publish, predicts a researcher’s most rational use of resources. Using this model, we are able to determine whether a given policy, such as our CET policy, can incentivize more reliable and reproducible research. We conclude that novel publication policies, such as the RR policy and our proposed CET policy, have the potential to better align scientists’ incentives with the goal of publishing reliable science.