Seminar

Network Vector Autoregression

We consider here a large-scale social network with a continuous response observed for each node at equally spaced time points. The responses from different nodes constitute an ultra-high dimensional vector, whose time series dynamic is to be investigated. In addition, the network structure is also taken into consideration, for which we propose a network vector autoregressive (NAR) model. The NAR model assumes each node’s response at a given time point as a linear combination of (a) its previous value, (b) the average of its connected neighbors, (c) a set of node-specific covariates, and (d) an independent noise. The corresponding coefficients are referred to as the momentum effect, the network effect, and the nodal effect respectively. Conditions for strict stationarity of the NAR models are obtained. In order to estimate the NAR model, an ordinary least squares type estimator is developed, and its asymptotic properties are investigated. We further illustrate the usefulness of the NAR model through a number of interesting potential applications. Simulation studies and an empirical example are presented.

Approach to incorporate Measurement Error Bias into the Modelling of Multivariate Data

It is common for notions of health and behaviour to be multidimensional. Structural equation modelling (SEM) incorporates ideas from regression, path-analysis and factor analysis. SEM allows the original predictors and outcomes summarised by underlying latent variables while also accounting for relationships between the latent variables. A Bayesian approach to SEM enables models that are easy to interpret and supported by data, while also portraying research hypotheses. The development and application of Bayesian approach to SEM is presented.

IAM-PIMS Public Lecture on Monday, January 22, 2018: An ODE to Statistics: Inference about Nonlinear Dynamics

This is a public lecture, hosted by the Institute of Applied Mathematics.

***

Data Science and Machine Learning at Ecoation

Ecoation is an AgTech company that provides early detection of pests, diseases and deficiencies for growers to reduce economic damage, increase crop value and decrease production costs. Data Science and Machine Learning plays a core part of product development at Ecoation. We analyze data from tens of thousands of plants in greenhouses everyday to help us understand the differences in data patterns across plant states, and provide results to growers. In the talk, we will further describe common tasks for data scientists at Ecoation. We will also provide examples of problems and challenges we have worked on in machine learning, data engineering, and other areas. 

UBC Statistics Seminar on Tue, Nov 14: Network Meta-Analysis of Disconnected Networks

Network meta-analysis is a methodology used to compare the efficacy and safety of multiple medical interventions by synthesizing data across clinical studies. The term "network" is coined because each medical intervention can be represented as a node in a network and any two nodes are linked when there is at least one study that compares the two medical interventions. Most of the current literature focuses on connected networks, which arise when there is at least one path that connects all the nodes. When this is not the case, the network is said to be disconnected. Although disconnected networks arise rather frequently, their analysis is not common practice because the standard (contrast-based) method (with fixed baseline effects) used in connected networks is thought not to work in disconnected networks. In this talk, we will provide theoretical confirmation that the standard contrast-based method with fixed baseline treatment effects does not work in disconnected networks. As an alternative to synthesizing evidence in disconnected networks, we explore empirically the suitability of using random baseline treatment effects in a contrast-based approach.

UBC Statistics Seminar on Tue, Dec 5: Multivariate One-sided Tests For Multivariate Normal and Mixed Effects Regression Models With Missing Data, Semi-continuous Data, and Censored Data

In many applications, statistical models for real data often have natural constraints or restrictions on some model parameters. For example, the growth rate of a child is expected to be positive, and patients receiving anti-HIV treatments are expected to exhibit a decline in their viral loads. Hypothesis testing for certain model parameters incorporating the natural constraints is expected to be more powerful than testing ignoring the constraints. Although constrained statistical inference, especially multi-parameter order-restricted hypothesis testing, has been studied in the literature for several decades, methods for models for complex longitudinal data are still very limited. We develop innovative multi-parameter order-restricted (or one-sided) hypothesis testing methods for modelling the following complex data: (1) multivariate normal data with non-ignorable missing values; (2) semi-continuous longitudinal data; and (3) left censored or truncated longitudinal data due to detection limits. We focus on testing mean parameters in the models, and the approaches are based on the likelihood methods. Some asymptotic results are obtained, and some computational challenges are discussed. Simulation studies are conducted to evaluate the proposed methods. Several real datasets are analyzed to illustrate the power advantages of the proposed new tests.

UBC Statistics Seminar on Tue, Nov 21: Bayesian latent variable models for understanding (pseudo-) time-series single-cell gene expression data

In the past five years biotechnological innovations have enabled the measurement of transcriptome-wide gene expression in single-cells. However, the destructive nature of the measurement process precludes genuine time-series analysis of e.g. differentiating cells. This has led to the pseudotime estimation (or cell ordering) problem: given static gene expression measurements alone, can we (approximately) infer the developmental progression (or "pseudotime") of each cell? In this talk I will introduce the problem from the typical perspective of manifold learning before re-casting it as a (Bayesian) latent variable problem. I will discuss approaches including nonlinear factor analysis and Gaussian Process Latent Variable Models, before introducing a new class of covariate-adjusted latent variable models that can infer such pseudotimes in the presence of heterogeneous environmental and genetic backgrounds.

UBC Statistics Seminar on November 2 at 4pm: Systems Monitoring and Personalized Health Management

Abstract:  Due to the advancement of computation power, sensor technologies, and data collection tools, the field of systems monitoring and health management have been evolved over the past several decades with different names under different application domains, such as statistical process control (SPC), process monitoring, health surveillance, prognostics and health management (PHM), engineering asset management (EAM), personalized medicine, etc. There are tremendous opportunities in interdisciplinary research of system monitoring through integration of SPC, system informatics, data analytics, PHM, and personalized health management. In this talk we will present our views and experience in the evolution of systems monitoring, challenges and opportunities, and applications in machine systems health management as well as human health management.

Bio: Kwok L Tsui is chair professor in the Department of Systems Engineering and Engineering Management at City University of Hong Kong. Prior to the current position, Dr. Tsui has been professor/associate professor in the School of Industrial and Systems Engineering at Georgia Institute of Technology in 1990-2011; and member of technical staff in the Quality Assurance Center at AT&T Bell Labs in 1986-1990. He received his Ph.D. in Statistics from the University of Wisconsin at Madison. Professor Tsui was a recipient of the National Science Foundation Young Investigator Award. He is Fellow of the American Statistical Association, American Society for Quality, International Society of Engineering Asset Management, and Hong Kong Institution of Engineers; elected council member of International Statistical Institute; and U.S. representative to the ISO Technical Committee on Statistical Methods. Professor Tsui was Chair of the INFORMS Section on Quality, Statistics, and Reliability and the Founding Chair of the INFORMS Section on Data Mining. Professor Tsui’s current research interests include data mining, surveillance in healthcare and public health, prognostics and systems health management, calibration and validation of computer models, process control and monitoring, and robust design and Taguchi methods.

UBC Statistics Seminar on Tuesday, November 7: Sequential decision model for inference and prediction on non-uniform hypergraphs with application to knot matching from computational forestry

We consider the knot matching problem arising in computational forestry. The knot matching problem is an important problem that needs to be solved to advance the state of the art in automatic strength prediction of lumber. We show that this problem can be formulated as a quadripartite matching problem and develop a sequential decision model that admits efficient parameter estimation along with a sequential Monte Carlo sampler on graph matching that can be utilized for rapid sampling of graph matching. We demonstrate the effectiveness of our methods on 30 manually annotated boards and present findings from various simulation studies to provide further evidence supporting the efficacy of our methods.

Introductory Statistics: the Flexible Learning Project

 

Introductory Statistics is taught in many units across UBC. Typically, instructional resources and expertise are not shared, resulting in duplication of efforts or underuse of valuable material. The Flexible Learning Introductory Statistics project has sought to address this issue via a collaboration among 10 Statistics instructors from 7 units and 3 faculties, including 5 members of our department. The project resulted in a suite of resources that address conceptually challenging topics in introductory Statistics, resources that are easy to use and are grounded in existing research on learning and Statistics. The project also produced an on-line open resource repository called StatSpace (https://statspace.elearning.ubc.ca/).  

We’ll tell the long history that led to this project. We’ll show you some of the resources and explain how they were developed and assessed. And we’ll discuss the opportunities these resources will provide to support student learning of fundamental concepts of Statistics.