Seminar

An Attempt at Robust and Consistent Estimation of Fixed Parameters in General State-Space Models

Please note:  This Talk will be by Video Conference

State-space models (SSMs) encompass a wide range of popular models encountered in various fields such as mathematical finance, control engineering and ecology. SSMs are essentially characterized by a hierarchical structure, with latent (unobserved) variables governed by Markovian dynamics. Fixed parameters in these models are traditionally estimated by maximum likelihood and typically include regression, auto-regression and scale parameters. The sensitivity of these estimates to deviations from the assumed model is problematic, all the more so as distributional assumptions about latent variables cannot be verified by the data analyst. Standard robust estimation techniques from generalized linear and time series models cannot be directly adapted to SSMs, and this mainly because of high-dimensional integrals that generally need to be approximated. We propose a robust estimating method by downweighting observations on the joint log-likelihood scale and by approximating the marginal log-likelihood by Laplace's method. Our attempt to compute a Fisher consistency correction term involves further approximations at the joint likelihood level to recover a typical M-functional form. Encouraging simulation results are presented to support this work in progress.

Joint work with Joanna Mills Flemming (Dalhousie University), Eva Cantoni (University of Geneva), Chris Field (Dalhousie University) and Ximing Xu (Nankai University).
 

Learning Analytics in Higher Education; Successes, Opportunities, and Challenges


Learning analytics, as defined by the Society for Learning Analytics research (SoLAR, https://solaresearch.org/), refers to "the measurement, collection, analysis and reporting of data about learners and their contexts, for purposes of understanding and optimizing learning and the environments in which it occurs.” Developments in education and learning technologies in recent decades mean that universities are now awash in data about learners and learning. Online teaching tools such as Learning Management Systems (e.g. UBC’s ‘Connect’ system), discussion forums, messaging and homework systems, simulations, peer feedback environments and audio/video tools used in flipped or blended courses all collect rich sets of data about learner activity, behaviour, course choices, and performance. As a result, we now have wealth of e-traces about learners, courses, and programs.

In this talk we will review the kinds of questions that learning analytics research typically seeks to address. These include descriptive questions about learner behaviours; predictive questions about anticipated outcomes; course-level questions about quality of instruction; and institutional questions about overall trends and patterns. We will give concrete examples from our work, and describe the current state of learning analytics in higher education generally, and at UBC in particular, as well as the prospects for future work. We will also investigate the relationship between learning analytics and statistics.

Some prediction problems in intensive care units (ICUs)

A primary goal for ICU patients is treating them to achieve positive patient outcomes (e.g., hospital discharge alive, improvement from in-hospital ailments, extended survival). A major analytical issue is the preponderance of information available at ICU entry (e.g., age, sex, co-morbidities, prescriptions, vital signs), and longitudinally (e.g., vital sign changes, dynamic renal function, in-ICU treatment). I will present a number of interesting analytic challenges in predictive modeling that my collaborators and I have encountered from a large ICU database, and discuss a few remedies that we have investigated, including implementation of a patient similarity step in an effort to improve predictive accuracy.

Multiomic Association Analysis with Kernal Machines

In the past, most studies focused on directly associating genetic variants to phenotypes without considering the intermediate layers. With recent advances in acquisition technology, we can now measure the genome, epigenome, and transcriptome among other genomic layers. Combining these data types bring about new statistical challenges. The shear dimensionality of the data introduces a serious multiple testing problem with conventional univariate analysis, especially if one is to examine the interactions between genomic layers. In this talk, we will discuss how kernel machines can be applied to reduce the data dimensionality in a biologically meaningful way. We will also describe how kernel machines can be used to analyze interactions between genomic layers. Further, we will present the concept of multikernel machines and how it enables mediation analysis to be performed.

Speaker Bio:  Bernard Ng is a postdoctoral fellow under the Department of Statistics at the University of British Columbia. Prior to his current position, he was a joint postdoctoral fellow at Stanford University and INRIA. His research focuses on statistical method development for neuroimaging and genomics applications. Bernard completed his MASc and PhD in Electrical Engineering at the University of British Columbia and his BASc in Electronics Engineering at Simon Fraser University.
 

Lower Confidence Limit of the Conditional Reliability for Weibull Lifetime with Type I Censored Data

For the data obtained from the tests with fixed stopping time (type I censored data), how to obtain the accurate lower confidence limit of the conditional reliability when the lifetime follows Weibull distribution is a difficult problem in reliability engineering. In this talk I will present a new method for calculating the lower confidence limit of the conditional reliability for Weibull lifetime with type I censored data based on the theory of ordering method in the sample space. The software will also be demonstrated.

Design of Dose-response Clinical Trials

Challenges in designing these studies include: selection of the dose
   frequency and the dose range, choice of clinical endpoints or biomarkers,
   and use of control(s), among others.  Consequences of bad Phase II study
   designs may lead to the delay of the entire clinical development program
   or the waste of R&D investment.  Misleading results obtained from poor
   designs could cause a Phase III program to confirm a wrong set of doses,
   or to stop developing a potentially useful drug.  Therefore, it is
   critical to consider an entire drug development plan, to make best use of
   all the available information, and to include all relevant experts in
   designing Phase II dose response clinical trials.  This presentation
   discusses some of these considerations.

   Biography of Speaker

   Naitee Ting is a Fellow of ASA. He is currently a Sr. Principal
   Biostatistician in the Biometrics Department of Boehringer-Ingelheim
   Pharmaceuticals Inc. (BI).  He joined BI in September of 2009, and before
   joining BI, he was at Pfizer Inc. for 22 years (1987-2009).  Naitee
   received his Ph.D. in 1987 from Colorado State University (major in
   Statistics).  He has an M.S. degree from Mississippi State University
   (1979, Statistics) and a B.S. degree from College of Chinese Culture
   (1976, Forestry).

   Naitee published articles in Technometrics, Statistics in Medicine, Drug
   Information Journal, Journal of Statistical Planning and Inference,
   Journal of Biopharmaceutical Statistics, Biometrical Journal, Statistics
   and Probability Letters, and Journal of Statistical Computation and
   Simulation.  His book ?Dose Finding in Drug Development? was published in
   2006 by Springer.  The book ?Fundamental Concepts for New Clinical
   Trialists?, co-authored with Scott Evans, was published by CRC in 2015.
   Naitee is an adjunct professor of Columbia University, University of
   Connecticut and University of Rhode Island.  Naitee has been an active
   member of both the American Statistical Association (ASA) and the
   International Chinese Statistical Association (ICSA).

    esig-bi-logo-sm.png <http://bit.ly/OJRIa8>

   Naitee Ting, Ph.D.
   Biostatistics
   Boehringer Ingelheim Pharmaceuticals, Inc.
   Ridgefield, Connecticut
   P: 203 798 4999

   naitee.ting@boehringer-ingelheim.com

   icon-twitter.gif <http://bit.ly/LO4IvZ>icon-linkedin.gif
   <http://linkd.in/LO4tAP> icon-youtube.gif <http://bit.ly/PJCrqF>

Challenges, Tools and Examples for Big Data Inference

The Opening Conference and Boot Camp of the Thematic Program on Statistical Inference, Learning, and Models for Big Data was held at the Fields Institute in Toronto from January 12th to January 23rd, 2015. A total of 35 scientific talks were presented, providing an overview of the main themes of the Program. Even if big data problems from numerous fields were covered, common challenges emerged and some tools were seen in many different contexts. A number of successful applications of big data inference were also presented.  In this talk, I will describe those challenges and tools that stood out frequently and will summarize some examples of application that were presented during Opening Conference and Boot Camp. This work is based on a paper (in press for ISI Review) that has been written by the postdoctoral fellows and long-term visitors of the Fields institute that participated in the Big Data Program.

From Pixels to Points: Using Tracking Data to Measure Performance in Professional Sports

From Pixels to Points: Using Tracking Data to Measure Performance in Professional Sports

Show Abstract

In this talk I will explore how players perform, both individually and as a team, on a basketball court. By blending advanced spatio-temporal models with geography-inspired mapping tools, we are able to understand player skill far better than either individual tool allows. Using optical tracking data consisting of hundreds of millions of observations, I will demonstrate these ideas by characterizing defensive skill and decision making in NBA players.

Some Recent Topics in Informative Sampling

In analytic studies, survey data are often obtained with complex sampling designs and the resulting analyses require special attention to handle the sampling design. When the distribution in the sample is different from the distribution in the population, the sampling design is called informative and the analytic inference under informative sampling becomes more complicated. In this talk, some recent topics on informative sampling are covered. Topics include optimal estimation, Bootstrap-based test, analysis of multilevel models, Bayesian inference, and multiple imputation under informative sampling.
 
The talk is sponsored by the CANSSI CRT project Statistical Inference for Complex Surveys with Missing Observations.

Two Statistics Co-op program presentations

Two 30-minute talks will be given to discuss the co-op experience.

Talk 1 (11:00am - 11:30am)

Speaker:  Derek Chiu, UBC Statistics Master's Student (Co-op)

Title:  Meta-consensus clustering algorithm for HGSC subtype discovery

Abstract:  High grade serous carcinoma (HGSC) is the most common form of ovarian cancer, and can be divided into several subtypes, each with distinct pathological properties. Cluster analysis can be used to group patient samples based on similarity in gene expression data. The goal is to classify patients into these subtypes for more targeted treatment. However, there are two sources of variability we need to take into account. First, there is variability between different runs of a clustering algorithm. Also, different clustering algorithms produce different results. We constructed a “meta-consensus” clustering method adapted from Monti et al. to handle some of these limitations. The approach pools together results from different clustering methods and replications before arriving at a final cluster assignment. We also used internal clustering indices to assess performance on the notions of stability, compactness, and separability. This talk describes analyses performed at OVCARE, BC Cancer Agency as part of my Co-op work term.

Talk 2 (11:30am - 12:00pm)

Speaker:  Xiaoting Ding, UBC Statistics Master's Student (Co-op)

Title:  The effect of different case definition of current smoking on the discovery of smoking -related blood gene expression signatures in COPD. 

Abstract:  Smokers and people exposed to secondhand smoke are at most important risk for chronic obstructive pulmonary disease (COPD). Some findings showed that high proportions of smoking patients with lung disease underestimates or denies smoking. Our goal was to find differences in blood gene expression between current smokers and former smokers, where smoking status was defined by self-reported status, objective measurement (exhaled carbon monoxide) and the combination of these two. Our hypothesis was to use the combination of self-reported and objective measurement to represent the smoking status was a better way to classify the phenotype, so that we could identify differential expressions more effectively.