Seminar

The long road to 0.075 ppm and beyond: The statistician’s perspective on the process for setting ozone standards

***A video of the below talk can now be found here. Thank you very much to PIMS (the Pacific Institute for the Mathematical Sciences) for recording this!***

The presentation will take us along the road to the ozone standard for the United States, announced in Mar 2008 by the US Environmental Protection Agency, and then the new proposal in 2014. That agency is responsible for monitoring that nation’s air quality standards under the Clean Air Act of 1970. I will describe how I, a Canadian statistician, came to serve on the US Clean Air Scientific Advisory Committee (CASAC) for Ozone that recommended the standard and my perspectives on the process of developing it.  I will introduce the rich cast of players involved including the Committee, the EPA staff,  “blackhats”, “whitehats”, “gunslingers”, politicians and an unrevealed character waiting in the wings who appeared onstage only as the 2008 standards had been formulated.  And we will encounter a couple of tricky statistical problems that arose along with approaches, developed by the speaker and his coresearchers, which could be used to address them.  The first was about how a computational model based on things like meteorology could be combined with statistical models to infer a certain unmeasurable but hugely important ozone level, the “policy related background level” generated by things like lightning, below which the ozone standard could not go. The second was about estimating the actual human exposure to ozone that may differ considerably from measurements taken at fixed site monitoring locations.  Above all, the talk will be a narrative about the interaction between science and public policy - in an environment that harbors a lot of stakeholders with varying but legitimate perspectives, a lot of uncertainty in spite of the great body of knowledge about ozone and above all, a lot of potential risk to human health and welfare.

Probabilistic models for identification and interpretation of somatic single nucleotide variants in cancer genomes

Human cancer progresses under Darwinian evolution where (epi)genetic variation alters molecular phenotypes in individual cells. Consequently, tumours at diagnosis often consist of multiple, genotypically distinct cell populations. Somatic single nucleotide variants (SNVs) are mutations resulting from the substitution of a single nucleotide in the genome of cancer cells relative to non-malignant cells. SNVs can contribute to the malignant phenotype of cancer cells, though many SNVs likely have negligible selective value. Because many SNVs are selectively neutral, their presence in a measurable proportion of cells is likely due to drift or genetic hitchhiking. This makes SNVs an appealing class of genomic aberrations to use as markers of clonal populations and ultimately tumour evolution. Advances in sequencing technology, in particular the development of high throughput sequencing (HTS) technologies, have made it possible to systematically profile SNVs in tumour genomes. This has created an opportunity to not only catalogue SNVs in tumours but also study clonal population structure through digital allele counting capacity. Probabilistic modelling provides an attractive means to analyse this data in a coherent and statistically sound manner. I will present several probabilistic models we have developed to identify SNVs, infer clonal populations structures and resolve clonal genotypes at single cell resolution. I will then discuss our recent work that has applied these models to study metastatic migration in ovarian cancer patients.

Identification of Worsening Subjects and Treatment Responders in Comparative Longitudinal Studies

We develop a new modelling approach to enhance a recently proposed
method to detect increases of contrast enhancing lesions (CELs) on
repeated magnetic resonance imaging, which have been used as an
indicator for potential adverse events. The method signals patients
with unusual increases in CEL activity by estimating the probability
of observing CEL counts as large as those observed on a patient's
recent scans conditional on the patient's CEL counts on previous
scans. This index, computed based on a mixed effect negative binomial
regression model, can vary substantially depending on the choice of
distribution for the patient-specific random effects. Therefore, we
relax this parametric assumption to model the random effects with an
infinite mixture of beta distributions, using the Dirichlet process,
which allows any form of distribution.  As our inference is in the
Bayesian framework, we adopt a meta-analytic approach to develop an
informative prior based on previous trials. This is particularly
helpful at the early stages of a trial.  We illustrate our method with
10 multiple sclerosis (MS) trial datasets, and assess it by simulation
studies.

Identification of treatment responders is a challenge in comparative
studies where a treatment efficacy is measured by various
longitudinally collected continuous and count outcomes. Existing
procedures often identify responders based on only a single outcome.
We propose to classify patients according to their posterior
probability of being a responder estimated based on a multiple outcome
mixture model. Our novel model assumes that, conditioning on a cluster
label, each longitudinal outcome is from the generalized linear mixed
effect model (GLMM), arguably the most popular longitudinal model. As
GLMM is a rich class of models,  our general procedure enables finding
responders comprehensively defined by multiple outcomes from various
distributions. We utilize the Monte Carlo expectation-maximization
algorithm to obtain the maximum likelihood estimates of our
high-dimensional model. We demonstrate the generality of our procedure
on two MS trial datasets. The simulation study shows that
incorporating multiple outcomes improves the responder identification
performance.

Ex-post Risk Premia Tests using Individual Stocks: The IV-GMM solution to the EIV problemTitle

This paper develops an IV-GMM approach that uses past beta estimates and firm characteristics as instruments for estimating ex-post risk premia while addressing the error in-variables problem in the two-pass cross-sectional regression method. The approach is developed in the context of large cross sections of individual stocks and short time series. We establish the N-consistency of the IV-GMM ex-post risk premia estimator and obtain its asymptotic distribution along with an estimator of its asymptotic variance-covariance matrix. These results are then used to develop new tests for asset pricing model implications.
Empirically, we examine a number of popular asset pricing models and fund support for the recent q-factor model
proposed by Hou, Xue, and Zhang (2015).

Coast Mountain Bus Company Ridership Data. Sampling Plan and Analysis.

About 20% of transit buses are equipped with an Automated Passenger Counting unit (APC) that provides data on the number of passengers who board and depart the bus at each stop along a route. By assigning APC buses to different routes each day, Coast Mountain Bus Company is able to sample the number of passengers using transit buses over the service area at various times of the day.  The cleaned data is used to estimate various features of the demand for bus service in Metro Vancouver. In 2012 SCARL was asked to review how the APC buses were assigned and in 2015 review the method used to estimate ridership. This talk describes those projects including the method in use at the time and the recommendations made by SCARL about possible improvements.

Identifying Collusion in English Auctions

We develop a fully nonparametric identification framework and a test of collusion in ascending bid auctions. Assuming efficient collusion, we show that the underlying distributions of values can be identified despite collusive behaviour when there is at least one bidder outside the cartel. We propose a nonparametric estimation procedure for the distributions of values and a bootstrap stochastic dominance test of the null hypothesis of competitive behaviour against the alternative of collusion. Our framework allows for asymmetric bidders, and the test can be performed on individual bidders. The test is applied to the Guaranteed Investment Certificate auctions conducted over the Internet. There have been allegations of collusion in this market. Our test, however, does not uncover collusion for the auctions conducted over the Internet. A plausible explanation of this is that the auction design involved very limited information disclosure.

Bi-cross-validation for factor analysis

Factor analysis is a core technique in applied statistics with implications for biology, education, finance, psychology and engineering. It represents a large matrix of data through a small number k of latent variables or factors.  Despite more than 100 years of use, it remains challenging to choose k from the data. Ad hoc and subjective methods are popular, but subject to confirmation bias and they do not scale to automatic uses. There are many recent tools in random matrix theory (RMT) that apply to the factor analysis setting, so long as the noise has constant variance.  Real data usually involves heteroscedasticity foiling those techniques. There are also tools in the econometrics literature, but those apply mostly to the strong factor setting unlike RMT which handles weaker factors.  The best published method is parallel analysis, but that is only justified by simulations. We propose a bi-cross-validation approach holding out some rows and some columns of the data matrix, predicting the held out data via a factor analysis on the held in data.  We also use simulations to justify the method, though our simulations are designed using recent findings from RMT.  The new approach outperforms previous methods that we found, as measured by recovery of a true underlying factor matrix. 

This is joint work with Jingshu Wang of Stanford University.

***************************************************************************************************************************************************************************
Biosketch: Art Owen is a professor of statistics at Stanford University. He is best known for developing empirical likelihood and randomized quasi-Monte Carlo. Empirical likelihood is an inferential method that uses a data driven likelihood without requiring the user to specify a parametric family of distributions. It yields very powerful tests and is used in econometrics. Randomized quasi-Monte Carlo sampling, is a quadrature method that can attain nearly O(n**-3) mean squared errors on smooth enough functions. It is useful in valuation of options and in computer graphics. His present research interests focus on large scale data matrices. Professor Owen's teaching is focused on doctoral applied courses including linear modeling, categorical data, and stochastic simulation (Monte Carlo).

This talk was supported by the van Eeden fund, the Department of Statistics, and PIMS. A video of the event (recorded by PIMS) and slides from the talk are available here.

Statistics in Drug Development: A Personal Perspective


In this talk, I want to give the audience an idea of what the work of a statistician in a pharmaceutical company might look like using my own experiences. I will give two examples of applied statistics during drug development: an innovative strategy for a dose-ranging trial (MCPMod) and the estimation of the probability of success of a phase 3 trial given the mechanism of action of a compound for cardiovascular risk reduction. Emphasis will be on the strategic context of statistical methods within clinical drug development and on desirable skill-sets that help statisticians in the industry to succeed. I will conclude with some general comments that might be helpful for graduate students considering to work in the pharmaceutical industry.

Analysis of Aggregated Functional Data from Mixed Populations with Application to Energy Consumption

Understanding the energy consumption patterns of different types of consumers is essential in any planning of energy distribution. However, obtaining consumption information for single individuals is often either not possible or too expensive. Therefore, we consider data from aggregations of energy use, that is, from sums of individuals’ energy use, where each individual falls into one of C consumer classes. Although energy companies do have the reported number of individuals in each consumer class, reporting is not always reliable—the exact number of individuals of each class may be unknown. We develop a methodology to estimate the true number of consumers in each class and the expected energy use of each class as a function of time. We apply our method to a data set and study our method via simulation.

This is joint work with Amanda Lenzia, Camila P.E. de Souza, Ronaldo Dias, and Nancy L. Garcia.

Ordering-Free Inference from Locally Dependent Data

This paper focuses on a situation where data exhibit local dependence but their dependence ordering is not known to the econometrician. Such a situation arises, for example, when many differentiated products are correlated locally and yet the correlation structure is not known, or when there are peer effects among students' actions but precise information about their reference groups or friendship networks is absent. This paper introduces an ordering-free local dependence measure which is invariant to any permutation of the observations, and can be used to express various notions of temporal and spatial weak dependence. The paper begins with the two-sided testing problem of a population mean, and introduces a randomized subsampling approach where one performs inference using U-statistic type (or V-statistic type) test statistics that are constructed from randomized subsamples. The paper shows that one can obtain inference whose validity does not require knowledge of dependence ordering, as long as local dependence is sufficiently "local" in terms of the ordering-free local dependence measure. The method is extended to models defined by moment restrictions. This paper provides results from Monte Carlo studies.

The paper is written with an econometrics journal in mind, but it is fundamentally concerned with a statistics problem.