Seminar

Sequential ED-design for Binary Dose-response Experiments

Dose-response experiments and subsequent data analyses are often carried out according to optimal designs for the purpose of accurately determining a specific effective dose (ED) level. If the interest is the dose-response relationship over a range of ED levels, many existing optimal designs are misaligned. In this dissertation, we propose a new design procedure, called two-stage sequential ED-design, which directly and simultaneously targets several ED levels. We use a small number of trials to provide a tentative estimation of the model parameters. The doses of the subsequent trials are then selected sequentially, based on the latest model information, to maximize the efficiency of the ED estimation over several ED levels.

Although the commonly used logistic and probit models are convenient summaries of the dose-response relationship, they can be too restrictive. We introduce and study a more flexible albeit slightly more complex three-parameter logistic dose-response model. We explore the effectiveness of the sequential ED-design and the D-optimal design under this model, and develop an effective model fitting strategy. We develop a two-step iterative algorithm to compute the maximum likelihood estimate of the model parameters. We prove that the algorithm iteration increases the likelihood value, and therefore will lead to at least a local maximum of the likelihood function. We also study the numerical solution to the D-optimal design for the three-parameter logistic model.  Interestingly, all our numerical solutions to the D-optimal design are three-support distributions.

We also discuss the use of the ED-design when experimental subjects become available in groups.  We introduce the group sequential ED-design, and demonstrate how to construct this design. The ED-design has a natural extension to more complex models, and can satisfy a broad range of the demands that may arise in applications.

Elicitability and backtesting: Perspectives for banking regulation

Conditional forecasts of risk measures play an important role in internal risk management of financial institutions as well as in regulatory capital calculations. In order to assess forecasting performance of a risk measurement procedure, risk measure forecasts are compared to the realized financial losses over a period of time and a statistical test of correctness of the procedure is conducted. This process is known as backtesting. Such traditional backtests are concerned with assessing some optimality property of a set of risk measure estimates. However, they are not suited to compare different risk estimation procedures. We investigate the proposal of comparative backtests, which are better suited for method comparisons on the basis of forecasting accuracy, but necessitate an elicitable risk measure. We argue that supplementing traditional backtests with comparative backtests will enhance the existing trading book regulatory framework for banks by providing the correct incentive for accuracy of risk measure forecasts. In addition, the comparative backtesting framework could be used by banks internally as well as by researchers to guide selection of forecasting methods. The discussion focuses on three risk measures, Value-at-Risk, expected shortfall and expectiles, and is supported by a simulation study and data analysis. 

UBC Statistics Seminar on Tue, Oct 17: Markov random fields, geostatistics, and matrix-free computation

Abstract:  Since their introduction in statistics through the seminal works of Julian Besag, Gaussian Markov random fields have become central to spatial statistics, with applications in agriculture, epidemiology, geology, image analysis and other areas of environmental science. Specified by a set of conditional distributions, these Markov random fields provide a very rich and flexible class of spatial processes, and their adaptability to fast statistical calculations, including those based on Markov chain Monte Carlo computations, makes them very attractive to statisticians. In recent years, new perspectives have emerged in connecting Gaussian Markov random fields with geostatistical models, and in advancing vast statistical computations. In this talk, I will briefly discuss the scaling limit of lattice-based Gaussian Markov random fields, namely,  the de Wijs process that originates in the famous work of George Matheron on gold mines in South Africa. I will then explore how this continuum limit connection holds out further possibilities to fit a wide range of new continuum models by using Gaussian Markov random fields. The main focus of the talk will be on matrix-free computation for these models. In particular, for spatial mixed linear models, I will present frequentist residual maximum likelihood inference via matrix-free h-likelihood computation. I will draw applications both from areal-unit and point-referenced spatial data. The work resulted from collaborations with (late) Julian Besag, and Ph.D.students Somak Dutta (former) and Chunxiao Wang (current).

Short bio:  Debashis Mondal is an associate professor at the Department of Statistics, Oregon State University.  Prior to joining Oregon State in 2014, Mondal was on the statistics faculty at the University of Chicago. He received his Ph.D. in Statistics from the University of Washington and both his Bachelor’s and Master’s degrees in Statistics from the Indian Statistical Institute in Kolkata, India. Mondal's research interests include spatial statistics, MCMC and timeseries. He is a recipient of the NSF career award, the young researcher award by the International Indian Statistical Association, and is an elected member of the International Statistical Institute.

Website:  http://stat.oregonstate.edu/people/mondal-debashis

Special Lecture: 150 Years or More of Data Analysis in Canada

Join us as David Bellhouse (a statistician, Statistics historian, and professor emeritus from the University of Western Ontario), discusses the past and future of data analysis and Statistics in Canada. Topics will include Statistics from Victorian times through the foundational work of Fisher, Neyman, and Pearson and then into the computer age, bringing us to the emergence of Data Science.

4:30pm - 5:30pm:  Registration and mingling in the AERL lobby (with light refreshments)

5:30pm - 6:30pm:  Lecture in AERL 120 (the main lecture hall)

***

RSVP to headsec@stat.ubc.ca

by Friday, October 13, 2017.

This event is cosponsored by the Canadian Statistical Sciences Institute, the Statistical Society of Canada, and the UBC Department of Statistics.

UBC Statistics Seminar on Tuesday, Oct 10: Understanding gene regulation through graph-based posterior regularization in structured probabilistic models

 

Abstract:  Despite having sequenced the human genome over fifteen years ago, much is still unknown about how it functions. With the advent of high-throughput genomics technologies, it is now possible to measure properties of the genome across the entire genome in a single experiment, such as measuring where a given protein binds to the DNA or what genes are expressed. However, the complexity and massive scale of these data sets--billions of base pairs with thousands of measurements each--pose challenges to their analysis. My research focuses on the development of new machine learning methods that address the challenges posed by genomics data sets. 

 

I will focus on a method for combining probabilistic models with graph-based methods for semi-supervised learning. Graph-based based methods have been successful in solving many types of semi-supervised learning problems by optimizing a graph smoothness criterion. This criterion states that data instances nearby in a given graph are likely to have similar properties. A graph smoothness criterion cannot be directly incorporated into a generative unsupervised model because it is usually not clear what probabilistic process generated the data instances with respect to the graph, and incorporating the graph directly into a factorizable (i.e. time-series) model would break the model's factorizable structure, making exact inference methods like belief propagation intractable. This method, called entropic graph-based posterior regularization (EGPR) provides a way to express a graph smoothness criterion in a probabilistic model by defining a regularization term on an auxiliary posterior distribution variable. We applied this approach to regulatory genomics data sets from the human genome, leading to the discovery of a new type of regulatory domain.

 

Note: the material in this talk will be distinct from my 2017-09-21 talk at VanBUG. 

 

Biography:  Maxwell Libbrecht is an Assistant Professor in Computing Science at Simon Fraser University. He received his PhD in 2016 from the Computer Science and Engineering department at University of Washington, advised by Bill Noble and Jeff Bilmes. He received his undergraduate degree in Computer Science from Stanford University, where he did research with Serafim Batzoglou. His research focuses on developing machine learning methods applied to high-throughput genomics data sets. He was the first author of a paper named one of ISCB's Top 10 Regulatory and Systems Genomics papers of 2015.

 

UBC Statistics Seminar on Tuesday, September 26: An Easy-to-implement Variable Selection Method for Models Following Heredity

Speaker's Page:  https://carlsonschool.umn.edu/faculty/william-li

Abstract:  In many practical regression problems, it is desirable to select important variables with heredity constraint satisfied. In other words, when an interaction term is selected, it is preferred to select all the corresponding ain effects as well. In this paper, we propose a general strategy to maintain heredity in variable selection through a novel heredity-induced data standardization. After the standardization, any variable selection method (including stepwise selection, lasso, SCAD and others) can be applied and the selected model is automatically guaranteed to satisfy the heredity constraint. Furthermore, the same procedure works for all types of regression including linear regression, generalized linear regression and regression with censored outcome. Therefore, our proposed strategy is easy (almost effortless) to implement in practice to maintain the heredity. Simulations and real examples are used to illustrate the merits of the proposed methods.

 

UBC Statistics Seminar on Tue, Oct 3 at 11am (ESB 4192): A statistical view of uncertainty quantification: interface between statistics and applied mathematics

Speaker's Page:  https://www.isye.gatech.edu/users/jeff-wu.

C. F. Jeff Wu is Professor and Coca-Cola Chair in Engineering Statistics at the School of Industrial and Systems Engineering, Georgia Institute of Technology.

He was elected a Member of the National Academy of Engineering (2004), and a Member (Academician) of Academia Sinica (2000). A Fellow of the Institute of Mathematical Statistics (1984), the American Statistical Association (1985), the American Society for Quality (2002), and the Institute for Operations Research and Management Sciences (2009). He received the COPSS (Committee of Presidents of Statistical Societies) Presidents' Award in 1987, which was given to the best researcher under the age of 40 per year and was commissioned by five statistical societies. His other major awards include the 2011 COPSS Fisher Lecture, the 2012 Deming Lecture (plenary lectures during the annual Joint Statistical Meetings), the Shewhart Medal (2008) from ASQ, and the Pan Wenyuan Technology Award (2008). In 2016 he received the (inaugural) Akaike Memorial Lecture Award.  He has won numerous other awards, including the Wilcoxon Prize, the Brumbaugh Award, the Jack Youden Prize (twice), and the Honoree of the 2008 Quality and Productivity Research Conference. He was the 1998 P. C. Mahalanobis Memorial Lecturer at the Indian Statistical Institutes and an Einstein Visiting Professor at the Chinese Academy of Sciences (CAS). He is an Honorary Professor at several institutions, including the CAS and National Tsinghua University. He received an honorary doctor (honoris causa) of mathematics at the University of Waterloo in 2008.

***

Abstract: 

Because of the advances in complex mathematical models and fast computer codes, computer experiments have become popular in engineering and scientific investigations. Statisticians have worked on the design, modeling and computation aspects of computer experiments. Applied mathematicians have approached a closely related class of problem called UQ (uncertainty quantification). Interface between the two approached is made in the talk. Two problems on the statistical side are presented to illustrate this interface.

1. Consider deterministic computer experiments with tuning parameters which determine the accuracy of the numerical algorithm (e.g., mesh density in finite element analysis). To efficiently integrate computer outputs with different tuning parameters, a class of nonstationary Gaussian process models consistent with the knowledge in numerical analysis is proposed to model the integrated output. Estimation is performed by using Bayesian computation. Numerical studies show the advantages of the proposed method over existing methods. A related problem is given to illustrate the interplay between modeling and design. For this and a broader class of models with multi-levels of fidelity, the nested space-filling designs are most suitable. Some examples are given and the underlying mathematics discussed.

2. Calibration parameters in deterministic computer experiments are those attributes that cannot be measured or available in physical experiments or observations. Kennedy-O’Hagan (2001) suggested an approach to estimation by using data from physical experiments and computer simulations. We show that a simplified version of the original KO method leads to asymptotically inconsistent calibration. This calibration inconsistency can be remedied by modifying the original estimation procedure. A novel calibration method, called the L2 calibration, is proposed and proven to be consistent and enjoys optimal convergence rate. A numerical example and some mathematical analysis are used to illustrate the source of inconsistency.

Two UBC Statistics Master's Students' Presentations - Tue, Aug 22 at 11am (ESB 4192)

11am - 11:30am:  Qiong Zhang

Title:  Small Area Quantile Estimation under Unit-Level Models

Abstract:  Sample surveys are widely used as a cost-effective way to collect information on variables of interest in target populations. In applications, we are generally interested in parameters such as population means, totals, and quantiles. Similar parameters for subpopulations or areas, formed by geographic areas and socio-demographic groups, are also of interest in applications. However, the sample size might be small or even zero in subpopulations due to the probability sampling and the budget limitation. There has been intensive research on how to produce reliable estimates for characteristics of interest for subpopulations for which the sample size is small or even zero. We call this line of research Small Area Estimation (SAE). 

In this talk, I present the work from my Master's thesis, which studies a number of unit-level model-based small area quantile estimators. Since the model-based estimates can be misleading and their mean squared errors can be underestimated if the model assumption is wrong, simulation studies have been conducted to investigate the performance of three small area quantile estimators in the literature. They are found not to be very robust in some likely situations. Based on the observations, a few attempts have been made to obtain more robust small area quantile estimators. My talk will cover: (1) a brief introduction to small area estimation, (2) motivation for our new approach, and (3) simulation study results. 

*************************************************************************

 

11:30am - 12:00pm:  Yiwei Hou

Title:  Risk Region Estimation for Light-tailed Multivariate Samples

Abstract:  Accurate assessments for the probabilities of extreme events in multivariate cases are of great importance in various applications. Assuming the data can be described by a multivariate probability density, we can define risk regions as the multivariate quantile regions that correspond to very small probabilities.  In applications, these risk regions can serve as  multivariate stress test scenarios in financial risk management or be used for flagging events of extreme aviation risk when assessing airline performances.  However, estimation for such risk regions is difficult since the probability level  may be so small that there is hardly any or no data in these regions. There is an ongoing development of sophisticated statistical methods to estimate multivariate risk regions under different assumptions in the literature.  We investigate the problem of risk region estimation for a particular class of distributions that have homothetic level sets of non-specified shape. Such distributions generalize the family of elliptical distributions, moving away from the elliptical symmetry. For this class of distributions, we propose a new inference framework that allows flexible model assumptions and assess its performance through simulation studies.

In this talk, I present the work from my Master's thesis, which discusses this new estimation method. My talk will cover 1) the motivation for the new method, 2) details about this method and 3) simulation results and data examples.

 

The application of Machine Learning to risk/need assessment instrument LS/CMI in the prediction of offender recidivism

Research Summary: The Level of Service/Case Management Inventory (LS/CMI) is an offender risk/need assessment inventory that is designed for use with adult offenders who are either in custody or are serving their sentence in the community. Its items are scored in a dichotomous manner and summed to generate a score that correlates moderately with offender recidivism. In this study, we apply machine learning algorithms to gain more predictive accuracy for recidivism. The optimized methods for training the machine learning algorithms in this article are drawing uniform samples. Among all algorithms studied, decision trees perform significantly better than others. Policy Implications: The combination of a strong criminological theory-driven instrument, like the LS/CMI, with machine learning algorithms and clustering offenders according to their LS/CMI-score result in a significant improvement in the prediction of recidivism. Correctional agencies should explore the use of these techniques to improve the predictive validity of the LSI/CMI or similar instruments with their own offender population.

Estimation in Functional Linear Quantile Regression - FRIDAY, August 18th at 11am in ESB 4192

We consider the estimation in functional linear quantile regression in which the dependent variable is scalar while the covariate is a function, and the conditional quantile for each fixed quantile index is modeled as a linear functional of the covariate. There are two common approaches for modeling the conditional mean as a linear functional of the covariate. One is to use the functional principal components of the covariates as basis to represent the functional covariate effect. The other one is to extend the partial least square to model the functional effect. The former belongs to unsupervised method and has been generalized to functional linear quantile regression. The latter is a supervised method and is superior to the unsupervised PCA method. In this talk, we propose to use partial quantile regression and its tensor approximation to estimate the functional effect in functional linear quantile regression. Asymptotic properties have been studied and show the virtue of our method in large sample. Simulation study is conducted to compare it with existing methods. Real data examples are analyzed and some interesting findings are discovered.