Seminar

Modelling under-reported data through INAR-hidden Markov chains - June 22, 2017 at 4pm in ESB 4192

The interest in the analysis of count time series has been growing in the past years, and many models have been considered in the literature (Al-Osh and Alzaid, 1987, J Time Series Analysis). The main reason for this increasing popularity is the limited performance of the classical time series analysis approach when dealing with discrete valued time series. With the introduction of discrete time series analysis techniques, several challenges appeared such as unobserved heterogeneity, periodicity, under-reporting, among others. Many efforts have been devoted in order to introduce seasonality in these models (Morina et al., 2011, Statistics in Medicine) and also coping with unobserved heterogeneity. However, the problem of under-reported data is still in a quite early stage of study in many different fields. This phenomenon is very common in many contexts such as epidemiological and biomedical research. It might lead to potentially biased inferences and may also invalidate the main assumptions of the classical models. Especially, in public health context it is well known that several diseases have been traditionally under-reported (occupational related diseases, food exposures diseases, ...). The model we will present in this work considers two discrete time series: the observed series of counts Y_t which may be under-reported, and the underlying series X_t with an INAR(1) structure X_n = alpha*X_{n-1}+W_n, where 0 < alpha < 1 is a fixed parameter and W_n are the innovations which are Poisson(lambda) distributed. The binomial thinning operator (or binomial subsampling) is defined as alpha*X_{n-1}= \sum_{i=1}^{X_{n-1}} Z_i(alpha); where Z_i are i.i.d Bernoulli random variables with probability of success equal to alpha. The way we allow the observed process Y_n to be under-reported is by defining that Y_n is X_n with probability 1-omega  or is q*X_n with probability omega. Obviously, this definition means that the observed Y_n coincides with the underlying series X_n, and therefore the count at time n is not under-reported with probability 1-omega. Several applications in the field of public health will be discussed, using real data regarding incidence and mortality attributable to diseases related to occupational and environmental exposures and known toxics and traditionally under-reported. Full details of the work can be found in Fernandez-Fontelo et al. (2016, Statistics in Medicine, v 35, pp 4875-4890).

 

2 UBC Statistics MSc Students on Tue, May 16 at 11am (ESB 4192)

11am - 11:30am

Speaker:  Md Rashedul (Rashed) Hoque

Title:  Approximation of the Formal Bayesian Model Comparison using the Extended Conditional Predictive Ordinate Criterion

Abstract:  The optimal method for Bayesian model comparison is the formal Bayes factor (BF), according to decision theory. The formal BF is computationally troublesome for more complex models. If predictive distributions under the competing models do not have a closed form, a cross-validation idea, called the conditional predictive ordinate (CPO) criterion, can be used. In the cross-validation sense, this is a “leave-out one” approach. CPO can be calculated directly from the Monte Carlo (MC) outputs, and the resulting Bayesian model comparison is called the pseudo Bayes factor (PBF). We can get closer to the formal Bayesian model comparison by increasing the “leave-out size”, and at “leave-out all” we recover the formal BF. But, the MC error increases with increasing `leave-out size'. In this study, we examine this for linear and logistic regression models.

Our study reveals that the Bayesian model comparison can favor a different model for PBF compared to BF when comparing two close linear models. So, larger “leave-out sizes” are preferred which provide result close to the optimal BF. On the other hand, MC samples based formal Bayesian model comparisons are computed with more MC error for increasing “leave-out sizes”; this is observed by comparing with the available closed form results. Still, considering a reasonable error, we can use “leave-out size” more than one instead of fixing it at one. These findings can be extended to logistic models where closed form solution is unavailable.

11:30am - 12:00pm

Speaker:  Derek Cho

Title:  Prediction of Indian reserve populations in Canada using administrative data sources

Abstract:  Statistics Canada has recently been interested in using administrative data to construct models for predictive purposes. One such application is to use data from the Indian Register to predict populations of Indian reserves where Census collection is not possible for the 2016 Census of Population. Using a naïve robust mixed effects model, we can predict the populations of non-enumerated Indian reserves. In addition, I will also go over some of my other work at Statistics Canada last summer including work done with the Immigration Database and Census coding.

Estimating Dependence Between Lumber Strength Properties with Copula Models

We propose a copula-based approach to estimate the dependence between a pair of lumber strength properties that cannot be observed simultaneously, i.e., MOR and UTS. We also present a graphical method to evaluate whether a model can accurately represent the damage effect due to proof load. By fitting copula models, we conclude that there is a strong dependence between MOR and UTS even conditioning on other related properties such as MOE. We also find that the reflected Gumbel copula is a more suitable model for the lumber data compared with the Gaussian copula.

2 UBC Statistics Students' MSc Presentations

11:00am - 11:30am:  Yijun Xie

Title:  A Flexible Inference Method for an Autoregressive Stochastic Volatility Model with an Application to Risk Management

Abstract:  Autoregressive Stochastic Volatility (ARSV) model is a discrete-time process that can model financial returns. Existing inference methods can only be applied to the classic ARSV model. In this talk, I present the work from my master's thesis, which discusses a new inference method that allows flexible model assumptions for the ARSV model. I also present an application to risk management and compare the ARSV model with another commonly used model for financial time series, namely the GARCH model. My talk will cover the following: (1) motivation for an extension of the classic ARSV model, (2) details about the new inference method, and (3) examples of using ARSV model in risk management.

************

11:30am - 12:00pm:  Jonathan Agyeman

Title:  On the choice of scoring functions for forecast comparisons

Abstract:  Forecasting of risk measures is an important part of risk management for financial institutions.  Value-at-Risk and Expected Shortfall are two commonly used risk measures and accurately predicting these risk measures enables financial institutions to plan adequately for possible losses. Point forecasts from different methods can be compared using consistent scoring functions, provided the underlying functional to be forecasted is elicitable. It has been shown that the choice of a scoring function from the family of consistent scoring functions does not influence the ranking of forecasting methods as long as the underlying model is correctly specified and nested information sets are used. However, in practice, these conditions do not hold, which may lead to discrepancies in the ranking of methods under different scoring functions.

We investigate the choice of scoring functions in the face of model misspecification, parameter estimation error and nonnested information sets. We concentrate on the family of homogeneous consistent scoring functions for Value-at-Risk and the pair of Value-at-Risk and Expected Shortfall and identify conditions required for existence of the expectation of these scoring functions. We also assess the finite-sample properties of the Diebold-Mario Test, as well as examine how these scoring functions penalize for over-prediction and under-prediction with the aid of simulation studies.

2 UBC Statistics Co-op Students' MSc Presentations

4:00pm - 4:30pm:  Annie Li

Title:  The Use of Genome-Wide Association and Mendelian Randomization in Epidemiology Studies

Abstract:  In observational epidemiology studies, one problem is the difficulty of identifying causal associations between modifiable traits and diseases. With the development of high-throughput genotyping technologies, Genome-wide association approach provides a powerful tool for identifying genetic variants associated with clinical conditions and traits. The identified genetic variants can then be used as a proxy in a Mendelian randomization framework for analysing causal associations. This talk is about my experience with genome-wide association studies and Mendelian randomization during my CO-OP work terms in the Centre for Heart and Lung Innovation, UBC and St. Paul’s Hospital.

*****************************

4:30pm - 5:00pm:  Ming Wan

Title:  Genotype Imputation and Association Testing in Genome-Wide Association Studies

Abstract:  In the past decade, Genome-Wide Association Studies (GWAS) have greatly advanced our understanding of the genetic bases on complex diseases. In this presentation, I will talk about the statistical methods applied in two GWA studies from the Denise Daley lab, St. Paul’s hospital. In particular, 1) The use of modified principal component analysis to better control for population stratification in the presence of related samples; 2) Imputation of ungenotyped allele variants using Hidden Markov Models to boost study power. Additionally, I will also touch on the computational issues that arise from large sample size association studies.

2 UBC Statistics Co-op Students' MSc Presentations

*Please note the corrected times for these 2 seminars.*

4:00pm - 4:30pm:  Jie Cui

Title:  My Co-op Experience at Scotiabank

Abstract:  Cross-selling refers to the practice of selling related or complimentary product/service to an already existing customer. In the face of a hyper-competitive retail banking market, banks tend to shift focus to cross-selling since the cost of acquiring a new customer is much higher than that of retaining an existing one. To thrive and succeed in the digital age, banks are relying more and more on data-driven campaign strategies. One project of my co-op placement was to build a predictive model for cross-selling among certain lines & loan customers. During this seminar, I will briefly talk about projects I was involved in and working experience at Scotiabank. 

4:30pm - 5:00pm:  Mary He

Title:  Working with High-dimensional Genomics Data at PROOF Center of Excellence

Abstract:  With the rapid development of sequencing technology, fast and affordable sequencing of large amount of nucleotides has become possible, producing expression data for tens of thousands of genes easily. However, in contrast to the large number of features measured, clinical experiments usually involve only a few samples, making the detection of differences between treatment groups or phenotypes a challenging task. In this seminar, I will briefly talk about the statistical tools I learned for differential expression analysis and biomarker discovery in the high-p-low-n settings, and share with you a few relevant projects that I was involved in.

An Introduction to Claims Development and Basic Reserving Techniques

*Please note the unusual location of this talk.*

Actuarial reserving is the practice of estimating the current value of all future liabilities an insurer has taken on through its policies, and is necessary in order to set aside an adequate reserve. This estimation process is challenging since it may take several years for all claims in a given policy period to be reported and closed. In addition, information pertaining to existing claims is often altered long after the end of a policy period. 

In order to develop an estimate of reserves to set aside, actuaries evaluate the development of reported and paid claims over time. In this seminar, we will first delve into the claims process, following a claim from its first report to the insurer to its ultimate settlement. Next we will introduce the concept of a loss development triangle, a way of displaying how a group of claims evolve over time. The development triangle is one of the most common tools used by actuaries to evaluate claims. We will then explore some basic techniques for estimating an insurance company’s unpaid claims by going through a detailed case study for an event insurance company with two lines of insurance, medical insurance and indemnity.

This talk is based on a case studies project that won first prize in the Actuarial Students National Association Case Competition 2017.

Statistical Methods for Big Tracking Data: An Application to Marine Mammal Tracking

Recent advances in technology have led to large sets of tracking data, which brings new challenges in statistical modeling and prediction. Built on recent developments in Gaussian process modeling for spatio-temporal data and stochastic differential equations (SDEs), we develop a sequence of new models and corresponding inferential methods to meet these challenges. We first propose Bayesian Melding (BM) and downscaling frameworks to combine observations from different sources. To use BM for big tracking data, we exploit the properties of the processes along with approximations to the likelihood to break a high dimensional problem into a series of lower dimensional problems. To implement the downscaling approach, the integrated nested Laplace approximation (INLA) is applied to fit a linear mixed effect model that connects the two sources of observations. We apply these two approaches in a case study involving the tracking of marine mammals. Both of our frameworks are shown to have superior predictive performance compared with traditional approaches in both cross-validation and simulation studies.

We further develop the BM frameworks with stochastic processes that can reflect the time varying features of the tracks. We propose a linear SDE with splines as its coefficients and call it a generalized Ornstein-Ulhenbeck (GOU) process. The GOU achieves flexible modeling of the tracks in both mean and covariance with a reasonably parsimonious parameterization. Its inferences and predictions can be computed via the Kalman filter and smoother. BM with the GOU achieves a smaller prediction error and better credibility intervals in cross-validation comparisons to the basic BM and downscaling models. Following the success with the GOU, we further study a special class of SDEs called the potential field (PF) models, which formulates the drift term as the gradient of another function. A mixture of PF model is shown to provide good interpretation of the animal’s track.

Dimension reduction using Independent Component Analysis with an application to a business psychology dataset

In my presentation I will review the methodology of Independent Component Analysis (ICA) and propose a method of dimension reduction using ICA. ICA is used for separating mixed signals into statistically independent additive subcomponents. The methodology extracts as many independent components as there are dimensions or features in the original dataset. Since not all these components may be of importance, a few solutions have been proposed to reduce the dimension of the data using ICA, most of which rely on prior knowledge or estimation of the number of independent components that are to be used in the model. In my talk I will discuss a methodology that addresses the problem of dimension reduction that best approximates the original dataset without the prior knowledge or estimation of the number of components to be retained.

Adaptive MCMC For Everyone

 

*A small pre-talk reception will be served just outside the venue at 3:30pm.*

 

2016-17 Constance van Eeden Lecture

Abstract: Markov chain Monte Carlo (MCMC) algorithms, such as the Metropolis Algorithm and the Gibbs Sampler, are an extremely useful and popular method of approximately sampling from complicated probability distributions. Adaptive MCMC attempts to automatically modify the algorithm while it runs, to improve its performance on the fly. However, such adaptation often destroys the ergodicity properties necessary for the algorithm to be valid. In this talk, we first illustrate MCMC algorithms using simple graphical Java applets. We then discuss adaptive MCMC, and present examples and theorems concerning its ergodicity and efficiency. We close with some recent ideas which make adaptive MCMC more widely applicable in broader contexts.

Speaker's Bio: Jeffrey Rosenthal is a professor in the Department of Statistics at the University of Toronto. He received his BSc from the University of Toronto at the age of 20, his PhD in Mathematics from Harvard University at the age of 24, and tenure at the University of Toronto at the age of 29. He received the 2006 CRM-SSC Prize, the 2007 COPSS Presidents' Award, the 2013 SSC Gold Medal, and teaching awards at both Harvard and Toronto. He is a fellow of the Institute of Mathematical Statistics and of the Royal Society of Canada. Rosenthal's book for the general public, Struck by Lightning: The Curious World of Probabilities, was published in sixteen editions and ten languages, and was a bestseller in Canada, leading to numerous media and public appearances, and to his work exposing the Ontario lottery retailer scandal. He has also dabbled as a computer game programmer, musical performer, and improvisational comedy performer, and is fluent in French. His web site is www.probability.ca.