Seminar

Modeling Operational Risk with Truncated Samples

Operational risk for a bank is the risk arising from execution of their business functions. The increasing sophistication of derivative securities enables banks to mitigate many risks associated with lending and investment activities, leaving operational risk as a growing portion of their overall risk portfolios. Regulations require banks set aside capital to cover operational losses, called regulatory capital. We discuss the loss distribution approach for estimating regulatory capital according to the advanced modeling approach discussed in Basel II. For a given business line and event type, our approach applies simple statistical and visual tools for choosing a loss severity distribution from a list of candidate parametric distributions. Candidate distribution parameters are estimated from operational loss data that are truncated below at a known minimum reporting threshold using the maximum likelihood method. This truncated approach assumes that unobservable losses below the reporting threshold follow the same distribution as observable losses. Evaluating estimated loss severity distributions at this minimum reporting threshold produces estimates of the proportion of losses that are unobserved, which we call implied probability. We use implied probability estimates in addition to AIC, QQ-plots, and Quantile Score when selecting candidate loss severity distributions and discuss some of the challenges associated with this process. We then simulate operational losses and use our advanced modeling approach to estimate regulatory capital.

UBC Statistics M.Sc. Co-op Student Presentations (2)

4:00pm - 4:30pm

Speaker:  Hiwot Tafessu

Title:  Healthcare utilization and associated costs among people living with HIV/AIDS with HCV co-infection and mental illness in British Columbia, Canada

Abstract:  Due to the widespread use of modern combination antiretroviral therapy in high-resource countries like Canada, HIV infection has become a chronic manageable disease. As a consequence, morbidity and mortality due to non-AIDS related comorbidities have become increasingly prevalent. Hepatitis C virus (HCV) co-infection represents the most prevalent comorbidity, particularly in people who inject drugs. It has also been shown that the majority of healthcare utilization among this population is due to non-AIDS related conditions, including mental illness. This presentation will be about my co-op experience at the BC Centre for Excellence in HIV/AIDS, where I analyzed trends in hospitalizations (2000-2014) and associated costs in this population group from the British Columbia Seek and Treat for Optimal Prevention of HIV/AIDS cohort.

*****

4:30pm - 5:00pm

Speaker:  Lisa Leung

Title:  Biostatistician Co-op Experience at Heart + Lung Institute in St. Paul’s Hospital

Abstract:  Chronic Obstructive Pulmonary Disease (COPD) is the third leading cause of death worldwide. It is characterized by reduced lung function measures as a person ages. Besides smoking, COPD is associated with a few factors in which genetic factors are being studied as causal risks. For this talk, I will briefly introduce the genomic studies and analyses of COPD at the Heart + Lung Institute in St. Paul’s Hospital, and talk about my experiences with Dr. Ma’en Obeidat's analysis team.

Modelling Microbial Data with Time, Treatment and/or Space Variation

This talk is motivated by issues arising from microbial oceanic data that biological researchers have been collecting to understand variation in response to environmental changes. The data typically consists of counts of OTUs (operational taxonomic units) from ocean samples varying either in time, treatment and/or space. At any given observation there will typically be at least 1000 OTUs of potential interest. In particular I will describe some data arising from an experimental study of microbial organisms in the oceans to assess the effect of enhanced carbon loading. I will indicate briefly our approach to modelling the data to take into account both the experimental and temporal variation. In our view, an essential first step is to carry out dimension reduction via clustering based on the results of Poisson generalized linear models. Then we can carry out the tests for any significant experimental effect. In this data, it is likely that only a small subset of the OTUs may show a response to the carbon loading. During the talk, it will be clear that there are still some unresolved statistical issues on which we welcome feedback. I should note that although the data I have is from the ocean, similar issues will arise with microbial data coming from many other environments such as the human gut.

Modelling Malaria in India: Statistical, Mathematical and Graphical approaches.

Malaria has existed in India since antiquity. Different periods of elimination and control policies have been adopted by the government for tackling the disease. Malaria parasite was discovered in India by Sir Ronald Ross who also developed the simplest mathematical model in early 1900. Malaria modelling has since come through many variations that incorporated various intrinsic and extrinsic/environmental factors to describe the disease progression in population. Collection of disease incidence and prevalence data, however, has been quite variable with both governmental and non-governmental agencies independently collecting data at different space and time scales. In this talk I will describe our work on modelling malaria prevalence using three different approaches. For monthly prevalence data, I will discuss (i) a regression-based statistical model based on a specific data-set, and (ii) a general mathematical model that fits the same data. For more coarse-grained temporal (yearly) data, I will show graphical analysis that uncovers some useful information from the mass of data tables. This presentation aims to highlight the suitability of multiple modelling methods for disease prevalence from variable quality data.

Development and validation of a clinically applicable RNA-based molecular classifier of HGSC ovarian carcinoma

High-grade serous carcinomas (HGSCs) account for approximately 70% of all epithelial ovarian cancers and patients presenting with these tumours have the lowest survival rates. Although HGSC tumours appear morphologically similar, they constitute different molecular subtypes. This presentation outlines my co-op project, which proposes a three-stage model development and validation approach to diagnose patients into these subtypes across gene expression platforms. These model development stages include subtype discovery, which is followed by model training and external validation using data on a similar and different gene expression platform. Finally, a chosen classifier will be used to assign HGSC molecular subtype classes to all available data using clinically applicable gene expression methods. 

UBC Statistics M.Sc. Co-op Student Presentations (2)

4:00pm - 4:30pm:  Shanshan Pi

Title:  Nice Co-op Experience in Pediatric Anesthesia Research Team at BCCHR

Abstract:  I will first give a brief introduction of the Pediatric Anesthesia Research Team at BCCHR and then will talk about my two projects during my Co-op time: (1). Effectiveness of the NeuroSENSE for Monitoring the Hypnotic Depth of Anesthesia. (2). Integrating intraoperative physiology data into outcomes analysis for ACS national surgical quality improvement program.

***

4:30pm - 5:00pm:  Weining Hu

Title:   Statistical application for software engineering job in developing Alexa

Abstract:  In recent years, virtual assistants, such as Siri, Allo, and Alexa, have attracted a lot of attention, and many of these virtual assistants have been enabled in devices such as smart speakers, smartphones, and even cars. The popularity of virtual assistants has led to interesting problems in both research and engineering. In this presentation, I will give a brief overview of my co-op work as a software engineer intern with the Amazon Alexa Communications team in Toronto. In particular, I will present the infrastructure of Alexa; introduce the team I was working with and the products we released; and share with you some work that will require statistics knowledge from an engineer's perspective.

Nonparametrics without tuning parameters: shape-constraints and mixture models

The talk is a glimpse back at the line of work that led from the research on shape-constrained statistical inference to opening new perspectives in mixture models, with outcomes in deconvolution and empirical Bayes prediction. Particular personal topics include s-unimodality in density estimation, shape-constrained aspects of empirical  Bayes prediction, and Kiefer-Wolfowitz nonparametric estimators of mixing distributions. Those unfamiliar with the subject may discover some data-analytic stimuli, new methods illustrated on the presented examples; those well in touch may recognize some recent developments. The unifying theme, and thus of an interest by itself, is the role played by the modern convex optimization methodology, both from the algorithmic and theoretical point of view.

2017-18 Constance van Eeden Lecture: Reproducibility of science: p-values, multiple testing and optional stopping

*A small pre-talk reception will be served just outside the venue in AERL at 4:30pm.*

Abstract: Three of the statistical causes for the lack of reproducibility of science will be discussed, along with a suggested cure. The first cause is the common misinterpretation of p-values; the second is the frequent lack of sufficient adjustment for multiple testing; the third is the common ignoring of optional stopping when testing. The suggested cure for all three is to use odds of hypotheses as the basic inference tool. Surprisingly, this can be done in a way that simultaneously accommodates both frequentist and Bayesian reasoning.

Speaker's Bio: Jim Berger has made fundamental contributions to the foundations of statistics and to statistical decision theory. He is one of the world’s leading figures in Bayesian statistics and has made seminal contributions to the areas of model selection, multiple inference, computer modeling and simulation. For his work, he has received many honours, including the COPSS Presidents' Award, a Guggenheim Fellowship, the R. A. Fisher Lectureship, and the Wilks Memorial Award from the ASA. He has supervised 36 Ph.D. dissertations, published over 190 papers, written or edited 16 books or special volumes, and is a founding editor of the Journal on Uncertainty Quantification. Berger has also made substantial interdisciplinary contributions in the fields of  astronomy, geophysics, medicine and the validation of complex computer models.

---------

This talk is supported by the van Eeden fund, the Department of Statistics, and PIMS.

Probabilistic models of splicing (dys)regulation in human disease

Abstract:  Splicing, the cellular process by which "junk" intronic regions are removed from precursor messenger RNA, is tightly regulated in healthy human development but frequently dysregulated in disease.  Massively parallel sequencing of RNA (RNA-seq) has become a ubiquitous technology in biology to assay the resulting “transcriptome”: the collection of messenger RNA molecules expressed from the genes of an organism. However, significant computational and statistical challenges remain to translate the resulting noisy, confounded RNA-seq data into meaningful understanding of the biological system or disease state under consideration. I will describe our use of probabilistic models to address such challenges: a novel approach to quantifying alternative splicing across different tissues/diseases and a neural-network model that predicts splicing from DNA sequence, improving interpretation of rare variants from exome or whole-genome sequencing studies.

Speaker's Biography:  Dr. Knowles studied Natural Sciences and Information Engineering at the University of Cambridge before obtaining an MSc in Bioinformatics and Systems Biology at Imperial College London. During his PhD studies in the Cambridge University Engineering Department Machine Learning Group under Zoubin Ghahramani he worked on Bayesian nonparametric models for factor analysis, hierarchical clusterings and network analysis, as well as on (stochastic) variational inference. He is currently a post-doctoral researcher at Stanford University with Sylvia Plevritis (Center for Computational Systems Biology/Radiology) and Jonathan Pritchard (Genetics/Biology) having previously worked with Daphne Koller (Computer Science). His work involves the application of statistical machine learning in functional genomics, with the occasional foray into imaging of biological systems. As of 2017 he is an O-1 Alien of Extraordinary Ability and has a T-shirt to prove it.

Design and analysis of computer experiments: assessing and advancing the stat of the art

Computer experiments have been widely used in practice as important supplements to traditional physical experiments in studying complex processes. However, a computer experiment is expensive in terms of its computational time. Hence, Gaussian process (GP) was proposed to be used as a statistical surrogate. The scope of the research is rather broad: we are concerned with design and analysis of computer experiments based on a GP. We use comprehensive assessment strategies to evaluate the effect of several factors on the prediction accuracy of the GP model. In addition, we propose new methods motivated by the assessment. Working with an engineering computer model, we also provide insights into issues faced by practitioners.