Seminar

Making a Training Dataset from Multiple Data Distributions

Over time we might accumulate lots of data from several different populations: e.g., the spread of a virus across different countries. Yet what we wish to model is not any one of these populations. One might want a model for the spread of the virus that is robust to the different countries, or is predictive on a new location we have only limited data for. We overview and formalize the objectives these present for mixing different distributions to make a training dataset, which have historically been hard to optimize. We show that by assuming we train models near "optimal" for our training distribution these objectives simplify to convex objectives, and provide methods to optimize these reduced objectives. Experimental results show improvements across language modeling, bio-assays, and census data tasks.

To join this seminar virtually, please request Zoom connection details from ea@stat.ubc.ca. 

Some new advances in precision medicine modeling

Earlier work has shown that similarity-based predictive models can improve upon predictive performance, as compared to using the entire training data to help build models, particular regarding model discrimination for binary responses. My collaborators and I have some updated results to share, regarding similarity-based modeling for joint consideration of model calibration and discrimination, as well as for dynamic prediction models. In addition, we have been developing transfer learning methods for targeted prediction. Properties of our methods will be investigated in comprehensive simulation studies, and we will demonstrate the methods through separate analyses of a publicly-available intensive care unit (ICU) database.

Collaborators: Keeley lsinghood, Minzee Kim, Tatiana Krikella, Subha Maity, Mengqi Xu Department of Statistics and Actuarial Science, and School of Public Health Sciences, University of Waterloo; Department of Statistical Sciences, University of Toronto

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca.

Event Photo
Professor Joel Dubin

Classification and diffusion-induced neural density estimators and simulators for generative AI

Neural network-based methods for conditional density estimation have recently gained substantial attention, as various neural density estimators have outperformed classical approaches in real-data experiments. Despite these empirical successes, implementation can be challenging due to the need to ensure non-negativity and unit-mass constraints, and theoretical understanding remains limited. In particular, it is unclear whether such estimators can adaptively achieve faster convergence rates when the underlying density exhibits a low-dimensional structure. This paper addresses these gaps by proposing a structure-agnostic neural density estimator, called the classification-induced neural density estimator and simulator (CINDES) that is straightforward to implement and provably adaptive, attaining faster rates when the true density admits a low-dimensional composition structure. Another key contribution of our work is to show that the proposed mator integrates naturally into generative sampling pipelines, most notably score-based diffusion models, where it achieves provably faster convergence when the underlying density is structured. We validate its performance through extensive simulations and a real-data application. We also prove the optimality of score-based diffusion models for density estimation when the target density admits a factorizable, low-dimensional, nonparametric structure in a separate work. The main challenge is that the low-dimensional, factorizable structure no longer holds for most diffused timesteps, and it is very difficult to show that these diffused score functions can be well approximated without a significant increase in the number of network parameters.

(Join works with Yihong Gu, Dehao Dai, Mukherjee, and Ximing Li)

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca.

 

Professor Jianqing Fan, a member of the US National Academy of Sciences, the Royal Academy of Belgium, and Academia Sinica, is the Frederick L. Moore’18 Professor of Finance, Professor of Operations Research and Financial Engineering, and Former Chairman of the Department of Operations Research and Financial Engineering at Princeton University, where he directs both Financial Econometrics and Statistics and Data Science labs. He received his Ph.D. from the University of California at Berkeley and held faculty positions at the University of North Carolina at Chapel Hill, University of California at Los Angeles, and the Chinese University of Hong Kong before joining Princeton University. He is a fellow of the American Association for the Advancement of Science, the Institute of Mathematical Statistics, the American Statistical Association, and the Society of Financial Econometrics. He has served as the president of the Institute of Mathematical Statistics and the International Chinese Statistical Association, and has been a joint editor of the Journal of the American Statistical Association, Annals of Statistics, Probability Theory and Related Fields, Econometrics Journal, Journal of Econometrics, Journal of Business and Economics Statistics, and Management Science (Finance Department editor). Awards include the COPSS Presidents' Award, the Morningside Gold Medal of Applied Mathematics, the Guggenheim Fellowship, the P.L. Hsu Prize, the Royal Statistical Society Guy medal in silver, the Noether Distinguished Scholar Award, Le Cam Award and Lecture, the Frontiers of Science Award, and the Wald Award and Lecture. His research interests include high-dimensional statistics, data science, machine learning, deep learning, mathematics of AI, financial economics, and computational biology. He coauthored 4 books and published over 300 highly cited papers, with over 100,000 citations.

Event Photo
Professor Jianqing Fan

Efficient Estimation and Closed-Form Uncertainty Quantification for Net Benefit of Algorithms that Predict Individualized Treatment Benefit

Treatment benefit predictors (TBPs) quantify the expected treatment benefit given individual characteristics. The net benefit function evaluates the expected gain in clinical utility from using a TBP to guide treatment decisions, relative to the default decisions of treating no one and treating everyone. The existing estimator for the net benefit of a given TBP implicitly assumes a 1:1 randomization design in the randomized controlled trial data. When this assumption is violated, the estimator can exhibit biased and unstable finite-sample behaviour. We identify the source of this implicit design restriction and propose a corrected congruent-based modification of the net benefit estimator that remains valid under arbitrary randomization schemes. We further introduce an alternative net benefit formulation based on the average treatment effect among individuals recommended for treatment by a TBP (ATT-based). The ATT-based estimator is asymptotically equivalent to the corrected congruent-based estimator and makes more efficient use of the full sample by estimating the proportion recommended for treatment using all individuals. Three uncertainty quantification procedures are developed for the ATT-based estimator: a large-sample variance approximation, an aggregated nonparametric bootstrap, and a Bayesian analysis. A Monte Carlo simulation study evaluates point estimation of net benefit curves using mean squared error, while uncertainty quantification is assessed through confidence interval length and coverage probabilities for 95% intervals constructed using asymptotic, bootstrap, and Bayesian methods. Simulation studies show that the corrected congruent-based estimator removes the bias under unequal randomization, while the ATT-based estimator demonstrated improved finite-sample efficiency. An empirical application to the GUSTO randomized controlled trial illustrates the ATT-based estimator and accompanying uncertainty intervals in a real-world setting. Overall, the proposed methodology enables valid estimation and uncertainty quantification of net benefit for treatment benefit predictors under arbitrary randomization schemes.

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca. 

Event Photo
Sasha Sharma

Topics in Trend Filtering with Poisson Loss

Many types of data come as counts — disease cases per day, website visits per hour, or pixel intensities in images. A common goal is to recover the smooth trend underlying these noisy counts. Trend filtering is a nonparametric method that fits flexible, piecewise polynomial curves which adapt automatically to abrupt changes in the signal without prespecifying where they occur. However, existing methods assume Gaussian noise, whereas count data follow Poisson-type models whose variability grows with the signal magnitude, making the effective noise heteroscedastic.

This dissertation develops scalable algorithms for trend filtering under Poisson loss and methodology for solving real-life applications. We propose two proximal algorithms that extend the estimator from simple time series to general graph structures. We apply this framework to epidemic surveillance, producing an R package (rtestim) that estimates time-varying reproduction numbers with principled, cross-validated tuning. We further identify and resolve a numerical instability in the linear system solvers that arise as inner subproblems of these algorithms, by recasting the system as a linear Gaussian state-space model, yielding a solver that is both stable and efficient. Finally, an ongoing work of ours shows that observation-dependent penalty weights can recover minimax optimal rates under the heteroscedastic noise inherent in exponential-family models.

Event Photo
Jiaping (Olivia) Liu

Extreme Value Theory: A Projection Estimator for the Angular Dependence Function

Extreme value theory provides a principal framework for modeling rare and extreme events, with applications in fields such as environmental science, finance, and engineering. Classical multivariate extreme value theory describes the limiting behaviour of normalised random vectors, particularly when the components are asymptotically dependent. However, many practical applications require a more flexible description of extremal dependence that can accommodate both asymptotic dependence and asymptotic independence. The Angular Dependence Function (ADF), introduced by Wadsworth and Tawn (2013), provides a flexible and interpretable way to characterize extremal dependence by describing how the rate of joint tail decay varies with the relative contribution of each component. Accurate estimation of the ADF is critical for understanding and modeling joint extremes.

While several estimators have been proposed for the ADF, they face significant limitations in finite samples, including high variability, irregular behavior across the domain, and violations of key theoretical constraints. The violations of the theoretical constraints are problematic and current approaches to addressing these issues are typically ad hoc, involving post hoc adjustments.

To resolve the issue of the violations, this work introduces a projection-based estimator for the ADF. Inspired by projection methods for the Pickands dependence function introduced in Fils-Villetard et al. (2008), the proposed method projects an initial non-parametric estimate onto a closed, convex set of admissible functions in L²([0, 1]). By construction, this estimator strictly enforces the theoretical upper and lower bounds. A simulation study across various Gaussian dependence levels demonstrates its ability to preserve validity and eliminate violations. This proposed method focuses on convex ADFs under positive quadrant dependence, reflecting the dependence structures most commonly observed in practice and other literature.

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca. 

Rolling Extrapolation of Censored Survival Data and Its Applications to Lifetime Outcome Estimation

Estimation of lifetime outcomes is a fundamental problem in biostatistics, epidemiology, and health economic evaluation. In many cohort studies, however, follow-up durations are limited and survival data are heavily censored, making direct estimation of lifetime survival, life expectancy, and cumulative disease burden impossible. Conventional approaches typically rely on parametric survival models, but long-term extrapolations are often highly sensitive to model misspecification and may produce substantial bias.

In this talk, I will present a novel statistical framework, termed the Rolling Extrapolation Algorithm (REA), for extrapolating censored survival data beyond the observed follow-up period. The method incorporates external population information through a matched reference cohort and models relative survival between the study and reference populations. A key observation is that the logit transformation of relative survival often exhibits approximate linearity under broad classes of excess hazard models. Rather than performing a single long-term extrapolation, REA fits a restricted cubic spline model and iteratively predicts one step ahead, updating the fitted model in a rolling fashion until a lifetime horizon is reached.

Simulation studies demonstrate that REA substantially improves extrapolation accuracy compared with conventional one-shot spline and parametric approaches under a variety of hazard patterns. The resulting lifetime survival estimates can be combined with longitudinal quality-of-life, disability, healthcare expenditure, and productivity data to estimate life expectancy, years of life lost, disability-adjusted life years, and lifetime economic burden. Applications will be illustrated using nationwide cohort studies in Taiwan, including analyses of long-term PM2.5 exposure and healthy lifestyle factors.

The talk will focus on the statistical principles underlying REA, its theoretical motivation, empirical performance, practical implementation, and remaining methodological challenges in survival extrapolation and lifetime outcome estimation.

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca.

Event Photo
Jing-Shiang Hwang

A novel class of mixed Poisson distributions and wastewater-based epidemiology

Mixed Poisson families are widely used to model count data with overdispersion, zero inflation, or heavy tails in a variety of applications including finance, biology, and the physical sciences. The mixing distribution assigned to the Poisson rate is typically restricted to have nonnegative support. Surprisingly, this assumption is unnecessary. For example, the Hermite distribution is analogous to mixing a Poisson with an untruncated Gaussian and can be derived using generating functions so long as constraints on the natural parameter are satisfied. I will give a general characterization of this unusual class as well as several concrete examples, including an apparently novel generalization of the discrete stable family. I will also briefly present some applied work in wastewater-based epidemiology examining spatiotemporal variation of the pepper mild mottle virus biomarker. 

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca.

Event Photo
Will Townes

Two MSc student presentations (Zhili Jiang & Zachary Lau)

Presentation 1

Time: 11:00am - 11:30am

Speaker: Zhili Jiang, UBC Statistics MSc student

Title: A Joint Model for Longitudinal and Survival Data with Nonlinear Trajectories and Interval-Censored Dropout, with Application to HIV Vaccine Studies

Abstract: Joint modeling of longitudinal biomarkers and time-to-event outcomes provides an important framework for understanding vaccine-induced immune responses and their relationship with clinical outcomes. In vaccine trials, dropout is often assumed to be non-informative and exactly observed, which may lead to biased inference when these assumptions are violated. In this study, we extend existing joint modeling approaches by incorporating a biologically motivated nonlinear mixed-effects model for longitudinal antibody trajectories and modeling dropout under both right- and interval-censored settings. The proposed framework provides a more realistic characterization of the association between immune dynamics and dropout through shared random effects. The method is applied to data from the VAX004 HIV-1 vaccine trial. The results suggest that dropout is associated with the underlying longitudinal antibody processes through shared random effects, supporting the presence of informative dropout under the proposed joint modeling framework. This association is consistently observed across Cox right-censored, Weibull right-censored, and Weibull interval-censored specifications. For the longitudinal component, the exponential-decay model provides a substantially better representation of antibody dynamics than linear and power-law alternatives. Simulation studies demonstrate reliable parameter estimation, although Hessian-based standard errors may underestimate uncertainty for parameters associated with the nonlinear component of the model. Overall, the proposed framework provides a flexible and biologically interpretable approach for joint modeling in vaccine studies, offering improved handling of realistic dropout mechanisms and the potential for extension to more complex longitudinal and survival settings. 

Presentation 2

Time: 11:30am – 12:00pm

Speaker: Zachary Lau, UBC Statistics MSc student

Title: Scalable Gaussian Processes and Active Learning for Emulator Design in Solar Wind Simulation

Abstract: In this work, we discuss emulator design for solar wind simulators. Our work focuses on two areas. Firstly, we focus on scaling Gaussian Process regression to work well on simulator grids with millions of points. We accomplish this by extending existing work on Kronecker product covariance based algorithms to work efficiently with a dataset larger than working memory. Secondly, we implement and experiment with existing acquisition functions for active learning in the large data regime found in simulators. We find encouraging, though not definitive, results in favour of the Expected Predictive Information Gain acquisition function, particularly when it targets a prior concentrated in a particular part of the search space. To the best of our knowledge, this work is the first time that Gaussian Process Regression has been applied at this scale in Solar Wind modelling, the first time that these acquisition functions have been implemented at this scale for Gaussian Process models, and the first time that active learning has been applied to the problem of emulator design for the solar wind.

To join these seminars virtually, please request Zoom connection details from hr.ops@stat.ubc.ca.

Manifold Sampling with Automatic Tuning

Many statistical and applied problems involve sampling from distributions constrained to curved lower-dimensional spaces, or manifolds. Standard MCMC methods are inapplicable in these settings because they do not naturally respect the constraint geometry, while existing manifold samplers can be highly sensitive to step-size tuning.

Our main contribution is an automatically tuned manifold sampler with a local step-size selection procedure that adapts to the geometry of the manifold. Under regularity conditions, we show that our method is invariant using the involutive MCMC framework. We further implement a contour-based sampling method with automatic tuning that achieves strong performance in terms of effective sample size per second while maintaining stable acceptance rates on several challenging target distributions. Empirical results show that automatic tuning can make manifold sampling more reliable and less sensitive to step-size choice for constrained and contour-based inference problems.

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca. 

Event Photo
Junsong Tang