Seminar

Making a Training Dataset from Multiple Data Distributions

Over time we might accumulate lots of data from several different populations: e.g., the spread of a virus across different countries. Yet what we wish to model is not any one of these populations. One might want a model for the spread of the virus that is robust to the different countries, or is predictive on a new location we have only limited data for. We overview and formalize the objectives these present for mixing different distributions to make a training dataset, which have historically been hard to optimize. We show that by assuming we train models near "optimal" for our training distribution these objectives simplify to convex objectives, and provide methods to optimize these reduced objectives. Experimental results show improvements across language modeling, bio-assays, and census data tasks.

To join this seminar virtually, please request Zoom connection details from ea@stat.ubc.ca. 

Is a spike-and-slab prior ever a good idea?

Statisticians who use Bayes factors often complain that, somewhat awkwardly, their hypothesis testing and estimation results can disagree: a Bayes factor may favour the null hypothesis while the 95% credible interval excludes the null value. In this talk, I show that this apparent contradiction disappears once testing and estimation are required to use the same "implied prior model odds." Upon careful consideration, we see that using a Bayes factor for decision-making implicitly commits one to a mixture model and, when one's null is a point null, implicitly specifies a spike-and-slab prior. Such priors are not without their challenges. The resulting model-averaged posterior has an atom at the null value, its CDF has a jump, and credible intervals at conventional levels such as 95% may not exist at all. Moreover, if the frequentist p-value is held fixed while the sample size grows, the range of probability levels at which a credible interval can be defined steadily shrinks. This is a lesser-known correlate of the Jeffreys–Lindley paradox, one that carries the paradox from testing into interval estimation. All of this leads to a simple question: is a spike-and-slab prior (or, equivalently, a Bayes factor against a point null) ever a good idea?  

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca.

Event Photo
Harlan Campbell

Trend Filtering Deconvolution Methods in Epidemiology

Reported cases provide a delayed and incomplete view of infection incidence and so make it difficult to recover the true timing and magnitude of an epidemic. To use such surveillance data to estimate latent infections, we develop a trend filtering deconvolution approach. We use this to retrospectively estimate daily COVID-19 infections across U.S. states prior to the Omicron period using seroprevalence data in an antibody prevalence model. We adapt it for the Omicron period by using viral shedding and wastewater concentration data that allow for infection estimation when case and seroprevalence surveillance data become less reliable. Then we study the statistical properties of the estimator by extending trend filtering theory to the case where the design arises from delay distributions as a Toeplitz convolution operator and develop a backfitting extension to handle multiple overlapping infection curves, such as those from different variants.

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca. 

Event Photo
Rachel Lobay

Bridging the Gap: How Practicing Data Scientists Use LLMs in the Wild and What It Means for Data Science Education

Since the widespread availability of generative artificial intelligence (GenAI), particularly Large Language Models (LLMs), fundamental questions have emerged about the future of coding in data science. Some predict that data scientists will no longer need traditional coding skills, while others question whether LLMs might replace data scientists entirely. However, these discussions have largely proceeded without empirical evidence of how practicing data scientists actually use these tools.

This study addresses this gap by surveying trained, practicing data scientists to understand if and how they integrate LLMs into their workflows, particularly for writing and editing code and performing other data science tasks. Building on our recent investigation of data science educators' perspectives on LLMs, this research examines real-world usage patterns among practitioners to bridge the gap between current practice and educational preparation.

Our findings will contribute to the data science community in two critical ways. First, by documenting how data scientists are actually working with LLMs four years after their initial release, we provide actionable insights that allow practitioners to learn and adopt effective strategies for integrating these tools into their work. Second, we inform data science education by evaluating whether current pedagogies adequately prepare students for this evolving landscape.

This research will help answer key pedagogical questions: Should coding education emphasize code reading, tracing, and editing over writing large amounts of de novo code? Should greater focus be placed on writing high-quality documentation and specifications, given their value as prompt context for LLMs? Should testing receive increased emphasis to enable verification of LLM-generated code? By grounding these questions in empirical evidence of practitioner behaviour, we aim to provide data-driven guidance for evolving data science curricula.

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca.

This talk is one of the Teaching and Learning in Statistics and Data Science Seminar Series. 

Event Photo
Dr. Tiffany Timbers
Tags

Veridical Data Science towards Trustworthy AI

Data science underpins modern AI and many advances in healthcare, yet human judgment permeates every stage of the data science life cycle. These judgment calls introduce hidden uncertainties that go well beyond sampling variability and drive many of the risks associated with AI.

We introduce veridical data science, grounded in three fundamental principles—Predictability, Computability, and Stability (PCS)—to make such uncertainties explicit and assessable and to aggregate reality-checked algorithms for better results. The PCS framework unifies and extends best practices in statistics and machine learning and is illustrated through healthcare applications, including identifying genetic drivers of heart disease, reducing cost of prostate cancer detection, improving uncertainty quantification beyond standard conformal prediction, and proposing, Green Shielding, a new user-centric framework for safeguarding users of AI. 

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca.

Event Photo
Professor Bin Yu

Computationally Efficient Design with Precision Criteria

Estimation frameworks for statistical inference are preferred to hypothesis testing when quantifying uncertainty and precise estimation are more valuable than binary decisions about statistical significance. Study design for estimation-based investigations often uses precision criteria to select sample sizes that control the length of interval estimates with respect to a sampling distribution. In this work, we define a distribution that characterizes the probability of obtaining a sufficiently narrow interval estimate as a function of the sample size. This distribution can be used to determine the smallest sample size needed to ensure an interval estimate is sufficiently narrow. We prove that this distribution is approximately normal in large-sample settings for many data generation processes. However, this approximate normality may not hold for studies with moderate sample sizes, particularly when incorporating prior information or obtaining asymmetric interval estimates. Thus, we also propose an efficient simulation-based approach to approximate the distribution for the sample size by estimating the sampling distribution of interval estimate lengths at only two sample sizes. Our methodology provides a unified framework for design with precision criteria in Bayesian and frequentist settings.

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca.

Event Photo
Dr. Luke Hagar

Some new advances in precision medicine modeling

Earlier work has shown that similarity-based predictive models can improve upon predictive performance, as compared to using the entire training data to help build models, particular regarding model discrimination for binary responses. My collaborators and I have some updated results to share, regarding similarity-based modeling for joint consideration of model calibration and discrimination, as well as for dynamic prediction models. In addition, we have been developing transfer learning methods for targeted prediction. Properties of our methods will be investigated in comprehensive simulation studies, and we will demonstrate the methods through separate analyses of a publicly-available intensive care unit (ICU) database.

Collaborators: Keeley lsinghood, Minzee Kim, Tatiana Krikella, Subha Maity, Mengqi Xu Department of Statistics and Actuarial Science, and School of Public Health Sciences, University of Waterloo; Department of Statistical Sciences, University of Toronto

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca.

Event Photo
Professor Joel Dubin

Classification and diffusion-induced neural density estimators and simulators for generative AI

Neural network-based methods for conditional density estimation have recently gained substantial attention, as various neural density estimators have outperformed classical approaches in real-data experiments. Despite these empirical successes, implementation can be challenging due to the need to ensure non-negativity and unit-mass constraints, and theoretical understanding remains limited. In particular, it is unclear whether such estimators can adaptively achieve faster convergence rates when the underlying density exhibits a low-dimensional structure. This paper addresses these gaps by proposing a structure-agnostic neural density estimator, called the classification-induced neural density estimator and simulator (CINDES) that is straightforward to implement and provably adaptive, attaining faster rates when the true density admits a low-dimensional composition structure. Another key contribution of our work is to show that the proposed mator integrates naturally into generative sampling pipelines, most notably score-based diffusion models, where it achieves provably faster convergence when the underlying density is structured. We validate its performance through extensive simulations and a real-data application. We also prove the optimality of score-based diffusion models for density estimation when the target density admits a factorizable, low-dimensional, nonparametric structure in a separate work. The main challenge is that the low-dimensional, factorizable structure no longer holds for most diffused timesteps, and it is very difficult to show that these diffused score functions can be well approximated without a significant increase in the number of network parameters.

(Join works with Yihong Gu, Dehao Dai, Mukherjee, and Ximing Li)

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca.

 

Professor Jianqing Fan, a member of the US National Academy of Sciences, the Royal Academy of Belgium, and Academia Sinica, is the Frederick L. Moore’18 Professor of Finance, Professor of Operations Research and Financial Engineering, and Former Chairman of the Department of Operations Research and Financial Engineering at Princeton University, where he directs both Financial Econometrics and Statistics and Data Science labs. He received his Ph.D. from the University of California at Berkeley and held faculty positions at the University of North Carolina at Chapel Hill, University of California at Los Angeles, and the Chinese University of Hong Kong before joining Princeton University. He is a fellow of the American Association for the Advancement of Science, the Institute of Mathematical Statistics, the American Statistical Association, and the Society of Financial Econometrics. He has served as the president of the Institute of Mathematical Statistics and the International Chinese Statistical Association, and has been a joint editor of the Journal of the American Statistical Association, Annals of Statistics, Probability Theory and Related Fields, Econometrics Journal, Journal of Econometrics, Journal of Business and Economics Statistics, and Management Science (Finance Department editor). Awards include the COPSS Presidents' Award, the Morningside Gold Medal of Applied Mathematics, the Guggenheim Fellowship, the P.L. Hsu Prize, the Royal Statistical Society Guy medal in silver, the Noether Distinguished Scholar Award, Le Cam Award and Lecture, the Frontiers of Science Award, and the Wald Award and Lecture. His research interests include high-dimensional statistics, data science, machine learning, deep learning, mathematics of AI, financial economics, and computational biology. He coauthored 4 books and published over 300 highly cited papers, with over 100,000 citations.

Event Photo
Professor Jianqing Fan

Efficient Estimation and Closed-Form Uncertainty Quantification for Net Benefit of Algorithms that Predict Individualized Treatment Benefit

Treatment benefit predictors (TBPs) quantify the expected treatment benefit given individual characteristics. The net benefit function evaluates the expected gain in clinical utility from using a TBP to guide treatment decisions, relative to the default decisions of treating no one and treating everyone. The existing estimator for the net benefit of a given TBP implicitly assumes a 1:1 randomization design in the randomized controlled trial data. When this assumption is violated, the estimator can exhibit biased and unstable finite-sample behaviour. We identify the source of this implicit design restriction and propose a corrected congruent-based modification of the net benefit estimator that remains valid under arbitrary randomization schemes. We further introduce an alternative net benefit formulation based on the average treatment effect among individuals recommended for treatment by a TBP (ATT-based). The ATT-based estimator is asymptotically equivalent to the corrected congruent-based estimator and makes more efficient use of the full sample by estimating the proportion recommended for treatment using all individuals. Three uncertainty quantification procedures are developed for the ATT-based estimator: a large-sample variance approximation, an aggregated nonparametric bootstrap, and a Bayesian analysis. A Monte Carlo simulation study evaluates point estimation of net benefit curves using mean squared error, while uncertainty quantification is assessed through confidence interval length and coverage probabilities for 95% intervals constructed using asymptotic, bootstrap, and Bayesian methods. Simulation studies show that the corrected congruent-based estimator removes the bias under unequal randomization, while the ATT-based estimator demonstrated improved finite-sample efficiency. An empirical application to the GUSTO randomized controlled trial illustrates the ATT-based estimator and accompanying uncertainty intervals in a real-world setting. Overall, the proposed methodology enables valid estimation and uncertainty quantification of net benefit for treatment benefit predictors under arbitrary randomization schemes.

To join this seminar virtually, please request Zoom connection details from hr.ops@stat.ubc.ca. 

Event Photo
Sasha Sharma

Topics in Trend Filtering with Poisson Loss

Many types of data come as counts — disease cases per day, website visits per hour, or pixel intensities in images. A common goal is to recover the smooth trend underlying these noisy counts. Trend filtering is a nonparametric method that fits flexible, piecewise polynomial curves which adapt automatically to abrupt changes in the signal without prespecifying where they occur. However, existing methods assume Gaussian noise, whereas count data follow Poisson-type models whose variability grows with the signal magnitude, making the effective noise heteroscedastic.

This dissertation develops scalable algorithms for trend filtering under Poisson loss and methodology for solving real-life applications. We propose two proximal algorithms that extend the estimator from simple time series to general graph structures. We apply this framework to epidemic surveillance, producing an R package (rtestim) that estimates time-varying reproduction numbers with principled, cross-validated tuning. We further identify and resolve a numerical instability in the linear system solvers that arise as inner subproblems of these algorithms, by recasting the system as a linear Gaussian state-space model, yielding a solver that is both stable and efficient. Finally, an ongoing work of ours shows that observation-dependent penalty weights can recover minimax optimal rates under the heteroscedastic noise inherent in exponential-family models.

Event Photo
Jiaping (Olivia) Liu