The frontier of Spatial Statistics: Methods, applications and scalability
ChairMarcos Prates · Universidade Federal de Minas Gerais
As data collection technologies advance, the demand for more realistic frameworks capable of handling complex dependencies in geospatial data has increased significantly. This thematic session showcases cutting-edge Bayesian frameworks designed to tackle the complex non-Gaussian structures, data irregularities, and computational bottlenecks inherent in modern spatial data. On the methodological and scalable front, a combination of nearest-neighbor mixture processes and Hausdorff distances offers elegant solutions for multi-resolution data fusion and atmospheric interpolation while ensuring algorithmic scalability. To address complex observational structures, semi-parametric approaches using Bernstein Polynomials are proposed to model zero-inflated, spatially correlated, recurring events, providing crucial insights into criminal recidivism trajectories. Finally, advanced geostatistical frameworks demonstrate how to robustly handle censored responses and detection limits in environmental contamination studies. In summary, this session illustrates innovative methods in spatial statistics that provide scalable, rigorous solutions to complex real-world phenomena.
Nearest-Neighbor Mixture Processes for Data Fusion
Bruno Sansó · University of California - Santa Cruz
Abstract
Spatial data for a given variable are often available at varying spatial resolutions. This issue is known as spatial misalignment. A specific instance of spatial misalignment in environmental sciences is that of observations that are collected by monitoring stations on the ground, as well as data retrieved by satellites. To obtain a unified description of the spatial field, some form of spatial data fusion needs to be performed. Common statistical approaches to this problem assume that the data observed at the coarse resolution correspond to the aggregation of spatial process indexed at the resolution of the fine scale data. Alternatively, the coarse resolution observations are "downscaled'' by regressing them on the fine resolution observations. In this work we consider an alternative approach that uses distances between sets to defined a unified model that encompasses both sources of information within the same stochastic process. Leveraging recent developments in the literature of spatial process, we develop a family of nearest neighbors processes defined on a space with the maximum inscribed circle Hausdorff distance. Our proposed models generalize the nearest-neighbor Gaussian processes, as well as the nearest neighbor mixture processes. This allows modeling of Gaussian and non-Gaussian spatial fields. The nearest neighbor structure induces sparsity in the likelihood, allowing for efficient computations, even for large datasets. We illustrate our methods with examples of spatial interpolation of atmospheric variables, Gaussian and non-Gaussian, observed over the state of California at ground based stations as well as remotely sensed from satellites.
The analysis of criminal recidivism: a hierarchical model-based approach for the analysis of zero-inflated, spatially correlated recurring event data
Marcos Oliveira Prates · Universidade Federal de Minas Gerais
Abstract
The life course perspective in criminology has become prominent in recent years, offering valuable insights into various patterns of criminal pathways. Noticeably, the study of criminal trajectories aims to understand crime's beginning, persistence, and desistance. Central to this analysis is the identification of patterns in the frequency of criminal victimization and recidivism, along with the factors that contribute to them. Specifically, this work introduces a new class of models that overcome limitations in traditional methods used to analyse criminal recidivism. The proposed models are designed for recurrent events data characterized by excess of zeros and spatial correlation. In addition to their parametric counterparts, we propose flexible semi-parametric versions approximating the intensity function using Bernstein Polynomials. The performance of these models was evaluated in a simulation study with various scenarios, and we applied them to analyse criminal recidivism data in the Metropolitan Region of Belo Horizonte, Brazil. The results provide a detailed analysis of high-risk areas for recurrent crimes and the behaviour of recidivism rates over time. This research significantly enhances our understanding of criminal trajectories, paving the way for more effective strategies in combating criminal recidivism. The author acknowledged CNPq and FAPEMIG for partial financial support.
Geostatistical Modeling of Censored Responses: A Bayesian Approach using Stan
Victor Hugo Lachos · University of Connecticut
Abstract
Geostatistical data in epidemiology, hydrology, and environmental studies frequently encounter detection limits, resulting in censored measurements (left-, right-, or interval-censored). Addressing these challenges requires specialized statistical methods for accurate inference and prediction. Although classical approaches such as the Expectation-Maximization algorithm, Monte Carlo EM, and Stochastic Approximation EM have been employed, this study presents a Bayesian geostatistical framework that leverages Stan's No-U-Turn Sampler to obtain posterior simulations for parameter estimation. The performance of the methodology was evaluated through both uncensored and censored simulation studies, and further validated with real contamination datasets: dioxin in Missouri and arsenic in Michigan. Our results yielded credible intervals for model parameters, enabling robust inference and the identification of optimal covariance structures.
TS2
Bayesian phylogenetics: where high-dimensional data and complicated models meet
ChairLuiz Max Carvalho · School of Applied Mathematics, Getulio Vargas Foundation
Phylogenetic trees are a central object in subjects as diverse as Linguistics, Medicine, Evolutionary Biology and Molecular Epidemiology. Over the last three decades, Bayesian methods for phylogenetic inference/reconstruction have become central to various areas of Science due to the increase in the availability of powerful computers and molecular (DNA) data. In recent years, however, we have witnessed an unprecedented increase in the amount of molecular data, especially for pathogens such as Ebola and SARS-Cov-2. Scientists have in turn built rather complicated, high dimensional models to try and extract insights from these data. While Bayesian inference for these models usually employs Markov chain Monte Carlo methods, developing efficient algorithms is hampered by the intractability of the likelihood under most commonly employed models. This intractability places limits on important statistical tasks such as model selection, sensitivity analysis and prior and posterior checks. While the gap between the data that we have and the questions we'd like to ask is widening, Bayesian methods have helped immensely in this task, as evidenced by two Mitchell prizes (2006 and 2011) being won by phylogenetic applications.
In this thematic session we will explore modern solutions to Bayesian phylogenetic problems such as inference for coalescent point processes, continuous-time Markov chains for infections diseases, and diagnostics for Markov chain Monte Carlo in phylogenetic space. I have included a backup talk in case anything happens the day of.
Coalescent point processes
Julia Adelia Palacios · Department of Statistics, Stanford University
Abstract
The coalescent is a central modeling framework in modern population genetics and evolutionary biology. Coalescent models provide robust approximations to the distribution of sampled genealogies under a variety of population dynamics and allow us to express genealogical tree priors in terms of population parameters of interest. Point process constructions of these models have proven to be an exceptionally fruitful mathematical approach. In this talk, I will present three new instances in which point process constructions facilitate the derivation of novel coalescent priors and simulation. In particular, we derive the coalescent distribution under (1) epidemiological disease dynamics with latent periods, (2) linear birth-death dynamics and, (3) when the TMRCA is bounded (bounded coalescent).
Bayesian computation for discrete-valued stochastic processes in infectious diseases
Marc Adam Suchard · Department of Biostatistics in the UCLA Fielding School of Public Health, University of California Los Angeles (UCLA)
Abstract
Bayesian computation remains onerous at scale for inference under many discrete-valued stochastic process-based models, while these statistical models remain ubiquitous across biology and public health. In this talk, I will explore how one can construct high-performance samplers for flexible continuous-time Markov chain (CTMC) models. CTMCs underpin the most popular models for learning about how rapidly evolving pathogens change over time and space to give rise to human infection, and the dimensionality of these problems are daunting. These approaches enable the introduction of novel non-parametric CTMC models with Gaussian process priors that capture biological realism previously missing. Applied to the analysis of early SARS-CoV-2 and Yellow Fever virus (YFV) genomes, such flexible models remove bias in inference of the location and timing of the SARS-CoV-2 split-over into humans and identify the non-linear effect of temperature on YFV spread in South America, while the statistical machinery is over an order of magnitude more time efficient than conventional approaches.
Nonparametric tests for the convergence of tree-valued Markov chains
Mauricio Darós · School of Applied Mathematics, Getulio Vargas Foundation
Abstract
Trees are a fundamental object in many applied fields, and phylogenetic trees in particular play a central role in Biology and Medicine. Bayesian inference for these objects often involves intricate Markov chains Monte Carlo (MCMC) algorithms, giving rise to tree-valued Markov processes, whose convergence needs to be checked in order to validate the results. Convergence assessment for tree-valued MCMC usually involves multidimensional scaling (MDS) procedures which produce high-dimensional cloud points for which only the first few dimensions are usually analysed. In this talk we exploit nonparametric tests based on graph colouring to develop a robust testing framework aimed at assessing whether a set of K independent MCMC runs can be said to have converged to the same target. Our results provide practitioners with a principled framework for assessing convergence of tree-valued MCMC.
TS3
Advances in Bayesian Nonparametric and Latent Variable Modeling: Recent Contributions from j-ISBA Members
ChairNicolas Bianco · Karlsruhe Institute of Technology
Modern statistical applications increasingly involve data that are high-dimensional, dependent, or non-Euclidean, requiring Bayesian models that are both flexible and computationally tractable. This session presents three contributions that develop latent variable and Bayesian nonparametric methods to address these challenges. The first introduces an adaptive shrinkage prior for high-dimensional regression by combining the Bayesian Lasso with Dirichlet process mixtures, improving variable selection and coefficient estimation. The second develops a state-space model for multivariate spherical time series, using latent Gaussian processes to capture temporal and cross-sectional dependence in directional data. The third proposes a Bayesian nonparametric framework for covariate-adjusted hypothesis testing in shape regression, enabling principled inference for non-Euclidean data with applications in biomedical imaging. Together, these works illustrate how latent representations, flexible prior distributions, and efficient posterior computation extend Bayesian inference to increasingly complex data structures. Beyond its scientific contributions, this session contributes to j-ISBA's broader commitment to supporting early-career researchers, complementing the dedicated Francisco Torres session focused on young scientists. Moreover, this initiative strengthens the connections between j-ISBA and the growing Bayesian research community in Latin America, reflecting a shared commitment to an inclusive and globally connected network of Bayesian researchers.
Adaptive shrinkage with a nonparametric Bayesian Lasso
Santiago Marin · Australian National University
Abstract
Modern approaches to perform Bayesian variable selection rely mostly on the use of shrinkage priors. That said, an ideal shrinkage prior should be adaptive to different signal levels, ensuring that small effects are ruled out, while keeping relatively intact the important ones. With this task in mind, we develop the nonparametric Bayesian Lasso, an adaptive and flexible shrinkage prior for Bayesian regression and variable selection, particularly useful when the number of predictors is comparable or larger than the number of available data points. We build on spike-and-slab Lasso ideas and extend them by placing a Dirichlet process prior on the shrinkage parameters. The result is a prior on the regression coefficients that can be seen as an infinite mixture of Laplace distributions, all offering different amounts of regularization, ensuring a more adaptive and flexible shrinkage. We also develop an efficient Markov chain Monte Carlo algorithm for posterior inference. Through extensive simulation studies and real-world data analyses, we illustrate that our proposed method leads to coefficient recovery, variable selection accuracy, and out-of-sample predictions that are comparable to or better than those from state-of-the-art shrinkage priors, highlighting the benefits of the nonparametric Bayesian Lasso over existing methods.
Multivariate Dynamic Projected Normal: A State-Space Model for Cross-Dependent Spherical Time Series
Maria Fernanda Pintado · CUNEF Madrid
Abstract
Directional data on the unit sphere arise throughout the sciences are frequently observed jointly across multiple related series through time. Existing Bayesian models for spherical time series, such as the projected dynamic linear model of Zito and Kowal (2024), are built for a single series, leaving cross-sectional dependence between jointly evolving spherical processes largely unaddressed. We propose the Multivariate Dynamic Projected Normal, a state-space model for pairs of time series on the unit sphere obtained by radially projecting a latent Gaussian vector whose mean is driven by a shared dynamic latent state. Temporal dependence is induced through this common state, while a contemporaneous cross-covariance block captures additional dependence between the two series. We develop a Bayesian data-augmentation scheme that extends the identifiability results of Hernandez-Stumpfhauser et al. (2017) to the joint setting, yielding tractable full conditional distributions for the marginal covariance structure together with a feasible sampling strategy for the cross-covariance block that preserves positive definiteness of the joint covariance matrix. The resulting Gibbs sampler enables fully Bayesian inference for flexible, dependent spherical time series models of arbitrary dimension.
Bayesian hypothesis testing for Shape Regression with Application to Autism Spectrum Disorder
Ingrid Johana Guevara Romero · Pontificia Universidad Católica de Chile
Abstract
Bayesian model comparison offers a rigorous framework for evaluating competing scientific hypotheses while effectively accounting for uncertainty. While Bayesian model selection has been extensively explored in the context of Euclidean regression models, there has been relatively little focus on hypothesis testing for non-Euclidean responses, such as shape data. In these instances, standard statistical methodologies present challenges, as observations reside in nonlinear spaces and relevant covariates may obscure the scientific question at hand. We introduce a Bayesian nonparametric framework for covariate-adjusted hypothesis testing involving two-dimensional shape data. Shape configurations are modeled using a Dependent Dirichlet Process mixture of complex Watson distributions, which allows for flexible estimation of multimodal shape distributions while simultaneously incorporating covariate effects. This methodology frames competing scientific hypotheses as Bayesian model comparison problems, resulting in posterior probabilities that directly quantify the evidence supporting either shared or distinct population shape distributions. Furthermore, we derive a representation of the marginal distributions of individual landmarks, facilitating visualization and interpretation of the fitted model. Simulation studies illustrate that our proposed methodology effectively distinguishes genuine distributional differences from apparent differences induced by confounding variables. We demonstrate the approach by analyzing the shape of the corpus callosum in children with autism spectrum disorder, adjusting for age and sex. The results yield substantial evidence in favor of a common shape distribution between diagnostic groups, suggesting that previously observed structural differences are more likely attributed to size rather than geometric shape. More broadly, this work demonstrates how Bayesian nonparametric model comparison can provide interpretable probabilistic inference for complex non-Euclidean regression challenges encountered in biomedical imaging and related fields.
TS4
Advances in Bayesian Computation for Latent Gaussian Models
ChairHaavard Rue · KAUST
In this session we highlight recent developments in Bayesian computation for the class of Latent Gaussian Models (LGMs). LGMs form a broad class of Bayesian hierarchical models in which Gaussian priors are assigned to the latent components entering the linear predictor. This enables the use of the Laplace method while dealing with the other model parameters, the hyper-parameters. By performing a Gaussian approximation (GA) for the conditional distribution of the latent field, conditional on the data and the hyper-parameters, it form a Laplace approximation to the (joint) distribution for the hyper-parameters.
The inference proceed using either simulation-based, with the GA as a proposal, or with analytic approximations for the latent field as well. The objective remains the same: obtaining posterior distributions for all unknown quantities in the model. The Integrated Nested Laplace approximations (INLA) one for each latent field, was originally proposed.
Moving further, the INLA software keep including new prior, likelihood, and latent models, as well providing tools for model evaluation. This session address recent developments in this direction, 1) on theory and aligned efficient computations to perform model evaluation considering leave-group-out, instead of just the leave-one-out approach, 2) on models for correlation matrices, and 3) on the simultaneous estimating equation models.
Recent developments in the R-INLA project
Haavard Rue · KAUST
Abstract
I will in this presentation give an overview of recent developments in the R-INLA project, which includes 1. the inclusion of a new parallel sparse matrix solver (sTiles), 2. our work on (group-)cross validation including log-Gaussian Point processes, 3. our on-going work on model validation and model-checking.
Prior models for graphical correlation matrices
Elias T Krainski · KAUST
Abstract
The correlation between two variables is an important summary of its dependency. With several variables, the correlation matrix summarizes the marginal linear relationships among each pair of them. In data learning a correlation matrix some challenges arise, including rapid growth in the number of parameters. We address this challenge by proposing a graph approach to model correlation matrices. Such a graph represent conditional distributions and we propose a model based on that with as many parameters as the number of conditional distributions. We further leverage expert knowledge by introducing a prior that penalizes divergence from a base correlation matrix. When combined, it provides an informative model-based prior for correlation matrices. We will illustrate the performance of such a prior against others for fair comparison cases and illustrate how ours work for prior knowledge elicitation with an application.
INLA for structural equation models
Haziq Jamil · KAUST
Abstract
Structural equation models (SEM) are widely used to study causal pathways, latent constructs, and measurement error. Full Bayesian estimation via Markov chain Monte Carlo (MCMC), however, is often too slow for the complexity of modern applications. An approximate Bayesian approach to SEM is presented, drawing on ideas from the integrated nested Laplace approximation (INLA) framework. A Laplace approximation to the joint posterior is computed first, and its mean is then shifted by a variational Bayes correction to better capture the posterior mass. Each marginal is estimated by a simplified Laplace approximation, which profiles the posterior density efficiently along each parameter direction while correcting for asymmetry, yielding a parametric skew-normal fit. An efficient Gaussian copula sampling scheme then delivers the essential quantities: factor scores, modelfit indices, and credible intervals for nonlinear derived parameters such as indirect effects. The approach achieves speeds close to maximum likelihood estimation, while retaining the inferential richness of full Bayesian analysis. The methodology is implemented in the R package INLAvaan, and its speed and accuracy are illustrated against MCMC benchmarks on simulated and real data.
TS5
Bayesian Methods for Modelling and Forecasting Life Tables in Actuarial Science
ChairMariane Branco Alves · DME/IM-UFRJ
Life tables and age-specific mortality curves are fundamental tools in actuarial science, demography, public policy and longevity-risk assessment. In recent decades, stochastic mortality models have played a central role in the construction, graduation and forecasting of life tables, particularly in contexts involving pension systems, insurance portfolios and population-level mortality surveillance. Bayesian methods provide a natural framework for this class of problems, allowing uncertainty quantification, the incorporation of structured dependence across age and time, and the development of flexible dynamic models for mortality evolution.
This session brings together three contributions on Bayesian modelling of life tables and mortality curves. The presentations share a common focus on dynamic and structured statistical modelling, while approaching the problem from complementary perspectives: mortality surfaces with local age–time dependence, state-space formulations of the Lee–Carter model, and Bayesian nonparametric dynamic hazard rates for evolutionary life tables. Together, the talks aim to highlight recent methodological developments and actuarial applications of Bayesian inference for mortality modelling, graduation and forecasting.
Dynamic Bayesian Mortality Surface Modelling with Local Age–Time Dependence
Mariane Branco Alves · DME-LabMA-LSE/IM-UFRJ
Abstract
Reliable mortality forecasts are central to the measurement and management of longevity risk in pension, insurance and annuity applications. Classical stochastic mortality forecasting is often based on Lee–Carter-type models, in which age-specific log-mortality rates are represented through a fixed age profile, one or more period factors and age-specific sensitivities to these factors. Although these models are parsimonious and widely used, they typically impose a low-dimensional global structure on mortality improvement: temporal changes are summarised by common latent factors, while local dependence among neighbouring ages is not explicitly modelled. The approach proposed here moves
away from this global-factor representation by modelling log-mortality as a latent age–time surface with structured dependence across ages and temporal evolution over calendar years. We investigate two related dynamic Bayesian formulations. The first is motivated by Lavine's general connection between conditionally Gaussian Markov random fields and dynamic linear model representations. Although this connection was not originally developed for mortality modelling, we adapt it to the age–time structure of mortality curves, providing a coherent Markovian mechanism for borrowing information simultaneously across neighbouring ages and adjacent years. The second formulation is based on matrix-variate dynamic linear models, in which a specific dependence structure is imposed
across ages. This structure allows the model to account for association among age-specific mortality rates while preserving the analytical tractability of sequential Bayesian updating. As a consequence, the resulting inference remains computationally efficient and suitable for forecasting the wholemortality surface. Discount factors are used to specify evolution variances and smoothing parameters indirectly, avoiding the computational burden associated with their repeated direct estimation. The state-space representation also provides a natural framework for intervention modelling, allowing temporal structural breaks in mortality regimes to be accommodated within the dynamic evolution of the curves. This is particularly relevant when mortality trajectories are affected by abrupt shifts due to exceptional events or changes in population-level risk patterns. The methodology is
illustrated using real age-specific mortality data. The proposed formulations are compared with Lee–Carter-type models used as reference models, with the empirical analysis suggesting advantages in terms of local smoothing, extrapolation, computational efficiency, intervention flexibility and predictive uncertainty quantification over the mortality surface. Overall, the proposed framework preserves the forecasting interpretation of state-space models while introducing explicit dependence structures for the joint modelling of age association, temporal evolution, mortality improvement and longevity-risk prediction.
*Joint work with: Luís Philipe Craveiro Mendes (IM–UFRJ) and Rafael Santos Erbisti (IME–UFF)
Dynamic Linear Lee-Carter Model
João Batista Pereira · DME-LSE/IM-UFRJ
Abstract
Understanding mortality patterns and forecasting mortality are fundamental objectives in many fields. For example, in demography, understanding mortality patterns can support the planning of public policies; in actuarial science, it enables the construction of life tables and insurance pricing; and in medicine, it may help explain or predict the impacts of a disease. An essential tool for understanding and forecasting mortality is the modeling of mortality rates. In this context, two key aspects must be taken into account: (i) mortality rates vary with age; and (ii) mortality patterns change over time. The predominant approach to modeling mortality rates is the Lee-Carter model. We propose a fully state-space representation of the Lee-Carter model as a dynamic linear model. Specifically, just as the time-specific parameters are assumed to evolve dynamically over time, the age-specific parameters are assumed to evolve dynamically over age, thereby inducing, a priori, a dependence structure among the observations not only over time but also over age. The proposed framework allows the mortality-related parameters to follow more flexible dynamic models, such as damped growth models, leading to more realistic mortality forecasts over both time and age. Inference for the proposed model is conducted within
the Bayesian framework. The proposed state-space representation of the Lee-Carter model allows traditional stochastic simulation methods, such as Markov chain Monte Carlo (MCMC), to be combined with sequential algorithms commonly used for inference in dynamic models. These algorithms can be used to sample not only the time-specific parameters but also the age-specific parameters. The methodology is illustrated through simulation studies and applications to widely used mortality datasets.
Bayesian nonparametric dynamic hazard rates in evolutionary life tables
Luis E. Nieto-Barajas · ITAM, Mexico
Abstract
In the study of life tables the random variable of interest is usually assumed discrete since mortality rates are studied for integer ages. In dynamic life tables a time domain is included to account for the evolution effect of the hazard rates in time. In this article we follow a survival analysis approach and use a nonparametric description of the hazard rates. We construct a discrete time stochastic processes that reflects dependence across age as well as in time. This process is used as a bayesian nonparametric prior distribution for the hazard rates for the study of evolutionary life tables. Prior properties of the process are studied and posterior distributions are derived. We present a simulation study, with the inclusion of right censored observations, as well as a real data analysis to show the performance of our model.
TS6
Bayesian Clustering for Complex Temporal and Functional Data
ChairGarritt Page · Brigham Young University
Recent advances in Bayesian model-based clustering have enabled increasingly sophisticated analyses of complex data characterized by high dimensionality and temporal dependence. This session brings together three complementary contributions that develop novel Bayesian model-based clustering techniques for discovering and interpreting hidden structure in longitudinal, multivariate, and functional data. Together, the talks illustrate how modern Bayesian modeling addresses challenging inferential problems arising in biomedical, financial, and biomechanics data applications while providing scientifically meaningful interpretations.
The session features an emerging researcher, an earlier/mid career researcher and a more established investigator. Bryan Tobar (PhD student at the Pontificia Universidad Católica de Chile) will present a Bayesian extension of Product Partition Models for detecting delayed structural changes in multivariate time series, allowing change points to propagate across related series with temporal lags. The methodology is illustrated using international financial market data. Max Russo (Assistant Professor at The Ohio State University) will present a Bayesian biclustering framework for high-dimensional longitudinal data that jointly identifies patient subgroups and symptom clusters, providing new insights into the progression of X-linked dystonia-parkinsonism. Garritt Page (Professor at Brigham Young University) will present work that seeks to explain the mechanisms underlying Bayesian model-based clustering of functional data by identifying the curve characteristics that drive cluster formation. Building on this framework, a new Bayesian approach to functional clustering is introduced that directly highlights the features of the curves responsible for the resulting clusters.
Collectively, these presentations demonstrate the breadth and flexibility of Bayesian methods for uncovering latent structure in modern complex data and should appeal to a broad audience interested in Bayesian methodology, computation, and applications.
Bayesian Modeling of Delayed Structural Changes in Multivariate Time Series
Bryan Tobar · Pontificia Universidad Católica de Chile
Abstract
Understanding how structural disruptions propagate across financial markets is essential for studying contagion and systemic risk. We propose a Bayesian extension of Product Partition Models for multivariate change-point detection that allows changes in one series to relate to the probability of changes in other series after a temporal delay. Dependence is introduced through a multivariate latent process with flexible temporal correlation structures. The model accommodates simultaneous and asynchronous changes while borrowing information across related series. Posterior inference is performed through an efficient MCMC algorithm. Simulation studies assess change-point recovery under different propagation patterns, and the methodology is illustrated using international financial market data.
Bayesian biclustering for temporally heterogeneous high-dimensional longitudinal data
Max Russo · Ohio State University
Abstract
X-linked dystonia-parkinsonism (XDP) is a rare genetic form of dystonia found almost entirely among males of Filipino descent, characterized by highly heterogeneous symptoms and unknown progression patterns. Identifying subtypes of XDP is pivotal in advancing our understanding of the disease and in providing effective targeted treatment for the affected
patients. However, analysis of existing longitudinal studies of XDP is complicated by the fact that (i) the patients are observed for only a short length of time at different stages of progression, and (ii) the disease symptoms are measured using a large number of interdependent scales. To overcome these challenges, we propose a novel Bayesian statistical model that simultaneously clusters individuals according to the trajectory of their progression and clusters variables into jointly relevant aspects of the disease. We apply the model to clinical XDP data and find that it reveals novel insights into the patterns of disease progression.
Interpretable Bayesian Functional Clustering via Decomposition of Curve Features
Garritt Page · Brigham Young University
Abstract
Functional clustering is commonly performed by first representing curves with a finite basis expansion and then clustering the resulting coefficient vectors using Bayesian model-based methods. While this strategy has proven effective across a wide range of applications, it often provides little insight into why curves are assigned to particular clusters. Moreover, the resulting clustering can depend strongly on the choice of basis representation and prior specification, making interpretation difficult. We propose a Bayesian functional clustering framework that decomposes spline representations into interpretable components corresponding to the intercept, slope, and curvature of each function. This decomposition allows cluster membership to be understood in terms of scientifically meaningful curve characteristics rather than distances between high-dimensional basis coefficients. Through simulated and real data examples, we demonstrate how the proposed approach both clarifies the mechanisms underlying Bayesian functional clustering and provides greater flexibility in identifying the aspects of functional observations that drive cluster formation.
TS7
Bayesian models for finite population prediction with special applications in small area estimation and rare clustered populations
ChairKelly Cristina Mota Gonçalves · DME-UFRJ
This session presents three original contributions to finite population prediction. Two papers focus on small area estimation, while the third addresses rare and clustered populations. Each work illustrates its theoretical framework with applications using real data.
How do we handle mass points in small area estimation? A Bayesian missing data solution
Anna Sikov · Universidad Nacional de Ingeniería, Peru
Abstract
This talk addresses the challenge of predicting small area characteristics when direct
estimates follow a mixed distribution exhibiting both a continuous component and a mass
point. While existing small area literature typically treats such mass points as true under-
lying structural values, we explore a distinct scenario where the mass point is a consequence of very high rates coupled with small sample sizes. This scenario presents a widespread challenge in small area estimation using survey data, as exemplified by district-level labor informality rates in Peru, where survey data for a substantial proportion of districts yield a direct estimate of exactly one.
We adopt a Bayesian framework to implement a novel modeling approach that addresses
this challenge by drawing an analogy to a missing data problem, treating areas with 100%
estimated informality as non-respondents. Analogous to methods for non-ignorable non-
response, we construct the full likelihood and decompose it into a product of terms that
can be modeled. Additionally, we discuss out-of-sample prediction under this framework.
The performance and utility of the proposed solution are demonstrated by estimating labor
informality rates across the districts of the Peruvian Amazonia region.
A Beta–Beta Prime model for rates and their precision for small area estimation: an application to Brazilian food insecurity index
Fernando Antônio da Silva Moura · DME-UFRJ
Abstract
Statistics bureaus around the world have been faced with an increasing need to provide reliable estimates of economic and social indices, such as proportions or rates, from socioeconomic survey data at small area levels. For example, in 2015, the Member States of the United Nations committed themselves to the 2030 Agenda for Sustainable Development
Goals, which requires the national statistical offices and other govern ment agencies of the Member States to provide high-quality, timely, and reliable national indicators at a disaggregated level. However, due to the relatively small sample size of these areas or domains, it is not viable to obtain estimates with an acceptable level of accuracy without using model-based approaches. Here, we propose to model the direct estimator
of the rates or proportions in small area domains as being beta distributed. The novelty is that we also model the sampling precision estimator as a beta prime distribution. An evaluation study with real data shows that there is an additional gain in jointly modeling the direct estimator and its sampling precision estimator. An application is also provided to
estimate the food insecurity index in small areas of a Brazilian state.
Model-based Inference for Rare and Clustered Populations from Adaptive Cluster Sampling using Auxiliary Variables
Izabel Nolau de Souza · FGV EMAp
Abstract
Rare populations, such as endangered animals and plants, drug users and individuals with rare diseases, tend to cluster in regions. Adaptive cluster sampling is generally applied to obtain information from clustered and sparse populations since it increases survey effort in areas where the individuals of interest are observed. This work aims to propose a unit-level model which assumes that counts are related to auxiliary variables, improving the sampling process, assigning different weights to the cells, besides referring them spatially. The proposed model fits rare and grouped populations, disposed over a regular grid, in a Bayesian framework. The approach is compared to alternative methods using simulated data and a real experiment in which adaptive samples were drawn from an African Buffaloes population in a 24,108 square kilometers area of East Africa. Simulation studies show that the model is efficient under several settings, validating the methodology proposed in this paper for practical situations.
TS8
Cutting-Edge Bayesian Models for Disease Mapping
ChairGuilherme Lopes de Oliveira · Centro Federal de Educação Tecnológica de Minas Gerais (CEFET-MG)
A Bayesian hierarchical model for disease mapping that accounts for scaling and heavy-tailed latent effects
Alexandra M. Schmidt · McGill University, Canada
Abstract
In disease mapping, the relative risk of a disease is commonly estimated across different areas within a region of interest. The number of cases in an area is often assumed to follow a Poisson distribution whose mean is decomposed as the product between an offset and the logarithm of the disease's relative risk. The log risk may be written as the sum of fixed effects and latent random effects. A modified Besag-York-Mollié (BYM2) model decomposes each latent effect into a weighted sum of independent and spatial effects. We build on the BYM2 model to allow for heavy-tailed latent effects and accommodate potentially outlying risks, after accounting for the fixed effects. We assume a scale mixture structure wherein the variance of the latent process changes across areas and allows for outlier identification. We propose two prior specifications for this scale mixture parameter. These are compared through various simulation studies and in the analysis of Zika cases from the first (2015-2016) epidemic in Rio de Janeiro city, Brazil. The simulation studies show that the proposed model always performs at least as well as an alternative available in the literature, and often better, both in terms of widely applicable information criterion, mean squared error and of outlier identification. In particular, the proposed parametrisations are more efficient, in terms of outlier detection, when outliers are neighbours. Our analysis of Zika cases finds 23 out of 160 districts of Rio as potential outliers, after accounting for the socio-development index. Our proposed model may help prioritise interventions and identify potential issues in the recording of cases.
Continuous-Space Disease Mapping for Areal Data: A Spatiotemporal Nearest-Neighbor Gaussian Process Model for Tuberculosis in Rio Grande do Sul
Lucas da Cunha Godoy · PUC Chile
Abstract
Tuberculosis remains a critical public health challenge in Brazil, particularly in the state of Rio Grande do Sul, where incidence rates exceed the national average. Designing effective intervention strategies requires reliable local risk estimates, but analyzing municipal data is complicated by zero-inflation, small population sizes, and complex spatiotemporal dependencies. Standard disease mapping models typically rely on adjacency-based spatial structures that overlook the specific geometry of regions, failing to account for the fact that large, sparsely populated municipalities exert different spatial influences than small, densely populated ones. To resolve these limitations, we propose a spatiotemporal generalized linear mixed model to estimate tuberculosis incidence across the 497 municipalities of Rio Grande do Sul from 2011 to 2021. We introduce a separable spatiotemporal extension of isotropic Gaussian processes tailored for areal data, utilizing the spatial covariance through the ball-Hausdorff distance to explicitly account for the shape and size of regions. To ensure computational scalability, we implement a nearest-neighbor approximation for the spatial component which we make available through the INLAspacetime package. This framework yields robust, risk-smoothed incidence rates, providing policymakers with a reliable tool to identify high-risk areas and target social determinants effectively. Joint work with Elias Teixeira Krainski (KAUST), Marcos Oliveira Prates (UFMG), Hans Montcho (KAUST), Sabrina da Cunha Godoy (Secretaria da Saúde do Estado do Rio Grande do Sul), and Jun Yan (University of Connecticut).
Directional Asymmetry in Edge-Based Spatial Models via a Skew-Normal Prior
Rosangela Helena Loschi · Universidade Federal de Minas Gerais, Brazil
Abstract
We introduce a skewed edge-based spatial prior, named ReNeGe-SN, that extends the Gaussian RENeGe framework by incorporating directional asymmetry through a skew–normal distribution. Skewness is defined on the edge graph and propagated to the node space, aligning asymmetric behavior with transitions across neighboring regions rather than with marginal node effects. The model is formulated within the skew–normal framework and employs identifiable hierarchical priors together with low-rank parameterizations to ensure scalability. The skew–normal's stochastic representation is considered to facilitate the computational implementation. Simulation studies show that ReNeGe-SN recovers compact, edge-aligned directional structure more accurately than symmetric Gaussian priors, while remaining competitive under irregular spatial patterns. An application to cancer incidence data in Southern Brazil illustrates how the proposed approach yields stable area-level estimates while preserving localized, directionally driven spatial variation. Joint work with Danna L. Cruz-Reyes (Universidad Nacional de Colombia), Renato M. Assunção (UFMG), and Reinaldo B. Arellano-Valle (PUC Chile).
TS9
Bayesian nonparametric modeling for complex data structures
ChairAmy Herring · Dept. of Statistical Science, Global Health Institute and Dept. of Biostatistics & Bioinformatics, Duke University
This session brings together recent developments in Bayesian nonparametrics designed to handle complex dependence, high-dimensional covariates, and non-standard data structures. Highlights include novel calibration frameworks for covariate-informed product partition models, spatially structured latent position models for network and spatial interaction data, foundational developments in stick-breaking processes, and flexible BNP approaches to missing data and causal inference. These talks demonstrate how flexible BNP priors enhance modern network analysis, spatial epidemiology, and causal modeling.
A Bayesian product partition model using weighted covariates
Sally Paganin · Dept. of Statistics, Ohio State University
Abstract
Product partition models with covariates (PPMx) incorporate auxiliary information into Bayesian nonparametric clustering, but calibrating how much the covariates should matter is not straightforward: overweighting them can override structure supported by the response variable, while underweighting wastes useful information. We propose a Bayesian PPM with a calibration mechanism that tunes covariate influence on the resulting partition, allowing interpolation between partitions driven primarily by the response and those driven by covariate similarity. We illustrate the methodology in stochastic block models for network community detection using several types of node-level covariates, including spatial location, and show that the resulting communities are both topologically and covariate-coherent.
Size-biased sampling in conditionally independent stick-breaking models
María Fernanda Gil Leyva Villa · Dept. of Probability and Statistics, IIMAS, UNAM
Spatially informed Bayesian latent position model for global migration data
Michela Frigeri · Dept. of Mathematics, Politecnico di Milano
Abstract
Commercial and social exchanges between countries are often represented as count network data. While international trade and migration flow data are widely available, modeling the dependence structure underlying these interactions remains challenging.
In particular, cross-country exchanges are shaped not only by node-specific tendencies, but also by latent similarities and by structured forms of proximity, such as geographical and economic closeness. Latent position models (LPMs) offer a convenient framework to this end by embedding countries into a latent space where proximity indicates a higher probability of interaction. However, many existing LPMs do not explicitly include information on countries' physical propinquity. To address this limitation, we propose here a Bayesian spatially informed latent position model. This framework accounts for node heterogeneity, modeling their propensity to import and export, combining it with a spatially informed characterization of the latent space. Latent positions are provided with a flexible marginal prior, allowing the data to determine the relevance of geographical proximity between countries within each dimension of the latent space. We implement a custom Markov chain Monte Carlo (MCMC) algorithm to efficiently handle the computational complexity of our approach. Finally, the model's performance is validated through simulation studies, considering varying levels of network granularity, density, and connectivity. We apply the proposed model to global migration data to characterize the latent dynamics governing migration flows.
Causal Bayesian nonparametric mixture model for missing data in per-treatment variables
Dafne Zorzetto · Biostatistics and Data Science Institute, Brown University