2026 Stata Sociology Virtual Symposium

20 October 2026

What is the Stata Sociology Virtual Symposium?

The 2026 Stata Sociology Virtual Symposium is a meeting of researchers in sociology from around the world discussing current theory and applied methods using Stata. The proceedings consist of invited talks by top Stata users in a virtual platform that allows you to experience this one-day event from wherever you are. The symposium is organized by StataNordic and Metrika Consulting.

2026 Speakers

Ben

Ben Jann

University of Bern

Maarten

Maarten Buis

University of Konstanz

Trenton

Trenton D. Mize

Purdue University

Emily

Emily Hannum

University of Pennsylvania

Tim

Tim Liao

University of Illinois

Agenda

Times shown in CEST, EDT, and CDT
11:00 CEST 05:00 EDT 04:00 CDT

Broadcast start and Welcome

11:15 CEST 05:15 EDT 04:15 CDT

Drawing maps in Stata using the geoplot command

Ben Jann, University of Bern

Abstract
geoplot is a powerful Stata command for drawing maps from shape files and other datasets. Multiple layers of elements such as regions, borders, lakes, roads, labels, and symbols can be freely combined and the look of elements (e.g. their color) can be varied depending on the values of variables. Compared to previous solutions in Stata, geoplot provides more user convenience, more functionality, and more flexibility. In this talk I will give an overview of the command and illustrate its use with examples.
12:00 CEST 06:00 EDT 05:00 CDT

Break

12:15 CEST 06:15 EDT 05:15 CDT

Agent based models in Mata: Modelling aggregate processes

Maarten Buis, University of Konstanz

Abstract
An Agent Based Model (ABM) is a simulation in which agents that each follow simple rules interact with one another and thus produce an often surprising outcome at the macro level. The purpose of an ABM is to explore mechanisms through which actions of the individual agents add up to a macro outcome by varying the rules that agents have to follow or varying with whom the agent can interact (for example, varying the network). These models have many applications, like the study of segregation of neighborhoods, the adoption of new technologies, or the spread of a disease. In this talk I will discuss the -abm- package. This package contains a number of Mata classes intended to help manage various aspects of an Agent Based model. Implementing a new ABM will always require that person developing the ABM does some programming, but many tasks will be similar across ABMs. For example, in many ABMs the agents live on a square grid (like a chessboard), and can only interact with their neighbours. The abm_grid class, which is part of the -abm- package contains a set of Mata functions that will do those tasks. Alternatively, the agents could live in a network, and the abm_nw class is there to handle that case. The talk will begin with a brief introduction to Mata and classes in Mata and than discuss how to use the classes in the -abm- package to create a ABM by going through various examples.
13:00 CEST 07:00 EDT 06:00 CDT

Long Break

14:00 CEST 08:00 EDT 07:00 CDT

Comparing Effects Within and Across Models Using Marginal Effects

Trenton D. Mize, Purdue University

Abstract
Comparing effects is a common task for the applied data analyst. For example, tests of interaction involve comparing the effects of one variable at multiple levels of another variable. Comparisons of effect sizes within models involve quantifying each focal variable?s effect and testing their equality. Cross-model comparisons are also common, for example when comparing processes across different groups. Tests of attenuation like mediation similarly involve comparisons across multiple models. Despite their ubiquity, such tests have challenges: coefficients that don?t quantify effects in the metric of interest, rescaling of coefficients in nonlinear/categorical models, and predictors on differing metrics. In this presentation, I detail the statistical underpinnings of within- and across-model comparisons, present a marginal effects framework that allows for comparisons across most any regression model or predictor type, and show how to automate these tests using the new suest2, mecompare, meinequality, and totalme commands. Compared to existing software, the new commands can be used for: (1) single or multiple models; (2) single-level, multilevel, and longitudinal models; (3) any amount of change for continuous predictors; (4) summary measures for nominal and ordinal variables; (5) comparisons across variables on different metrics; (6) comparisons across different model types; and (7) custom and nonstandard tests of effects. I demonstrate the utility of the approach and the new commands using publicly available social science data.
14:45 CEST 08:45 EDT 07:45 CDT

Break

15:00 CEST 09:00 EDT 08:00 CDT

Approximating Childrens Potential Environmental Exposures with Cross-Sectional Surveys in Stata

Emily Hannum, University of Pennsylvania

Abstract
Messy matching across space and time with MICS and EM-DAT Childrens environmental exposures accumulate across the life course. Yet longitudinal, geocoded data that could capture cumulative exposures are scarce in low- and middle-income countries (LMICs), where most of the world?s children live. More commonly available in such settings are cross-sectional surveys one-time-point snapshots of each child. This presentation uses a cross-national survey of children in LMICs to demonstrate how to approximate potential environmental exposure: aligning each childs life course with dated environmental hazard events, summing exposure over meaningful developmental windows, and assessing the reasonableness of the approach. Matching survey data to external hazard records can be imprecise. Spatial matching can be messy if place names are labeled inconsistently, geocoding is coarse or uneven in scope, or assumptions about where a child has lived must be made. Temporal matching can be messy if caregiver-reported birth dates and the start and end dates of hazard events are inexact. Approximate exposures are cumulative but potential rather than observed. The presentation will discuss steps to match round-six Multiple Indicator Cluster Survey (MICS6) data from seven Asian LMICs to hazard data from the Emergency Events Database (EM-DAT). The product is a child-level panel that aggregates exposure over the first 1,000 days, in recent years, and in the interval between. The presentation will also demonstrate an option available in MICS6 for assessing a central threat to this design: the assumption that children have not moved. The steps outlined here facilitate research on environmental exposures and their implications for children in the absence of geocoded longitudinal data. Adding more detailed location histories to cross-sectional surveys could extend the approach to longer windows and to adult populations. Empirical paper used in this illustration: Zhang, Yujie, Jere R. Behrman, Emily Hannum, Minhaj Uddin Mahmud, and Fan Wang. 2025. Are Natural Disasters Disastrous for Education? Evidence from Seven Asian Countries. Report No. 1492. Asian Development Bank Institute. https://doi.org/10.56506/RYSP7608.
15:45 CEST 09:45 EDT 08:45 CDT

Break

16:00 CEST 10:00 EDT 09:00 CDT

xtvfreg: A Stata Command for Modeling Mean and Variance in Panel Data

Tim Liao, University of Illinois

Abstract
Standard panel data models assume homoscedasticity, estimating only how covariates shift the mean of an outcome while treating variance as constant across observations. This assumption can obscure important dimensions of inequality, particularly when the dispersion of outcomes-not just their average level-differs systematically across groups or individuals. This presentation introduces xtvfreg, a user-written Stata command that jointly estimates mean and variance equations for panel data using an iterative weighted generalized least squares (GLS) procedure. Building on the framework proposed and demonstrated for answering a substantive question in Mooi-Reci and Liao (2025, European Sociological Review), xtvfreg allows researchers to model not only who has higher or lower outcomes (via the mean equation) but also who experiences greater or lesser dispersion in those outcomes (via the variance equation), separately by group when necessary. The presentation demonstrates the commands syntax, walks through its iterative estimation algorithm, and provides a step-by-step application using the NLSY womens wage panel data. Attendees will leave with the knowledge to implement xtvfreg in their own research and to think more carefully about variance as a substantive outcome of interest, not merely a nuisance parameter..
16:45 CEST 10:45 EDT 09:45 CDT

Adjourn

Register

Registration is free. Seats are limited. You will receive a Zoom link before the symposium starts. Registration closes on October 16.

Name: *
Email: *
Affiliation: