Fredrik Sävje
Fredrik Savje
- Professor
- Econometrics
- Uppsala University
Working papers
-
Network experiments are used throughout the social and medical sciences to investigate causal effects under the presence of interference. While a large body of work has developed improved statistical procedures, the fundamental limits of statistical estimation in these settings is less well understood. In this paper, we develop and investigate a design-based theory of minimax risk for network experiments under an arbitrary neighborhood interference model. Our notion of minimax risk describes the optimal precision among all statistical procedures for investigating a particular causal effect on the observed interference network. We show that the minimax risk is a function of the corresponding conflict graph, which captures inherent unobservability of estimand-relevant potential outcomes given the observed interference network. Our main contribution is a series of upper and lower bounds on the minimax rate in terms of local and global connectivity properties of the conflict graph. To illustrate their utility, we apply these general results to obtain minimax analyses for two commonly studied effects: the direct treatment effect and global average treatment effect.
-
We describe a new family of experimental designs that extends the principle of stratified randomization to settings with continuous, constrained multivariate, and other irregular treatment spaces. Our approach is to first match units into homogeneous groups, then use Monte Carlo couplings to assign within-group treatments to be highly dispersed over the treatment space. We show that ensuring similar units receive dissimilar treatments improves estimation efficiency. The efficiency gains are proportional to the product of dispersion and match quality, where dispersion measures how spread out the assignments are relative to independent randomization. We develop a new spectral analysis showing how efficiency depends on alignment between the smoothness and shape of the estimator's influence function and the coupling's principal directions. We illustrate these designs with examples from development, behavioral, and labor economics. In particular, our empirical application uses data from a real experiment allocating savings monitors using their position within village social networks.
-
Researchers use interference models based on exposure mappings to facilitate estimation of causal effects in randomized experiments with interference. To test the veracity of such models, researchers can use specification tests that aim to detect departures from the stipulated model. However, existing tests suffer from poor power and are often unable to detect important model violations. The main result in this paper is to show that the specification testing problem for exposure mapping models is inherently difficult, and the poor power of existing tests is inescapable. In particular, the worst-case Type I and Type II error rates must sum to one for any specification test of such models, ruling out the existence of a uniformly consistent test. This is the worst-case overall error rate achieved by a naive test that discards all data and arbitrarily rejects the null at random; the testing problem is in this sense impossible. This negative result holds true for all exposure mappings, all sample sizes, for uniformly bounded outcomes, and for alternatives that are maximally separated from the null. While some tests can detect some type of departures from the null model, there will always be relevant departures from the null that are undetectable. Informative specification tests must therefore restrict the alternative model against which they seek to attain power for, beyond the restrictions imposed by the exposure mappings alone. We illustrate this by providing a uniformly consistent test for differentiating no-interference from a network-linear-in-means model.
-
We describe a new design-based framework for drawing causal inference in randomized experiments. Causal effects in the framework are defined as linear functionals evaluated at potential outcome functions. Knowledge and assumptions about the potential outcome functions are encoded as function spaces. This makes the framework expressive, allowing experimenters to formulate and investigate a wide range of causal questions. We describe a class of estimators for estimands defined using the framework and investigate their properties. The construction of the estimators is based on the Riesz representation theorem. We provide necessary and sufficient conditions for unbiasedness and consistency. Finally, we provide conditions under which the estimators are asymptotically normal, and describe a conservative variance estimator to facilitate the construction of confidence intervals for the estimands.
Recent publications
-
Journal of the American Statistical Association (2026), in print.
Unbiased and consistent variance estimators generally do not exist for design-based treatment effect estimators because experimenters never observe more than one potential outcome for any unit. The problem is exacerbated by interference and complex experimental designs. Experimenters must accept conservative variance estimators in these settings, but they can strive to minimize conservativeness. In this paper, we show that the task of constructing a minimally conservative variance estimator can be interpreted as an optimization problem that aims to find the lowest estimable upper bound of the true variance given the experimenter's risk preference and knowledge of the potential outcomes. We characterize the set of admissible bounds in the class of quadratic forms, and we demonstrate that the optimization problem is a convex program for many natural objectives. The resulting variance estimators are guaranteed to be conservative regardless of whether the background knowledge used to construct the bound is correct, but the estimators are less conservative if the provided information is reasonably accurate. Numerical results show that the resulting variance estimators can be considerably less conservative than existing estimators, allowing experimenters to draw more informative inferences about treatment effects.
-
Observational Studies (2025), 11(1), 3–16.
We argue that randomized controlled trials (RCTs) are special even among settings where average treatment effects are identified by a nonparametric unconfoundedness assumption. This claim follows from two results of Robins and Ritov (1997): (1) with at least one continuous covariate control, no estimator of the average treatment effect exists which is uniformly consistent without further assumptions, (2) knowledge of the propensity score yields a uniformly consistent estimator and honest confidence intervals that shrink at parametric rates with increasing sample size, regardless of how complicated the propensity score function is. We emphasize the latter point, and note that successfully-conducted RCTs provide knowledge of the propensity score to the researcher. We discuss modern developments in covariate adjustment for RCTs, noting that statistical models and machine learning methods can be used to improve efficiency while preserving finite sample unbiasedness. We conclude that statistical inference has the potential to be fundamentally more difficult in observational settings than it is in RCTs, even when all confounders are measured.
-
Journal of the American Statistical Association (2024), 119(548), 2934–2946.
The design of experiments involves a compromise between covariate balance and robustness. This paper provides a formalization of this trade-off and describes an experimental design that allows experimenters to navigate it. The design is specified by a robustness parameter that bounds the worst-case mean squared error of an estimator of the average treatment effect. Subject to the experimenter’s desired level of robustness, the design aims to simultaneously balance all linear functions of potentially many covariates. Less robustness allows for more balance. We show that the mean squared error of the estimator is bounded in finite samples by the minimum of the loss function of an implicit ridge regression of the potential outcomes on the covariates. Asymptotically, the design perfectly balances all linear functions of a growing number of covariates with a diminishing reduction in robustness, effectively allowing experimenters to escape the compromise between balance and robustness in large samples. Finally, we describe conditions that ensure asymptotic normality and provide a conservative variance estimator, which facilitate the construction of asymptotically valid confidence intervals.
Software
-
Julia package with a fast implementation of the Gram-Schmidt Walk for balancing covariates in randomized experiments (also R wrapper).
-
R package with tools for distance metrics.
-
Quick Generalized Full Matching in R.
-
Quick Threshold Blocking in R.
-
C library for size-constrained clustering.
Last updated August 17, 2026.