What it is
A survey experiment embeds random assignment directly into a questionnaire. Different respondents receive different versions of a question, a piece of information, or a described scenario, and the difference in their responses estimates the causal effect of the variation. Because assignment is random, the researcher obtains clean identification without a separate field intervention — the survey is the experiment.
The logic is the same as a randomised controlled trial: random assignment ensures that arms differ only in what they were shown, so any average difference in outcomes is attributable to the manipulation. Survey experiments are now the workhorse method for studying information provision, framing, question-wording effects, and stated decisions over described scenarios.
When to use it
Survey experiments are appropriate when the quantity of interest is the causal effect of information, framing, or context on beliefs, attitudes, or stated choices — and when fielding an external intervention is either impossible or unnecessary to answer the question.
Typical applications include: testing whether correcting misperceptions about a policy parameter changes support for the policy (Kuziemko, Norton, Saez & Stantcheva, 2015; Cruces, Perez-Truglia & Tetaz, 2013; Haaland, Roth & Wohlfart, 2023); estimating how the description of a beneficiary affects willingness to donate or support redistribution; measuring whether question order or wording moves reported behaviour (a method-validation question); and isolating the causal effect of a single attribute — gender, caste, religion, age — by randomising it inside a described vignette.
A note on scope. Survey experiments in the sense used here are within-instrument manipulations — the experimental variation is delivered inside the questionnaire itself. Field experiments that deliver an information intervention outside the survey (e.g., Jensen, 2010; Dhar, Jain & Jayachandran, 2022) share the identification logic but face different design constraints around take-up, spillover, and attrition, and belong to the field-RCT literature rather than this guide.
Survey experiments are less appropriate when the outcome of interest is realised behaviour rather than a stated response. Stated outcomes systematically over- or under-estimate behavioural effects in domains with social-desirability content (Pager & Quillian, 2005); treat survey-experimental effects as estimates of belief or attitude shifts, not predictions of action, unless a behavioural complement is appended (Hainmueller, Hangartner & Yamamoto, 2015).
How it works
The survey instrument contains a manipulation — a randomly assigned variation in question content, stimulus, or context. The manipulation can take several forms:
Information provision. A randomly chosen subset of respondents is shown a fact, statistic, or message before answering a question; the other arm is not. The difference in responses estimates the effect of the information on beliefs or stated preferences.
Question wording / framing. Different respondents receive the same question with different wording — for instance, a policy described as having a “90% survival rate” versus a “10% mortality rate.” Average response differences measure the framing effect (Tversky & Kahneman, 1981; Schwarz, 1999).
Vignette manipulation. Respondents are presented with a described scenario — a hiring decision, a loan application, a policy choice — in which one or more characteristics are randomly varied. For single-vignette manipulations with one or two varied attributes, this guide applies. For multi-attribute vignettes with several factors crossed at once, see the Factorial Survey / Vignette Experiments guide; for forced-choice profile comparisons, see the Conjoint / Discrete Choice Analysis guide.
Split-sample design. The full sample is randomly divided into arms that receive different questionnaire modules or different question orderings. This is used to test question-order effects, to evaluate alternative measurement instruments, or to vary the salience of a concept before a key outcome question.
After data collection, the analysis compares mean responses across arms using OLS regression with treatment indicators. Heteroskedasticity-robust (HC2) standard errors are standard; cluster-robust SEs are required when the manipulation varies at a level coarser than the respondent (enumerator, village, household).
Identification assumptions. The estimator is the ATE on the assigned-respondent population, and it is unbiased under:
- Random assignment. Guaranteed by design; verified by balance checks and a joint F-test.
- SUTVA / no interference between respondents. Plausible for cross-sectional online or F2F surveys with isolated respondents; non-trivial for network samples, panels, or settings where respondents discuss the survey.
- No carryover within respondent. Relevant for within-subject and order experiments; addressed by between-subject designs or by counterbalancing and respondent fixed effects.
- No pre-treatment-question contamination. Baseline questions asked before the manipulation can sensitise respondents to the construct and shift their responses to the treatment (Krupnikov & Levine, 2014). Where possible, draw covariates from a pre-survey wave or administrative records rather than from items asked immediately before the manipulation.
- No differential attrition between arms. A longer information vignette can change the drop-out rate; track completion by arm and bound the ATE (Lee, 2009) when arms attrit differently.
Key decisions
Number of arms. A two-arm design (treatment vs. control) is the minimum and gives the highest per-comparison power for a given total sample. Adding arms multiplies comparisons and reduces per-arm n. As an anchor: for a 0.2 SD ATE on a continuous outcome, n ≈ 393 per arm gives 80% power at α = 0.05; for a 5pp ATE on a binary outcome at baseline 50%, n ≈ 1,570 per arm. Survey experiments routinely under-power because researchers anchor on psychology-style samples (n = 100–300 per arm) and end up only able to detect very large effects.
Embedding vs. standalone. An embedded manipulation lives inside a larger instrument and is harder for respondents to read as experimental; a standalone task makes the manipulation explicit and admits richer comprehension checks. The trade-off is not only demand effects: embedding the manipulation deep in a long instrument compounds order effects and respondent fatigue, while embedding it too early invites pre-treatment-question contamination. Rule of thumb: place the manipulation late enough that no intervening item primes the outcome construct, but early enough that fatigue has not degraded attention.
Within-person vs. between-person variation. Most survey experiments vary the stimulus between respondents to avoid contamination and demand effects. Within-subject designs — where the same respondent sees multiple versions — buy precision through respondent fixed effects but require counterbalancing, and they risk respondents inferring the manipulation. If within-subject, the analytic spec must include respondent fixed effects and cluster SEs at the respondent level; the between-subject estimator is wrong.
Outcome measure. Pre-specify the primary outcome and ensure it is directly downstream of the manipulation. If the manipulation is information about returns to education, the outcome should be beliefs about returns or enrolment intentions — not a distally related attitude. Pre-register the manipulation, primary outcome, multiple-testing rule, subgroup interactions, and attention-check exclusion rule before fielding (Olken, 2015) on AsPredicted, OSF, or the AEA RCT Registry. For information experiments, consider pairing the stated outcome with a behavioural complement — a real donation, a clickthrough, an incentivised choice — to validate stated against revealed (Hainmueller, Hangartner & Yamamoto, 2015).
Balance checks. Run per-variable regressions for diagnostic detail, then a single joint F-test (regressing the treatment indicator on all baseline covariates) as the primary balance verdict. With k covariates, expect roughly 5% of per-variable tests to flag at p < 0.05 by chance; the joint F-test absorbs this. Below n ≈ 200 per arm, imbalance is meaningfully likely even under correct randomisation.
Manipulation and attention checks. Build in a manipulation check (a question after the stimulus asking the respondent to recall it — e.g., “What number did the previous screen show?”) and an attention check / instructional manipulation check (an IMC: a benign question whose prompt instructs the respondent to ignore the visible response options; Berinsky, Margolis & Sances, 2014). Pre-register whether inattentive respondents are dropped — dropping post-randomisation breaks ITT and creates post-treatment bias. Standard practice is to report ITT as primary and per-protocol as secondary.
Covariate adjustment. Use the Lin (2013) estimator — covariates demeaned and interacted with treatment — not naive ANCOVA. Lin’s estimator is unbiased in finite samples and weakly dominates the unadjusted estimator on precision; naive ANCOVA (treatment plus uninteracted covariates) is biased toward zero in small samples and can hurt precision. R’s estimatr::lm_lin() implements it directly; in Python, mean-centre covariates and interact them with treatment.
Generalisability across samples. Modern survey experiments are mostly fielded on online convenience panels (MTurk, Prolific, Lucid, CloudResearch, Dynata) rather than probability samples. The replication evidence converges on the conclusion that treatment-effect signs and ranks generalise well across sample sources, but levels and heterogeneity less so (Mullinix, Leeper, Druckman & Freese, 2015; Coppock, Leeper & Mullinix, 2018; Coppock, 2019; Coppock & McClellan, 2019; Berinsky, Huber & Lenz, 2012). Report which panel was used, the screener and quota structure, and treat extrapolation to other populations as a substantive claim, not a default.
Caveats & common mistakes
Demand effects. The modern empirical update is that demand effects are typically small in online survey experiments, even when respondents are told the hypothesis (Mummolo & Peterson, 2019). Where the concern is load-bearing — for example, when the outcome is itself a sensitive belief or behaviour — use the de Quidt, Haushofer & Roth (2018) bounding design: include explicit experimenter-demand treatments that signal the hypothesised direction, and use the resulting effects to bound the true ATE. In face-to-face administration the concern is sharper because the enumerator can see the arm and inadvertently signal the expected response; deliver the manipulation via tablet self-administration or audio through headphones to keep the enumerator blind to assignment.
Pre-treatment-question contamination. Baseline questions asked just before the manipulation can sensitise respondents to the construct and shift treatment responses (Krupnikov & Levine, 2014). Where possible, draw baseline covariates from a separate pre-survey wave or from administrative records rather than from items asked within the same instrument immediately before the manipulation.
Inattention and weak manipulations. If respondents do not read or process the stimulus they cannot respond to it. If fewer than 80% in the treatment arm pass the manipulation check, the stimulus is too weak or too long — rewrite. For text-heavy stimuli read aloud by enumerators, variation in delivery is itself noise; tablet self-administration or recorded audio standardises the stimulus across enumerators.
Differential attrition. A longer or harder stimulus can change drop-out. If treatment-arm completion is more than 5 percentage points below control, attrition is differential; report Lee (2009) bounds alongside the point estimate.
Stated vs. revealed. Survey experiments measure stated attitudes and intentions, not realised behaviour. The effect of information on stated support may not transfer to actual voting, adoption, or compliance. Where a behavioural outcome is feasible, append one (a donation, a clickthrough, an incentivised choice) to validate.
Multiple comparisons. Multi-arm and multi-outcome designs generate many comparisons. Specify primary comparisons in advance; correct secondary comparisons for multiple testing — Benjamini-Hochberg FDR for exploratory work, Romano-Wolf FWER for confirmatory secondaries, and Anderson (2008) inverse-covariance-weighted indices for summarising correlated outcomes. See the Multiple Hypothesis Testing Correction guide.
Heterogeneous treatment effects. Subgroup analyses are typically severely under-powered relative to the main effect — a 0.2 SD interaction needs roughly four times the sample of a 0.2 SD main effect (Gelman, 2018). Pre-register subgroup analyses; treat post-hoc subgroup discoveries as hypothesis-generating, not confirmatory. For principled exploration of heterogeneity, use causal-forest / generic-ML methods (Athey & Imbens, 2017).
Mediation. “What’s the mechanism?” usually cannot be answered by adding a mediator regression: identification of the indirect effect requires sequential ignorability, which is typically violated in observational mediator settings (Bullock, Green & Ha, 2010). Credible mediation usually requires a separate experiment that manipulates the proposed mediator directly (Imai, Tingley & Yamamoto, 2013).
Analysis Guide
# 1. Balance check: regress each baseline covariate on the treatment indicator and
# collect the results into one tidy table — per-variable p-values are diagnostic,
# not the primary verdict. With k covariates, expect ~5% to flag at p < 0.05 by
# chance. HC2 SEs are the principled small-sample choice (MacKinnon & White 1985).
import pandas as pd
import numpy as np
import statsmodels.formula.api as smf
from statsmodels.stats.multitest import multipletests
balance_vars = ['age', 'female', 'education', 'log_hh_income']
rows = []
for v in balance_vars:
fit_v = smf.ols(f'{v} ~ treat', data=df).fit(cov_type='HC2')
rows.append({'var': v, 'beta': fit_v.params['treat'],
'se': fit_v.bse['treat'], 'p': fit_v.pvalues['treat']})
print(pd.DataFrame(rows))
# 2. Joint balance test: the primary verdict on randomisation — regress the
# treatment indicator on all baseline covariates and read the F p-value.
# A small p (e.g., < 0.05) flags non-random assignment and requires investigation.
joint = smf.ols('treat ~ age + female + education + log_hh_income',
data=df).fit(cov_type='HC2')
print('Joint F p-value:', joint.f_pvalue)
# 3. Lin (2013) covariate-adjusted ATE: demean covariates and interact each with
# treatment. This estimator is unbiased in finite samples (naive ANCOVA is not)
# and weakly dominates the unadjusted ATE on precision (Lin 2013;
# Egami, Hamilton & Imai 2021). Without covariates, drop the interactions.
for v in ['age', 'female', 'education', 'log_hh_income']:
df[f'{v}_c'] = df[v] - df[v].mean()
fit = smf.ols('outcome ~ treat * (age_c + female_c + education_c + log_hh_income_c)',
data=df).fit(cov_type='HC2')
print(fit.summary())
# 4. Cluster-robust SEs when the manipulation or sampling is clustered (treatment
# by enumerator, village, or household; multiple respondents per cluster).
# Standard OLS SEs understate uncertainty when units within a cluster are
# correlated; cluster-robust SEs are the principled fix.
fit_cl = smf.ols('outcome ~ treat', data=df).fit(
cov_type='cluster', cov_kwds={'groups': df['enumerator_id']})
print(fit_cl.summary())
# 5. Manipulation / attention check and per-protocol estimate: ITT (full sample)
# preserves random assignment and is the primary estimand. Per-protocol (subset
# passing the manipulation check) is reported as secondary — exclusion can
# correlate with potential outcomes and break random assignment.
itt = smf.ols('outcome ~ treat', data=df).fit(cov_type='HC2')
pp = smf.ols('outcome ~ treat',
data=df[df['passed_mcheck'] == 1]).fit(cov_type='HC2')
print('ITT:', itt.params['treat'], 'PP:', pp.params['treat'])
# 6. Pre-registered heterogeneous treatment effect: an interaction isolates whether
# the ATE differs by a pre-specified subgroup. Post-hoc subgroup searches inflate
# Type I error and are hypothesis-generating, not confirmatory.
# Threshold below is illustrative; pre-specify before seeing the data.
df['low_ed'] = (df['education'] <= 6).astype(int)
fit_het = smf.ols('outcome ~ treat * low_ed + age_c + female_c + log_hh_income_c',
data=df).fit(cov_type='HC2')
print(fit_het.summary())
# 7. Multiple-testing correction across secondary outcomes — Benjamini-Hochberg FDR
# is appropriate for exploratory secondaries; pre-register the family of outcomes
# before running, and report adjusted p-values alongside raw.
pvals = [...] # collected from per-outcome treat regressions
reject, padj, *_ = multipletests(pvals, alpha=0.05, method='fdr_bh')
print(pd.DataFrame({'p_raw': pvals, 'p_bh': padj, 'reject': reject})) SurveyCTO / XLSForm
Place the randomisation at the start of the survey, before any item that could be affected by the manipulation, and place demographic / baseline covariates before the randomisation so they cannot be contaminated by the treatment.
For a two-arm design, use a calculate field with once(if(random() < 0.5, 0, 1)) — wrapped in once() so the draw is fixed at first evaluation and not recomputed on every navigation. For k arms, use once(int(random() * k)). Store the result in a field named treatment_arm. The bare once(random()) returns a continuous uniform draw and is not a usable assignment.
| Item | Type | Calculation / Note |
|---|---|---|
treatment_arm | calculate | once(if(random() < 0.5, 0, 1)) (two-arm) or once(int(random() * 3)) (three-arm) |
info_block | note | relevant: ${treatment_arm} = 1 — shown to treatment arm only |
mcheck | select_one | ”What number/figure did the previous screen show?” — asked of all respondents; the manipulation check |
imc | select_one | An attention check whose prompt instructs the respondent to ignore the visible options (Berinsky, Margolis & Sances, 2014) |
outcome | integer/decimal/select_one | The primary outcome, placed immediately after the manipulation check |
For stratified randomisation (by region, gender, or baseline category), generate the assignment in advance using a pre-randomised stratified file and look it up with pulldata() keyed on respondent ID. In-form random() cannot guarantee within-stratum balance for small cells.
For face-to-face administration, enumerator-blinding is critical: deliver the manipulation via tablet self-administration or via headphone audio so the enumerator cannot see which arm the respondent received and cannot signal the expected response.
Keep all arms in the same form file rather than fielding separate form versions — separate versions complicate merging and version control. Always export treatment_arm, the manipulation check, and the attention check to the final dataset.
Reading the output
- The
treatcoefficient is the average treatment effect on the assigned-respondent population. Report the point estimate, HC2 robust SE, 95% CI, and p-value together — CI width is the more honest summary of precision than the p-value alone. - For a binary outcome it is a percentage-point difference; for a continuous outcome it is a scale-point difference.
- Joint balance F p-value > 0.10 supports random assignment. A p-value < 0.05 flags non-random assignment — investigate the form logic before interpreting the ATE.
- If the
treatcoefficient changes by more than 25% of its uncontrolled value or by more than one robust SE when Lin-style covariates are added, inspect balance and check whether n < 200 per arm. - If the share passing the manipulation check is below 80% in the treatment arm, the stimulus is too weak or too long — rewrite it; do not patch the analysis.
- If treatment-arm attrition exceeds control-arm attrition by more than 5 percentage points, attrition is differential — report Lee (2009) bounds alongside the point estimate.
- If the effect direction reverses under a de Quidt-style explicit-demand check, demand effects are first-order — report bounded estimates rather than the raw ATE.
References
Anderson, M. L. (2008). Multiple inference and gender differences in the effects of early intervention: A reevaluation of the Abecedarian, Perry Preschool, and Early Training projects. Journal of the American Statistical Association, 103(484), 1481–1495. https://doi.org/10.1198/016214508000000841
Berinsky, A. J., Huber, G. A., & Lenz, G. S. (2012). Evaluating online labor markets for experimental research: Amazon.com’s Mechanical Turk. Political Analysis, 20(3), 351–368. https://doi.org/10.1093/pan/mpr057
Berinsky, A. J., Margolis, M. F., & Sances, M. W. (2014). Separating the shirkers from the workers? Making sure respondents pay attention on self-administered surveys. American Journal of Political Science, 58(3), 739–753. https://doi.org/10.1111/ajps.12081
Bullock, J. G., Green, D. P., & Ha, S. E. (2010). Yes, but what’s the mechanism? (Don’t expect an easy answer). Journal of Personality and Social Psychology, 98(4), 550–558. https://doi.org/10.1037/a0018933
Coppock, A. (2019). Generalizing from survey experiments conducted on Mechanical Turk: A replication approach. Political Science Research and Methods, 7(3), 613–628. https://doi.org/10.1017/psrm.2018.10
Coppock, A., Leeper, T. J., & Mullinix, K. J. (2018). Generalizability of heterogeneous treatment effect estimates across samples. Proceedings of the National Academy of Sciences, 115(49), 12441–12446. https://doi.org/10.1073/pnas.1808083115
Coppock, A., & McClellan, O. A. (2019). Validating the demographic, political, psychological, and experimental results obtained from a new source of online survey respondents. Research & Politics, 6(1). https://doi.org/10.1177/2053168018822174
Cruces, G., Perez-Truglia, R., & Tetaz, M. (2013). Biased perceptions of income distribution and preferences for redistribution: Evidence from a survey experiment. Journal of Public Economics, 98, 100–112. https://doi.org/10.1016/j.jpubeco.2012.10.009
de Quidt, J., Haushofer, J., & Roth, C. (2018). Measuring and bounding experimenter demand. American Economic Review, 108(11), 3266–3302. https://doi.org/10.1257/aer.20171330
Druckman, J. N., & Green, D. P. (Eds.). (2021). Advances in experimental political science. Cambridge University Press.
Haaland, I., Roth, C., & Wohlfart, J. (2023). Designing information provision experiments. Journal of Economic Literature, 61(1), 3–40. https://doi.org/10.1257/jel.20211658
Hainmueller, J., Hangartner, D., & Yamamoto, T. (2015). Validating vignette and conjoint survey experiments against real-world behavior. Proceedings of the National Academy of Sciences, 112(8), 2395–2400. https://doi.org/10.1073/pnas.1416587112
Hainmueller, J., Hopkins, D. J., & Yamamoto, T. (2014). Causal inference in conjoint analysis: Understanding multidimensional choices via stated preference experiments. Political Analysis, 22(1), 1–30. https://doi.org/10.1093/pan/mpt024
Imai, K., Tingley, D., & Yamamoto, T. (2013). Experimental designs for identifying causal mechanisms. Journal of the Royal Statistical Society: Series A, 176(1), 5–51. https://doi.org/10.1111/j.1467-985X.2012.01032.x
Krupnikov, Y., & Levine, A. S. (2014). Cross-sample comparisons and external validity. Journal of Experimental Political Science, 1(1), 59–80. https://doi.org/10.1017/xps.2014.7
Kuziemko, I., Norton, M. I., Saez, E., & Stantcheva, S. (2015). How elastic are preferences for redistribution? Evidence from randomized survey experiments. American Economic Review, 105(4), 1478–1508. https://doi.org/10.1257/aer.20130360
Lee, D. S. (2009). Training, wages, and sample selection: Estimating sharp bounds on treatment effects. Review of Economic Studies, 76(3), 1071–1102. https://doi.org/10.1111/j.1467-937X.2009.00536.x
Lin, W. (2013). Agnostic notes on regression adjustments to experimental data: Reexamining Freedman’s critique. Annals of Applied Statistics, 7(1), 295–318. https://doi.org/10.1214/12-AOAS583
MacKinnon, J. G., & White, H. (1985). Some heteroskedasticity-consistent covariance matrix estimators with improved finite sample properties. Journal of Econometrics, 29(3), 305–325. https://doi.org/10.1016/0304-4076(85)90158-7
Mullinix, K. J., Leeper, T. J., Druckman, J. N., & Freese, J. (2015). The generalizability of survey experiments. Journal of Experimental Political Science, 2(2), 109–138. https://doi.org/10.1017/XPS.2015.19
Mummolo, J., & Peterson, E. (2019). Demand effects in survey experiments: An empirical assessment. American Political Science Review, 113(2), 517–529. https://doi.org/10.1017/S0003055418000837
Mutz, D. C. (2011). Population-based survey experiments. Princeton University Press.
Olken, B. A. (2015). Promises and perils of pre-analysis plans. Journal of Economic Perspectives, 29(3), 61–80. https://doi.org/10.1257/jep.29.3.61
Pager, D., & Quillian, L. (2005). Walking the talk? What employers say versus what they do. American Sociological Review, 70(3), 355–380. https://doi.org/10.1177/000312240507000301
Schwarz, N. (1999). Self-reports: How the questions shape the answers. American Psychologist, 54(2), 93–105. https://doi.org/10.1037/0003-066X.54.2.93
Sniderman, P. M., & Grob, D. B. (1996). Innovations in experimental design in attitude surveys. Annual Review of Sociology, 22, 377–399. https://doi.org/10.1146/annurev.soc.22.1.377
Tversky, A., & Kahneman, D. (1981). The framing of decisions and the psychology of choice. Science, 211(4481), 453–458. https://doi.org/10.1126/science.7455683