Metter. / Mixtapes / Methods Mixtape / Survey & Elicitation Methods

08 · Survey & Elicitation Methods

Factorial Survey / Vignette Experiments

A survey-experimental method in which respondents read short narrative descriptions of people, situations, or policies — with key attributes assigned at random — and report a judgment, an attitude, or a behavioural intention. Randomisation identifies the causal effect of each attribute.


What it is

A factorial survey experiment — also called a factorial vignette experiment (Rossi & Anderson, 1982; Auspurg & Hinz, 2015) — presents respondents with short narrative descriptions of people, situations, or policies, with key attributes within the description assigned at random across respondents or across vignettes shown to the same respondent. Respondents read or hear the vignette and report a judgment, attitude, or behavioural intention. Because the attributes are randomised, differences in responses across vignette versions identify the causal effect of each attribute on the judgment — the same Average Marginal Component Effect (AMCE) estimand used in conjoint analysis (Hainmueller, Hopkins & Yamamoto, 2014).

What this guide covers and what it does not. The terms “stated preference” and “hypothetical scenario” are used loosely across fields. This guide covers narrative-vignette designs specifically. For the related methods:

  • Tabular trade-off tasks with multi-attribute profiles → the Conjoint / Discrete Choice guide.
  • Monetary valuation of non-market goods using payment scenarios → the Contingent Valuation guide.
  • Calibration vignettes for correcting cross-group response-scale heterogeneity → the Anchoring Vignettes guide (same word, different method).
  • Question-wording / framing manipulations and split-sample experiments embedded in a survey → the Survey Experiments / Split-Sample Design guide.

What is “stated” here is stated, not behavioural. A factorial vignette measures the causal effect of an attribute on a stated response — not on an actual behaviour. Stated judgments correlate imperfectly with real behaviour, and the gap can be large for socially sensitive outcomes. Pager and Quillian (2005) compared employer interview statements about discrimination to behavioural audit-study results and found that employers stated non-discriminatory intentions but discriminated in the actual hiring process. Plan the design with this gap in mind, and discuss it in the results (more in Caveats).

When to use it

Factorial vignettes are appropriate when the research question concerns how specific characteristics of a person, situation, or policy affect a judgment — and when direct questioning would be biased by social desirability, sensitivity, or lack of an obvious counterfactual to compare against.

Examples across fields:

  • Hiring and HR research — varying applicant characteristics (name, gender, education, employment gap) to test for evaluator bias
  • UX research — varying product features in a use-case scenario to test stated preference for design changes
  • Public health — measuring stigma around disease status, mental health, or disability by varying the characteristics of a vignette patient
  • Policy support research — varying implementation details (eligibility, cost, who delivers it) to identify which features drive support
  • Marketing — varying brand, price, and use-context in product narratives
  • Sociology and political science — eliciting social norms (e.g., what is “fair” in distribution scenarios) and testing how political cues shape evaluation
  • Development economics — discrimination in lending, schooling, and service delivery; norms around domestic work, child marriage, and labour supply

The inferential logic is the same as a field audit study (Bertrand & Mullainathan, 2004): hold everything constant except the attribute of interest, randomise that attribute, attribute differences in outcomes to the attribute. The difference is that a factorial vignette asks for a stated response, while an audit study observes a behavioural one. The audit study has stronger external validity; the vignette is cheaper, more flexible, and allows interior attributes (income, occupation) and many-attribute factorials that audits cannot run.

The method is less suitable when: a direct question would produce honest answers (the indirection adds noise without value); the attributes cannot be plausibly varied without straining realism; or the research requires revealed rather than stated behaviour and an audit-style field design is feasible.

How it works

The researcher identifies the judgment to be measured (“Would you hire this person?”, “Is this behaviour acceptable in your community?”, “Would you support this policy?”) and specifies the vignette attributes hypothesised to affect it.

Design choice — between vs within respondent. In a between-person design, each respondent sees one vignette. In a within-person design, each respondent evaluates multiple vignettes (typically 3–6) drawn from the factorial. Within-person designs are more efficient (each respondent is their own control) but introduce carryover and order risks (see Caveats).

Estimator. OLS with the judgment as outcome and dummies for each vignette attribute level. The coefficient on a level dummy is the AMCE — the average causal effect of changing that attribute level relative to the omitted baseline, averaged over the experimental joint distribution of the other attributes and over respondents (Hainmueller, Hopkins & Yamamoto, 2014). Cluster standard errors at the respondent level when the design is within-person.

Identification assumptions. Five conditions are required for the AMCE to identify a causal effect:

  1. Random assignment of attribute levels — each attribute level is independent of every other attribute and of respondent characteristics. Guaranteed by a randomised factorial design; verifiable by regressing each attribute on the others and confirming coefficients are near zero.
  2. No profile-order (carryover) effects — for within-person designs, the judgment on vignette t does not depend on vignettes seen earlier. Randomise vignette order across respondents and include a vignette-order indicator in the regression as a diagnostic (Hainmueller, Hopkins & Yamamoto, 2014; Dafoe, Zhang & Caughey, 2018).
  3. No interference across respondents — respondents do not discuss vignettes with each other during the survey.
  4. Profile equivalence — respondents read each attribute the same way regardless of the other attributes in the vignette. Violations include respondents inferring an unstated attribute from a presented one. Dafoe et al. (2018) show that information equivalence — making sure attribute meanings are constant across profiles — is a separate condition beyond randomisation.
  5. Atypical profiles handled honestly — dropping logically impossible combinations changes the estimand. Either restrict to realistic combinations from the start (and report the restriction) or randomise across the full factorial and discuss atypicality in the results.

Where the estimand lives. The AMCE is an average over the experimental joint distribution of attributes — typically uniform — which is rarely the real-world joint distribution. If the policy question is “what would the average response be in the field where attribute X is concentrated at level x?”, post-stratify or weight the AMCEs to the target distribution (de la Cuesta, Egami & Imai, 2022). For subgroup comparisons across respondent characteristics, Marginal Means are baseline-free and the recommended estimand (Leeper, Hobolt & Tilley, 2020).

Key decisions

Vignette format. Attribute values are embedded in a short narrative — for example, “Consider a 28-year-old applicant with a degree in computer science and three years of experience, applying for a junior developer role” — rather than presented as a tabular list. The narrative format is more natural in oral and field-survey settings and is better suited to social judgments than to preference ranking. Use a single canonical term for the vignette throughout your instrument (we use “vignette” in this guide).

Number of attributes and levels. Three to six attributes is typical. More attributes allow more causal questions per instrument but increase cognitive load and risk generating implausible combinations. Each attribute should carry genuine theoretical weight — adding attributes to bulk out the design degrades data quality (Auspurg & Hinz, 2015, ch. 5).

Within-person vs between-person. Within-person designs (each respondent evaluates 3–6 vignettes) are about 3–5× more efficient than between-person designs for a given total sample (Stefanelli & Lukac, 2020). The cost is carryover, fatigue, and the possibility that respondents detect the design and respond strategically. Four to six vignettes per respondent is the standard compromise; beyond about eight, response quality declines (Sauer et al., 2011). For sensitive topics or short instruments, prefer between-person.

Deck construction. Two routes:

  • Full factorial random draws — with k attributes and L levels each, the universe has L^k unique vignettes. For small designs (e.g., 4 attributes × 3 levels = 81), draw vignettes uniformly at random from the full set. This is the simplest approach and gives orthogonality in expectation.
  • D-efficient fractional factorial — for larger designs or when you need restricted combinations, generate a D-efficient design using R’s AlgDesign::optFederov() or support.CEs::Lma.design(). This guarantees balance and minimises variance per attribute.

Use a pre-generated CSV of vignettes loaded into the form via pulldata() rather than constructing vignettes in-form — see the SurveyCTO section.

Response scale. Binary scales (acceptable/not, would hire/not) produce easily interpretable differences. Ordinal scales (1–5 agreement, 1–7 likelihood) capture more variation and improve efficiency. Match the scale to how the judgment is actually formed in real decisions.

Sample size. Stefanelli and Lukac (2020) give concrete power calculations for factorial vignettes. For an AMCE of 0.1 standard deviations at 80% power, α = 0.05:

  • Between-person, 4 attributes × 3 levels: approximately 1,600 respondents.
  • Within-person, 5 vignettes per respondent, intra-respondent correlation ρ ≈ 0.3: approximately 320 respondents.
  • Add a multiplicative factor of 1 + (m̄ − 1)ρ for further cluster-sampled designs.

For pre-specified subgroup heterogeneity tests, multiply the base sample by 2–4 to support the interaction term.

Profile-order randomisation and diagnostics. Randomise vignette order across respondents in within-person designs. Include a vignette-order indicator in the regression as a balance check; significant order effects flag carryover that must be reported and ideally controlled for.

Survey weights. If the inference target is a population (rather than a sample average), estimate AMCEs with survey weights. The bare regression averages over the sample distribution; population-relevant claims need weighting.

Caveats & common mistakes

Stated vs revealed — the central caveat. Factorial vignettes measure the effect of attributes on stated responses, not behaviour. The gap can be large for socially sensitive outcomes — Pager and Quillian (2005) found that employer self-reports of non-discrimination did not predict their behavioural discrimination in audit experiments. Report the AMCE as a stated-response effect, not a behavioural prediction. Where possible, validate against an external behavioural measure (administrative data, an audit study sub-sample) and discuss the gap.

Hypothetical bias. Stated responses to hypothetical scenarios are systematically more pro-social, more generous, and more discrimination-free than real behaviour. The bias varies by domain — Murphy et al. (2005) and List & Gallet (2001) find median hypothetical-to-real ratios of 1.3–3× in stated-preference valuation studies. Cheap-talk scripts (Cummings & Taylor, 1999) and consequentiality framing (Vossler, Doyon & Rondeau, 2012) reduce but do not eliminate the bias. Include both where possible: a short note before the vignette block telling respondents that this is hypothetical and asking them to answer as they would in a real decision, plus a statement that results will be used in a real decision-making context.

Demand effects. Respondents who infer the research hypothesis from the vignette structure may adjust their answers in the direction they believe the researcher wants. Mummolo and Peterson (2019) provide the contemporary empirical estimate — demand effects in survey experiments are typically small in magnitude — but de Quidt, Haushofer and Roth (2018) show they can be substantial when the design is transparent. Mitigation: embed the experimental vignettes inside a longer, mixed-topic instrument; vary attributes that are not of interest as decoys; frame questions in third-person rather than directing them at the respondent.

Framing, anchoring, and question form. The same content elicits different responses depending on how it is framed (Tversky & Kahneman, 1981; Kahneman & Tversky, 1979). The question form, scale labels, and surrounding instrument text all shape the answer (Schwarz, 1999). For factorial vignettes specifically, the order and anchoring of attribute presentation within the narrative can shift effects. Pilot the wording carefully and consider counter-balancing attribute presentation order across versions.

Preference construction, not preference revelation. Stated preferences for unfamiliar or complex hypothetical scenarios are often constructed in the moment rather than retrieved from a stable underlying preference (Lichtenstein & Slovic, 1971, 2006). The same respondent can give a different answer to logically equivalent questions framed differently. This is not a bug to fix; it is a property of the elicitation. For familiar judgments (whether to hire someone with a CV that fits the local labour market) the construction problem is mild; for unfamiliar judgments (whether to support a policy the respondent has never considered) it can be severe.

Within-person contamination. When respondents evaluate multiple vignettes, later responses may be affected by earlier ones — contrast effects (judging vignette 2 relative to vignette 1) or consistency effects (aligning later answers with earlier ones). Randomise vignette order across respondents; vary multiple attributes simultaneously (rather than one at a time); test for order effects in the regression. Where order effects are detected, report results both with and without an order control.

Atypical and impossible profiles. Dropping combinations that look implausible breaks the orthogonal factorial design and changes the estimand. Two defensible options: (a) restrict the design to realistic combinations from the start and report the restriction; (b) keep the full factorial and discuss atypicality in the results. Either is acceptable; silently dropping inconvenient combinations is not.

Cognitive piloting is not optional. Run cognitive interviews (think-aloud protocols; Willis, 2005; Beatty & Willis, 2007) with 15–20 respondents from the target population on draft vignettes. The interviews routinely surface incoherent combinations, ambiguous wording, and unintended inferences that the researcher cannot anticipate.

Multiple-testing correction. With many attributes × levels × subgroup interactions, the number of tests can exceed 30. Pre-register a primary contrast for which no correction is applied. For secondary item-level tests within one outcome, Benjamini-Hochberg false-discovery-rate control is acceptable. For cross-subgroup heterogeneity claims, Romano-Wolf step-down (Clarke, Romano & Wolf, 2020 — implemented as rwolf2 in Stata; bootstrap-based in R) is the more defensible standard.

Subgroup comparisons via Marginal Means. AMCEs depend on the baseline level — two analysts picking different baselines can reach opposite conclusions about subgroup heterogeneity. Leeper, Hobolt and Tilley (2020) recommend Marginal Means (the unconditional response rate at each attribute level) for cross-group comparisons. The cregg R package implements both AMCEs and MMs.

Differential treatment of stated-norm questions. When the goal is to elicit a social norm (what is considered appropriate behaviour), the Krupka-Weber (2013) coordination-game method often yields more reliable estimates than direct vignette judgments — respondents are paid to coordinate their answer with what others say, which incentivises norm-relevant reporting. For inferred valuations of behaviour (what others would do), Lusk and Norwood (2009) is an alternative.

Analysis Guide

import pandas as pd
import numpy as np
import statsmodels.formula.api as smf
from statsmodels.stats.multitest import multipletests

# 1. Randomisation balance check — under correct attribute assignment, each
#    attribute should be independent of the others; regress one on the rest
#    and confirm coefficients near zero. A non-trivial coefficient signals a
#    bug in the deck-generation script
bal = smf.ols('C(attr1) ~ C(attr2) + C(attr3) + C(attr4)', data=df).fit()
print(bal.f_test('C(attr2)[T.1]=0, C(attr3)[T.1]=0'))   # joint test

# 2. AMCE — between-person design (one vignette per respondent). HC3 is the
#    small-sample heteroskedasticity-robust SE recommended by Long & Ervin
#    (2000); HC1 (the Stata default) tends to under-cover at small N
fit = smf.ols('judgment ~ C(attr1) + C(attr2) + C(attr3) + C(attr4)',
            data=df).fit(cov_type='HC3')
print(fit.summary())

# 3. AMCE — within-person design (multiple vignettes per respondent).
#    statsmodels cov_type='cluster' is CR1, not CR2; for CR2 small-sample-
#    corrected cluster SEs use pyfixest or the R tab. With many clusters
#    (e.g., >100 respondents) CR1 is acceptable; with few clusters use
#    wild-cluster bootstrap (Cameron, Gelbach & Miller 2008)
fit_cl = smf.ols('judgment ~ C(attr1) + C(attr2) + C(attr3) + C(attr4)',
               data=df).fit(cov_type='cluster',
                            cov_kwds={'groups': df['respondent_id']})
print(fit_cl.summary())

# 4. Profile-order diagnostic — within-person designs are vulnerable to
#    carryover; include the order indicator and joint-test that order
#    coefficients are zero; significant order effects mean carryover is
#    contaminating the AMCE
fit_ord = smf.ols('judgment ~ C(attr1) + C(attr2) + C(attr3) + C(attr4) + C(vignette_order)',
                data=df).fit(cov_type='cluster',
                             cov_kwds={'groups': df['respondent_id']})
print(fit_ord.f_test('C(vignette_order)[T.2]=0, C(vignette_order)[T.3]=0'))

# 5. Pre-specified heterogeneity test — the interaction term is the
#    additional effect of attr1 for female=1 relative to female=0; the
#    main attr1 coefficient is now the effect for female=0 only. Do not
#    interpret the main effect as an average effect once an interaction
#    is included; recover marginal AMCEs via predict() at fixed female values
fit_het = smf.ols('judgment ~ C(attr1) * C(female) + C(attr2) + C(attr3)',
                data=df).fit(cov_type='cluster',
                             cov_kwds={'groups': df['respondent_id']})
print(fit_het.summary())

# 6. Multiple-testing correction across attribute-level contrasts. Pre-register
#    one primary contrast (no correction). For secondary tests, apply BH for
#    exploratory work or Romano-Wolf step-down for confirmatory subgroup claims
pvals = fit.pvalues.drop('Intercept')
reject, p_adj, _, _ = multipletests(pvals, method='fdr_bh', alpha=0.05)
print(pd.DataFrame({'p_raw': pvals, 'p_bh': p_adj, 'reject_bh': reject}))

SurveyCTO / XLSForm

The recommended pattern is pre-randomisation off-form, loaded into the survey via pulldata(). This makes the assignment deterministic, reproducible from the seed file, and supports stratification by enumerator or region. Avoid in-form random()-based attribute generation — it cannot be reproduced for audit and can shift between form recomputes.

1. Generate the deck file in R or Python before fieldwork:

vignettes.csv:
respondent_id,vignette_order,attr1_text,attr2_text,attr3_text,attr4_text,attr1,attr2,attr3,attr4
R0001,1,28-year-old,male,three years of experience,a degree in computer science,1,1,2,3
R0001,2,42-year-old,female,seven years of experience,a degree in business,2,2,3,2
R0002,1,...

One row per (respondent_id × vignette_order). The text columns are the substrings inserted into the vignette narrative; the numeric columns store the realised attribute values for analysis.

2. Form structure:

type                 name              calculation                                         label
calculate            attr1_text        pulldata('vignettes','attr1_text','key',
                                                concat(${respondent_id},'_',${v_order}))
calculate            attr2_text        pulldata('vignettes','attr2_text','key', ...)
... (one calculate per attribute text and one per stored numeric value) ...
note                 vignette_text                                                          Consider a ${attr1_text} ${attr2_text} applicant with ${attr3_text} and ${attr4_text}, applying for a junior developer role.
select_one yes_no    judgment                                                               Would you be willing to interview this applicant?

The note field assembles the narrative using concat() or in-line ${...} substitution. Each calculate field that holds an attribute value (the numeric columns) must be persisted to the submission — set appearance=hidden and ensure the field is included in the form bundle.

3. Within-person designs. Use a repeat loop that iterates vignette_order from 1 to N, recomputing the pulldata key on each iteration so a fresh vignette is drawn per loop pass.

4. Stratified randomisation. When you need balance within strata (region, enumerator route, baseline-survey subgroup), build the strata into the deck file’s key structure — concat(${stratum},'_',${respondent_id},'_',${v_order}) — and ensure the pre-generation script balances attribute frequencies within each stratum.

5. Mandatory persisted fields. Per submission, store: respondent_id, vignette_order, all numeric attribute values, the judgment response, and any attention or manipulation check responses. Without these the analysis cannot proceed.

6. Pilot the pulldata round-trip end-to-end. Confirm that the values stored in the submission match the deck file. A common failure: the form pulls correctly but appearance=hidden calculates do not get persisted to the submission, and the analyst receives blank attribute columns.

Reading the output

  • Each AMCE coefficient is the average causal effect of that attribute value on the judgment, relative to the omitted baseline, in the units of the response scale. The estimand is averaged over the experimental distribution of the other attributes — not the real-world distribution. If real-world relevance matters, weight to the target distribution (de la Cuesta, Egami & Imai, 2022).
  • Binary judgment (0/1): an AMCE of 0.12 means the attribute level raises the probability of a positive judgment by 12 percentage points relative to baseline. |AMCE| ≥ 0.15 is a strong driver of stated response; |AMCE| ≤ 0.03 is typically below practical relevance even if statistically significant.
  • Ordinal scale (1–5): interpret as a scale-point shift; divide by (max − min) to express as a proportion of the full scale.
  • Position / order check — the joint test on vignette_order indicators should be non-significant (p > 0.05). A significant test indicates carryover; report results both with and without the order control and discuss the gap.
  • Randomisation balance check — the F-test from the balance regression in step 1 should be non-significant. A material coefficient flags a deck-generation bug.
  • Interaction interpretation — once an interaction is in the model, the main effect on the interacted attribute is the effect for the reference level of the interacting variable, not an average effect. Use the cregg mm workflow (R) or predict() at fixed values (Python) to recover marginal AMCEs.
  • Heterogeneity claims — for cross-group claims, prefer Marginal Means (baseline-free) over subgroup AMCEs (baseline-dependent). Report subgroup AMCEs only with caveats and pre-specification.
  • Multiple-testing correction — pre-register a single primary contrast. For secondary item-level tests, BH controls FDR; for confirmatory subgroup claims, Romano-Wolf step-down is the more defensible standard. Report both corrected and raw p-values.
  • The result is stated, not behavioural. Discuss the gap between stated and revealed when reporting, and avoid policy claims framed as “people will do X” — the correct framing is “people say they will do X”.

References

Alexander, C. S., & Becker, H. J. (1978). The use of vignettes in survey research. Public Opinion Quarterly, 42(1), 93–104. https://doi.org/10.1086/268432

Auspurg, K., & Hinz, T. (2015). Factorial survey experiments. Sage Publications.

Bansak, K., Hainmueller, J., Hopkins, D. J., & Yamamoto, T. (2021). Conjoint survey experiments. In J. N. Druckman & D. P. Green (Eds.), Advances in experimental political science (pp. 19–41). Cambridge University Press.

Beatty, P. C., & Willis, G. B. (2007). Research synthesis: The practice of cognitive interviewing. Public Opinion Quarterly, 71(2), 287–311. https://doi.org/10.1093/poq/nfm006

Bertrand, M., & Mullainathan, S. (2004). Are Emily and Greg more employable than Lakisha and Jamal? A field experiment on labor market discrimination. American Economic Review, 94(4), 991–1013. https://doi.org/10.1257/0002828042002561

Cameron, A. C., Gelbach, J. B., & Miller, D. L. (2008). Bootstrap-based improvements for inference with clustered errors. Review of Economics and Statistics, 90(3), 414–427. https://doi.org/10.1162/rest.90.3.414

Clarke, D., Romano, J. P., & Wolf, M. (2020). The Romano-Wolf multiple-hypothesis correction in Stata. Stata Journal, 20(4), 812–843. https://doi.org/10.1177/1536867X20976314

Cummings, R. G., & Taylor, L. O. (1999). Unbiased value estimates for environmental goods: A cheap talk design for the contingent valuation method. American Economic Review, 89(3), 649–665. https://doi.org/10.1257/aer.89.3.649

Dafoe, A., Zhang, B., & Caughey, D. (2018). Information equivalence in survey experiments. Political Analysis, 26(4), 399–416. https://doi.org/10.1017/pan.2018.9

de la Cuesta, B., Egami, N., & Imai, K. (2022). Improving the external validity of conjoint analysis: The essential role of profile distribution. Political Analysis, 30(1), 19–45. https://doi.org/10.1017/pan.2020.40

de Quidt, J., Haushofer, J., & Roth, C. (2018). Measuring and bounding experimenter demand. American Economic Review, 108(11), 3266–3302. https://doi.org/10.1257/aer.20171330

Hainmueller, J., Hopkins, D. J., & Yamamoto, T. (2014). Causal inference in conjoint analysis: Understanding multidimensional choices via stated preference experiments. Political Analysis, 22(1), 1–30. https://doi.org/10.1093/pan/mpt024

Jasso, G. (2006). Factorial survey methods for studying beliefs and judgments. Sociological Methods & Research, 34(3), 334–423. https://doi.org/10.1177/0049124105283121

Kahneman, D., & Tversky, A. (1979). Prospect theory: An analysis of decision under risk. Econometrica, 47(2), 263–291. https://doi.org/10.2307/1914185

Krupka, E. L., & Weber, R. A. (2013). Identifying social norms using coordination games: Why does dictator game sharing vary? Journal of the European Economic Association, 11(3), 495–524. https://doi.org/10.1111/jeea.12006

Leeper, T. J., Hobolt, S. B., & Tilley, J. (2020). Measuring subgroup preferences in conjoint experiments. Political Analysis, 28(2), 207–221. https://doi.org/10.1017/pan.2019.30

Lichtenstein, S., & Slovic, P. (1971). Reversals of preference between bids and choices in gambling decisions. Journal of Experimental Psychology, 89(1), 46–55. https://doi.org/10.1037/h0031207

Lichtenstein, S., & Slovic, P. (Eds.). (2006). The construction of preference. Cambridge University Press. https://doi.org/10.1017/CBO9780511618031

List, J. A., & Gallet, C. A. (2001). What experimental protocol influence disparities between actual and hypothetical stated values? Environmental and Resource Economics, 20(3), 241–254. https://doi.org/10.1023/A:1012791005641

Long, J. S., & Ervin, L. H. (2000). Using heteroscedasticity consistent standard errors in the linear regression model. American Statistician, 54(3), 217–224. https://doi.org/10.1080/00031305.2000.10474549

Lusk, J. L., & Norwood, F. B. (2009). An inferred valuation method. Land Economics, 85(3), 500–517. https://doi.org/10.3368/le.85.3.500

MacKinnon, J. G., Nielsen, M. Ø., & Webb, M. D. (2023). Cluster-robust inference: A guide to empirical practice. Journal of Econometrics, 232(2), 272–299. https://doi.org/10.1016/j.jeconom.2022.04.001

Mummolo, J., & Peterson, E. (2019). Demand effects in survey experiments: An empirical assessment. American Political Science Review, 113(2), 517–529. https://doi.org/10.1017/S0003055418000837

Murphy, J. J., Allen, P. G., Stevens, T. H., & Weatherhead, D. (2005). A meta-analysis of hypothetical bias in stated preference valuation. Environmental and Resource Economics, 30(3), 313–325. https://doi.org/10.1007/s10640-004-3332-z

Mutz, D. C. (2011). Population-based survey experiments. Princeton University Press.

Pager, D., & Quillian, L. (2005). Walking the talk? What employers say versus what they do. American Sociological Review, 70(3), 355–380. https://doi.org/10.1177/000312240507000301

Rossi, P. H., & Anderson, A. B. (1982). The factorial survey approach: An introduction. In P. H. Rossi & S. L. Nock (Eds.), Measuring social judgments: The factorial survey approach (pp. 15–67). Sage.

Sauer, C., Auspurg, K., Hinz, T., & Liebig, S. (2011). The application of factorial surveys in general population samples: The effects of respondent age and education on response times and response consistency. Survey Research Methods, 5(3), 89–102. https://doi.org/10.18148/srm/2011.v5i3.4625

Schwarz, N. (1999). Self-reports: How the questions shape the answers. American Psychologist, 54(2), 93–105. https://doi.org/10.1037/0003-066X.54.2.93

Stefanelli, A., & Lukac, M. (2020). Subjects, trials, and levels: Statistical power in conjoint experiments. SocArXiv. https://doi.org/10.31235/osf.io/spkcy

Tversky, A., & Kahneman, D. (1981). The framing of decisions and the psychology of choice. Science, 211(4481), 453–458. https://doi.org/10.1126/science.7455683

Vossler, C. A., Doyon, M., & Rondeau, D. (2012). Truth in consequentiality: Theory and field evidence on discrete choice experiments. American Economic Journal: Microeconomics, 4(4), 145–171. https://doi.org/10.1257/mic.4.4.145

Willis, G. B. (2005). Cognitive interviewing: A tool for improving questionnaire design. Sage Publications.

Last updated: 5 June 2026