What it is
Endorsement experiments measure the credibility or influence of an actor — a political party, government agency, armed group, NGO, regulator, brand, religious authority, expert, celebrity, or anyone else whose name carries weight — by randomising whether that actor’s endorsement of a policy or proposal is shown to respondents. Both treatment and control groups rate the same set of statements; the treatment group is additionally told that the actor endorses each one. The endorsement effect is the difference in mean reported support between arms. The method identifies the net shift in survey responses caused by attaching the actor’s name, not necessarily a shift in underlying attitudes or behaviour (more on this distinction in Caveats).
A negative endorsement effect — reduced support when the actor is named — signals that the actor is distrusted or politically toxic in the study population. A positive effect signals credibility or alignment. Because respondents are asked only about policies, not about the actor directly, the method reduces (but does not eliminate) social desirability bias relative to a direct question about a sensitive political figure or group. Demand effects, item-level idiosyncrasies, and congruence/identity channels remain potential sources of bias.
When to use it
Endorsement experiments are appropriate in two related contexts: when measuring direct attitudes toward politically sensitive actors is likely to produce biased responses due to fear or social pressure, and when the research question concerns the credibility or influence of those actors rather than the frequency of some behaviour.
The method’s most cited applications are in politics. Lyall, Blair and Imai (2013) used it in Afghanistan to measure support for the Taliban, ISAF, and the Afghan National Army as endorsers of reconstruction policies. Nicholson (2012) used it to identify polarising cue effects from US political leaders — even uncontroversial actors produce systematic endorsement shifts. Blair, Imai and Lyall (2014) combined endorsement experiments with the item count technique to cross-validate estimates of sensitive political support. The same design works outside politics: studying when consumers shift toward or away from a product endorsed by a celebrity, how doctors react to a treatment endorsed by a professional society, or how employees respond to a policy endorsed by leadership versus the rank-and-file. Anywhere a source’s name might carry weight, the method isolates how much.
The method is less suitable when: the actor is not known to a substantial share of the sample (endorsement effects are attenuated toward zero for unrecognised actors — measure recognition in the main survey and analyse heterogeneous effects, see Caveats); the research question requires prevalence estimation rather than credibility measurement; or items are chosen such that baseline support is near 0 or 1, leaving no room for the endorsement to move the response. Pilot item baseline-support distributions and drop or rewrite items that cluster at the extremes before fielding.
How it works
Respondents are randomly assigned to a treatment or control condition, typically with equal allocation. Both groups receive the same set of policy statements and rate their support — commonly on a binary (support/oppose) or short ordinal scale (1–4). The treatment group is additionally told that each policy is endorsed by the actor of interest. The endorsement effect for a given item is the difference in mean support between treatment and control for that item.
Two design rules matter for valid inference. The actor must be named on every item in the treatment arm, not a random subset. If only some items in the treatment arm carry the endorsement, respondents may infer the actor’s position on the other items and contaminate the estimates. The treatment and control prompts must be matched in length and register — the standard convention is to attach a neutral attribution sentence in the control arm (e.g., “A policy currently under discussion proposes…”) of comparable length to the treatment attribution (“[Actor] has endorsed a policy that proposes…”). Without this match, the endorsement manipulation is partly confounded with prompt length and tone.
With multiple items, the averaged endorsement effect — the per-respondent mean across all items, regressed on treatment — is the primary estimand. It pools across items and is more stable than any single-item estimate. Imai, Park and Greene (2015) build a structural extension using item response theory (IRT): the model places both respondents and the actor on a single latent dimension of support, and the actor’s position on that dimension is the credibility estimate. The IRT version is more efficient but rests on three assumptions worth knowing before you commit: a single latent trait drives responses across all items (so items from substantively different policy areas can break the model), item responses are independent given the latent trait (violated when items are thematically clustered), and probability of support rises with latent support. Four to five items is the usual minimum for stable estimation.
Key decisions
Number of items. More items produce more stable averaged endorsement effects and are necessary for the structural IRT model. Imai, Park and Greene (2015) recommend at least four to five items. There is an inherent tension: items spanning different policy domains broaden the aggregate effect to capture general credibility rather than single-issue alignment, but the IRT model assumes unidimensionality — a single latent dimension drives responses across all items. If items load on substantively different ideological dimensions, the IRT estimate is meaningless and the descriptive averaged effect is the safer summary. Check unidimensionality empirically (e.g., factor analysis on control-arm responses) before relying on the IRT model.
Response scale. Binary scales are easier to implement and straightforward to analyse. Ordinal scales (e.g., 1–4) capture more variation and improve statistical efficiency, but require greater cognitive effort and are harder to administer in low-literacy or oral interview settings. The scale should be consistent across all items in the experiment. For ordinal outcomes, OLS on the raw scale is the conventional working model — defensible for treatment-effect interpretation under randomisation — but ordered logit or probit is the formally correct analysis, and at least one robustness check using the ordered model should be reported.
Multiple actors. A single instrument can test multiple endorsers by assigning different treatment versions — one per actor — across respondents. Between-person designs are strongly preferred. Within-person designs — in which each respondent evaluates policies with and without an endorser, or with multiple endorsers in sequence — are nominally more efficient but compromise the manipulation: once a respondent sees the same policy twice with and without an endorsement, the design is transparent and the response shift is more likely to reflect demand effects than genuine attitude change. If testing multiple actors between-person, sample size scales linearly with the number of actors.
Item framing. Items should be policy statements that are plausibly endorsable by all actors being tested. An item about road construction or school building is more credibly endorsable by both a government and an armed group than an item about electoral procedures. Items that are obviously partisan make the endorsement manipulation less credible and risk triggering demand effects. Pilot items for perceived plausibility and choose items where the endorser’s position is not the default expectation — this reduces (but does not separate) the congruence channel relative to the credibility channel.
Sample size. The endorsement effect is a difference in means, so power follows the same logic as any two-arm experiment. Rule of thumb: n per arm ≈ 16/δ² for 80% power at α = 0.05, where δ is the effect size in standardised units. For a target effect of 0.1 SD this works out to 800 respondents per arm, or 1,600 per actor. For binary outcomes with baseline support around 0.5 and a target shift of 5 percentage points, the same calculation gives roughly 1,600 per actor as well. Use pwr.t.test() in R or power twomeans in Stata for precise figures. Cluster-randomised designs need a design-effect adjustment — multiply n by 1 + (m̄ − 1) × ρ, where m̄ is the mean cluster size and ρ is the intra-cluster correlation. The IRT model has different power properties; see Imai, Park and Greene (2015) for guidance on that estimator.
Caveats & common mistakes
Credibility, congruence, and identity are entangled by design. A positive endorsement effect mixes at least three things: respondents trust the actor and update their views (credibility), respondents already think the actor and policy belong together (congruence), or respondents see the actor as part of their in-group (identity). The basic design cannot separate them, and the IRT model does not separate them either — its latent parameter is a composite of all three. To get at mechanism, you need mediation analysis (Bullock, Green & Ha, 2010) or paired source-vs-argument experiments (Bullock, 2011). Reporting an endorsement effect as “pure credibility” is a common overclaim.
Demand effects and manipulation checks. Respondents in the treatment condition may adjust their answers not because of genuine attitude updating but because they perceive the survey as asking them to evaluate the actor. This risk is lower when the endorsement is embedded naturally within a longer multi-topic questionnaire and higher when the endorsement is the clearly focal element of the interview. Include a post-treatment manipulation check — for example, after the policy items, ask in the treatment arm “Did any specific group or person endorse the policies in this section?” with the actor’s name as one of several response options. Respondents who fail the check did not register the manipulation; pre-specify a robustness analysis restricted to those who passed (Berinsky, Margolis & Sances, 2014).
Cueing vs. argument effects. The endorsement experiment measures a cue effect: the actor’s name is the entire treatment, with no policy argument attached. Bullock (2011) shows that argument content often dominates source cues among more informed respondents — i.e., the same actor’s effect can be much larger in a uniformed sample than in an informed one. Endorsement experiments are not substitutes for richer frame experiments (Brader, Valentino & Suhay, 2008) that vary both source and argument.
Validation against behaviour. The endorsement effect identifies a survey-response shift in a specific elicitation context. It does not necessarily predict voting, compliance, donation, or other behavioural outcomes attributable to the actor. Treat the estimate as one measure of stated political credibility, not as a behavioural prediction.
Low actor recognition — measure, don’t gate-keep. The experiment only works if respondents know who the actor is. Effects are attenuated when they don’t. The wrong move is to drop the actor based on a pilot recognition rate. The right move is to measure recognition for every respondent in the main survey and pre-specify two analyses: the full-sample estimate (the conservative population number) and a heterogeneous-effects analysis restricted to recognisers (the policy-relevant magnitude). When the recognition question itself is sensitive, embed the actor’s name in a list of names to lower the disclosure cost. The recognisers-only estimate is descriptive — recognition is not random, so it does not carry the same causal weight as the full-sample effect.
Cross-respondent contamination. Endorsement experiments are often fielded face-to-face by enumerators. If enumerators discuss the manipulation between respondents in the same village or office — even informally — the treatment-control comparison is contaminated. Train enumerators that the manipulation is confidential, and consider varying which item carries the endorsement across enumerator routes so that any leakage is harder to act on.
Multiple comparisons. Testing multiple items and multiple actors generates many hypothesis tests. The pre-analysis plan should specify a primary estimand — typically the averaged endorsement effect — for which no correction is needed. For secondary item-level tests within one actor, Benjamini-Hochberg false-discovery-rate control is acceptable if pre-specified. For comparisons across actors used to claim that a specific actor is distrusted (a policy-relevant substantive claim), family-wise error rate control via Holm or Romano-Wolf step-down is the stronger standard.
Null effects are not null findings. A zero endorsement effect means the actor’s named association with a policy does not shift reported support in a survey context. It does not mean the actor has no political influence through other channels. The endorsement experiment is a narrow test of one specific causal pathway — attitude change via disclosed association.
Analysis Guide
import pandas as pd
import numpy as np
import statsmodels.formula.api as smf
from statsmodels.miscmodels.ordinal_model import OrderedModel
items = ["item1", "item2", "item3", "item4", "item5"]
# 1. Item-level endorsement effects with SE and CI — looping over items and storing
# only the coefficient is the most common mistake in this design; without SE you
# cannot interpret magnitude. Collect estimates into a dataframe for a coefficient
# plot or downstream multiple-testing correction
item_effects = []
for y in items:
fit = smf.ols(f"{y} ~ treatment", data=df).fit(cov_type="HC1")
item_effects.append({
"item": y,
"estimate": fit.params["treatment"],
"std_error": fit.bse["treatment"],
"ci_low": fit.conf_int().loc["treatment", 0],
"ci_high": fit.conf_int().loc["treatment", 1],
"p_value": fit.pvalues["treatment"],
})
print(pd.DataFrame(item_effects))
# 2. Averaged endorsement effect on complete cases only — using pandas .mean(skipna=True)
# on the row silently mixes per-respondent denominators (one respondent's mean is
# over 3 items, another's over 5); restrict to complete-case rows so every averaged
# value is comparable
complete = df[items].notna().all(axis=1)
df["mean_support"] = np.nan
df.loc[complete, "mean_support"] = df.loc[complete, items].mean(axis=1)
fit_avg = smf.ols("mean_support ~ treatment", data=df).fit(cov_type="HC1")
print(fit_avg.summary())
# 3. With covariates and pre-specified recognition heterogeneity — covariates only
# reduce variance under randomisation; the treatment × recognition interaction is
# the pre-specified heterogeneous-effects test; the recognisers-only effect is the
# policy-relevant magnitude but is descriptive, not causal w.r.t. recognition
fit_het = smf.ols(
"mean_support ~ treatment * recognised + age + C(female) + hh_size",
data=df,
).fit(cov_type="HC1")
print(fit_het.summary())
# Marginal treatment effect at recognised = 0 and recognised = 1
b_t = fit_het.params["treatment"]
b_tr = fit_het.params["treatment:recognised"]
print(f"dy/d(treatment) | recognised=0: {b_t:.3f}")
print(f"dy/d(treatment) | recognised=1: {b_t + b_tr:.3f}")
# 4. Ordered model robustness check for ordinal outcomes — OLS on a 1-4 scale is a
# working approximation; ordered probit is the formally correct analysis. Run it
# on at least one item and report alongside the OLS result
ord_data = df[["item1", "treatment", "age", "female", "hh_size"]].dropna()
fit_ord = OrderedModel(
ord_data["item1"],
ord_data[["treatment", "age", "female", "hh_size"]],
distr="probit",
).fit(method="bfgs", disp=False)
print(fit_ord.summary())
# 5. Cluster-robust SE for cluster-randomised endorsement experiments — if the
# randomisation was by village or enumerator route rather than respondent,
# cluster the SE accordingly (and adjust power calculations for the design effect)
fit_cluster = smf.ols(
"mean_support ~ treatment + age + C(female) + hh_size",
data=df,
).fit(cov_type="cluster", cov_kwds={"groups": df["psu"]})
print(fit_cluster.summary())
# 6. Structural IRT model (Imai, Park & Greene 2015) — no Python equivalent of the
# R 'endorse' package exists. The hierarchical IRT model for endorsement
# experiments (latent credibility/ideal-point with item parameters and covariates)
# is not implemented in pystan, pymc, or statsmodels as a turnkey procedure. Use
# the R tab (endorse::endorse) for this analysis, or call it from Python via rpy2:
# from rpy2.robjects.packages import importr
# endorse = importr('endorse')
# fit = endorse.endorse(...) SurveyCTO / XLSForm
Assign treatment in a calculate field at the start of the form using once(if(random() < 0.5, 0, 1)). The outer once() is critical — without it the expression re-evaluates on form revisits and the assignment changes mid-interview. Persist the assignment to a stored field that is included in the submission so the analyst can recover which arm each respondent was in.
Create two versions of each policy item — one with the endorsement attribution (“[Actor] has endorsed a policy that proposes…”) and one with a neutral attribution of comparable length (“A policy currently under discussion proposes…”) — and use relevant conditions to display the correct version. The policy statement itself must be worded identically across versions and the surrounding attribution sentences must match on length and register; any difference in phrasing or prompt length confounds the endorsement effect.
For stratified randomisation (e.g., balancing treatment within region, gender, or recognition tertile), pre-randomise the assignments on the server and load them via pulldata() against an uploaded case management file rather than randomising in the form. The random() approach cannot guarantee balance within strata at small sample sizes.
Include a post-treatment manipulation check, presented after the policy items in both arms with the actor’s name embedded in a list of plausible alternatives: “Did any of the following groups or individuals endorse the policies you just read about? [list].” Respondents in the treatment arm who fail to identify the actor did not register the manipulation; pre-specify a robustness analysis restricted to those who passed.
Also measure actor recognition separately (“Have you heard of [actor]?”) earlier in the survey, before the experimental items — embedded in a list of names rather than as a direct standalone question if the recognition itself is sensitive. This recognition measure feeds the heterogeneous-effects analysis in step 3 of the Analysis Guide.
Reading the output
- The coefficient on
treatmentin each item regression is the endorsement effect for that item — the percentage-point (binary) or scale-point (ordinal) shift in support caused by disclosing the actor’s name. Always report alongside the SE and 95% CI; the point estimate alone is uninterpretable. - A negative coefficient means the actor’s association reduced support; a positive coefficient means it increased it. Effects whose 95% CI crosses zero are not distinguishable from no effect at conventional thresholds — but a wide CI is a precision problem, not evidence of zero effect.
- The averaged endorsement effect (from the
mean_supportregression on complete cases) is the primary estimand. A one-item result can reflect idiosyncratic item-actor congruence rather than the actor’s general credibility. - For binary support scales, a coefficient of 0.05 = a 5 percentage-point shift — small but politically meaningful in close electoral or compliance contexts. For 1–4 ordinal scales, divide the coefficient by the scale range to judge magnitude on a 0–1 metric, then sanity-check against the ordered probit/logit robustness check.
- Recognisers-only effect vs. full-sample ITT: report both. If the ITT is near zero but the recognisers-only effect is meaningfully non-zero, the actor has a real effect among those who know them and the full-sample result is diluted by non-recognition. If both are near zero, the actor has no detectable credibility signal in this population.
- Manipulation-check failure rate above ~20%: the treatment was not reliably received and the ITT is a lower bound on the true effect among those who registered the manipulation. Report the restricted-sample effect as a robustness check and discuss the failure rate.
- IRT posterior summary (from
endorse): thelambdaparameter is the actor’s location on the latent support dimension — positive values indicate net credibility, negative values indicate net distrust. The 95% credible interval and Rhat convergence diagnostic must both be reported. Remember the latent parameter is a composite of credibility, congruence, and identity channels — not pure credibility.
References
Berinsky, A. J., Margolis, M. F., & Sances, M. W. (2014). Separating the shirkers from the workers? Making sure respondents pay attention on self-administered surveys. American Journal of Political Science, 58(3), 739–753. https://doi.org/10.1111/ajps.12081
Blair, G., Imai, K., & Lyall, J. (2014). Comparing and combining list and endorsement experiments: Evidence from Afghanistan. American Journal of Political Science, 58(4), 1043–1063. https://doi.org/10.1111/ajps.12086
Brader, T., Valentino, N. A., & Suhay, E. (2008). What triggers public opposition to immigration? Anxiety, group cues, and immigration threat. American Journal of Political Science, 52(4), 959–978. https://doi.org/10.1111/j.1540-5907.2008.00353.x
Bullock, J. G. (2011). Elite influence on public opinion in an informed electorate. American Political Science Review, 105(3), 496–515. https://doi.org/10.1017/S0003055411000165
Bullock, J. G., Green, D. P., & Ha, S. E. (2010). Yes, but what’s the mechanism? (Don’t expect an easy answer). Journal of Personality and Social Psychology, 98(4), 550–558. https://doi.org/10.1037/a0018933
Imai, K., Park, B., & Greene, K. F. (2015). Using the item response theory to measure the endorsement experiment. American Journal of Political Science, 59(1), 55–75. https://doi.org/10.1111/ajps.12098
Lyall, J., Blair, G., & Imai, K. (2013). Explaining support for combatants during wartime: A survey experiment in Afghanistan. American Political Science Review, 107(4), 679–705. https://doi.org/10.1017/S0003055413000403
Mutz, D. C. (2011). Population-based survey experiments. Princeton University Press.
Nicholson, S. P. (2012). Polarizing cues. American Journal of Political Science, 56(1), 52–66. https://doi.org/10.1111/j.1540-5907.2011.00541.x
Sniderman, P. M., & Grob, D. B. (1996). Innovations in experimental design in attitude surveys. Annual Review of Sociology, 22(1), 377–399. https://doi.org/10.1146/annurev.soc.22.1.377