What it is
When respondents from different groups, regions, or countries are asked to rate a subjective construct — political freedom, health status, corruption, satisfaction with public services — they may not interpret or use the response scale in the same way. A respondent in a low-accountability context who rates government responsiveness as “3 out of 5” may be describing worse objective conditions than a respondent in a high-accountability context who gives the same rating. This is a form of differential item functioning (DIF) — specifically, heterogeneity in how respondents map a latent construct onto the response categories — and it makes direct cross-group comparisons unreliable even when the survey instrument is identical (King et al., 2004).
Anchoring vignettes address this by asking respondents to rate a set of hypothetical scenarios — the vignettes — alongside their self-assessment. Each vignette describes a researcher-defined level of the underlying construct. Under the assumption that all respondents interpret each vignette as representing the same level (the vignette equivalence assumption — discussed below), variation in how respondents rate the vignettes can only reflect variation in their use of the scale, not differences in what they are describing. These ratings are used to calibrate — or “anchor” — each respondent’s self-assessment, producing corrected estimates that are more comparable across groups.
When to use it
Anchoring vignettes are appropriate when: subjective outcomes are being compared across groups that may plausibly differ in how they interpret response scales; cross-national or cross-regional comparisons of self-reported wellbeing, health, governance perceptions, or similar constructs are central to the analysis; or a treatment in an RCT might plausibly affect not just the underlying outcome but also how respondents report it — in which case raw mean comparisons between arms would be confounded.
King et al. (2004) introduced the method to address cross-national comparisons of health and political freedom. The dominant empirical application has been comparing self-reported health and disability across age, gender, and socioeconomic groups in large panel studies such as SHARE, ELSA, and the WHO World Health Survey (Bago d’Uva, O’Donnell & van Doorslaer, 2008; van Soest et al., 2011). Applications in development economics and field research include corruption and government responsiveness perceptions (where local reference points vary substantially), subjective poverty and wellbeing comparisons across treatment arms, and quality-of-service assessments across geographically dispersed sites.
The method is less useful when the outcome of interest is objective rather than subjective, when sample sizes are small (CHOPIT estimation typically requires N ≥ 500–1,000 per comparison group for stable identification; King & Wand, 2007), or when it is not plausible to construct vignettes that will be understood consistently across the study population.
How it works
The survey presents respondents with a self-assessment question and a set of vignette questions using the same response scale. A self-assessment might ask: “How much freedom do people in your country have to express political opinions?” The vignettes then describe specific hypothetical individuals — “Consider [Name], who can express political opinions freely in private but faces arrest if they speak publicly. How much freedom does [Name] have?” — and the respondent rates each on the same scale.
Two identifying assumptions. The entire correction stands or falls on two conditions. Vignette equivalence (VE) — every respondent reads each vignette as representing the same underlying level (e.g., everyone agrees that the vignette person who is “arrested if they speak publicly” has limited freedom). Response consistency (RC) — each respondent uses the same internal thresholds when rating themselves as when rating the vignettes. Both are partially testable in practice; the diagnostics live in the Caveats section.
How the correction works. Think of each respondent as having their own internal thresholds for switching from one rating to the next — the point at which they shift from “3” to “4” sits at a different place on the underlying scale for different people. The vignettes give us a fixed reference: because everyone is rating the same hypothetical situations, their ratings of the vignettes reveal where their personal thresholds sit. Once we know those thresholds, we can re-read each respondent’s self-assessment on a common scale and compare across groups without the noise of differing scale use.
CHOPIT model (King & Wand, 2007). CHOPIT is a hierarchical ordered probit that does exactly that. It fits two equations at once: one for the self-assessment (latent construct as a function of covariates) and one for the vignettes (with researcher-fixed levels). Both equations share the same set of respondent-specific thresholds, and those thresholds are themselves allowed to vary with respondent covariates. Vignette ratings pin down the thresholds; the self-assessments are then placed on a common scale. CHOPIT is implemented in the anchors R package (Wand, King & Lau, 2011).
Rank-order (non-parametric) method (Hopkins & King, 2010). Instead of estimating thresholds, this approach simply asks where each respondent’s self-rating falls relative to their vignette ratings — for example, “self ranks below all vignettes” or “self falls between vignettes 2 and 3.” Each respondent ends up in an ordinal category, and the population result is a range rather than a single point estimate. Ties (when a respondent rates themselves the same as a vignette) widen the range further. The trade-off is the usual one: the rank-order method asks less of the data and is more robust to wrong-shape assumptions, but the resulting estimates are less precise. It is implemented as method = "B" in the anchors package.
Key decisions
Number of vignettes. Three to five vignettes per construct is standard; King et al. (2004) used three, SHARE typically uses two to three per domain. More vignettes give better identification of the cutpoints but increase survey length and respondent fatigue. Each vignette should represent a clearly distinct level of the construct, and together they should span the full range of the response scale. With fewer than three vignettes, the cutpoints are weakly identified and CHOPIT estimation becomes unstable.
Number of response categories. Five-point Likert scales are most common in this literature. Seven-point scales improve identification of cutpoints (more category boundaries to estimate) but add noise; three-point scales rarely give enough cutpoint variation to identify the model. Match the scale to the substantive task and use the same scale across self-assessment and all vignettes.
Vignette construction. This is the most consequential design decision and the most easily underestimated. Vignettes must be:
- Unambiguous about the level of the construct they represent
- Interpreted consistently across the study population (the VE assumption)
- Ordered correctly by the researcher in terms of the underlying construct
- Free of cues that might prompt respondents to infer the “right” answer rather than apply their genuine scale
Pre-testing with cognitive interviews is essential. A vignette that seems clear to the researcher may be interpreted very differently by respondents with different reference experiences.
Vignette person attributes. Whether the hypothetical person in each vignette should be matched to the respondent on gender, age, or occupation is a real design choice (Hopkins & King, 2010). Matching can improve response consistency by making the comparison more cognitively natural for the respondent, but it means different respondents are effectively rating slightly different stimuli, complicating pooling and weakening VE in its strictest form. Most applied studies use a single fixed set of vignette persons for everyone; matched-vignette designs are an option when there is strong theoretical reason to expect that gender or age changes how the construct should be interpreted.
Vignette order randomization. Randomize the order in which the vignettes are presented across respondents. Without order randomization, systematic order effects (carry-over, anchoring on the first vignette, fatigue on the last) confound the cutpoint estimates with vignette identity.
Placement of vignettes relative to self-assessment. The literature is unsettled. Placing vignettes after the self-assessment is the most common practice and avoids hypothetical descriptions priming how respondents evaluate themselves. Placing vignettes before the self-assessment is argued by some (Kristensen & Johansson, 2008; Hopkins & King, 2010) to give respondents a calibration reference that improves response consistency. Both placements appear in published work. Pick one, pre-specify it, and acknowledge the choice.
Analysis approach. CHOPIT is the standard parametric approach (Wand, King & Lau, 2011). It requires N ≥ 500–1,000 per comparison group for stable estimation and assumes a normal latent construct with covariate-shifted cutpoints. The rank-order method is robust to those distributional assumptions but produces set-valued estimates (interval bounds on group prevalences rather than point estimates) — fine for ordinal subgroup comparisons, less suitable when a single corrected score per respondent is needed. Note that anchoring vignettes do not improve the precision of construct estimates — they correct bias at the cost of variance. The standard errors on corrected group means are typically wider than on raw means.
Subgroup comparisons. The corrected estimates are most informative when compared across groups — treatment vs. control, regions, socioeconomic categories. The value of the correction depends on how much response scale heterogeneity actually exists in the data. Test whether the correction materially changes substantive conclusions; in some applications it does not, and reporting both raw and corrected comparisons is the right thing to do.
Caveats & common mistakes
Vignette equivalence (VE). The method assumes that all respondents interpret each vignette as representing the same level of the underlying construct. VE is not directly testable but is partially testable: examine whether vignette ratings differ systematically across observable groups (region, education, age) in ways that cannot plausibly reflect scale usage alone — and, where available, compare vignette ratings to objective characteristics described in the vignette (King et al., 2004, §4; Bago d’Uva et al., 2008; van Soest et al., 2011). Cognitive piloting on the actual study population is the first-line check. If respondents from different groups systematically read the same vignette description as representing different levels — because local context gives the descriptions different meaning — the correction will itself be biased.
Response consistency (RC). The method also assumes that respondents apply the same cutpoints when rating the hypothetical vignette person as when rating themselves. RC is testable against auxiliary objective measures of the construct (van Soest et al., 2011; Kapteyn, Smith & van Soest, 2007 use grip strength, vision tests, and clinical disability assessments as objective benchmarks against self-reported health). Evidence is mixed: RC holds well for some constructs (work disability) and less well for others (general life satisfaction). The direction of any RC violation is difficult to predict without domain-specific knowledge; where auxiliary objective measures exist, run the test and report the result.
Survey length. Adding three to five vignettes per construct, each requiring a rating on the same scale as the self-assessment, adds meaningful length to the interview. In field settings with respondent fatigue, this can affect data quality on subsequent questions. The method is most practical when one or two constructs are the primary focus.
Interpretation of corrected estimates. The corrected estimates are not directly observable quantities — they are model-based transformations of the raw responses. The units of the corrected scale are latent and cannot be interpreted in the same way as, say, a count or a probability. This limits how they can be communicated to non-technical audiences and requires care when reporting effect sizes.
Null corrections are informative. If the corrected estimates agree closely with the raw self-assessments, this is a substantive finding — it suggests response scale heterogeneity is not large enough to materially affect conclusions in this sample. This result should be reported, not discarded.
Analysis Guide
import pandas as pd
import numpy as np
from statsmodels.miscmodels.ordinal_model import OrderedModel
# Note: there is no Python implementation of CHOPIT. The R 'anchors' package
# (Wand, King & Lau 2011) is the canonical implementation and is what we recommend
# for the formal CHOPIT analysis (and for the rank-order method "B"). Below is a
# working approximation in Python using statsmodels' OrderedModel — an ordered
# probit with covariate-shifted cutpoints, the non-hierarchical core of CHOPIT.
# For the full hierarchical model with respondent-specific thresholds, use the R
# tab or call anchors via rpy2.
# Pre-modelling diagnostic for vignette ordering — under the researcher-intended
# ordering, mean vignette ratings should be monotone across vignettes; a non-monotone
# pattern flags that respondents do not perceive the vignettes in the intended order
print(df[["vign1", "vign2", "vign3"]].mean())
# 1. Stack vignette and self responses into long format — one row per respondent x
# item (3 vignettes + 1 self); item dummies enter as fixed effects on the latent
# level (the vignettes' researcher-defined levels), and covariates enter on the
# cutpoints to capture scale usage differences
df["id"] = np.arange(len(df))
long = df.melt(
id_vars=["id", "age", "female", "edu"],
value_vars=["self", "vign1", "vign2", "vign3"],
var_name="item",
value_name="rating",
).dropna(subset=["rating"])
# 2. Fit ordered probit with researcher-fixed vignette levels and covariates that
# shift the cutpoints — item dummies identify vignette positions relative to a
# respondent's own (self) position; including age/female/edu as covariates lets
# the cutpoints depend on respondent characteristics, mimicking the CHOPIT
# scale-usage mechanism. statsmodels OrderedModel has no native cluster-robust
# SE option for respondent clustering — bootstrap by respondent if SEs matter
item_dummies = pd.get_dummies(long["item"], prefix="item", drop_first=True)
X = pd.concat(
[item_dummies, long[["age", "female", "edu"]].reset_index(drop=True)],
axis=1,
).astype(float)
y = long["rating"].astype(int).reset_index(drop=True)
fit = OrderedModel(y, X, distr="probit").fit(method="bfgs", disp=False)
print(fit.summary())
# 3. Recover the respondent-level latent score — predict the linear index for each
# respondent's "self" row; this is the corrected self-rating on the common latent
# metric. Group-mean comparisons of this score are the substantive output;
# do not compare it to the raw self-rating (different units)
self_rows = long["item"] == "self"
xb = fit.model.exog @ fit.params[: fit.model.exog.shape[1]]
corrected = pd.Series(xb[self_rows.values], index=long.loc[self_rows, "id"].values)
# Compare distributions of corrected across groups (e.g., treatment arms, regions) SurveyCTO / XLSForm
Implement the self-assessment and each vignette as select_one items sharing a single choice list (the same response scale across all items). Present each vignette as a note field containing the vignette text, followed immediately by the select_one field that records the rating — keeping each vignette text-and-rating as a tight pair reduces cognitive load.
Randomize vignette order across respondents using a stable per-form draw: a calculate field with once(random()) generates a permutation index, and a repeat or branching pattern displays the vignettes in the order indicated by the index. Persist the realized order as a stored variable so the analyst can confirm balanced randomization and test for residual order effects.
Allow “don’t know” and refusal as separate response codes on each item; pre-specify how these are handled (typically dropped for the CHOPIT model, retained as a separate ordinal level for the rank-order method).
If running the recommended robustness checks for VE and RC, include any auxiliary objective measures (e.g., a brief functional-health measure for a health-construct study) in the same instrument so the tests can be run on the same sample.
Reading the output
- The CHOPIT corrected score and the rank-order category are the model-estimated latent level of the construct, on a standardised metric. Neither is directly comparable to the raw self-assessment — do not report them side by side as if they share units.
- Compare the distribution of corrected scores across groups (treatment vs. control, regions, socioeconomic categories) rather than comparing corrected to raw scores for the same individual. Report both raw and corrected group comparisons; if they agree, response-scale heterogeneity is not a binding constraint in this sample — a substantive, reportable finding.
- Pre-modelling diagnostic: the sample means of the vignette ratings should be monotone in the researcher-intended ordering of vignette latent levels. If mean(vign1) > mean(vign2), most respondents perceive vignette 1 as representing a higher level than vignette 2 — the intended ordering does not match respondent perception, and VE is in doubt before any model is fit.
- CHOPIT convergence warning signs: failure to converge after 200 iterations, estimated cutpoints that are not strictly increasing within a respondent, or a Hessian that is not positive-definite. Any of these indicates VE or RC is failing badly or the sample is too small (typical threshold N < 500 per group). Drop back to the rank-order method.
- Rank-order output: the interval-valued estimates collapse when many respondents tie with vignettes; if more than 20% of respondents tie with one specific vignette, that vignette is poorly placed on the latent scale and should be re-piloted.
- Where auxiliary objective measures exist, run the RC test (e.g., regress objective measure on the corrected score within self-rating category; significant coefficient ⇒ RC violated). Report the test result alongside the main estimates.
References
Bago d’Uva, T., O’Donnell, O., & van Doorslaer, E. (2008). Differential health reporting by education level and its impact on the measurement of health inequalities among older Europeans. International Journal of Epidemiology, 37(6), 1375–1383. https://doi.org/10.1093/ije/dyn146
Grol-Prokopczyk, H., Freese, J., & Hauser, R. M. (2011). Using anchoring vignettes to assess group differences in general self-rated health. Journal of Health and Social Behavior, 52(2), 246–261. https://doi.org/10.1177/0022146510396713
Hopkins, D. J., & King, G. (2010). Improving anchoring vignettes: Designing surveys to correct interpersonal incomparability. Public Opinion Quarterly, 74(2), 201–222. https://doi.org/10.1093/poq/nfq011
Kapteyn, A., Smith, J. P., & van Soest, A. (2007). Vignettes and self-reports of work disability in the United States and the Netherlands. American Economic Review, 97(1), 461–473. https://doi.org/10.1257/aer.97.1.461
King, G., Murray, C. J. L., Salomon, J. A., & Tandon, A. (2004). Enhancing the validity and cross-cultural comparability of measurement in survey research. American Political Science Review, 98(1), 191–207. https://doi.org/10.1017/S000305540400108X
King, G., & Wand, J. (2007). Comparing incomparable survey responses: New tools for anchoring vignettes. Political Analysis, 15(1), 46–66. https://doi.org/10.1093/pan/mpl011
Kristensen, N., & Johansson, E. (2008). New evidence on cross-country differences in job satisfaction using anchoring vignettes. Labour Economics, 15(1), 96–117. https://doi.org/10.1016/j.labeco.2006.11.001
Murray, C. J. L., Tandon, A., Salomon, J. A., & Mathers, C. D. (2002). Enhancing cross-population comparability of survey results. GPE Discussion Paper No. 35. World Health Organization.
van Soest, A., Delaney, L., Harmon, C., Kapteyn, A., & Smith, J. P. (2011). Validating the use of anchoring vignettes for the correction of response scale differences in subjective questions. Journal of the Royal Statistical Society Series A, 174(3), 575–595. https://doi.org/10.1111/j.1467-985X.2011.00694.x
Wand, J., King, G., & Lau, O. (2011). anchors: Software for anchoring vignette data. Journal of Statistical Software, 42(3), 1–25. https://doi.org/10.18637/jss.v042.i03