Metter. / Mixtapes / Methods Mixtape / Survey & Elicitation Methods

05 · Survey & Elicitation Methods

Conjoint / Discrete Choice Analysis

A method for estimating how people value and trade off the attributes of a policy, product, or candidate by presenting randomised choice tasks and observing which options they select.


What it is

Conjoint analysis — also called discrete choice experiments (DCE) — estimates how respondents value and trade off the characteristics of a choice option. Respondents are presented with a series of hypothetical choice tasks, each showing two or more options described by a set of attributes at varying levels. By observing choices across many tasks, the researcher can estimate how much each attribute contributes to the probability of selection and what tradeoffs respondents are willing to make between them.

The primary estimand in modern survey-based conjoint analysis is the Average Marginal Component Effect (AMCE): the average effect of changing one attribute level on the probability that a respondent selects a profile, averaging over the distribution of other attributes and over respondents (Hainmueller, Hopkins & Yamamoto, 2014). Because attribute levels are randomly assigned across profiles and tasks, the AMCE has a clean causal interpretation under a small set of design-based assumptions (stated formally in “How it works”) — it does not require the parametric utility specifications that earlier random utility models impose, but it does require its own structural conditions.

Two interpretive cautions are essential to flag at the outset (and discussed in Caveats): (i) the AMCE is not “the average respondent’s preference” — it is a weighted average where preference intensity confounds preference direction, so a majority-preferred option can carry a negative AMCE (Abramson, Koçak & Magazinnik, 2022); and (ii) for subgroup comparisons the AMCE is the wrong target because it depends on the choice of baseline level — Marginal Means (MMs) are baseline-free and the recommended estimand for cross-group comparison (Leeper, Hobolt & Tilley, 2020).

When to use it

Conjoint analysis is appropriate when the research question concerns the relative importance of multiple attributes simultaneously, rather than the effect of any single attribute in isolation. It is well suited to:

  • Preferences for public services with multiple quality dimensions (proximity, provider qualifications, cost, waiting time)
  • Voting behaviour and candidate evaluation (where candidates differ on many dimensions at once)
  • Agricultural input or technology adoption decisions (price, yield potential, risk profile, compatibility)
  • Policy evaluation where respondents must weigh competing objectives

It is less appropriate when: a single attribute is the focus and other attributes can be held constant experimentally; the outcome of interest is revealed rather than stated preference (conjoint estimates hypothetical choices, not actual ones); or cognitive demands exceed respondent capacity — low-literacy contexts, designs with six or more attributes, or attributes that require numerical comparisons across multiple units (currency, distance, time) are warning signs and should be pilot-tested with cognitive interviews before fielding.

How it works

The researcher specifies a set of attributes — the dimensions that describe each option — and for each attribute, a set of levels that it can take. Attributes might include cost (low / medium / high), distance to service (under 1 km / 1–5 km / over 5 km), and provider type (public / private / NGO). A profile is one specific combination of attribute levels. In each task, a respondent is shown two profiles — each fully described — and asked to choose which they prefer.

What the AMCE is. For a given attribute level, the AMCE is the average percentage-point change in the probability that a profile is chosen when that level is shown, compared to when the baseline level is shown, averaged across all other attributes and across respondents. It is the OLS coefficient on the level’s dummy in a regression of the binary choice indicator on all attribute-level dummies (baseline omitted), with standard errors clustered at the respondent.

Precision and sample size. Each respondent contributes 2 × T observations — two profiles per task, across T tasks — and the cluster-robust SE accounts for the correlation across both. Precision improves with more respondents and more tasks per respondent, until satisficing sets in (see Caveats). As a planning rule for a five-attribute, two-profile, ten-task design, 300–500 respondents are enough to detect AMCEs of 3–5 percentage points at 80% power; 1,000+ is the usual figure when precise subgroup comparisons are the target. Use cregg’s simulation utilities or design-specific Monte Carlo for tighter planning, especially when the design is restricted (Egami & Imai, 2019) or attribute prevalences are uneven.

What has to be true for the AMCE to mean what you think it means. Five conditions: attribute levels really are assigned at random; a respondent’s choice in any one task does not depend on profiles they saw in earlier tasks; left/right placement within a task does not influence choice (so position must be randomised and controlled for or balance-checked); respondents read each attribute level the same way regardless of the other attributes in the profile (Ono & Burden, 2018 document cases where this fails); and respondents do not communicate about the profiles during the survey. If any of these breaks, the AMCE you estimate is not the AMCE you think you are estimating.

Other estimators. AMCE-OLS is the modern standard for causal effects on choice probability. When the question is about market shares or willingness to pay rather than causal effects, conditional logit (McFadden, 1974) is the right tool — clogit / asclogit in Stata, mlogit or apollo in R. For individual-level preference heterogeneity, mixed logit with random coefficients fits the question — mixlogit in Stata, gmnl or apollo in R. Both logit approaches buy efficiency at the cost of stronger distributional assumptions; AMCE-OLS is more transparent for causal interpretation (Train, 2009).

Key decisions

Number of attributes. Five to seven attributes per task is the typical range. Beyond seven or eight, cognitive burden increases and respondents begin to satisfice — focusing only on the one or two attributes they care about most and ignoring the rest. Bansak et al. (2019) show that conjoint task quality degrades meaningfully as the number of attributes grows beyond this range. Include only attributes that are genuinely central to the research question.

Levels per attribute. Two to five levels per attribute is standard. More levels give finer-grained estimates of the attribute’s effect but reduce the frequency with which each level appears and therefore reduce precision per level. Levels should be realistic and recognisably distinct — small numerical differences are difficult for respondents to process reliably.

Number of tasks per respondent. Five to ten tasks is the norm for in-person and tablet-based surveys. Fewer tasks reduce statistical power but limit respondent fatigue; more tasks can introduce learning or anchoring effects across the sequence. Bansak et al. (2018) tested designs up to 30 tasks and found that fatigue and satisficing effects on AMCE estimates are surprisingly modest for engaged respondents through about ten tasks — this is the empirical basis for the “ten tasks” guideline. Robustness checks that include task order as a covariate are good practice.

Profile generation. With even modest attribute counts the full factorial space (e.g., 3⁵ = 243 combinations) is large, but this is not a problem — random independent draws sample from that space and produce orthogonal attributes in expectation. The standard approach for survey-experimental conjoint is independent random draws for each profile in each task. Two complications matter in practice:

  • Restricted designs. Some combinations are implausible (“MD degree” + “age 18”, or “low-cost private school” + “high teacher pay”) and should be excluded. Restricted designs break the independence-of-attributes condition; the conditional AMCE (Egami & Imai, 2019) is the correct estimand and cregg supports it.
  • D-efficient fractional factorial designs. For small samples (below ~500 respondents × 10 tasks) or designs with many attributes, a D-efficient fractional factorial design produced via R’s support.CEs / idefix or Stata’s dcreate (Hole) improves precision substantially over random independent draws. Random draws are fine at typical sample sizes; below that, build the design.

Profile position within a task. Within each task, randomise which profile appears on the left vs the right. Position effects are well-documented in the conjoint literature (a small left-side bias is typical), and an unrandomised position confounds position with attribute composition. Store the realised position alongside the attribute values so the analyst can include it as a covariate or balance-check it.

Forced choice versus opt-out. Most survey-based conjoint experiments use forced choice — respondents must select one of the two profiles shown. Adding a “neither” or “no preference” option changes the estimand: instead of “given a choice, which is preferred”, the estimand becomes “would the respondent choose either of these at all”, which is closer to revealed-preference logic and requires a multinomial or conditional-logit formulation with a third “none” alternative. Forced choice is appropriate when the research context involves genuine decisions where opting out is not realistic; opt-out is appropriate when the decision to engage at all is part of the research question.

Analysis approach. AMCE-OLS with respondent-clustered SEs is the modern standard for population-average causal effects (Hainmueller et al., 2014). Marginal Means (MMs) are the recommended estimand for any cross-group comparison because they are baseline-free (Leeper, Hobolt & Tilley, 2020). Conditional logit and mixed logit are appropriate when welfare estimates (willingness-to-pay) or individual-level heterogeneity are the primary target, at the cost of stronger distributional assumptions. State the estimand the analysis is targeting before choosing the estimator.

Caveats & common mistakes

AMCE is not “the average respondent’s preference.” Abramson, Koçak and Magazinnik (2022) show that the AMCE is a weighted average in which strongly-felt minority preferences can swamp weakly-felt majority preferences. A majority-preferred option can have a negative AMCE; an option with a large positive AMCE need not be the modal choice. So do not write “the average respondent prefers X” on the basis of a positive AMCE. Report Marginal Means alongside AMCEs so the reader sees the unconditional choice frequency at each level, and describe the AMCE as the average causal effect on choice probability rather than as a statement about what respondents actually want.

For subgroup comparisons, use Marginal Means — not subgroup AMCEs. Leeper, Hobolt and Tilley (2020) show that comparing AMCEs across subgroups is misleading. The AMCE depends on which level you pick as the baseline, so two analysts who pick different baselines can reach opposite conclusions about subgroup heterogeneity from the same data. Marginal Means don’t have this problem — they describe the unconditional probability of choice at each level, no baseline required. Use cj(estimate = "mm") in R or margins, over(...) in Stata. The pooled AMCE remains a valid summary of the average causal effect; just don’t lean on subgroup AMCEs to claim that preferences differ across groups.

Bimodal preferences masked by the pooled estimate. Even after correcting for baseline-dependence, a near-zero pooled AMCE can mask sharply bimodal preferences — for example, half the sample strongly favours attribute level X and half strongly opposes it, averaging to zero. Check with MMs by subgroup or by fitting a mixed logit with random coefficients (gmnl, apollo in R; mixlogit in Stata) and inspecting the dispersion of individual-level coefficients.

Hypothetical bias. Conjoint experiments elicit stated preferences for hypothetical profiles. These may differ from the choices respondents would make with real consequences. Hypothetical bias is context-dependent and hard to quantify without a revealed-preference benchmark. Standard remedies include cheap-talk scripts (Cummings & Taylor, 1999) that explicitly remind respondents of the gap between hypothetical and real-stakes choices, consequentiality framing (Vossler, Doyon & Rondeau, 2012) that connects the task to a downstream decision the respondent has reason to care about, and incentive compatibility where the choice format is constructed so that truthful reporting is the dominant strategy. Hainmueller, Hangartner and Yamamoto (2015) validate conjoint estimates against a natural experiment on Swiss naturalisation and find that, for that case, the conjoint correctly predicted real-world choices — useful baseline evidence, but not a general guarantee.

Within-task position (left/right) effects. Respondents tend to choose the left-side profile slightly more often than chance. Position must be randomised at the form level and either controlled for as a covariate or balance-checked in the analysis. A non-trivial position coefficient indicates the form did not randomise position correctly.

Cognitive overload and satisficing. Respondents facing complex profiles may simplify the task by focusing on one or two attributes and ignoring the rest. This produces estimates that understate the marginal effects of neglected attributes. Diagnose with: variance in choice across tasks within respondent (very low variance suggests always-left/always-right), time-per-task distribution (under 5 seconds suggests not reading), and a screener task with a dominant profile that flags respondents who fail to pick the obvious choice. Short, clear attribute labels and level descriptions are essential, especially in low-literacy field settings.

Profile similarity and dominance. If two profiles in a task happen to be very similar on all attributes, respondents may respond idiosyncratically rather than meaningfully. If one profile clearly dominates the other on all attributes, the task provides no information about tradeoffs. Neither requires special treatment in large samples — they average out — but both are worth monitoring.

Order effects. Preferences estimated from early tasks in the sequence may differ from those estimated from later tasks, as respondents learn the task format, fatigue, or anchor on earlier profiles. Include task order as a covariate in robustness checks.

Analysis Guide

import pandas as pd
import numpy as np
import statsmodels.formula.api as smf
from statsmodels.stats.multitest import multipletests

# Note: there is no Python equivalent of cregg. The OLS-with-dummies-and-
# respondent-cluster-SE approach below is the standard route for AMCEs and is
# equivalent to cregg::cj(estimate = "amce") in R. Marginal Means can be computed
# directly from primitives (group means of the choice indicator) — shown below —
# but cregg's full MM/MM-differences feature set (proper SEs, plotting, formal
# tests of subgroup heterogeneity) has no maintained Python port; use the R tab
# for confirmatory MM analyses. For welfare/WTP, use linearmodels or biogeme; for
# mixed-coefficient models, xlogit (PyPI). For fractional-factorial design
# generation, use R's support.CEs / idefix.

attrs = ["attr1_level", "attr2_level", "attr3_level", "attr4_level", "attr5_level"]

# 1. Randomisation balance — regress each attribute on the others; under correct
#    independent random draws the F-test should be non-significant. Any strong
#    relationship signals a coding bug in profile generation. As a frequency-based
#    alternative, inspect pd.crosstab(df.attr1_level, df.attr2_level) for parity
for a in attrs:
  others = [x for x in attrs if x != a]
  formula = f"C({a}) ~ " + " + ".join(f"C({x})" for x in others)
  fit = smf.ols(formula, data=df).fit()
  print(f"{a}: F = {fit.fvalue:.3f}, p = {fit.f_pvalue:.3f}")

# 2. Pooled AMCE — cluster at respondent (NOT task) because each respondent
#    contributes 2 x T correlated observations across both tasks and the two
#    profiles per task. Include C(position) to absorb left/right position effects
amce_formula = "choice ~ " + " + ".join(f"C({a})" for a in attrs) + " + C(position)"
fit_amce = smf.ols(amce_formula, data=df).fit(
  cov_type="cluster", cov_kwds={"groups": df["respondent_id"]}
)
print(fit_amce.summary())

# 3. Marginal Means via groupby — for each attribute level, the unconditional mean
#    of the choice indicator (averaged over all other attributes and respondents).
#    This is the baseline-free estimand recommended for any cross-group comparison
#    (Leeper, Hobolt & Tilley 2020). Subgroup AMCEs are baseline-dependent and can
#    flip sign when the baseline level is changed
for a in attrs:
  mm = df.groupby([a, "treatment_arm"])["choice"].mean().unstack("treatment_arm")
  mm["difference"] = mm.iloc[:, 1] - mm.iloc[:, 0]
  print(f"\nMarginal Means for {a}:")
  print(mm)
# Note: SEs on these MMs and a formal test of MM-differences require either a
# manual delta-method computation (clustered by respondent) or cregg in R

# 4. Subgroup AMCE via interactions — fit the full model with each attribute
#    interacted with the subgroup variable; the interaction coefficients are the
#    differences in conditional AMCEs. Use only for descriptive comparison alongside
#    MMs; do not lean on subgroup AMCEs alone for heterogeneity claims
het_formula = (
  "choice ~ (" + " + ".join(f"C({a})" for a in attrs) + ") * C(female) + C(position)"
)
fit_het = smf.ols(het_formula, data=df).fit(
  cov_type="cluster", cov_kwds={"groups": df["respondent_id"]}
)
print(fit_het.summary())

# 5. Multiple-testing correction — with many attribute levels x subgroups the number
#    of hypotheses can easily exceed 50; BH for exploratory analysis, Holm for
#    confirmatory subgroup tests. Romano-Wolf step-down has no maintained Python
#    implementation; use 'rwolf' in Stata or the wildrwolf package in R for
#    confirmatory analyses
pvals = fit_amce.pvalues.drop("Intercept").values
_, p_bh, _, _ = multipletests(pvals, method="fdr_bh")
_, p_holm, _, _ = multipletests(pvals, method="holm")
print("BH-adjusted p-values:", p_bh)
print("Holm-adjusted p-values:", p_holm)

SurveyCTO / XLSForm

Conjoint tasks are generated before fieldwork using a pre-generation script (R or Python) and loaded into the form via pulldata() from a server-side dataset. Do not regenerate attribute assignments using random() at runtime — runtime randomisation cannot be reproduced and is harder to verify during analysis.

Pre-generation script. Produce a profiles.csv file keyed on respondent_id × task_id × profile_position with one column per attribute. For typical sample sizes (N ≥ 500 respondents × ~10 tasks), independent random draws per row are sufficient and ensure orthogonality in expectation. For smaller samples or designs with many attributes, generate a D-efficient fractional factorial design using R’s support.CEs or idefix, or Stata’s dcreate (Hole). Randomise profile position (left vs right) within each task at this stage and store it in the CSV.

Load into the form via pulldata(). Attach profiles.csv to the form server-side and load each attribute with a key composed of respondent, task, and position:

calculation: pulldata('profiles', 'attr1', 'key',
                      concat(${respondent_id},'_',${task_id},'_',${profile_position}))

The form then displays the loaded attribute values using ${attr1} etc. Persist all attribute values, the task number, and the profile position to the submitted data — they are required for estimation.

Attention/screener check. Include one task with a dominant profile (one alternative clearly better on a single salient attribute) so the analyst can flag respondents who fail to pick the obvious choice. Drop or down-weight failers in a pre-specified robustness analysis. Bansak et al. (2018) document that attention varies substantially even in well-administered surveys.

Cheap-talk script (optional but recommended where hypothetical bias is a concern). Add a short note before the first task acknowledging that the choices are hypothetical and asking the respondent to answer as they would in a real decision. Cummings and Taylor (1999) and Vossler, Doyon and Rondeau (2012) document non-trivial reductions in hypothetical bias from this and related framings.

Reading the output

  • Each AMCE coefficient is the percentage-point change in the probability of a profile being chosen relative to the omitted baseline level for that attribute, averaged over all other attributes and respondents. On a forced-binary choice task the AMCE is bounded in [−0.5, 0.5].
  • A coefficient of 0.08 means profiles with that level were chosen 8 percentage points more often than profiles with the baseline level, on average. |AMCE| ≥ 0.15 indicates a strong driver of choice; |AMCE| ≤ 0.03 is typically below practical relevance even if statistically significant in large samples.
  • A near-zero pooled AMCE does not mean the attribute is unimportant. It may mean the specific levels tested do not span a meaningful range, the level was weighed similarly to the omitted baseline, or — critically — that bimodal preferences are cancelling out (Abramson, Koçak & Magazinnik 2022). Check with MMs by subgroup or mixed-logit individual-level dispersion before concluding the attribute is unimportant.
  • Position coefficient (from the regression with i.position): should be small in absolute terms (|AMCE_position| < 0.03 is the usual benchmark). A larger position effect indicates the form did not randomise position correctly or there is a systematic left-side bias in the population.
  • Randomisation balance check (cregg estimate = "freq"): empirical attribute frequencies should match the design draws. Substantial deviation flags a profile-generation bug.
  • Marginal Means subgroup comparison: report MMs by subgroup and the differences in MMs across groups; do not report subgroup AMCEs as the basis for substantive heterogeneity claims (they are baseline-dependent).
  • Satisficing diagnostics: flag respondents who (a) take under 5 seconds per task, (b) show very low variance in choice across tasks (e.g., always left or always right), or (c) fail the screener / dominant-profile task. Pre-specify whether the primary analysis includes or excludes failers.
  • Multiple-testing correction: with many attribute levels × subgroups, apply BH (FDR) for exploratory results and Holm or Romano-Wolf for confirmatory subgroup claims. Report both corrected and uncorrected p-values.

References

Abramson, S. F., Koçak, K., & Magazinnik, A. (2022). What do we learn about voter preferences from conjoint experiments? American Journal of Political Science, 66(4), 1008–1020. https://doi.org/10.1111/ajps.12714

Bansak, K., Hainmueller, J., Hopkins, D. J., & Yamamoto, T. (2018). The number of choice tasks and survey satisficing in conjoint experiments. Political Analysis, 26(1), 112–119. https://doi.org/10.1017/pan.2017.40

Bansak, K., Hainmueller, J., Hopkins, D. J., & Yamamoto, T. (2019). Beyond the breaking point? Survey satisficing in conjoint experiments. Political Science Research and Methods, 7(1), 121–135. https://doi.org/10.1017/psrm.2017.40

Bansak, K., Hainmueller, J., Hopkins, D. J., & Yamamoto, T. (2021). Conjoint survey experiments. In J. N. Druckman & D. P. Green (Eds.), Advances in experimental political science (pp. 19–41). Cambridge University Press.

Cummings, R. G., & Taylor, L. O. (1999). Unbiased value estimates for environmental goods: A cheap talk design for the contingent valuation method. American Economic Review, 89(3), 649–665. https://doi.org/10.1257/aer.89.3.649

de la Cuesta, B., Egami, N., & Imai, K. (2022). Improving the external validity of conjoint analysis: The essential role of profile distribution. Political Analysis, 30(1), 19–45. https://doi.org/10.1017/pan.2020.40Author page (preprint)

Egami, N., & Imai, K. (2019). Causal interaction in factorial experiments: Application to conjoint analysis. Journal of the American Statistical Association, 114(526), 529–540. https://doi.org/10.1080/01621459.2018.1476246

Hainmueller, J., Hangartner, D., & Yamamoto, T. (2015). Validating vignette and conjoint survey experiments against real-world behavior. Proceedings of the National Academy of Sciences, 112(8), 2395–2400. https://doi.org/10.1073/pnas.1416587112

Hainmueller, J., Hopkins, D. J., & Yamamoto, T. (2014). Causal inference in conjoint analysis: Understanding multidimensional choices via stated preference experiments. Political Analysis, 22(1), 1–30. https://doi.org/10.1093/pan/mpt024

Leeper, T. J. (2023). cregg: Simple conjoint tidying, analysis, and visualization [R package]. CRAN. https://CRAN.R-project.org/package=cregg

Leeper, T. J., Hobolt, S. B., & Tilley, J. (2020). Measuring subgroup preferences in conjoint experiments. Political Analysis, 28(2), 207–221. https://doi.org/10.1017/pan.2019.30

McFadden, D. (1974). Conditional logit analysis of qualitative choice behavior. In P. Zarembka (Ed.), Frontiers in econometrics (pp. 105–142). Academic Press. — Author page (PDF)

Ono, Y., & Burden, B. C. (2018). The contingent effects of candidate sex on voter choice. Political Behavior, 41(3), 583–607. https://doi.org/10.1007/s11109-018-9464-6

Train, K. E. (2009). Discrete choice methods with simulation (2nd ed.). Cambridge University Press. — Author page (free PDF)

Vossler, C. A., Doyon, M., & Rondeau, D. (2012). Truth in consequentiality: Theory and field evidence on discrete choice experiments. American Economic Journal: Microeconomics, 4(4), 145–171. https://doi.org/10.1257/mic.4.4.145

Last updated: 5 June 2026