What it is
Ambiguity aversion is a preference for known risks over unknown ones — a distinction first articulated by Knight (1921). A risk-averse person dislikes variance in outcomes. An ambiguity-averse person dislikes not knowing the probability distribution itself — they prefer a gamble with specified probabilities over one without, even when expected values are identical.
The canonical demonstration is the Ellsberg (1961) two-urn paradox. Urn R contains 100 balls in a known 50/50 mix of red and black. Urn A contains 100 balls in unknown proportions. Asked to bet on red, most respondents pick Urn R; asked to bet on black, they pick Urn R again. This pattern violates Savage’s Sure-Thing Principle and is inconsistent with any subjective probability the respondent could hold over Urn A’s composition: under expected-utility theory with subjective beliefs, the two preferences imply both P(red in A) < 0.5 and P(black in A) < 0.5, which sum to less than 1. The systematic preference is for the known distribution over the unknown one. Across published samples, roughly 60–75% of respondents exhibit this pattern (Trautmann & van de Kuilen, 2015; Dimmock, Kouwenberg, Mitchell & Peijnenburg, 2016).
Field implementations often use smaller numbers of balls (10 or 20 per bag) for tangibility, with the understanding that the conceptual structure — known versus unknown distribution — is what carries the elicitation. The behavioural correlates of ambiguity aversion include lower stock market participation, lower insurance demand, lower portfolio diversification, and slower technology adoption (Dimmock, Kouwenberg, Mitchell & Peijnenburg, 2016; Dimmock, Kouwenberg & Wakker, 2015).
When to use it
Ambiguity aversion is worth measuring when the primary research question concerns decision-making under genuine uncertainty — as opposed to calculable risk — or when the study evaluates an intervention that changes the information environment around a decision. The empirical case for elicitation is strong: in nationally representative samples, ambiguity aversion predicts stock market non-participation, lower insurance demand, and lower portfolio diversification, over and above risk aversion (Dimmock, Kouwenberg, Mitchell & Peijnenburg, 2016).
Engle-Warnick, Escobal and Laszlo (2007) used Ellsberg-style tasks alongside risk preference measures in a study of technology adoption in Peru, finding that ambiguity aversion predicts adoption decisions over and above risk aversion alone. Akay, Martinsson, Medhin and Trautmann (2012) document ambiguity attitudes among rural Ethiopian farmers, providing a developing-country benchmark. Sutter, Kocher, Glätzle-Rützler and Trautmann (2013) link experimental ambiguity measures to adolescents’ field behaviour.
The method is less useful when the elicitation procedure cannot establish credible ambiguity — when respondents either distrust the procedure (suspecting the ambiguous urn is stacked) or interpret the ambiguous urn as a compound lottery (treating “unknown 0–100% red” as “uniform over 0–100%”). Both possibilities are detectable; both need separate handling — see Caveats.
How it works
There are two main families of elicitation. Choose deliberately based on what parameter you want to estimate.
Binary Ellsberg classification (the simpler method). Two paired choice tasks. In each, the respondent is shown two urns:
- Urn R (risky): A known mix of red and black, typically 50/50 (50/50 with 100 balls in the original; 10 balls is the standard field variant).
- Urn A (ambiguous): Unknown proportions of red and black, total ball count known.
Task 1: “Which urn do you want to draw from? If you draw a red ball, you win [amount].” Task 2: “Which urn do you want to draw from? If you draw a black ball, you win [amount].”
A respondent who picks Urn R in both is classified ambiguity-averse; Urn A in both is ambiguity-seeking; one each is consistent with a subjective prior over Urn A. The output is a categorical label, not a numerical parameter. This is the workhorse classification when the analyst needs a binary indicator for downstream regressions.
Matching probability (Dimmock, Kouwenberg & Wakker, 2015; Dimmock et al., 2016). The recommended modern method when a numerical ambiguity premium is wanted. The respondent reports the probability p* in Urn R at which they are indifferent between betting on Urn R at p* and betting on the winning colour from Urn A. The ambiguity premium is then:
b = 0.5 − m_low
where m_low is the matching probability for the low-payoff colour. Under ambiguity neutrality b = 0. The insensitivity index a = 1 − (m_high − m_low) captures likelihood insensitivity — the tendency to treat probabilities as more similar than they are (Baillon, Huang, Selim & Wakker, 2018). Together (a, b) decompose ambiguity attitude into aversion and insensitivity.
Source-based methods (Baillon et al., 2018). The frontier methodology generalises Ellsberg-style elicitation to natural events — climate outcomes, sports results, financial-market movements — instead of artificial urns. For field work on technology adoption under climate uncertainty or health-product uptake under disputed efficacy, source-based methods often have stronger external validity than the urn task.
The three theoretical models. The estimand depends on which decision-theoretic model is being applied. The three main families:
- Maxmin expected utility (MEU) — Gilboa and Schmeidler (1989). The respondent has a set of priors and acts as if maximising the worst-case expected utility. The binary RR classification approximates MEU under a strong-aversion assumption.
- Choquet expected utility (CEU) — Schmeidler (1989). The respondent uses a non-additive capacity instead of a probability measure.
- Smooth ambiguity model — Klibanoff, Marinacci and Mukerji (2005). Ambiguity attitude is captured by a second-order utility function over expected utilities; the parameter ρ analogous to risk aversion governs the smooth ambiguity premium.
The matching probability and ambiguity premium b are model-light summary statistics interpretable under all three; the structural parameters (α in α-MEU, ρ in smooth ambiguity) require richer designs.
Identification — the load-bearing RCLA problem (Halevy, 2007). Halevy’s (2007) experiment showed that respondents who satisfy reduction of compound lotteries (RCLA) in a control task are largely ambiguity-neutral, and respondents who fail RCLA account for most of measured ambiguity aversion. In other words, a great deal of what looks like ambiguity aversion in Ellsberg-style tasks is a failure to reduce a compound lottery, not a primitive preference. This is the load-bearing identification problem for any urn-based elicitation. The recommended diagnostic is a parallel two-stage compound-lottery task (Halevy’s “Task 4”); respondents who fail it should be flagged and the ambiguity-aversion estimate reported both including and excluding them.
Identification assumptions for the binary classification. Five conditions must hold for the RR pattern to identify ambiguity aversion:
- Sure-Thing Principle — Savage’s axiom holds for the respondent. Halevy’s RCLA failures are a violation of (a weaker but related) compound lottery reduction.
- Symmetric beliefs over Urn A — the respondent doesn’t have a strong prior that one colour is more likely. The two-task structure tests this.
- No distrust — the respondent doesn’t suspect the ambiguous urn is stacked. See Key Decisions for the second-mover procedure that addresses this.
- Comprehension of the difference between a known and unknown distribution.
- Risk attitude held constant across tasks — the same risk aversion governs both choices; this lets the difference in choice identify ambiguity attitude rather than confounding risk.
The matching-probability method holds the prize fixed and varies only the p* in the risky urn, which controls cleanly for risk attitude. The binary method does not.
Key decisions
Binary classification vs matching probability vs source-based. The first design decision. The binary Ellsberg classification is cheap and field-friendly but information-poor — it gives a 4-cell categorical outcome. The matching-probability method gives a numerical premium (Dimmock et al., 2015; Dimmock et al., 2016) and is the modern survey workhorse. Source-based elicitation (Baillon et al., 2018) replaces urns with natural events and is more externally valid when the substantive uncertainty is about a real-world source (climate, prices, treatment efficacy). Default recommendation: matching probability for new studies with respondent capacity; binary for very short batteries or low-numeracy settings; source-based when the research question is domain-specific.
Physical vs described urns. The urn metaphor is standard but may be unfamiliar in some field contexts. Physical urns — opaque bags with actual coloured balls — work better than verbal descriptions: the respondent watches the risky bag being filled in known proportions and sees the ambiguous bag filled without the proportions being revealed. Spinner wheels or coloured token draws are alternatives.
Separating ambiguity from distrust — second-mover procedure. In field settings, respondents may choose Urn R not because they are ambiguity-averse but because they suspect the ambiguous bag has been stacked against them. The standard fix in the literature is a second-mover procedure (Charness, Karni & Levin, 2013): seal the ambiguous bag in advance, then let the respondent themselves choose which colour pays after the bag is sealed. With the colour decision in the respondent’s hands, stacking the bag would have to anticipate their choice — implausibly. This is the strongest mitigation; weaker alternatives (community member fills the bag, public draw) leave residual concerns.
Varying the risky urn composition for matching probability. For the matching probability method, vary the risky urn across several compositions — 5/5, 4/6, 3/7, 2/8 red/black — keeping the ambiguous urn fixed. The switching point identifies an interval for m. Standard implementations use 8–10 risky-urn compositions in a Holt-Laury-style list. Multiple switching is a data-quality concern (typically 5–15% of respondents; Charness, Gneezy & Imas, 2013); use the first switching point and flag for sensitivity.
Incentivising the task. Real stakes produce more meaningful responses than hypothetical, but the evidence base for hypothetical bias in ambiguity tasks is thinner than for risk (Trautmann & van de Kuilen, 2015, §3.5) — direction and magnitude are less established. Pay a salient amount (one to two hours of local wage) and use a random-task-selection mechanism if embedded in a battery so wealth effects from prior tasks do not contaminate ambiguity choices.
Task-order randomisation. Order matters and is empirically substantial (Trautmann & van de Kuilen, 2015). Randomise: (a) which colour pays first (red-first vs black-first within the ambiguity task) to neutralise within-task order, and (b) whether the ambiguity battery runs before or after the risk battery (the Holt-Laury and Eckel-Grossman tasks). Store the order indicator and include it as a covariate in robustness checks.
Embedding in a battery. Ambiguity tasks are commonly administered alongside risk preference tasks. The cross-task contamination concerns are the same as in the risk guides: wealth effects (use end-of-session random selection of one paid task) and framing effects (randomise task order across respondents).
Sample size. For a binary classification outcome with expected RR prevalence of ~0.60, detecting a 10-percentage-point treatment effect at 80% power and α = 0.05 needs N ≈ 388 per arm under no clustering. For the matching probability with expected SD of b ≈ 0.15, detecting a 0.05 difference in mean b needs N ≈ 140 per arm. For cluster-randomised designs, multiply by 1 + (m̄ − 1)ρ.
Caveats & common mistakes
Compound lottery reduction is the load-bearing identification problem (Halevy, 2007). Some respondents treat the ambiguous urn as a compound lottery — “the bag was filled somehow, so there is some distribution over its composition.” Under this interpretation, a risk-averse respondent prefers the known urn because compound lotteries have higher effective variance, not because of a primitive ambiguity aversion. Halevy’s (2007) experimental finding is stronger than “much of what looks like ambiguity aversion”: respondents who satisfy reduction of compound lotteries (RCLA) in a control task are largely ambiguity-neutral, and respondents who fail RCLA are responsible for most of measured ambiguity aversion in his sample. Treat this as the central identification challenge for Ellsberg-style elicitation, not as a side caveat. Include a parallel compound-lottery diagnostic task and report ambiguity-aversion shares both including and excluding RCLA-failers.
Consistency across colours. A respondent who chooses Urn R for red but Urn A for black is consistent with a subjective belief that black is more likely than red in Urn A (a prior of ≥ 0.5 for black). The symmetric pattern reveals the opposite prior. Coding respondents as ambiguity-averse from a single task misses this; the two-task structure is the minimum for valid classification.
Competence hypothesis (Heath & Tversky, 1991). People are less ambiguity-averse in domains where they feel competent. This has direct implications for field design: an urn task measures ambiguity attitude over an abstract source of uncertainty about which no one is “competent”. If the substantive research question is about decisions where respondents feel knowledgeable — farmers about local weather, healthcare workers about a familiar treatment — the abstract urn measure may understate the relevant ambiguity attitude. Where possible, complement the urn task with a domain-specific source-based elicitation (Baillon et al., 2018).
Multi-switching in the matching-probability extended design. Typically 5–15% of respondents switch back and forth in the multi-row matching-probability list (Charness, Gneezy & Imas, 2013). Use the first switching point, flag multi-switchers as a reliability indicator, run sensitivity excluding them. Rates above 20% warrant reviewing the comprehension procedure.
Small classification cells in the binary design. In the four-cell binary classification (RR / RA / AR / AA), AA (ambiguity-seeking) is typically 12–18% in US samples (Dimmock et al., 2016) and lower in developing-country samples. RR (ambiguity-averse) is the modal cell at 60–75%. Rates outside this range warrant a comprehension check.
Gender differences. Borghans, Heckman, Golsteyn and Meijers (2009) document a gender gap in ambiguity aversion — women more ambiguity-averse than men — but smaller than the gender gap for risk aversion. Include gender as a control where relevant; do not extrapolate the magnitude between risk and ambiguity gaps.
Hypothetical bias evidence is thinner for ambiguity than for risk. The literature on hypothetical-vs-real elicitation for ambiguity is much smaller than for risk preferences (Trautmann & van de Kuilen, 2015, §3.5). Direction and magnitude of hypothetical bias for ambiguity are not well established. Pilot the procedure under real stakes where possible; if hypothetical is unavoidable, report the choice and the limitation.
Analysis Guide
import numpy as np
import pandas as pd
import statsmodels.formula.api as smf
from statsmodels.stats.proportion import proportion_confint
# Conventions:
# urn_red: 1 = chose risky urn R (known), 0 = chose ambiguous urn A, when red pays
# urn_black: 1 = chose risky urn R, 0 = chose ambiguous urn A, when black pays
# For the matching-probability extended design, switch_row_m holds the first
# row at which the respondent preferred the risky urn at p* over the ambiguous
# urn; the per-row p* values feed an external lookup table to recover m
# 0. Filter to respondents with both binary tasks observed — NA on either task
# must be excluded before classification, otherwise np.select will mis-bin
df = df.dropna(subset=['urn_red', 'urn_black'])
# 1. Classify the binary Ellsberg pattern — only RR (risky in both) is strictly
# ambiguity-averse and violates Savage's Sure-Thing Principle. RA / AR are
# consistent with an asymmetric subjective prior. AA is ambiguity-seeking
df['ambig_type'] = np.select(
[(df['urn_red'] == 1) & (df['urn_black'] == 1),
(df['urn_red'] == 0) & (df['urn_black'] == 0)],
['ambiguity-averse', 'ambiguity-seeking'],
default='subjective-prior')
print(df['ambig_type'].value_counts(normalize=True).mul(100).round(1))
# 2. Prevalence with a binomial 95% CI — the headline number for most studies.
# Trautmann & van de Kuilen 2015 give a 60-75% reference range for RR
df['ambig_averse'] = ((df['urn_red'] == 1) & (df['urn_black'] == 1)).astype(int)
n_aa = df['ambig_averse'].sum()
n_tot = len(df)
lo, hi = proportion_confint(n_aa, n_tot, alpha=0.05, method='wilson')
print(f'Ambiguity-averse share: {n_aa / n_tot:.2%} 95% CI [{lo:.2%}, {hi:.2%}]')
# 3. Distribution diagnostics — share above ~85% RR or below ~30% RR warrants a
# comprehension review. AA above 20% suggests respondents misunderstood or
# distrusted the risky urn
print('Share AA (ambiguity-seeking):', (df['ambig_type'] == 'ambiguity-seeking').mean())
print('Share AR + RA (subjective prior):', (df['ambig_type'] == 'subjective-prior').mean())
# 4. Regress the LPM on covariates with cluster-robust SEs — most ambiguity tasks
# are enumerator- or session-clustered, so HC1/HC2 understate uncertainty. For
# a clean LPM with predicted probabilities outside [0,1], report a logit
# alternative as a robustness check
fit = smf.ols('ambig_averse ~ age + female + log_hh_expenditure + treatment',
data=df).fit(cov_type='cluster',
cov_kwds={'groups': df['enumerator_id']})
print(fit.summary())
# Optional: logit alternative to handle LPM out-of-range fitted values
# fit_logit = smf.logit('ambig_averse ~ age + female + log_hh_expenditure + treatment',
# data=df).fit(cov_type='cluster',
# cov_kwds={'groups': df['enumerator_id']})
# 5. Matching-probability ambiguity premium (for the extended design only) —
# compute the implied m from the switch row in the multi-row task and the
# ambiguity premium b = 0.5 - m. m_table maps switch_row_m to the implied m
# df = df.merge(m_table, on='switch_row_m', how='left')
# df['ambiguity_premium'] = 0.5 - df['m_low']
# print(df['ambiguity_premium'].describe()) XLSForm / SurveyCTO
Task-order randomisation. Draw a binary indicator at the start of the ambiguity battery to determine whether red-pays or black-pays is asked first:
type name calculation
calculate task_order once(if(random() < 0.5, 0, 1))
Persist task_order to the submission so it can be included as a covariate in the analysis. The outer once() is critical — without it the value re-randomises on every form recompute.
Binary classification implementation. Use two consecutive select_one questions, one per colour, with a relevant condition gating them on task_order. Label the response options clearly: “Bag with 5 red and 5 black balls” vs “Bag with unknown mix”. Display each question on its own screen to prevent the respondent from changing their first answer after seeing the second.
Second-mover procedure for distrust. Seal the ambiguous bag in advance and let the respondent choose which colour pays after the bag is sealed. Store the respondent-chosen colour to a persisted field.
Compound-lottery diagnostic. Include a parallel control task — a clearly-described two-stage compound lottery with known second-stage probabilities — and persist the respondent’s reduction-of-compound-lotteries (RCLA) compliance to the submission. Halevy (2007) shows that RCLA failers account for most measured ambiguity aversion; analysing the data with and without them is the standard diagnostic.
Matching-probability extended design. Pre-randomise the multi-row matching list off-form and load via pulldata():
type name calculation
calculate composition_n pulldata('mp_design','composition','key',
concat(${respondent_id},'_',${row_idx}))
Build mp_design.csv with one row per (respondent_id × row_idx) holding the risky-urn composition for each row. The form loops over row_idx from 1 to N using a repeat block and presents the binary choice for each composition; switching point is identified post-hoc.
Persisted fields. At minimum: urn_red, urn_black, task_order, rcla_passed (compound-lottery control task), respondent-chosen-colour. For the extended design, also persist each row’s choice and the computed switch_row_m.
Reading the output
- Expected RR share: 60–75% in published samples (Trautmann & van de Kuilen, 2015). RR above 85% suggests respondents are using a “safe choice” heuristic without engaging with the ambiguous-urn distinction; below 30% suggests they misunderstand the task or accept the ambiguous urn as a 50/50 compound lottery. Either warrants a comprehension review.
ambig_type == "ambiguity-averse"(RR) is the strict Ellsberg classification — chose the risky urn in both tasks. This pattern violates Savage’s Sure-Thing Principle and is inconsistent with any subjective prior the respondent could hold.ambig_type == "subjective-prior"(RA or AR) — chose different urns across tasks. Consistent with a respondent who assigns a non-50/50 belief to the ambiguous urn (e.g., believes black is more likely if they choose Urn R for red and Urn A for black). Not irrationality.ambig_type == "ambiguity-seeking"(AA) — always chose the ambiguous urn. Expect 12–18% in US samples (Dimmock et al., 2016), lower in developing-country samples. Above 20% may indicate misunderstanding or distrust of the risky urn.- RCLA diagnostic. Report the ambiguity-aversion share both including and excluding respondents who fail the compound-lottery control task (Halevy, 2007). A large gap is the central identification concern; a small gap means RCLA failures are not driving the result.
- Regression interpretation. The LPM coefficient on a covariate is the percentage-point change in the probability of ambiguity aversion. A positive treatment coefficient of 0.10 means treatment increased the ambiguity-aversion share by 10 percentage points. Report the logit alternative as a robustness check given the LPM’s out-of-range fitted-value issue. Use cluster-robust SEs at the enumerator or session level; HC1/HC2 understate uncertainty when sessions are correlated.
- Matching-probability output. The ambiguity premium b = 0.5 − m_low is positive for ambiguity-averse respondents and zero for ambiguity-neutral. The insensitivity index a = 1 − (m_high − m_low) is positive when respondents treat extreme probabilities as more similar than they are. Both should be reported with bootstrap SEs.
References
Akay, A., Martinsson, P., Medhin, H., & Trautmann, S. T. (2012). Attitudes toward uncertainty among the poor: An experiment in rural Ethiopia. Theory and Decision, 73(3), 453–464. https://doi.org/10.1007/s11238-011-9250-y
Baillon, A., Huang, Z., Selim, A., & Wakker, P. P. (2018). Measuring ambiguity attitudes for all (natural) events. Econometrica, 86(5), 1839–1858. https://doi.org/10.3982/ECTA14370
Borghans, L., Heckman, J. J., Golsteyn, B. H. H., & Meijers, H. (2009). Gender differences in risk aversion and ambiguity aversion. Journal of the European Economic Association, 7(2–3), 649–658. https://doi.org/10.1162/JEEA.2009.7.2-3.649
Camerer, C. F., & Weber, M. (1992). Recent developments in modeling preferences: Uncertainty and ambiguity. Journal of Risk and Uncertainty, 5(4), 325–370. https://doi.org/10.1007/BF00122575
Charness, G., Gneezy, U., & Imas, A. (2013). Experimental methods: Eliciting risk preferences. Journal of Economic Behavior & Organization, 87, 43–51. https://doi.org/10.1016/j.jebo.2012.12.023
Charness, G., Karni, E., & Levin, D. (2013). Ambiguity attitudes and social interactions: An experimental investigation. Journal of Risk and Uncertainty, 46(1), 1–25. https://doi.org/10.1007/s11166-012-9157-1
Dimmock, S. G., Kouwenberg, R., Mitchell, O. S., & Peijnenburg, K. (2016). Ambiguity aversion and household portfolio choice puzzles: Empirical evidence. Journal of Financial Economics, 119(3), 559–577. https://doi.org/10.1016/j.jfineco.2016.01.003
Dimmock, S. G., Kouwenberg, R., & Wakker, P. P. (2015). Ambiguity attitudes in a large representative sample. Management Science, 62(5), 1363–1380. https://doi.org/10.1287/mnsc.2015.2198
Ellsberg, D. (1961). Risk, ambiguity, and the Savage axioms. Quarterly Journal of Economics, 75(4), 643–669. https://doi.org/10.2307/1884324
Engle-Warnick, J., Escobal, J., & Laszlo, S. (2007). Ambiguity aversion as a predictor of technology choice: Experimental evidence from Peru (CIRANO Working Paper 2007s-01). https://cirano.qc.ca/files/publications/2007s-01.pdf
Gilboa, I., & Schmeidler, D. (1989). Maxmin expected utility with non-unique prior. Journal of Mathematical Economics, 18(2), 141–153. https://doi.org/10.1016/0304-4068(89)90018-9
Halevy, Y. (2007). Ellsberg revisited: An experimental study. Econometrica, 75(2), 503–536. https://doi.org/10.1111/j.1468-0262.2007.00755.x
Heath, C., & Tversky, A. (1991). Preference and belief: Ambiguity and competence in choice under uncertainty. Journal of Risk and Uncertainty, 4(1), 5–28. https://doi.org/10.1007/BF00057884
Klibanoff, P., Marinacci, M., & Mukerji, S. (2005). A smooth model of decision making under ambiguity. Econometrica, 73(6), 1849–1892. https://doi.org/10.1111/j.1468-0262.2005.00640.x
Knight, F. H. (1921). Risk, uncertainty, and profit. Houghton Mifflin.
Schmeidler, D. (1989). Subjective probability and expected utility without additivity. Econometrica, 57(3), 571–587. https://doi.org/10.2307/1911053
Sutter, M., Kocher, M. G., Glätzle-Rützler, D., & Trautmann, S. T. (2013). Impatience and uncertainty: Experimental decisions predict adolescents’ field behavior. American Economic Review, 103(1), 510–531. https://doi.org/10.1257/aer.103.1.510
Trautmann, S. T., & van de Kuilen, G. (2015). Ambiguity attitudes. In G. Keren & G. Wu (Eds.), The Wiley Blackwell handbook of judgment and decision making (pp. 89–116). Wiley-Blackwell.
Wakker, P. P. (2010). Prospect theory: For risk and ambiguity. Cambridge University Press.