What it is
Standard surveys ask people what they did or what they want; subjective expectations elicitation asks what they think will happen. Respondents are asked to assign a probability — typically on a 0–100 or 0–10 scale, often using a visual aid — to a future event: the chance that a child who completes secondary school will find formal employment, the probability of a good harvest next season, the likelihood of falling ill in the coming year.
The motivation is that beliefs and preferences are both required to explain behaviour, but conventional surveys confound the two. A farmer who does not adopt a new seed variety may lack information about its yield potential, may be risk-averse, or both. Eliciting subjective expectations separates the belief component and lets the analyst test whether behaviour is consistent with stated beliefs, and whether changing beliefs changes behaviour (Manski, 2004). For risk and time preferences — the other half of the belief–preference decomposition — see the Holt-Laury MPL and Time Preference / Discount Rate guides.
When to use it
Subjective expectations are worth eliciting when the research question is why people make a decision, not just whether they do. If a decision model predicts that behaviour responds to beliefs about returns or risks, directly measuring those beliefs tests whether the model is correct and whether belief updating is a plausible mechanism for changing behaviour.
Attanasio and Kaufmann (2014) elicit Mexican mothers’ and youths’ expectations about the returns to schooling and show that beliefs about returns predict college attendance, holding income constant. Jensen (2010) provides Dominican secondary-school boys with information on the actual returns to schooling and shows that corrected beliefs raise schooling attainment by about 0.2–0.35 years over the next four years — implying that prior underestimation was a binding constraint. Attanasio and Augsburg (2016) use the same elicitation toolkit to estimate the subjective income process in rural India. Delavande, Giné and McKenzie (2011) survey the state of the method in developing countries across health, agriculture, migration, and labour markets; Delavande (2014) updates that survey with more recent evidence.
The method is less useful when the relevant future event is so distant or unfamiliar that respondents have no meaningful prior — asking smallholder farmers in 2005 about the probability of digital payment adoption in five years would have produced noise, not beliefs.
How it works
The basic elicitation asks: “What is the chance, on a scale of 0 to 100, that [event] will happen?” A score of 0 means the respondent believes the event is impossible; 100 means certain; 50 means they think it is equally likely to happen or not.
In practice, three formats are common and most field-credible work prefers a tactile 0–10 token format with non-specialist respondents:
Balls in a bin (preferred for low-numeracy populations). The respondent is given 10 tokens and a two-compartment tray. They place tokens in the “will happen” compartment until the allocation reflects their belief. The number of tokens on the “will happen” side is the probability × 10. Works well with low-numeracy populations, has the lowest focal-point clustering in head-to-head tests, and shares the same scale as the wheel.
Probability wheel. A circular spinner or pie chart divided into 10 or 20 equal segments. The enumerator fills in segments to represent the respondent’s stated probability — 6 out of 10 filled segments represents a 60% chance. Visual, tactile, and does not require numerical fluency (Delavande, Giné & McKenzie, 2011).
Numerical probability scale (0–100). The enumerator explains the scale using anchor examples — “0 means you are certain it will not happen, 100 means you are certain it will, 50 means you think it is just as likely to happen as not” — and asks the respondent to give a number. Simple to administer but requires numeracy and produces the most focal-point clustering.
For continuous outcomes — expected crop yield, expected earnings — the standard approach is to ask for the probability that the outcome falls below a series of thresholds (e.g., “what is the chance your harvest will be less than 10 bags? Less than 20 bags? Less than 30 bags?”). This traces out points on a subjective cumulative distribution function. Fitting a parametric family (beta, triangular, log-normal) to the elicited (threshold, cumulative probability) pairs recovers a continuous CDF from which the mean, variance, and quantiles follow (Dominitz & Manski, 1997; Attanasio, 2009; Delavande, Giné & McKenzie, 2011).
Identification assumptions. What can be recovered from elicited probabilities depends on five assumptions:
- The respondent holds a well-defined subjective probability over the event. A degenerate prior (no model of the event) cannot be recovered; the “50 means don’t know” pattern is the empirical signature of this failure (Bruine de Bruin, Fischhoff, Millstein & Halpern-Felsher, 2000).
- The elicitation format faithfully translates the internal belief into the response. Format effects across percent-chance, fractional, and population-frame frames are non-trivial (Delavande, Giné & McKenzie, 2011); pilot-test before fielding.
- No strategic distortion — neither demand effects nor aspirational over-reporting on highly valued outcomes (child health, schooling success).
- Rounding is benign or modelled. Manski & Molinari (2010) treat focal-point responses as informative about an interval of underlying belief and recover partial-identification bounds. Discarding focal responses as “noise” is one option; modelling them as bounds is the principled alternative.
- Validation against realised outcomes is performed where feasible (Delavande & Rohwedder, 2008; Hurd, 2009; Delavande, 2014). A belief-on-behaviour regression cannot distinguish a calibrated believer from a uniformly over-optimistic one — both produce positive coefficients. Realised-outcome calibration is the discriminating test.
Key decisions
Scale resolution. The 0–100 scale allows fine-grained responses but invites clustering at round numbers (0, 25, 50, 75, 100). A 0–10 scale is coarser but easier to use with visual aids and reduces focal-point responses. For most field surveys with non-specialist respondents, a 0–10 scale with physical tokens or a wheel is the default.
Visual aid vs. verbal elicitation. Verbal probability questions reliably produce mass at focal points — particularly 50, which respondents often use to mean “I don’t know” rather than “I think it’s equally likely either way” (Bruine de Bruin et al., 2000). Visual aids reduce this by making the scale tangible and requiring active allocation rather than a verbal response. When the population is not familiar with probability language, visual aids are strongly preferred.
What to elicit. The choice of events and reference periods should match the decision being studied. For returns to education, elicit the probability of formal employment at a realistic time horizon. For agricultural technology adoption, elicit expected yield under the new technology and under the status quo — both are needed to infer the expected gain. Eliciting central tendency and variance (through three to five threshold probabilities) is more informative than a single point estimate, but the recovered variance is highly sensitive to functional-form assumptions on the subjective distribution (Attanasio, 2009); report the choice and at least one sensitivity check.
Reference class and frame. “Percent chance for you personally” and “out of 100 households like yours, how many will…” are different objects (Manski, 2004, sections 4–5): the first is own-case probability, the second is a population-frame estimate. They are correlated but not identical; use the frame the decision model calls for, not whichever is more comfortable for the respondent.
Distinguishing uncertainty from ignorance. A response of 50 may represent genuine uncertainty (the respondent has thought about it and thinks both outcomes equally likely) or epistemic ignorance (they have no idea). Following up a 50 response with “are you unsure, or do you think it’s really just as likely to happen as not?” helps distinguish these (Bruine de Bruin et al., 2000). Code the two cases distinctly in the dataset; do not collapse them.
Incentivising accuracy. Asking for subjective probabilities is hypothetical by default. Proper scoring rules (Brier for binary, CRPS or log score for continuous) reward truthful reporting under risk-neutrality but are not incentive-compatible under risk-aversion — risk-averse respondents shade toward 50. The binarized scoring rule (Hossain & Okui, 2013) restores incentive compatibility under any monotone preferences and is the canonical fix; see the Belief Elicitation / Binarized Scoring Rule guide for the full mechanism.
Sample size. For a single elicited probability, n ≈ 400 gives a sampling standard error of roughly 2.5 percentage points on the mean (sd ≈ 30 on a 0–100 scale). For threshold elicitation that recovers a full subjective distribution, simulate the design: draw plausible (a, b) beta parameters, sample threshold responses with rounding, refit, and check the bias-and-coverage of the recovered moments at the proposed n.
Caveats & common mistakes
Focal point clustering. Responses cluster at 0, 50, and 100, especially in verbal elicitation. A focal share above 30–40% is diagnostic of comprehension failure or scale unfamiliarity. Two competing responses to focal points: (a) treat them as noise and exclude or downweight; (b) treat them as informative about an interval of underlying belief and recover partial-identification bounds (Manski & Molinari, 2010). The second is the more honest default when the focal share is non-trivial.
“Fifty means don’t know.” The most common interpretive trap (Bruine de Bruin et al., 2000; Fischhoff & Bruine de Bruin, 1999). Respondents who cannot form a probability default to 50, which inflates apparent uncertainty and makes the 50 response uninterpretable without follow-up. A separate “don’t know” option, distinct from the probability scale, forces the conceptual distinction.
Reference class sensitivity. Elicited probabilities depend on how the event is defined. “The probability of a good harvest” is vague; “the probability that your maize yield will exceed 10 bags per acre” is specific. More specific reference classes produce more meaningful and reliable responses but require the respondent to be familiar with the units. Pilot-test the question phrasing.
Beliefs vs. hopes. In domains where outcomes are highly valued — child health, schooling success — respondents may report aspirational beliefs rather than genuine expectations. The standard mitigation is not to switch frames mid-question (own-case and population-frame elicitations recover different objects, Manski 2004) but to add an explicit “we are asking what you think will happen, not what you hope will happen” preface and, where the decision model permits, run both frames and compare.
Probabilistic coherence is the exception. Even median respondents do not assign coherent probabilities to complementary or mutually exclusive events (Tversky & Koehler, 1994; Manski, 2004). Consistency checks (complementary probabilities sum to 100) are useful for flagging the worst comprehension failures, but a coherence threshold cannot be set so high that most respondents fail. Code violations as a covariate; do not silently drop.
Beliefs are not decision weights. Under rank-dependent utility or cumulative prospect theory, the probability that enters the decision is a transformation of the stated probability, not the probability itself. For risk-domain applications, flag this as a substantive caveat.
Analysis Guide
# 1. Recode don't-know sentinel to NaN before any statistic — leaving 999 in the
# data inflates the mean and distorts the distribution. Track the don't-know
# rate as a data-quality covariate: a high rate signals the elicitation format
# was not understood by this population.
# prob_employ: stated probability (0-100) that the child will be employed after school
import numpy as np, pandas as pd
import statsmodels.formula.api as smf
import matplotlib.pyplot as plt
from scipy.optimize import least_squares
from scipy.stats import beta
df['prob_employ'] = df['prob_employ'].replace(999, np.nan)
print('Don-know rate:', df['prob_employ'].isna().mean())
# 2. Inspect the full distribution with histogram bins aligned to focal points so
# the spike at 50 is not smeared across two bins. Compare mean and median —
# a large gap signals skew or aspirational over-reporting near the top.
print(df['prob_employ'].describe(percentiles=[.1, .25, .5, .75, .9]))
df['prob_employ'].hist(bins=range(0, 105, 5))
plt.title('Distribution of employment probability beliefs')
# 3. Compute the focal-point share on the NON-MISSING subsample only. Using
# .isin([0,50,100]) on a Series with NaNs treats NaN as non-focal and inflates
# the denominator, understating the focal share — restrict to notna() first.
non_na = df['prob_employ'].notna()
focal_share = df.loc[non_na, 'prob_employ'].isin([0, 50, 100]).mean()
print('Focal-point share:', focal_share) # > 0.30 flags comprehension failure
# 4. Coherence check: ask the complementary event (prob_unemploy) and check
# that the two sum to ~100. Flag rows outside [90, 110] as incoherent;
# code as a quality covariate rather than silently dropping.
df['psum'] = df['prob_employ'] + df['prob_unemploy']
df['incoherent'] = ((df['psum'] < 90) | (df['psum'] > 110)).astype(int)
print('Incoherent share:', df['incoherent'].mean())
# 5. Fit a beta CDF to elicited (threshold, cumulative-probability) pairs to
# recover the full subjective distribution for continuous outcomes. NOTE:
# scipy.stats.beta.fit does MLE on raw observations, NOT CDF-fitting on
# fractiles — use least_squares against beta.cdf at the elicited thresholds.
thresholds = np.array([0.20, 0.40, 0.60, 0.80]) # normalised to [0,1]
cum_probs = np.array([0.10, 0.35, 0.70, 0.92]) # respondent i's responses
res = least_squares(lambda ab: beta.cdf(thresholds, ab[0], ab[1]) - cum_probs,
x0=[2.0, 2.0], bounds=([0.1, 0.1], [50, 50]))
a_hat, b_hat = res.x
subj_mean = a_hat / (a_hat + b_hat)
subj_var = a_hat * b_hat / ((a_hat + b_hat) ** 2 * (a_hat + b_hat + 1))
print('Subjective mean:', subj_mean, 'variance:', subj_var)
# 6. Belief-on-behaviour regression — DESCRIPTIVE, not causal. Beliefs are not
# randomly assigned, so the coefficient is an association consistent with the
# decision model, not an effect estimate. HC2 SEs match the R block; cluster
# by village (or sampling cluster) if the design is clustered.
fit = smf.ols('enrolled ~ prob_employ + age + C(female) + log_hh_expenditure',
data=df).fit(cov_type='HC2')
print(fit.summary())
# 7. Belief updating under one-sided non-compliance (information offered but not
# all assigned-treatment respondents actually engaged with it): use the
# treatment assignment as an instrument for engagement to recover LATE.
# See the Encouragement Design guide for the full estimator and assumptions.
fit_update = smf.ols(
'prob_employ_endline ~ treat + prob_employ_baseline + age + C(female)',
data=df).fit(cov_type='HC2')
print(fit_update.summary()) SurveyCTO / XLSForm
For the balls-in-a-bin or wheel format, use an integer field constrained to 0..10 (or 0..100 for percent-chance). Attach the visual aid via the media::image column on the question row (e.g., media::image = bins_visual.png). Always store the raw probability value — do not recode to categories in the instrument.
For threshold elicitation of continuous outcomes, create one integer field per threshold and use a constraint on each to enforce monotonicity of the cumulative distribution. For threshold k:
constraint: . >= ${prob_below_threshold_k-1}
constraint_message: This must be at least as large as the previous answer,
because the chance of being below a higher amount cannot
be smaller than the chance of being below a lower one.
Randomise the order of threshold questions across respondents (using a calculate field with once(int(random()*k))) to detect bracketing effects (Delavande, Giné & McKenzie, 2011). Store a separate don't know field as a select_one yes_no adjacent to the probability field — do not encode don’t-know inside the probability scale itself, since this collapses uncertainty into ignorance.
Reading the output
- Focal-point share is the primary data-quality indicator. Compute via
df.loc[df['prob_employ'].notna(), 'prob_employ'].isin([0,50,100]).mean()in Python orprop.table(table(df$focal[!is.na(df$prob_employ)]))in R. Above 30–40% flags comprehension failure; with visual aids and trained enumerators rates below 20% are achievable. - Mean vs. median divergence signals skew. Use
df['prob_employ'].describe(percentiles=[.1,.25,.5,.75,.9])(Python) orsummary(df$prob_employ)(R). A mean well above the median often indicates aspirational over-reporting near the top of the scale. - Incoherent share (complementary probabilities summing outside [90, 110]) above 25% indicates the elicitation is asking too much of the population; consider switching to a tactile format or simpler reference class.
- Belief-on-behaviour regression: a significant positive coefficient on
prob_employis consistent with beliefs being a binding constraint, not proof — beliefs were not randomly assigned. - Belief-updating regression: the coefficient on
treatis the treatment effect on endline beliefs, controlling for baseline. Benchmark against Jensen (2010): a 25–30 percentage-point belief shift after an information intervention is what the canonical study produced; shifts above 10 pp are substantively meaningful, shifts below 5 pp typically do not move downstream behaviour. - Subjective distribution fit: report a goodness-of-fit residual (sum of squared deviations from the elicited cumulative probabilities) alongside the recovered mean and variance. Above ~0.05 in residual sum-of-squares the parametric family does not fit well; report bounds rather than point estimates of the moments.
References
Attanasio, O. P. (2009). Expectations and perceptions in developing countries: Their measurement and their use. American Economic Review, 99(2), 87–92. https://doi.org/10.1257/aer.99.2.87
Attanasio, O., & Augsburg, B. (2016). Subjective expectations and income processes in rural India. Economica, 83(331), 416–442. https://doi.org/10.1111/ecca.12192
Attanasio, O. P., & Kaufmann, K. M. (2014). Education choices and returns to schooling: Mothers’ and youths’ subjective expectations and their role by gender. Journal of Development Economics, 109, 203–216. https://doi.org/10.1016/j.jdeveco.2014.04.003
Bruine de Bruin, W., Fischhoff, B., Millstein, S. G., & Halpern-Felsher, B. L. (2000). Verbal and numerical expressions of probability: “It’s a fifty–fifty chance.” Organizational Behavior and Human Decision Processes, 81(1), 115–131. https://doi.org/10.1006/obhd.1999.2868
Bruine de Bruin, W., Manski, C. F., Topa, G., & van der Klaauw, W. (2011). Measuring consumer uncertainty about future inflation. Journal of Applied Econometrics, 26(3), 454–478. https://doi.org/10.1002/jae.1239
Delavande, A. (2014). Probabilistic expectations in developing countries. Annual Review of Economics, 6, 1–20. https://doi.org/10.1146/annurev-economics-072413-105148
Delavande, A., Giné, X., & McKenzie, D. (2011). Measuring subjective expectations in developing countries: A critical review and new evidence. Journal of Development Economics, 94(2), 151–163. https://doi.org/10.1016/j.jdeveco.2010.01.008
Delavande, A., & Rohwedder, S. (2008). Eliciting subjective probabilities in internet surveys. Public Opinion Quarterly, 72(5), 866–891. https://doi.org/10.1093/poq/nfn064
Dominitz, J., & Manski, C. F. (1997). Using expectations data to study subjective income expectations. Journal of the American Statistical Association, 92(439), 855–867. https://doi.org/10.1080/01621459.1997.10474041
Fischhoff, B., & Bruine de Bruin, W. (1999). Fifty-fifty = 50%? Journal of Behavioral Decision Making, 12(2), 149–163. https://doi.org/10.1002/(SICI)1099-0771(199906)12:2%3C149::AID-BDM314%3E3.0.CO;2-J
Hossain, T., & Okui, R. (2013). The binarized scoring rule. Review of Economic Studies, 80(3), 984–1001. https://doi.org/10.1093/restud/rdt006
Hurd, M. D. (2009). Subjective probabilities in household surveys. Annual Review of Economics, 1, 543–562. https://doi.org/10.1146/annurev.economics.050708.142955
Jensen, R. (2010). The (perceived) returns to education and the demand for schooling. Quarterly Journal of Economics, 125(2), 515–548. https://doi.org/10.1162/qjec.2010.125.2.515
Manski, C. F. (2004). Measuring expectations. Econometrica, 72(5), 1329–1376. https://doi.org/10.1111/j.1468-0262.2004.00537.x
Manski, C. F., & Molinari, F. (2010). Rounding probabilistic expectations in surveys. Journal of Business & Economic Statistics, 28(2), 219–231. https://doi.org/10.1198/jbes.2009.08098
Tversky, A., & Koehler, D. J. (1994). Support theory: A nonextensional representation of subjective probability. Psychological Review, 101(4), 547–567. https://doi.org/10.1037/0033-295X.101.4.547