What it is
The Becker-DeGroot-Marschak (BDM) mechanism elicits willingness to pay (WTP) for a good through a single-respondent auction (Becker, DeGroot & Marschak, 1964). Under expected-utility preferences, stating the true maximum price you would pay is theoretically incentive-compatible — you cannot do better by misreporting. The mechanism involves a real transaction with real stakes: whether you receive the good and at what price depends on what you stated.
The mechanism works as follows. The respondent is shown the good and states the maximum price they are willing to pay. A price is then randomly drawn from a pre-specified distribution. If the stated value is at least as high as the drawn price, the respondent purchases the good at the drawn price (not at the price they stated). If not, no transaction occurs. The respondent pays the drawn price, not the stated price — so stating anything other than the true reservation value can only hurt: understating risks missing the chance to buy at a price you would have accepted; overstating risks being obligated to buy at a price above your reservation value.
The theoretical IC result is conditional. Karni and Safra (1987) prove that no mechanism that takes preferences over lotteries as input can elicit preferences over riskless outcomes if respondents are not expected-utility maximisers. Horowitz (2006) shows BDM is not IC even for non-random goods under fairly weak conditions, and Cason and Plott (2014) provide canonical experimental evidence that respondents systematically fail to recognise BDM’s game form and treat it as a first-price auction or similar. In practice, subject comprehension is the central design problem, not the theoretical proof.
The BDM has been widely used to study demand for health products — bed nets, water purification tablets, oral rehydration salts, improved cookstoves — and for product valuations in marketing and behavioural economics, where identifying the price at which adoption falls off is the central question.
When to use it
The BDM is appropriate when: the good can be physically provided at the point of interview at a range of possible prices; real monetary transactions are feasible; and the incentive-compatibility property is likely to hold — that is, respondents understand the mechanism well enough to respond to its logic.
The core advantage over contingent valuation is that BDM is a real-stakes mechanism — by construction, the hypothetical bias that inflates CV estimates does not apply. BDM is often used as the empirical benchmark against which hypothetical methods are validated (Cummings & Taylor, 1999; Vossler, Doyon & Rondeau, 2012). Dupas (2014) used BDM-elicited WTP as the primary measure of demand for health products and is the standard contemporary application; Berry, Fischer and Guiteras (2020) and Burchardi, de Quidt, Gulesci, Lerva and Tripodi (2021) provide rigorous field comparisons of BDM against alternatives.
The method is less appropriate when: the good cannot be physically made available at the interview; respondents cannot reliably follow the mechanism (Berry et al., 2020 document this in low-numeracy field settings; Cason & Plott, 2014 document it in lab settings); the range of prices to be tested is so wide that a single BDM draw provides little information at policy-relevant price points; or the research question is about market demand at scale, where social context and availability matter as much as individual reservation prices.
How it works
The standard implementation has four steps:
Step 1 — Comprehension training. Before the real elicitation, run a practice round with a known-value item (a common branded good with a posted retail price; a familiar consumer item with widely-known value) and a scored comprehension test. The test asks the respondent to predict what would happen under several hypothetical scenarios — “If your stated price is 50 and the drawn price is 30, what happens?” Respondents who fail the test should be re-trained or, under a pre-specified rule, excluded from the analysis sample. Berry et al. (2020) and Cason and Plott (2014) show that without this step, a non-trivial share of respondents do not play the dominant strategy.
Step 2 — Valuation. The enumerator shows the respondent the good — physically present and of standardised quality — explains what it does, and asks for the maximum price they would pay. The respondent states a value in local currency.
Step 3 — Price draw. The enumerator draws a random price from a pre-specified distribution using a spinner, lottery cards, or a tablet randomiser. The drawn price and the distribution it was drawn from are visible to the respondent. This transparency is what makes the mechanism credible.
Step 4 — Transaction. If the stated WTP is at or above the drawn price, the respondent buys the good at the drawn price. If not, no transaction occurs. Payment is made immediately and the good is handed over.
Identification assumptions. For BDM to recover true WTP, six conditions must hold:
- Game-form recognition — the respondent understands that the stated price acts only as a threshold and the drawn price is what they pay. Empirically the most consequential and most often violated (Cason & Plott, 2014).
- Expected utility (or probabilistic sophistication) — the IC property assumes EU preferences over the price-distribution lottery (Karni & Safra, 1987; Horowitz, 2006).
- No reference dependence on the announced price distribution — respondents should value the good on its merits, not anchor on the range of the announced distribution.
- No binding liquidity constraint — the respondent has enough cash on hand that stated WTP reflects valuation, not budget. First-order in low-income field settings.
- No enumerator demand effects — the field protocol gives no signals about a “right” answer.
- Standardised good — the physical good is identical across respondents in quality, packaging, and condition.
Analysis output. The distribution of stated WTP is the primary output. Demand at a given price is the proportion of respondents with stated WTP at or above that price — a non-parametric demand curve. The empirical CDF of stated maxima gives the full population WTP distribution. For parametric inference, fit a censored regression (Tobit) or a survival model treating zero-WTP as left-censoring. Regression analysis can relate WTP to respondent characteristics or treatment assignment, but should use censored or quantile methods given that WTP is bounded below at zero with mass at the boundary.
Key decisions
Comprehension test and exclusion rule. Run a scored comprehension test before the live elicitation, with a pre-specified pass threshold and exclusion rule. This is not optional. Without it, a non-trivial share of stated WTPs are not generated by the BDM logic at all and the analysis is contaminated (Cason & Plott, 2014; Berry et al., 2020). Pre-register the pass threshold and exclusion rule before fieldwork.
Price distribution. Three design considerations:
- Support coverage. The distribution must cover the support of true WTP in the study population. If the maximum drawn price is below some respondents’ true WTP, those respondents are indifferent between all bids at or above the maximum and IC fails at the top (Horowitz, 2006). If the minimum is above zero, respondents with very low valuations are pooled. Set the support from pilot data.
- Uniform unless there’s a reason. A uniform distribution gives constant marginal incentives across the support; non-uniform distributions create gradients in the cost of misreporting at different points. Use uniform unless the study question explicitly requires concentrating draws around a policy-relevant price.
- Continuous vs discrete draws. Continuous draws give a single observation per respondent at the realised price. Discrete draws (drawing from a fixed list of price points) effectively bracket WTP into intervals and need interval-regression analysis. Discrete is simpler to administer in the field.
The distribution should be set from prior evidence, not post-hoc. The distribution range can also anchor WTP through reference dependence — respondents shown a range that tops out at 500 may state higher WTP than respondents shown a range topping out at 200, independently of valuation (Bohm, Lindén & Sonnegård, 1997; Drichoutis & Lusk, 2014).
Liquidity constraints. In low-income field settings, stated BDM bids are bounded above by cash on hand, not by valuation. If respondents do not have the cash to back a high bid, they will state a lower WTP — not because they value the good less, but because they cannot pay. Either (a) provide a small show-up payment large enough to cover the maximum possible drawn price, or (b) measure cash on hand at the start of the session and analyse WTP conditional on a non-binding budget. Berry et al. (2020) discuss this in detail.
Multiple price list (MPL) alternative. For low-numeracy settings, the multiple price list (the Holt-Laury MPL guide covers the broader MPL methodology) presents a series of yes/no purchase decisions at increasing prices. The respondent’s switching point reveals a WTP interval. MPL is simpler to administer and verify but reveals only a bracket; the analysis needs interval regression rather than OLS, and the format is susceptible to multiple switching and midpoint anchoring. Berry et al. (2020) recommend MPL over open-ended BDM in low-numeracy contexts; Burchardi et al. (2021) provide a comparative field evaluation.
Single round vs multiple goods. The standard BDM elicits WTP for a single good per session. Eliciting WTP for multiple goods simultaneously generates wealth effects and reference-point shifts that distort individual valuations. If multiple goods must be valued, use a random device to select which good is transacted at the end of the session.
WTP vs WTA framing. BDM can be implemented as a buyer auction (the respondent states the maximum they would pay) or as a seller auction (the respondent is endowed with the good and states the minimum they would sell for). The two yield systematically different values — the WTA/WTP gap is large, typically 2–4× — and the choice of framing has major substantive consequences. Plott and Zeiler (2005) argue that much of the observed gap reflects subject misconceptions rather than loss aversion; whatever the mechanism, do not mix WTP and WTA responses in the same analysis.
Position in the family of incentive-compatible elicitation mechanisms. BDM is one of several mechanisms with the dominant-strategy property. The second-price (Vickrey) auction (Vickrey, 1961) is its multi-respondent analogue; the random nth-price auction (Shogren, Margolis, Koo & List, 2001) is a hybrid that empirically often outperforms BDM on engagement; multiple price lists give a bracket rather than a point estimate. Lusk and Schroeder (2004) provide the standard comparison. Choose BDM when single-respondent administration is the constraint; choose alternatives when group elicitation is feasible or comprehension is a binding concern.
Caveats & common mistakes
Subject misconceptions — the central practical problem. Cason and Plott (2014) provide the canonical experimental evidence that subjects systematically fail to recognise BDM’s game form: they treat it as a first-price auction, anchor on the stated price as if it were what they would pay, or otherwise misunderstand the incentives. Plott and Zeiler (2005) argue that the famous “endowment effect” in BDM data is largely an artefact of these misconceptions, not loss aversion. Diagnostic signals: bid distributions that cluster at round numbers (>40% of bids at multiples of 25 or 50 is a red flag — use a digit-preference test, see the Digit Preference / Heaping Detector guide), pile up at the maximum of the announced price distribution, or show no relationship with respondent characteristics. The fix is comprehension training with a known-value item and a scored test, plus an exclusion rule for respondents who fail (see How it works).
The IC result requires expected utility. Karni and Safra (1987) prove that BDM is not incentive-compatible under non-expected-utility preferences (very common empirically). Horowitz (2006) extends this to weaker conditions. If respondents are loss-averse, probability-weighting, or otherwise non-EU, stated WTP does not necessarily equal true reservation value even with full comprehension. The size of the gap is typically small but not zero.
Liquidity constraints. In low-income field settings, stated BDM bids are bounded above by cash on hand rather than by valuation. A respondent with 200 in their pocket cannot credibly bid 400, regardless of how much they value the good. This is a first-order concern in development applications and routinely under-discussed. Either provide a show-up payment large enough to cover the maximum drawn price, or measure cash on hand and analyse WTP conditional on a non-binding budget.
Reference dependence on the announced distribution. Respondents may anchor on the range of the announced price distribution rather than on their preferences for the good. Bohm, Lindén and Sonnegård (1997) and Drichoutis and Lusk (2014) document empirical evidence of this in BDM data. Pilot two different distributions and check whether the WTP distribution shifts with the announced range; if it does, reference dependence is contaminating the elicitation.
Experimenter demand effects. Field enumerators conducting BDM sessions are not neutral. Enumerators who reveal enthusiasm for the good, rush through the mechanism explanation, or vary delivery across respondents introduce non-random noise. Standardised scripts, enumerator training, and audio monitoring are important quality controls. Mummolo and Peterson (2019) and de Quidt, Haushofer and Roth (2018) document the empirical magnitude of demand effects in survey-experimental settings.
Supply and quality standardisation. The physical good must be identical across respondents in quality, packaging, and condition. Variation across sessions confounds WTP comparisons. Spot-checking the goods distributed is a standard quality-control step.
External validity of WTP estimates. BDM-elicited WTP predicts individual purchasing behaviour better than contingent valuation, but the relationship to market demand is imperfect. Demand estimated in a research context may not translate to a commercial market, where social context, availability, marketing, and the absence of an enumerator all matter. Berry et al. (2020) and Burchardi et al. (2021) provide the strongest field evidence on this gap.
Analysis Guide
import numpy as np
import pandas as pd
import statsmodels.formula.api as smf
import matplotlib.pyplot as plt
# 1. Comprehension test pass rate and balance — non-trivial fail rates contaminate
# the analysis; balance check on age/expenditure across pass/fail rules out
# selection on observables driving differences in stated WTP
print('Pass rate:', df['passed_comprehension'].mean())
df.groupby('passed_comprehension')[['age', 'log_hh_expenditure']].mean()
# 2. Falsification check for IC failure — regress stated WTP on the realised
# drawn price; under correct BDM understanding the coefficient should be zero
# because the draw is independent of the bid. A non-zero coefficient is
# evidence that respondents update their bid in response to the draw, which
# means IC has failed (Cason & Plott 2014)
falsif = smf.ols('wtp_stated ~ drawn_price', data=df).fit(cov_type='HC1')
print(falsif.params['drawn_price'], falsif.pvalues['drawn_price'])
# 3. Inspect the WTP distribution; check median and look for clustering at
# round numbers (0, 25, 50, 100) — >40% of bids at multiples of 25 in a
# continuous-bid design is a red flag for anchoring rather than valuation
wtp_clean = df['wtp_stated'].dropna()
print(wtp_clean.describe(percentiles=[.1, .25, .5, .75, .9]))
heap_share = ((wtp_clean % 25) == 0).mean()
print(f'Share of bids at multiples of 25: {heap_share:.2%}')
wtp_clean.hist(bins=range(0, int(wtp_clean.max()) + 10, 10))
plt.title('Distribution of stated WTP'); plt.xlabel('WTP (local currency)')
# 4. Non-parametric demand curve — for each candidate price, the share of
# respondents with WTP at or above that price is the predicted demand; this
# is model-free, makes no functional-form assumption on the WTP distribution,
# and is directly interpretable for policy
prices = [0, 25, 50, 75, 100, 150, 200]
demand = pd.Series({p: (df['wtp_stated'] >= p).mean() for p in prices},
name='share_willing')
print(demand)
# Binomial SE per price point — useful for confidence bands on the curve
demand_se = pd.Series({p: np.sqrt(d * (1 - d) / df['wtp_stated'].notna().sum())
for p, d in demand.items()}, name='se')
# 5. Censored regression for WTP on characteristics — WTP is bounded below at
# zero with mass at the boundary; OLS biases coefficients. Tobit is the
# standard parametric alternative; quantile regression is the non-parametric
# alternative when the distributional assumption is suspect
from statsmodels.miscmodels.tmodel import TLinearModel # alternative: use
# linearmodels.iv or scipy-based tobit; or qreg via statsmodels.formula.api smf.quantreg
qfit = smf.quantreg('wtp_stated ~ age + C(female) + log_hh_expenditure + C(treatment)',
data=df).fit(q=0.5)
print(qfit.summary()) Implementation checklist
Before fielding a BDM study, confirm:
- The good is physically available in sufficient quantity at every interview location
- Enumerators have practiced the mechanism explanation to the point that they can deliver it consistently
- A comprehension check or practice round is included in the script before the main elicitation
- The price distribution is recorded in the data and can be verified against stated WTP values
- Payment float is available at each interview location to make change for any transaction
Reading the output
- Comprehension pass rate. Below ~70% suggests training was inadequate or the population struggles with the mechanism — re-train or switch to MPL. The pre-registered exclusion rule should drop or down-weight failers in a sensitivity analysis.
- Falsification check. The coefficient on
drawn_pricein the regression ofwtp_statedondrawn_priceshould be statistically zero. A meaningfully non-zero coefficient (|β| > 0.1 in standardised units, p < 0.05) is evidence that respondents are updating bids in response to the realised draw — IC has failed and the WTP estimates cannot be trusted. - Heaping diagnostic. More than 40% of bids at multiples of 25 (in a design with continuous bids) or strong piling at the maximum of the announced price distribution flags anchoring rather than valuation. See the Digit Preference / Heaping Detector guide for a formal test.
- Demand curve shape. A steeply downward-sloping curve indicates high price sensitivity — small price increases produce large demand drops. A flat curve indicates low sensitivity in the tested range. Report binomial SEs at each price point so adjacent shares with overlapping CIs are not reported as different.
- Sample-size and precision. For a target half-width of ±5 percentage points on the share willing at a target price, the binomial SE is √(p(1−p)/n); at p = 0.5 this needs n ≈ 100 per arm; at p = 0.2 it needs n ≈ 250. For a 10pp difference between arms at α = 0.05, β = 0.2, plan for ~400 per arm.
- WTP regression interpretation. OLS on
wtp_statedis biased when WTP is censored at zero with mass at the boundary. Report OLS, Tobit, and median regression side by side; if they diverge materially, the censoring or the distributional assumptions are doing the work. A positive coefficient ontreatmentindicates higher stated WTP in the treatment arm — interpretable as a treatment effect on demand under random assignment. - WTP at the bounds. If stated WTP values pile at 0 or at the maximum of the announced price distribution, the distribution was poorly specified for this population. The estimate is weakly identified and re-piloting the price range is the appropriate response.
References
Becker, G. M., DeGroot, M. H., & Marschak, J. (1964). Measuring utility by a single-response sequential method. Behavioral Science, 9(3), 226–232. https://doi.org/10.1002/bs.3830090304
Berry, J., Fischer, G., & Guiteras, R. (2020). Eliciting and utilizing willingness to pay: Evidence from field trials in northern Ghana. Journal of Political Economy, 128(4), 1436–1473. https://doi.org/10.1086/705374
Bohm, P., Lindén, J., & Sonnegård, J. (1997). Eliciting reservation prices: Becker-DeGroot-Marschak mechanisms vs. markets. The Economic Journal, 107(443), 1079–1089. https://doi.org/10.1111/j.1468-0297.1997.tb00008.x
Burchardi, K. B., de Quidt, J., Gulesci, S., Lerva, B., & Tripodi, S. (2021). Testing willingness to pay elicitation mechanisms in the field: Evidence from Uganda. Journal of Development Economics, 152, 102701. https://doi.org/10.1016/j.jdeveco.2021.102701
Cason, T. N., & Plott, C. R. (2014). Misconceptions and game form recognition: Challenges to theories of revealed preference and framing. Journal of Political Economy, 122(6), 1235–1270. https://doi.org/10.1086/677254
Cummings, R. G., & Taylor, L. O. (1999). Unbiased value estimates for environmental goods: A cheap talk design for the contingent valuation method. American Economic Review, 89(3), 649–665. https://doi.org/10.1257/aer.89.3.649
de Quidt, J., Haushofer, J., & Roth, C. (2018). Measuring and bounding experimenter demand. American Economic Review, 108(11), 3266–3302. https://doi.org/10.1257/aer.20171330
Drichoutis, A. C., & Lusk, J. L. (2014). Judging statistical models of individual decision making under risk using in- and out-of-sample criteria. PLOS ONE, 9(7), e102269. https://doi.org/10.1371/journal.pone.0102269
Dupas, P. (2014). Short-run subsidies and long-run adoption of new health products: Evidence from a field experiment. Econometrica, 82(1), 197–228. https://doi.org/10.3982/ECTA9508
Horowitz, J. K. (2006). The Becker-DeGroot-Marschak mechanism is not necessarily incentive compatible, even for non-random goods. Economics Letters, 93(1), 6–11. https://doi.org/10.1016/j.econlet.2006.03.033
Karni, E., & Safra, Z. (1987). “Preference reversal” and the observability of preferences by experimental methods. Econometrica, 55(3), 675–685. https://doi.org/10.2307/1913606
Lusk, J. L., & Schroeder, T. C. (2004). Are choice experiments incentive compatible? A test with quality differentiated beef steaks. American Journal of Agricultural Economics, 86(2), 467–482. https://doi.org/10.1111/j.0092-5853.2004.00592.x
Mummolo, J., & Peterson, E. (2019). Demand effects in survey experiments: An empirical assessment. American Political Science Review, 113(2), 517–529. https://doi.org/10.1017/S0003055418000837
Noussair, C., Robin, S., & Ruffieux, B. (2004). Revealing consumers’ willingness-to-pay: A comparison of the BDM mechanism and the Vickrey auction. Journal of Economic Psychology, 25(6), 725–741. https://doi.org/10.1016/j.joep.2003.06.004
Plott, C. R., & Zeiler, K. (2005). The willingness to pay–willingness to accept gap, the “endowment effect,” subject misconceptions, and experimental procedures for eliciting valuations. American Economic Review, 95(3), 530–545. https://doi.org/10.1257/0002828054201387
Shogren, J. F., Margolis, M., Koo, C., & List, J. A. (2001). A random nth-price auction. Journal of Economic Behavior & Organization, 46(4), 409–421. https://doi.org/10.1016/S0167-2681(01)00165-2
Vickrey, W. (1961). Counterspeculation, auctions, and competitive sealed tenders. Journal of Finance, 16(1), 8–37. https://doi.org/10.1111/j.1540-6261.1961.tb02789.x
Vossler, C. A., Doyon, M., & Rondeau, D. (2012). Truth in consequentiality: Theory and field evidence on discrete choice experiments. American Economic Journal: Microeconomics, 4(4), 145–171. https://doi.org/10.1257/mic.4.4.145