Metter. / Mixtapes / Methods Mixtape / Survey & Elicitation Methods

07 · Survey & Elicitation Methods

Willingness to Pay — Contingent Valuation

A stated preference method for estimating what people would pay for goods or services not traded in markets — health interventions, public infrastructure, environmental quality — using hypothetical payment scenarios.


What it is

Contingent valuation (CV) is a stated-preference method that estimates willingness to pay (WTP) for goods and services with no market price. Respondents read a detailed hypothetical scenario describing the good and the payment mechanism, then say whether they would pay a specified amount, or the most they would pay. The method is called “contingent” because the valuation is contingent on the hypothetical market described in the scenario.

CV has its roots in environmental economics — early uses estimated WTP for wilderness preservation, clean air, and water quality. After the 1989 Exxon Valdez oil spill, the National Oceanic and Atmospheric Administration (NOAA) convened a panel of economists (chaired by Kenneth Arrow and Robert Solow) to assess whether CV could produce defensible damage estimates for natural-resource litigation. The resulting NOAA Panel report (Arrow et al., 1993) concluded that CV can produce valid WTP estimates under specific methodological conditions and laid down a set of design recommendations that remain the canonical standard — see “Following the NOAA Panel” below. CV has since spread well beyond environmental valuation: health economics (Mühlbacher et al., 2016), transport, marketing, regulatory cost-benefit analysis, and development applications such as bed nets, water connections, and improved sanitation (Carson, 2012; Whittington, 2010; Bishop & Boyle, 2019).

The Diamond-Hausman critique (Diamond & Hausman, 1994; Hausman, 2012) and the rebuttals from Carson (2012) and Kling, Phaneuf & Zhao (2012) remain the central methodological debate. For many non-market goods CV is the only available tool; for others, incentive-compatible alternatives like the BDM mechanism substantially reduce hypothetical bias and should be preferred when feasible.

When to use it

CV is appropriate when no market price exists for the good of interest, revealed-preference methods (inferring value from actual choices) are infeasible, and the research question requires a monetary welfare estimate rather than a simple preference ranking.

Typical applications: WTP for bed nets, water connections, sanitation, vaccine programmes, and health services in development and global-health contexts; for ecosystem services, recreation, biodiversity, and pollution control in environmental policy; for new product features, brand attributes, and price elasticities in market research; for transport time savings; for new regulations in cost-benefit analysis. Whittington (1998, 2010) discusses CV design for low-income settings; Bishop and Boyle (2019) is the current general textbook reference.

CV is the wrong tool when an incentive-compatible mechanism is feasible. If respondents have a real chance of receiving the good and their stated value determines whether they pay, the Becker-DeGroot-Marschak (BDM) mechanism is incentive-compatible in principle and substantially reduces hypothetical bias in practice. Use CV when real transactions are not feasible or when estimating aggregate demand from a population sample rather than eliciting individual reservation prices.

How it works

The researcher builds a scenario describing the good, the payment mechanism (a user fee, a tax, a one-time contribution, a monthly fee), and the consequences of payment and non-payment. The scenario must be detailed enough to be credible and simple enough to be understood in an interview. The response is elicited using one of several formats (see Key Decisions).

The estimator depends on the format. Open-ended responses give continuous WTP values; the basic estimate is the sample mean or median. Dichotomous-choice (DC) responses — yes/no to a single bid — are binary. The probability of acceptance is modelled as a function of the bid using probit or logit (parametric) or a step function (non-parametric, Turnbull). WTP is recovered from the fitted model, not from raw responses.

Mean vs median WTP. WTP distributions are almost always right-skewed, so mean and median differ. The median is the bid at which 50% would accept — robust to distributional misspecification and the policy-relevant estimand for majority-rule referenda. The mean is the area under the demand curve and feeds aggregate benefit calculations, but is sensitive to the right tail and the assumed distribution. Report both.

Parametric WTP recovery. Hanemann (1984, 1989) shows that under each distributional choice, mean and median WTP have a closed-form expression in the coefficients of a probit or logit regression of acceptance on bid:

  • Normal or logistic (symmetric distribution): mean = median = −α/β, where α is the intercept and β is the bid coefficient.
  • Log-normal or log-logistic (the standard CV choice — constrains WTP to be positive): median = exp(−α/β); mean involves the scale parameter and can be undefined when the distribution’s tail is heavy enough. Use the formulas in Hanemann (1984) or the DCchoice R package’s built-in extraction.

The Krinsky-Robb procedure (Krinsky & Robb, 1986) computes confidence intervals on these recovered WTP values: draw repeatedly from the asymptotic multivariate-normal distribution of the regression coefficients, transform each draw into a WTP value, and take the 2.5/97.5 percentiles of the resulting distribution.

Non-parametric WTP recovery — Turnbull. The Turnbull (1976) estimator makes no distributional assumption. It enforces monotonicity (acceptance probability cannot rise with bid) using the pool-adjacent-violators algorithm, then computes a lower bound on mean WTP from the resulting step function (Haab & McConnell, 1997). Useful as a conservative robustness check alongside the parametric estimate.

Welfare interpretation. DC CV recovers Hicksian compensating surplus (the income transfer that would leave the respondent indifferent between having the good at the proposed price and not having it at no cost), not Marshallian consumer surplus. For most policy applications the distinction is small; for goods that are a large share of income it can matter.

Identification assumptions. For the parametric DC estimator to recover the population WTP, four conditions must hold: (i) the bid is assigned independently of respondent characteristics (satisfied by random assignment); (ii) the assumed distribution of WTP is approximately correct in the study population; (iii) for double-bounded DC, the same WTP distribution governs both responses — no anchoring or framing effects in the follow-up; (iv) respondents understand the scenario as describing a genuine payment for the genuine good. Violations of (ii) are addressed by trying multiple distributions and reporting sensitivity; violations of (iii) by the bivariate-probit specification (Cameron & Quiggin, 1994); violations of (iv) by cognitive piloting, debriefing questions, and consequentiality framing (see Caveats).

Key decisions

Elicitation format. The choice of format is the most consequential design decision.

  • Open-ended: “What is the maximum you would pay for X?” Cognitively simple, but prone to strategic responses, zeros, and outliers. Use only with sophisticated respondents in familiar markets.
  • Payment card: shows a range of values and asks the respondent to mark their maximum WTP. Reduces extreme values but introduces anchoring on the displayed range. Analyse as interval data (the marked value as a lower bound, the next-higher card value as an upper bound) using interval regression (Cameron, 1988), not as a continuous outcome.
  • Single-bounded dichotomous choice (SBDC): a single randomly assigned bid, yes/no response. NOAA Panel default. Mirrors a take-it-or-leave-it purchase. Main cost: sample size — binary responses identify the WTP distribution less efficiently than continuous answers.
  • Double-bounded DC (DBDC): follow-up bid (higher if accepted, lower if refused). About 30–50% more efficient than SBDC (Hanemann, Loomis & Kanninen, 1991), but the follow-up can anchor the second response. When the data fail consistency tests, use the bivariate-probit specification (Cameron & Quiggin, 1994; Alberini, Kanninen & Carson, 1997).
  • One-and-a-half-bound DC (Cooper, Hanemann & Signorello, 2002): respondents see a bid range from the start and answer yes/no on the lower bound; those who accept then answer on the upper bound. Roughly matches DBDC efficiency without the anchoring of a surprise follow-up. Increasingly the preferred design.

Following the NOAA Panel. For a CV study that may face publication, regulatory, or litigation scrutiny, the Arrow et al. (1993) recommendations remain the canonical standard. The key items:

  1. Personal interviews preferred over phone, mail, or self-administered web. Mode effects in CV are large.
  2. Dichotomous-choice format (with follow-up debriefing — DC alone is not enough).
  3. Remind respondents of their budget constraint immediately before the WTP question.
  4. Remind respondents of available substitute goods to prevent embedding.
  5. Follow-up debriefing for “don’t know”, “no”, and protest responses; ask why.
  6. Treat “don’t know” responses conservatively (originally: as no; modern practice usually offers them as an explicit option and analyses them separately).
  7. Scenario should describe the good clearly and concretely, with realistic provision and payment mechanisms.
  8. Include a no-vote option; respondents should be able to opt out without social pressure.

Document compliance with each item in the study protocol and report deviations transparently.

Bid design for DC formats. Bid amounts must span the plausible WTP distribution. Two practical rules:

  • Use 4–8 distinct bid levels. Fewer reduces identification of the distribution’s tails; more spreads the sample too thinly per bid.
  • Cap the share assigned to any single bid at 30–40%. Optimal designs (Kanninen, 1995; Alberini, 1995) place bids at the 0.20, 0.40, 0.60, 0.80 quantiles of the prior WTP belief from pilot data.

Payment vehicle. The payment mechanism (user fee, tax, monthly bill, one-time contribution) affects responses. A payment vehicle that respondents distrust (a tax in a low-trust setting; a fee for a service traditionally provided free) elicits protest zeros. The vehicle should be realistic, familiar, and aligned with how the good would actually be provided.

Scope test design. To test scope sensitivity, run a split-sample: assign one arm a small quantity of the good (e.g., access for 1,000 households), the other a larger quantity (e.g., 10,000 households). The pass criterion: mean WTP should be statistically larger in the larger-quantity arm, and approximately proportional to the quantity ratio for plausible values. Failure (similar WTP across arms) is evidence of “warm glow” rather than marginal valuation (Kahneman & Knetsch, 1992).

Sample size. For single-bounded DC with 4–6 bid levels and a target half-width of ±20% on mean WTP, plan for 600–800 respondents. Double-bounded or one-and-a-half-bound designs reduce this by roughly 30–50% (Hanemann, Loomis & Kanninen, 1991). For scope tests, double the sample to support the split-arm comparison. For cluster-sampled designs, multiply by 1 + (m̄ − 1)ρ where m̄ is mean cluster size and ρ is the intra-cluster correlation.

Scenario realism and cognitive piloting. Respondents must understand what they are valuing. Cognitive interviews on draft scenarios — at minimum 15–20 respondents from the target population — are standard practice and surface ambiguities the analyst cannot anticipate.

Caveats & common mistakes

Hypothetical bias. Stated WTP from CV typically exceeds actual WTP in incentive-compatible experiments. Murphy, Allen, Stevens and Weatherhead (2005) — the standard meta-analysis — find a median ratio of stated-to-actual values of about 1.35 (mean 2.6, with substantial heterogeneity). The earlier List and Gallet (2001) meta-analysis gives a mean closer to 3. Two design-side remedies have empirical support:

  • Cheap talk (Cummings & Taylor, 1999): a short script before the WTP question that explicitly names the gap between hypothetical and actual payment and asks respondents to answer as they would in a real decision. Reduces but does not eliminate the bias.
  • Consequentiality framing (Vossler, Doyon & Rondeau, 2012): connect the survey explicitly to a downstream decision the respondent has reason to care about (“the results will be used to advise a regulator on whether to build X”). Post-2010 this is the central methodological response to hypothetical bias and arguably more important than cheap talk for current practice.

Include both in any new CV instrument, and report the share of respondents reporting they viewed the survey as consequential.

Scope insensitivity (the embedding effect). Respondents often state similar WTP for very different quantities of a good — saving 2,000 versus 200,000 birds (Diamond & Hausman, 1994), providing water to 100 versus 10,000 households (Kahneman & Knetsch, 1992). When mean WTP does not rise with the quantity of the good, the values reflect a generic “warm glow” rather than marginal valuation. Build a scope test into the design (split-sample with small and large quantities; see Key Decisions) and report the result. Scope insensitivity is most severe for unfamiliar public goods; tangible, local, personally relevant goods typically pass scope tests (Carson & Hanemann, 2005).

WTP versus WTA — the gap matters. Asking willingness-to-accept (WTA, the compensation a respondent would demand to give something up) yields systematically larger values than WTP for the same good. The WTA/WTP ratio is commonly 2–4× and can be larger for public or non-substitutable goods (Horowitz & McConnell, 2002; Plott & Zeiler, 2005). Switching framings without realising it switches welfare measures is one of the most common errors in stated-preference work. Pick the framing that matches the policy question — typically WTP for a gain, WTA for a loss — and stick to it.

Protest zeros. Some respondents state zero WTP because they object to the payment mechanism, distrust the implementing agency, or regard provision as a government obligation. Identify protests through follow-up debriefing questions (“Why did you say no?”). Then choose one of three documented approaches:

  1. Drop and report bounds: exclude protests and report mean WTP both including and excluding them as bounds on the population estimate.
  2. Selection correction: treat the protest decision as a selection model (Heckman) and estimate WTP for non-protesters with a Mills-ratio correction.
  3. Sensitivity bounds: assign protests the lowest and highest plausible WTP values and report the resulting range.

Report the protest rate alongside the main estimate; protest rates above ~15% suggest the payment vehicle should be reconsidered.

Payment vehicle bias. WTP estimates can differ materially across payment vehicles for the same good. When comparing across studies, payment-vehicle differences are a common and under-appreciated source of heterogeneity.

Survey mode effects. The NOAA Panel preferred in-person interviews. Modern CV is increasingly done online; mode effects are large and not fully understood. Web surveys often produce lower WTP than in-person, more “don’t know” responses, and weaker scope sensitivity. If using web mode, pilot against an in-person sub-sample and report any divergence.

Strategic behaviour. Open-ended formats invite respondents to overstate (to push provision) or understate (to avoid being charged). DC formats reduce but do not eliminate strategic incentives.

Aggregation from individual WTP to population benefit. To scale mean (or median) WTP to a total population benefit, multiply by the relevant population. Use sample weights if the sample is not representative. Do not extrapolate WTP to populations whose covariates fall outside the sample range. For cost-benefit purposes, report sensitivity to the choice of mean versus median.

Critiques and ongoing debate. Diamond and Hausman (1994) and Hausman (2012) argue that scope insensitivity, embedding, and large WTA-WTP gaps indicate CV measures something other than welfare-relevant valuation. Carson (2012) and Kling, Phaneuf and Zhao (2012) argue these problems are tractable under NOAA Panel design. The debate is unresolved. For a 2026 reader: treat NOAA-compliant DC CV with consequentiality framing as the defensible state of the art; treat departures from those design choices as evidence to discount the resulting estimates.

Analysis Guide

import numpy as np
import pandas as pd
import statsmodels.formula.api as smf

# 1. Single-bounded DC probit — acceptance as a function of the bid; probit link
#    keeps predicted probabilities in [0, 1]; the bid coefficient must be negative
#    (higher bid -> lower acceptance) for the model to identify a WTP distribution.
#    For cluster-randomised bid assignment or cluster-sampled designs, add
#    cov_type='cluster', cov_kwds={'groups': df['psu']} for clustered SEs
fit = smf.probit('wtp_response ~ bid_amount', data=df).fit(cov_type='HC1')
print(fit.summary())

# 2. Median WTP from a probit — for a symmetric distribution (normal/logistic),
#    -intercept / bid_coef is both mean and median. Under the log-logistic model
#    that R uses for skewed WTP, this formula gives the MEDIAN, not the mean.
#    Report both labels carefully
b0, b_bid = fit.params['Intercept'], fit.params['bid_amount']
median_wtp = -b0 / b_bid
print('Median WTP (probit, symmetric):', median_wtp)

# 3. Krinsky-Robb confidence interval — draw 5000 parameter vectors from the
#    asymptotic multivariate-normal of the estimates, compute median WTP for each,
#    take 2.5/97.5 percentiles. More reliable than the delta method for ratios
np.random.seed(42)
draws = np.random.multivariate_normal(fit.params.values, fit.cov_params().values, size=5000)
wtp_draws = -draws[:, 0] / draws[:, 1]
ci_lo, ci_hi = np.percentile(wtp_draws, [2.5, 97.5])
print(f'Median WTP 95% CI: ({ci_lo:.2f}, {ci_hi:.2f})')

# 4. WTP at a respondent profile — include ALL covariates in the profile; omitting
#    any covariate gives a wrong WTP. Do not extrapolate to profiles outside the
#    sample range
fit2 = smf.probit('wtp_response ~ bid_amount + age + C(female) + log_hh_expenditure',
                data=df).fit(cov_type='HC1')
p = fit2.params
xb_profile = (p['Intercept']
            + p['age'] * 30
            + p['C(female)[T.1]']
            + p['log_hh_expenditure'] * np.log(60000))
median_wtp_profile = -xb_profile / p['bid_amount']

# 5. Raw acceptance proportions by bid — the empirical demand curve. This is NOT
#    Turnbull (which additionally enforces monotonicity via pool-adjacent-violators
#    and integrates the step function for a lower-bound mean). For proper Turnbull,
#    see the R tab — DCchoice::turnbull.sb is the canonical implementation
p_yes = df.groupby('bid_amount')['wtp_response'].mean().reset_index()
print(p_yes)

SurveyCTO / XLSForm

Bid assignment. Pre-randomise bid assignments to respondents on the server (in a CSV keyed by respondent ID) and load with pulldata() rather than randomising in the form. This makes assignment deterministic, supports stratification on baseline covariates, and avoids the failure modes of in-form random(). The pre-randomisation script should distribute respondents evenly across the 4–8 bid levels with caps on each bid level’s share.

For an in-form fallback when pre-randomisation is not feasible, use once(if(random() < 0.5, 0, 1)) for a simple two-arm design or once(int(random() * N)) cast to a 1..N index into a bid vector. The outer once() is critical — without it the expression re-evaluates on every form recompute.

Persisted fields. Store at minimum: bid_assigned (the first bid shown), wtp_response (1/0 yes/no), protest_reason_code (categorical, from the debriefing question below). For double-bounded DC, also persist bid_followup (the actually-presented follow-up amount, which depends on the first response) and wtp_response2.

Scenario presentation. The scenario text should remind respondents of (i) the good or service being offered, (ii) the payment vehicle and frequency, (iii) their budget constraint, (iv) substitutes available without the offered good, and (v) the no-vote option. Display all five before the bid question, not split across screens.

Cheap-talk and consequentiality. Include a short cheap-talk note before the bid question (“People sometimes say they would pay more than they actually would when no money is at stake — please answer as if you would actually have to pay.”) and a consequentiality statement (“Results from this survey will be used to advise [decision-maker] on whether to [proceed with the project].”). Both improve estimate quality (Cummings & Taylor, 1999; Vossler, Doyon & Rondeau, 2012).

Debriefing for “no” responses. Follow every “no” with: “You said you would not pay [amount]. Is it because (a) you cannot afford it; (b) you do not value this service; (c) you object to the payment mechanism; (d) you think someone else should pay; (e) other?” Codes (c) and (d) are the standard protest categories.

Reading the output

  • The expression −α/β (Python: -fit.params['Intercept'] / fit.params['bid_amount']; R: extracted by DCchoice::summary(fit_sb)) gives median WTP under the assumed distribution, in the same currency units as the bid. For symmetric distributions (probit, logit on raw bid) mean = median; under log-logistic or log-normal the mean differs and must be computed separately.
  • Report both mean and median WTP, with Krinsky-Robb 95% confidence intervals on both. If they diverge materially (more than ~20%), the WTP distribution is heavily skewed — favour the median.
  • Turnbull (R, turnbull.sb) gives a non-parametric lower bound on mean WTP. Use as a robustness check; if the parametric estimate is substantially higher than the Turnbull bound and the distributional fit looks poor, prefer the Turnbull.
  • The raw acceptance proportion at each bid (Python step 5) is the empirical demand curve. A steep drop between two adjacent bids signals that the median WTP falls between them — useful for bid-vector calibration in a pilot.
  • The bid coefficient must be negative and statistically significant. If it is not, the model is unidentified — the bid vector is too narrow, sample is too small, or respondents are not responding to bids. Do not compute WTP from an unidentified model.
  • Protest rate above ~15%: the payment vehicle should be reconsidered. Report the rate alongside the main WTP estimate and the sensitivity to the protest-handling rule.
  • Scope test: mean WTP in the larger-quantity arm should be statistically larger than in the smaller-quantity arm. If they are not distinguishable, report the failure — the result is informative about the limits of the elicitation.

References

Aizaki, H., Nakatani, T., & Sato, K. (2014). Stated preference methods using R. CRC Press.

Alberini, A. (1995). Optimal designs for discrete choice contingent valuation surveys: Single-bound, double-bound, and bivariate models. Journal of Environmental Economics and Management, 28(3), 287–306. https://doi.org/10.1006/jeem.1995.1019

Alberini, A., Kanninen, B., & Carson, R. T. (1997). Modeling response incentive effects in dichotomous choice contingent valuation data. Land Economics, 73(3), 309–324. https://doi.org/10.2307/3147170

Arrow, K., Solow, R., Portney, P. R., Leamer, E. E., Radner, R., & Schuman, H. (1993). Report of the NOAA panel on contingent valuation. Federal Register, 58(10), 4601–4614. https://www.federalregister.gov/documents/1993/01/15

Bishop, R. C., & Boyle, K. J. (Eds.). (2019). A primer on nonmarket valuation (2nd ed.). Springer. https://doi.org/10.1007/978-94-007-7104-8

Bishop, R. C., & Heberlein, T. A. (1979). Measuring values of extramarket goods: Are indirect measures biased? American Journal of Agricultural Economics, 61(5), 926–930. https://doi.org/10.2307/3180348

Cameron, T. A. (1988). A new paradigm for valuing non-market goods using referendum data: Maximum likelihood estimation by censored logistic regression. Journal of Environmental Economics and Management, 15(3), 355–379. https://doi.org/10.1016/0095-0696(88)90008-3

Cameron, T. A., & James, M. D. (1987). Efficient estimation methods for “closed-ended” contingent valuation surveys. Review of Economics and Statistics, 69(2), 269–276. https://doi.org/10.2307/1927234

Cameron, T. A., & Quiggin, J. (1994). Estimation using contingent valuation data from a “dichotomous choice with follow-up” questionnaire. Journal of Environmental Economics and Management, 27(3), 218–234. https://doi.org/10.1006/jeem.1994.1035

Carson, R. T. (2012). Contingent valuation: A practical alternative when prices aren’t available. Journal of Economic Perspectives, 26(4), 27–42. https://doi.org/10.1257/jep.26.4.27

Carson, R. T., & Hanemann, W. M. (2005). Contingent valuation. In K.-G. Mäler & J. R. Vincent (Eds.), Handbook of environmental economics (Vol. 2, pp. 821–936). North-Holland.

Carson, R. T., & Mitchell, R. C. (1989). Using surveys to value public goods: The contingent valuation method. Resources for the Future.

Cooper, J. C., Hanemann, M., & Signorello, G. (2002). One-and-one-half-bound dichotomous choice contingent valuation. Review of Economics and Statistics, 84(4), 742–750. https://doi.org/10.1162/003465302760556549

Cummings, R. G., & Taylor, L. O. (1999). Unbiased value estimates for environmental goods: A cheap talk design for the contingent valuation method. American Economic Review, 89(3), 649–665. https://doi.org/10.1257/aer.89.3.649

Diamond, P. A., & Hausman, J. A. (1994). Contingent valuation: Is some number better than no number? Journal of Economic Perspectives, 8(4), 45–64. https://doi.org/10.1257/jep.8.4.45

Haab, T. C., & McConnell, K. E. (1997). Referendum models and negative willingness to pay: Alternative solutions. Journal of Environmental Economics and Management, 32(2), 251–270. https://doi.org/10.1006/jeem.1996.0968

Hanemann, W. M. (1984). Welfare evaluations in contingent valuation experiments with discrete responses. American Journal of Agricultural Economics, 66(3), 332–341. https://doi.org/10.2307/1240800

Hanemann, W. M. (1989). Welfare evaluations in contingent valuation experiments with discrete response data: Reply. American Journal of Agricultural Economics, 71(4), 1057–1061. https://doi.org/10.2307/1242685

Hanemann, W. M., Loomis, J., & Kanninen, B. (1991). Statistical efficiency of double-bounded dichotomous choice contingent valuation. American Journal of Agricultural Economics, 73(4), 1255–1263. https://doi.org/10.2307/1242453

Hausman, J. (2012). Contingent valuation: From dubious to hopeless. Journal of Economic Perspectives, 26(4), 43–56. https://doi.org/10.1257/jep.26.4.43

Horowitz, J. K., & McConnell, K. E. (2002). A review of WTA/WTP studies. Journal of Environmental Economics and Management, 44(3), 426–447. https://doi.org/10.1006/jeem.2001.1215

Kahneman, D., & Knetsch, J. L. (1992). Valuing public goods: The purchase of moral satisfaction. Journal of Environmental Economics and Management, 22(1), 57–70. https://doi.org/10.1016/0095-0696(92)90019-S

Kanninen, B. J. (1995). Bias in discrete response contingent valuation. Journal of Environmental Economics and Management, 28(1), 114–125. https://doi.org/10.1006/jeem.1995.1009

Kling, C. L., Phaneuf, D. J., & Zhao, J. (2012). From Exxon to BP: Has some number become better than no number? Journal of Economic Perspectives, 26(4), 3–26. https://doi.org/10.1257/jep.26.4.3

Krinsky, I., & Robb, A. L. (1986). On approximating the statistical properties of elasticities. Review of Economics and Statistics, 68(4), 715–719. https://doi.org/10.2307/1924536

List, J. A., & Gallet, C. A. (2001). What experimental protocol influence disparities between actual and hypothetical stated values? Environmental and Resource Economics, 20(3), 241–254. https://doi.org/10.1023/A:1012791005641

Mühlbacher, A. C., Kaczynski, A., Zweifel, P., & Johnson, F. R. (2016). Experimental measurement of preferences in health and healthcare using best-worst scaling: An overview. Health Economics Review, 6, 2. https://doi.org/10.1186/s13561-015-0079-x

Murphy, J. J., Allen, P. G., Stevens, T. H., & Weatherhead, D. (2005). A meta-analysis of hypothetical bias in stated preference valuation. Environmental and Resource Economics, 30(3), 313–325. https://doi.org/10.1007/s10640-004-3332-z

Plott, C. R., & Zeiler, K. (2005). The willingness to pay–willingness to accept gap, the “endowment effect,” subject misconceptions, and experimental procedures for eliciting valuations. American Economic Review, 95(3), 530–545. https://doi.org/10.1257/0002828054201387

Turnbull, B. W. (1976). The empirical distribution function with arbitrarily grouped, censored and truncated data. Journal of the Royal Statistical Society Series B, 38(3), 290–295. https://doi.org/10.1111/j.2517-6161.1976.tb01597.x

Vossler, C. A., Doyon, M., & Rondeau, D. (2012). Truth in consequentiality: Theory and field evidence on discrete choice experiments. American Economic Journal: Microeconomics, 4(4), 145–171. https://doi.org/10.1257/mic.4.4.145

Whittington, D. (1998). Administering contingent valuation surveys in developing countries. World Development, 26(1), 21–30. https://doi.org/10.1016/S0305-750X(97)00125-3

Whittington, D. (2010). What have we learned from 20 years of stated preference research in less-developed countries? Annual Review of Resource Economics, 2(1), 209–236. https://doi.org/10.1146/annurev.resource.012809.103908

Last updated: 5 June 2026