What it is
The Holt-Laury Multiple Price List (MPL) is the dominant laboratory method for eliciting risk preferences in economics (Holt & Laury, 2002). Respondents make ten binary choices, each between a relatively safe lottery (Option A) and a riskier lottery (Option B). Across the rows, the probability of the high payoff increases from 1/10 to 10/10 for both options. Option A offers a moderate high payoff and a small-gap low payoff; Option B offers a large high payoff and a near-zero low payoff.
A risk-neutral respondent should switch from A to B at row 5, where the expected values cross (making 4 safe choices in rows 1–4). Respondents who switch earlier — preferring the risky lottery even when its expected value is lower — are risk-seeking. Those who switch later are risk-averse, requiring a larger expected-value premium to accept the riskier option. The number of safe choices maps directly to a range of the Constant Relative Risk Aversion (CRRA) coefficient under expected utility with power utility (Holt & Laury, 2002, Table 3).
Holt and Laury (2002) found an asymmetric scale effect that has shaped the field’s incentivisation standard: as stakes rise, real-payoff choices reveal substantially more risk aversion, while hypothetical choices do not. The 2005 follow-up (Holt & Laury, 2005) confirmed the scale effect under randomised order. Hypothetical elicitation therefore tends to make populations look closer to risk-neutral than they are. For large-N surveys where incentivised MPL is impractical, the single-item risk question developed in the German Socioeconomic Panel (Dohmen et al., 2011) and used in the Global Preference Survey (Falk et al., 2018) is the established validated alternative.
When to use it
The Holt-Laury MPL is appropriate when risk aversion is a primary outcome, a moderator of treatment effects, or a control variable, and when respondents are sufficiently numerate to process probability-based choices.
Applications in development economics include baseline measurement of risk attitudes as a predictor of technology adoption, insurance take-up, and financial behaviour; heterogeneous treatment effects analysis where risk aversion is a plausible moderator; and cross-site comparisons of risk preference distributions. Tanaka, Camerer and Nguyen (2010) adapt the HL framework for Vietnamese household surveys and link elicited risk preferences to observed asset accumulation and farming decisions.
The method is less appropriate for low-numeracy populations who struggle with probability fractions — a well-documented field challenge. For those settings, the Eckel-Grossman single-gamble method sacrifices precision for comprehensibility and is often the better choice. Charness, Gneezy and Imas (2013), Dave et al. (2010), and Crosetto and Filippin (2016) provide systematic methodological comparisons; Holzmeister and Stefan (2021) is the most recent across-method consistency study. For very large samples or where incentivised laboratory protocols cannot be implemented, the Dohmen et al. (2011) single-item survey question — “How willing are you to take risks, in general? (0 = not at all, 10 = fully)” — has been validated against incentivised choices and is the workhorse measure in the Global Preference Survey (Falk et al., 2018).
How it works
Each row of the MPL presents two lotteries with a shared probability structure. Using the original Holt-Laury low-payoff design as illustration:
| Row | Option A | Option B |
|---|---|---|
| 1 | 1/10: $2.00, 9/10: $1.60 | 1/10: $3.85, 9/10: $0.10 |
| 2 | 2/10: $2.00, 8/10: $1.60 | 2/10: $3.85, 8/10: $0.10 |
| … | … | … |
| 10 | 10/10: $2.00, 0/10: $1.60 | 10/10: $3.85, 0/10: $0.10 |
Two equivalent summary statistics: the switching row (the first row at which the respondent picks B) and the number of safe choices n_safe (count of Option A choices across the 10 rows). Under monotone preferences (single switch) the two are deterministically related — switching row = n_safe + 1 — but n_safe is robust to multiple switching and is the canonical summary used in published implementations and in the structural estimation literature (Andersen, Harrison, Lau & Rutström, 2008). Compute and report both.
Under expected utility with power utility (the CRRA functional form), the number of safe choices maps to an interval for the risk-aversion coefficient r. The canonical mapping from Holt and Laury (2002, Table 3) for the low-payoff design is:
| n_safe | Switching row | CRRA range | Classification |
|---|---|---|---|
| 0–1 | 1–2 | r < −0.95 | Highly risk-loving |
| 2 | 3 | −0.95 to −0.49 | Very risk-loving |
| 3 | 4 | −0.49 to −0.15 | Risk-loving |
| 4 | 5 | −0.15 to 0.15 | Risk-neutral |
| 5 | 6 | 0.15 to 0.41 | Slightly risk-averse |
| 6 | 7 | 0.41 to 0.68 | Risk-averse |
| 7 | 8 | 0.68 to 0.97 | Very risk-averse |
| 8 | 9 | 0.97 to 1.37 | Highly risk-averse |
| 9–10 | 10 / never | r > 1.37 | Stay-safe |
The risk-neutral boundary (r = 0) is at 4 safe choices, i.e., switching at row 5 — this is where the expected values of the two options cross.
Joint risk-time estimation. The CRRA parameter elicited here is the standard input to the CRRA-adjusted discount factor in joint risk-time estimation (Andersen, Harrison, Lau & Rutström, 2008; see the Time Preference / Discount Rate guide). Studies that elicit both should fit the joint model rather than using a CRRA point estimate from this task and a separate discount rate from a time MPL.
Incentivisation. For incentivised sessions, the standard mechanism is random row payment: one of the 10 rows is selected at random after all choices are made, the respondent’s choice at that row determines which lottery they play, and the lottery is then resolved for real payment. Random row payment is incentive-compatible under expected utility (Holt, 1986; Charness, Gneezy & Halevy, 2017 discuss the assumptions). Paying all rows generates wealth effects across choices; paying only one fixed row removes incentives on the unrewarded rows.
Key decisions
Real vs hypothetical payoffs. Holt and Laury (2002) showed that hypothetical choices produce substantially more risk-neutral responses than incentivised ones, and that the scale effect (more aversion at larger stakes) appears only under real payoffs. Real payoffs are strongly preferred when feasible. When they are not feasible, three fallback rules: (1) use modest real stakes calibrated to local wages even if smaller than the ideal — small real stakes still outperform hypothetical; (2) if fully hypothetical, frame as concretely as possible and expect underestimation of risk aversion; (3) for very large surveys, switch to the Dohmen et al. (2011) single-item alternative rather than running a 10-row hypothetical MPL whose CRRA categories are biased.
Scaling effects. Holt and Laury (2002, 2005) found that CRRA estimates rise systematically with the magnitude of stakes — the same respondent looks more risk-averse under high real stakes than low real stakes. CRRA categories from studies using different payoff levels are therefore not directly comparable. Document the exact payoff amounts used and report them.
Adapting payoffs for the field context. Replace the original dollar amounts with local currency equivalents calibrated to local wage rates, typically one to three days’ wages for the high payoff in Option B. Preserve the ratio between Option A and Option B payoffs: the risky option (B) must offer a substantially higher top payoff to ensure the switching range spans the realistic risk-preference distribution. Pilot the payoff levels with a small sample first; Tanaka, Camerer and Nguyen (2010) document this calibration in Vietnamese villages and combine the basic structure with prospect-theory parameters for probability weighting and loss aversion.
Presenting probabilities. In field settings with low numeracy, replace fractions (2/10, 5/10) with visual aids — coloured balls in a bag, a pie chart, or a physical spinner divided into sectors. The number of balls in the bag or sectors on the spinner should correspond directly to the probability (e.g., a bag of 10 balls, 2 of which are red for the high payoff). Visual aids substantially reduce comprehension-driven noise.
Handling multiple switchers. Holt and Laury (2002) report ~13% multiple switchers in their original low-stakes design; Charness, Gneezy and Imas (2013) cite a range from 10–55% depending on instrument and population. Multiple switching is inconsistent with monotone expected utility but in practice reflects inattention, incomplete understanding, or genuine non-EU preferences (probability weighting, loss aversion). Three options: report n_safe (robust to multi-switching) as the primary summary; for switching-row analyses, use the first switching point and flag multi-switchers as a reliability indicator; exclude them in a sensitivity analysis if the rate is below ~15%. If the rate is above ~20%, the implementation or comprehension procedure needs review before treating the data as informative about preferences.
Number of rows. The standard 10-row design is well validated. Shorter versions (6 or 7 rows) reduce respondent burden but provide coarser classification; Lévy-Garboua et al. (2012) and others propose specific modified designs. Longer versions add rows at the extremes but provide little additional information for typical risk-preference distributions.
Sample size. For detecting a mean difference of half a row in switching point at 80% power and α = 0.05, plan for approximately N ≈ 200 per group. For regressions on covariates with reasonable signal-to-noise, 300–500 respondents per group is standard. Cluster-randomised designs need a design-effect adjustment of 1 + (m̄ − 1)ρ.
Caveats & common mistakes
The CRRA assumption — and probability weighting. Interpreting the switching row as a CRRA coefficient requires assuming respondents maximise expected utility under a power utility function. Under cumulative prospect theory (Tversky & Kahneman, 1992), choices in the Holt-Laury MPL are jointly affected by utility curvature and probability weighting — the MPL varies probability across rows, so it does not separately identify the two. The CRRA recovered from MPL choices is biased upward when respondents overweight small probabilities and downward when they underweight (Andersen et al., 2008; Harrison & Rutström, 2008). The switching row remains a valid non-parametric rank of risk aversion; the CRRA interval interpretation is a parametric layer that the CPT extension contests. For studies where the CPT separation matters, use Tanaka, Camerer and Nguyen (2010) or a structural design that varies probability and stake size independently.
Probability comprehension. The most common failure mode in field implementations is that respondents do not reliably understand the probability structure. Symptoms include a large share who never switch (switch_row == 11 in the code), who switch on the first row (switch_row == 1), or who multi-switch. If the combined share of “always-B” plus “never-switch” respondents exceeds about 10%, the payoff scale is mis-calibrated or comprehension is failing — review payoffs and the comprehension check before treating the data as informative about preferences. A scored comprehension exercise before the main task is mandatory, not advisable (see SurveyCTO section).
Functional-form alternatives. CRRA is the standard but not the only choice. CARA (constant absolute risk aversion) and expo-power forms are used in some structural papers (Harrison & Rutström, 2008). Switching-row classifications under different functional forms give different bin boundaries; the CRRA mapping in “How it works” is specific to power utility. Report the functional form assumption explicitly.
Gender and elicitation method. Filippin and Crosetto (2016) show that the gender gap in risk-taking is robust in the Holt-Laury MPL but attenuated or absent in some other elicitation tasks (Eckel-Grossman, BRET). When comparing risk preferences across genders, report the elicitation method clearly and interpret the gap as conditional on the method.
Domain specificity. Risk aversion measured over monetary lotteries does not necessarily predict risk-taking in health, agricultural, or social domains. Respondents who are highly risk-averse in the HL task may take significant health risks in daily life, and vice versa (Dohmen et al., 2011, document this domain-specificity). If the research question concerns domain-specific risk behaviour, supplementary domain-specific measures are more informative than the HL measure alone.
Order effects. When the MPL is embedded in a larger battery — especially alongside time preference or other economic-game tasks — responses can be affected by task order and the stakes of preceding tasks. Holt and Laury (2005) document that scale effects persist under randomised order. Randomise task order across respondents and use a single random-row payment procedure across the battery to reduce but not eliminate these effects.
Analysis Guide
import numpy as np
import pandas as pd
import statsmodels.formula.api as smf
# Choice matrix: rows = respondents, cols = choice_1 ... choice_10 (0 = A safe, 1 = B risky)
choice_cols = [f'choice_{i}' for i in range(1, 11)]
M = df[choice_cols].to_numpy()
valid = ~df[choice_cols].isna().any(axis=1) # drop respondents with any missing choice
# 1. n_safe = number of safe (A) choices across the 10 rows. Robust to multi-switching
# and the canonical summary mapped to CRRA in Holt-Laury (2002) Table 3
df['n_safe'] = np.where(valid, (M == 0).sum(axis=1), np.nan)
# 2. First switching row under monotone preferences; switch_row = n_safe + 1.
# Compute separately so the diagnostics in step 3 work; respondents with missing
# choices get switch_row = NaN to exclude them from category assignment
first_b = np.where(M == 1, np.arange(1, 11), 99).min(axis=1)
df['switch_row'] = np.where(valid,
np.where(first_b == 99, 11, first_b),
np.nan)
# 3. Monotonicity diagnostic — n_switches counts all A->B and B->A transitions; under
# monotone EU preferences this should be 0 (all-A or all-B) or 1 (single switch).
# Values > 1 are monotonicity violations; report the share alongside the estimate
df['n_switches'] = np.where(valid, np.abs(np.diff(M, axis=1)).sum(axis=1), np.nan)
df['multiple_switcher'] = (df['n_switches'] > 1).astype('Int64')
# 4. Distribution diagnostics — share at boundaries flags comprehension failure;
# combined share of always-B (switch_row == 1) and never-switch (switch_row == 11)
# above ~10% signals miscalibrated payoffs or comprehension problems
share_always_B = (df['switch_row'] == 1).mean()
share_never = (df['switch_row'] == 11).mean()
share_multi = (df['multiple_switcher'] == 1).mean()
print(f'always-B: {share_always_B:.2%} never-switch: {share_never:.2%} '
f'multi: {share_multi:.2%} boundary combined: {share_always_B + share_never:.2%}')
# 5. CRRA category from Holt-Laury (2002) Table 3 indexed by n_safe — this is the
# canonical mapping; switch_row = n_safe + 1 under monotone choices. Treat as
# approximate intervals, not point estimates, and report the functional-form
# assumption (power utility) alongside
df['crra_category'] = pd.cut(
df['n_safe'],
bins=[-1, 1, 2, 3, 4, 5, 6, 7, 8, 10],
labels=['highly risk-loving', 'very risk-loving', 'risk-loving',
'risk-neutral', 'slightly risk-averse', 'risk-averse',
'very risk-averse', 'highly risk-averse', 'stay-safe'])
# 6. Regress switch_row on characteristics, excluding multi-switchers in a sensitivity
# analysis. A positive coefficient means later switching = more risk aversion.
# For inference on the CRRA parameter itself (not switch_row), interval regression
# on the implied CRRA bounds is the standard alternative (Andersen et al. 2008).
clean = df[df['multiple_switcher'] == 0]
fit = smf.ols('switch_row ~ age + C(female) + log_hh_expenditure + C(treatment)',
data=clean).fit(cov_type='HC2')
print(fit.summary()) XLSForm / SurveyCTO
Display the 10 rows as 10 individual select_one yes_no fields (one per row) or as a single select_multiple with the rows labelled. Use a physical randomisation device — a spinner or a bag of 10 coloured balls — to convey probabilities, rather than text fractions alone. The balls/sectors should correspond one-to-one with the probability (a bag of 10 balls, 2 red for a 2/10 high-payoff probability).
Comprehension check before the main task is mandatory. The standard single-probability question (“If I have 10 balls and 3 are red, what are the chances of drawing a red ball?”) tests numeracy on a single probability but not the joint structure of the MPL. A stronger check has two parts: (a) present row 1 (1/10 of the high payoff for both options) and ask which option pays more on average; (b) present row 10 (a certain payoff in both options) and ask which pays more. Only respondents who answer both correctly proceed; the others are re-trained or excluded under a pre-specified rule.
Random row payment. At the end of the task, draw a random integer in [1, 10] and pay the choice the respondent made at that row. The canonical SurveyCTO pattern stores the draw to a hidden field so it cannot be re-randomised on form recompute:
type name calculation
calculate payment_row once(int(random() * 10) + 1)
Reference ${payment_row} downstream in the payment-display calculation; never call random() a second time. Store all 10 choices in the submitted data — not just the switching row — so multiple switchers can be identified in cleaning. Also store payment_row and the realised lottery outcome.
Reading the output
- Primary summaries. Report both
n_safe(count of safe choices, 0–10) andswitch_row(1–11; 11 = never switched).n_safeis robust to multi-switching and is the canonical input to the CRRA mapping;switch_rowis the more intuitive descriptor. - CRRA mapping. From Holt-Laury (2002) Table 3,
n_safe = 0–1is highly risk-loving (r < −0.95),n_safe = 4is risk-neutral (r ∈ [−0.15, 0.15], the row where expected values cross), andn_safe = 9–10is the stay-safe band (r > 1.37). Always report the functional form (power utility) and the payoff scale alongside the CRRA category — categories are not comparable across studies with different payoff levels (Holt & Laury, 2005). - Multi-switching diagnostic.
multiple_switcher == 1flags respondents whose choices violate monotone EU. A rate above ~15% suggests comprehension or implementation problems; above ~20% warrants stopping and reviewing the procedure before treating the data as informative. - Boundary-mass diagnostic. Combined share of always-B (
switch_row == 1) plus never-switch (switch_row == 11) above ~10% signals miscalibrated payoffs or comprehension failure — review the payoff scale and the comprehension check before interpreting CRRA distributions. - Regression interpretation. A positive coefficient on a covariate means later switching (greater risk aversion); a negative coefficient means earlier switching. OLS on
switch_rowtreats the discrete ordinal outcome as continuous — a useful approximation but coarse. For CRRA-parameter inference, interval regression on the implied CRRA bounds is the standard structural alternative (Andersen et al., 2008). - Joint with time preferences. When the design also elicits a time MPL, the CRRA estimated here is the standard input to the CRRA-adjusted discount factor δ_adj = (LL/SS)^(−1/ρ) (Andersen et al., 2008). Fit the joint risk-time model rather than substituting a point estimate.
References
Andersen, S., Harrison, G. W., Lau, M. I., & Rutström, E. E. (2006). Elicitation using multiple price list formats. Experimental Economics, 9(4), 383–405. https://doi.org/10.1007/s10683-006-7055-6
Andersen, S., Harrison, G. W., Lau, M. I., & Rutström, E. E. (2008). Eliciting risk and time preferences. Econometrica, 76(3), 583–618. https://doi.org/10.1111/j.1468-0262.2008.00848.x
Charness, G., Gneezy, U., & Halevy, Y. (2017). The (im)possibility of incentivizing the elicitation of economic preferences. European Economic Review, 100, 51–66. https://doi.org/10.1016/j.euroecorev.2017.07.005
Charness, G., Gneezy, U., & Imas, A. (2013). Experimental methods: Eliciting risk preferences. Journal of Economic Behavior & Organization, 87, 43–51. https://doi.org/10.1016/j.jebo.2012.12.023
Crosetto, P., & Filippin, A. (2016). A theoretical and experimental appraisal of four risk elicitation methods. Experimental Economics, 19(3), 613–641. https://doi.org/10.1007/s10683-015-9457-9
Dave, C., Eckel, C. C., Johnson, C. A., & Rojas, C. (2010). Eliciting risk preferences: When is simple better? Journal of Risk and Uncertainty, 41(3), 219–243. https://doi.org/10.1007/s11166-010-9103-z
Dohmen, T., Falk, A., Huffman, D., Sunde, U., Schupp, J., & Wagner, G. G. (2011). Individual risk attitudes: Measurement, determinants, and behavioral consequences. Journal of the European Economic Association, 9(3), 522–550. https://doi.org/10.1111/j.1542-4774.2011.01015.x
Falk, A., Becker, A., Dohmen, T., Enke, B., Huffman, D., & Sunde, U. (2018). Global evidence on economic preferences. Quarterly Journal of Economics, 133(4), 1645–1692. https://doi.org/10.1093/qje/qjy013
Filippin, A., & Crosetto, P. (2016). A reconsideration of gender differences in risk attitudes. Management Science, 62(11), 3138–3160. https://doi.org/10.1287/mnsc.2015.2294
Harrison, G. W., & Rutström, E. E. (2008). Risk aversion in the laboratory. Research in Experimental Economics, 12, 41–196. https://doi.org/10.1016/S0193-2306(08)00003-3
Holt, C. A., & Laury, S. K. (2002). Risk aversion and incentive effects. American Economic Review, 92(5), 1644–1655. https://doi.org/10.1257/000282802762024700
Holt, C. A., & Laury, S. K. (2005). Risk aversion and incentive effects: New data without order effects. American Economic Review, 95(3), 902–904. https://doi.org/10.1257/0002828054201459
Holzmeister, F., & Stefan, M. (2021). The risk elicitation puzzle revisited: Across-methods (in)consistency? Experimental Economics, 24(2), 593–616. https://doi.org/10.1007/s10683-020-09674-8
Lévy-Garboua, L., Maafi, H., Masclet, D., & Terracol, A. (2012). Risk aversion and framing effects. Experimental Economics, 15(1), 128–144. https://doi.org/10.1007/s10683-011-9293-5
Tanaka, T., Camerer, C. F., & Nguyen, Q. (2010). Risk and time preferences: Linking experimental and household survey data from Vietnam. American Economic Review, 100(1), 557–571. https://doi.org/10.1257/aer.100.1.557
Tversky, A., & Kahneman, D. (1992). Advances in prospect theory: Cumulative representation of uncertainty. Journal of Risk and Uncertainty, 5(4), 297–323. https://doi.org/10.1007/BF00122574