What it is
The trust game, introduced by Berg, Dickhaut and McCabe (1995), is the standard incentivised tool for measuring trust and trustworthiness as distinct behavioural constructs. Player 1 (the investor) receives an endowment and chooses how much to send to Player 2 (the trustee). The amount sent is multiplied by a factor k — typically 3 in the original design — and Player 2 then chooses how much of the augmented amount to return. The game is one-shot: subjects play exactly once, with no future interaction.
Player 1’s send is read as trust: the willingness to accept vulnerability in exchange for a possible larger return. Player 2’s return is read as trustworthiness: the willingness to share gains when there is no obligation to do so. The two are empirically distinct (Ashraf, Bohnet & Piankov 2006). Both readings are bounded by identification assumptions discussed below — neither send nor return is a clean construct without auxiliary controls.
When to use it
The trust game is appropriate when trust, trustworthiness, or reciprocal exchange is a primary outcome, a treatment mechanism, or a baseline measure of social capital. It is also used as a component of broader pro-social preference batteries.
Glaeser, Laibson, Scheinkman & Soutter (2000) administer the investment game to Harvard students alongside a survey, and find that survey questions about past trusting behaviour (lending money, leaving doors unlocked) predict Player 1’s send. The standard GSS attitudinal trust question (“most people can be trusted”) does not predict the send — but does predict Player 2’s return. Their conclusion: the standard survey trust measure is closer to a measure of trustworthiness than of trust.
Karlan (2005) administers the investment game to FINCA Peru microfinance borrowers and finds that Player 2 trustworthiness predicts loan repayment a year later; Player 1 send rates predict savings choices but are weaker predictors of repayment. This is the canonical demonstration that the two components of the trust game pick up distinct field-relevant behaviours.
For comparative work on identity and trust, Fershtman & Gneezy (2001) is the seminal in-group/out-group trust game in Israeli-Jewish settings. The design template — randomising counterpart identity (real or described) — has been adapted to caste, religion, and community in South Asian field settings.
The trust game is less appropriate when pure altruism is the quantity of interest — the dictator game better isolates altruism by removing Player 2’s response. It is also less appropriate when the relevant construct is generalised trust at the population level: Falk, Becker, Dohmen, Enke, Huffman & Sunde (2018) Global Preference Survey scale is the standard survey alternative. External validity of incentivised lab games for field behaviour is contested (Levitt & List 2007); pairing the trust game with a real-stakes field outcome strengthens the inferential chain.
How it works
Stage 1 — Investment. Player 1 receives an endowment E (Berg, Dickhaut & McCabe used $10) and chooses an amount s ∈ [0, E] to send. They keep E − s. The sent amount is multiplied by k before reaching Player 2.
Stage 2 — Return. Player 2 receives ks and chooses an amount r ∈ [0, ks] to return. They keep ks − r.
Outcomes. Trust is the share sent: s / E. Trustworthiness is the share returned: r / (ks). Player 1’s net gain from sending is r − s; this is positive when r / (ks) > 1/k, i.e., when Player 2 returns more than 1/k of the multiplied amount. With k = 3, the break-even return share is 1/3.
Identification assumptions. For Player 1’s send to identify trust and Player 2’s return to identify trustworthiness, six conditions must hold.
- Independence. Each subject plays one role in exactly one pair, with no repeated interaction. Violated in repeated trust games and any design where subjects play multiple rounds.
- No spillover across pairs within a session. Even when each subject plays once, if other pairs play first and outcomes are visible, behaviour is contaminated. Run all pairs simultaneously or in private sequence.
- Anonymity is operative. Subjects must believe their identity is unlinked to their decision — both to the counterpart and to the enumerator. Field designs frequently fail this through enumerator-visible tablet screens or post-session matching.
- Comprehension. Subjects understand the multiplication factor, the return mechanism, and the one-shot nature. Failure produces noise indistinguishable from low trust; comprehension is the single biggest source of measurement error in low-literacy field deployments.
- Framing matches interpretation. The intended interpretation is a one-shot anonymous exchange. If subjects interpret the transfer as a gift, a debt, or an investment with a future return, the measurement is contaminated by the connotation.
- No enumerator-mediated information leakage. Enumerators must not transmit information between Player 1 and Player 2 — including subtle differences in framing across pairs.
The subgame-perfect Nash equilibrium under standard self-interest is (s = 0, r = 0): a rational Player 2 returns nothing, so a rational Player 1 sends nothing. Empirical send and return rates above zero are therefore the object of measurement; the meta-analytic averages (Johnson & Mislin 2011) are roughly s/E ≈ 0.50 and r/(ks) ≈ 0.37.
Key decisions
Multiplication factor k. The original Berg–Dickhaut–McCabe design uses k = 3, and this is the modal choice in the literature, but Johnson & Mislin’s (2011) meta-analysis of 162 replications tabulates k = 2 and k = 4 as common alternatives. Higher k raises the joint surplus and lowers the trustworthiness threshold for Player 1 to break even. State k clearly to both players before play; report it in the dataset for analysis.
Endowment size. Stakes should be meaningful in local context — half a day’s to a full day’s wage is a common anchor. Charness, Gneezy & Halladay (2016) meta-analyse stakes effects across lab games and find behaviour is broadly robust to stake variation within typical ranges, but very small stakes generate noise and very large stakes shift the decision toward financial-transaction framing.
Strategy method vs direct response. In the direct-response design, Player 2 sees the actual amount sent before deciding how much to return. The strategy method asks Player 2 to specify a return for every possible send before learning what Player 1 actually sent — yielding a full return function per subject. Bellemare & Kröger (2007) and the Brandts & Charness (2011) meta-analysis find that the strategy method produces results broadly similar to direct response in trust games, with some attenuation of negative reciprocity. The strategy method adds cognitive load and benefits from a comprehension check during pilot.
Counterpart identity. Pairing with a co-ethnic, co-religionist, caste-in-group member, or co-villager typically raises both send and return relative to out-group pairing. Identity can be made salient (real partner, identified by group) or described (hypothetical partner, group attribute stated). Real-partner designs are more credible but logistically harder. Randomising identity within subject (each subject plays multiple rounds against different counterparts) recovers within-person sensitivity but violates the one-shot independence assumption — choose accordingly.
Anonymity protocol. Double-blind anonymity (subjects do not know each other’s identity; the enumerator does not know which decision belongs to which subject) is the design ideal. de Quidt, Haushofer & Roth (2018) show that experimenter demand effects can shift behaviour even under nominal anonymity; designing the protocol so that no demand cue points subjects toward a “right” answer matters as much as physical separation. In field settings, sessions where players are co-located but visually separated, with private decision entry, approximate the necessary conditions.
Sample size. Power calculations should be calibrated against the empirical standard deviation of trust shares: Johnson & Mislin (2011) report SD ≈ 0.25 for the share sent across studies. Detecting a 0.05 share-sent effect at 80% power with α = 0.05, equal allocation, and no clustering requires n ≈ 786 (using the standard two-sample formula n = 2 × (z_critical + z_power)² × σ² / δ², with z_critical = 1.96 and z_power = 0.84). Sessions are the realistic cluster (Fréchette 2012); inflating by the design effect 1 + (m − 1) × ρ for m = 20 subjects per session and ρ = 0.05 (an empirically reasonable trust-game ICC) gives a multiplier of 1.95, so n ≈ 1530 individuals or roughly 77 sessions of 20. Pilot the ICC if it is unknown; the session-level multiplier is the single biggest determinant of total sample.
Comprehension checks. A pre-decision quiz — testing that the subject understands the multiplication factor, who decides what, and that the interaction is one-shot — reduces noise and protects identification assumption (4). Pilot with re-routing on incorrect answers.
Caveats & common mistakes
Trust is empirically separable from risk preferences. A common framing claims that Player 1’s send is a gamble on Player 2’s response and therefore confounded with risk aversion. The literature has tested this and rejected the strong form: Eckel & Wilson (2004) and Houser, Schunk & Winter (2010) show that trust behaviour and risky-bet behaviour diverge even when stakes and probabilities are matched; Fairley, Sanfey, Vyrastekova & Weitzel (2016) reinforces. Risk preferences contribute to the send, but trust is not reducible to a risk preference. Elicit a separate risk measure (the Holt-Laury or Eckel-Grossman guides) as a control rather than a substitute.
Trustworthiness is contaminated by altruism, inequality aversion, and norms — and the dictator-game subtraction is not a complete fix. Cox (2004) is the canonical methodological answer: a triadic design runs the trust game alongside two dictator-game controls (one with Player 1’s endowment, one with Player 2’s received amount). Subtracting the matched dictator-game allocations from the trust-game send and return identifies the components attributable to trust and reciprocity respectively. The simple “trust game return minus dictator game gift” comparison the field has historically relied on is incomplete because it does not match the endowment Player 2 controls in the two settings. Run the triadic design if separating reciprocity from altruism is essential.
Selection in the trustworthiness regression. Trustworthiness is only defined when Player 1 sent something. Restricting the trustworthiness regression to p1_send > 0 is conditioning on an endogenous outcome: Player 1’s decision is itself a function of unobserved characteristics that may also affect Player 2’s return (especially in within-session designs with common shocks). The OLS coefficient on a covariate in the conditional regression is not the population effect on trustworthiness. Heckman-style sample-selection models (with an exclusion restriction) or Lee (2009) sharp bounds give credible alternatives.
Session-level dependence. Behavioural games are run in sessions. Subjects within a session share enumerator framing, hear the instructions together, may observe each other’s choices, and absorb a common shock. Fréchette (2012) is the standard reference for session-level dependence in lab experiments. Heteroskedasticity-robust standard errors (HC1/HC2) treat each subject as independent and are wrong when sessions are the unit at which design randomness is consumed. Cluster at the session level (or the village/day level in field deployments) for both Player 1 and Player 2 regressions.
Sequential structure and the strategy method. When Player 2 sees the actual send before responding, the return is a function of the level sent, not just the proportion. Disentangling an “amount effect” (more tokens in hand) from a “proportion effect” (larger send signals more trust) requires the strategy method or controlled variation in k. With direct response and a single value of k, the two effects are not identified separately.
Demand effects under nominal anonymity. “Anonymity reduces demand effects” is a common claim and is partially right: physical separation and double-blind protocols help. But de Quidt, Haushofer & Roth (2018) document non-trivial demand effects in incentivised games even under standard anonymity. Demand-effect bounds (asking subjects what they think the experimenter wants) are a useful diagnostic to include in pilot.
Cross-cultural framing. The implicit framing as an investment or loan may carry different connotations across settings — debt obligations, social exchange norms, religious proscriptions on interest. Ashraf, Bohnet & Piankov (2006) compare trust and trustworthiness across three countries and find substantial heterogeneity in both levels and the trust–trustworthiness relationship. Pilot the comprehension and the dominant interpretation before deployment.
One-shot vs repeated. Berg–Dickhaut–McCabe is strictly one-shot. Repeated trust games measure reputation, learning, and equilibrium selection rather than first-meeting trust. State the design explicitly to subjects: “you will play this game once.”
Analysis Guide
# p1_send: tokens Player 1 sent (0 to endowment)
# p2_return: tokens Player 2 returned (0 to k * p1_send)
# endowment, k: design parameters; session_id, pair_id, counterpart_group: design variables
import pandas as pd
import numpy as np
import statsmodels.api as sm
import statsmodels.formula.api as smf
# 1. Normalise both decisions as shares — trust is the fraction of the endowment
# sent (the construct), trustworthiness is the fraction of the multiplied amount
# returned (the construct), undefined when p1_send is zero because there is no
# investment to reciprocate
df["trust"] = df["p1_send"] / df["endowment"]
df["trustworthy"] = np.where(df["p1_send"] > 0,
df["p2_return"] / (df["k"] * df["p1_send"]),
np.nan)
# 2. Inspect the bounded distribution — look at the share of zeros (no-trust),
# the share of full senders (trust = 1), and the SD; heavy mass at the bounds
# motivates fractional logit over OLS; SD around 0.25 is the J&M (2011) anchor
print(df[["trust", "trustworthy"]].describe())
print("share zero (trust):", (df["trust"] == 0).mean(),
" share one (trust):", (df["trust"] == 1).mean())
# 3. Net gain to Player 1 conditional on sending — only meaningful on the
# subsample that invested; the unconditional mean confounds the no-trust mass
# with the trustworthiness norm
sent = df[df["p1_send"] > 0].copy()
sent["net_gain_p1"] = sent["p2_return"] - sent["p1_send"]
print((sent["net_gain_p1"] > 0).mean()) # share of investors made better off
# 4. Trust regression with Papke-Wooldridge fractional logit (Binomial family,
# logit link, no need to drop boundary 0/1 observations); session-clustered SEs
# because the session is the level at which subjects share enumerator framing,
# instructions, and common shocks (Frechette 2012)
fit_trust = smf.glm("trust ~ age + C(female) + log_hh_expenditure + C(treatment)",
data=df, family=sm.families.Binomial()).fit(
cov_type="cluster", cov_kwds={"groups": df["session_id"]})
print(fit_trust.summary())
# 5. Trustworthiness regression — same estimator, restricted to investors;
# NOTE: this conditions on an endogenous variable (p1_send > 0) so the
# coefficients describe the conditional population, not the marginal
# population effect on trustworthiness — see step 6 for the bounds
fit_tw = smf.glm("trustworthy ~ age + C(female) + log_hh_expenditure + C(treatment)",
data=sent, family=sm.families.Binomial()).fit(
cov_type="cluster", cov_kwds={"groups": sent["session_id"]})
print(fit_tw.summary())
# 6. Lee (2009) sharp bounds for trustworthiness under selection on p1_send > 0 —
# trim the larger selected group from each tail of the trustworthy distribution
# to bound the treatment effect; here we estimate bounds on E[trustworthy | T=1]
# minus E[trustworthy | T=0] when the treatment group has higher selection
def lee_bounds(d, y, t):
p1 = (d[t]==1).mean(); p0 = (d[t]==0).mean()
s1 = (d.loc[d[t]==1, "p1_send"] > 0).mean()
s0 = (d.loc[d[t]==0, "p1_send"] > 0).mean()
q = (s1 - s0) / s1 if s1 > s0 else (s0 - s1) / s0
g = d.loc[(d[t]==1) & (d["p1_send"]>0), y].dropna() if s1>s0 else d.loc[(d[t]==0) & (d["p1_send"]>0), y].dropna()
lo, hi = np.quantile(g, q), np.quantile(g, 1-q)
return (g[g <= hi].mean() - d.loc[(d[t]==0)&(d["p1_send"]>0), y].mean(),
g[g >= lo].mean() - d.loc[(d[t]==0)&(d["p1_send"]>0), y].mean())
print(lee_bounds(df, "trustworthy", "treatment"))
# 7. In-group premium — the coefficient on counterpart_group measures how much
# more Player 1 sends to an in-group counterpart relative to out-group, in
# log-odds of the trust share; back out marginal effects with .get_margeff()
fit_ig = smf.glm("trust ~ C(counterpart_group) + age + C(female)",
data=df, family=sm.families.Binomial()).fit(
cov_type="cluster", cov_kwds={"groups": df["session_id"]})
print(fit_ig.summary())
print(fit_ig.get_margeff().summary()) XLSForm / SurveyCTO
A field-deployable trust game has four moving parts: pair matching, anonymity, comprehension, and the return-elicitation mechanism. The canonical patterns:
- Pair matching via pre-randomised file. Build a roster of pairs offline (paired on whatever stratification you want — village, gender, baseline trust score) and load it as a server dataset. In the form:
pulldata("pairs", "partner_id", "id", ${respondent_id})retrieves each subject’s partner. Do not pair on-the-fly withrandom()inside the form — it is not reproducible and produces unbalanced pairs. - Counterpart-identity randomisation. If counterpart group (in-group / out-group) is randomly assigned, use the canonical SurveyCTO pattern:
once(if(random() < 0.5, 0, 1)). Theonce()wrapper persists the assignment across form recomputes; the inner conditional guarantees a fair Bernoulli draw. Store the resultingcounterpart_groupas a persisted integer field. - Multiplication. After Player 1’s
p1_send, acalculatefield computesp1_multiplied = ${p1_send} * ${k}. Display this to Player 2 before the return decision. - Strategy method via repeat group. For Player 2, define a
repeatgroup of lengthendowment / step + 1(e.g., 11 iterations for E = 100, step = 10). Each iteration asks “if Player 1 sent${step * position(..)}, you would return ___” with anintegerconstraint of[0, ${k} * ${step} * position(..)]. Persist all responses as a wide table for analysis. - Comprehension check. Before the decision, run a
select_onequiz: “If you send 30 and the multiplication factor is 3, how much does your partner receive before deciding what to return?” With arelevancerule that re-routes incorrect answers back to the instructions screen. Log the number of attempts as a quality flag. - Anonymity on the tablet. Use private entry mode for the decision fields. The enumerator should not see the screen at the moment of decision; either hand the tablet over or use a privacy screen. Do not display the partner’s ID, decision, or any link between Player 1 and Player 2 in any visible field.
- Persisted fields for analysis. Store
respondent_id,pair_id,partner_id,session_id,enumerator_id,endowment,k,counterpart_group,p1_send,p2_return(or the full strategy-method vector),comprehension_attempts, and any treatment arm. The analyst needs every one of these.
Reading the output
trustis the share of endowment sent (0–1). Johnson & Mislin (2011) meta-analytic mean is ≈ 0.50 across 162 replications; SD across studies ≈ 0.25. A sample mean below 0.30 or above 0.70 is a signal to inspect the framing and comprehension before drawing substantive conclusions.trustworthyis the share of the multiplied amount returned (0–1), defined only whenp1_send > 0. Johnson & Mislin’s meta-analytic mean is ≈ 0.37. The break-even threshold for Player 1 is 1/k (0.33 when k = 3); a sample mean below 1/k means Player 1 investors lost money on average.(net_gain_p1 > 0).mean()is the conditional share of investors made better off by investing; meta-evidence puts this in the 0.50–0.60 range. Materially below 0.50 indicates a weak reciprocity norm in your population.share zero (trust)andshare one (trust)flag the boundary mass. If either exceeds 0.20, fractional logit is the correct estimator; OLS will produce predictions outside the [0, 1] interval.- Cluster-robust SEs at the session level should be visibly larger than HC1/HC2 — by a factor of √(1 + (m − 1)ρ) for m subjects per session and ICC ρ. If they are not larger, your sessions are likely too small or ρ is genuinely near zero; report the session count and average size alongside the SE.
- The
counterpart_groupcoefficient infit_igis in log-odds because of the logit link. Use.get_margeff()(Python) ormargins::margins()(R) to convert to a marginal-effect interpretation in shares. - Lee bounds for the trustworthiness treatment effect should be reported alongside the conditional regression coefficient when treatment changed Player 1’s selection probability. Tight bounds (narrow interval, both sides same sign) support the conditional estimate; wide bounds or sign-flipping bounds mean the selection issue is binding.
References
Ashraf, N., Bohnet, I., & Piankov, N. (2006). Decomposing trust and trustworthiness. Experimental Economics, 9(3), 193–208. https://doi.org/10.1007/s10683-006-9122-4
Bellemare, C., & Kröger, S. (2007). On representative social capital. European Economic Review, 51(1), 183–202. https://doi.org/10.1016/j.euroecorev.2006.03.006
Berg, J., Dickhaut, J., & McCabe, K. (1995). Trust, reciprocity, and social history. Games and Economic Behavior, 10(1), 122–142. https://doi.org/10.1006/game.1995.1027
Brandts, J., & Charness, G. (2011). The strategy versus the direct-response method: A first survey of experimental comparisons. Experimental Economics, 14(3), 375–398. https://doi.org/10.1007/s10683-011-9272-x
Camerer, C. F. (2003). Behavioral game theory: Experiments in strategic interaction. Princeton University Press.
Cassar, A., d’Adda, G., & Grosjean, P. (2014). Institutional quality, culture, and norms of cooperation: Evidence from behavioral field experiments. Journal of Law and Economics, 57(3), 821–863. https://doi.org/10.1086/678331
Charness, G., Gneezy, U., & Halladay, B. (2016). Experimental methods: Pay one or pay all. Journal of Economic Behavior & Organization, 131, 141–150. https://doi.org/10.1016/j.jebo.2016.08.010
Cox, J. C. (2004). How to identify trust and reciprocity. Games and Economic Behavior, 46(2), 260–281. https://doi.org/10.1016/S0899-8256(03)00119-2
de Quidt, J., Haushofer, J., & Roth, C. (2018). Measuring and bounding experimenter demand. American Economic Review, 108(11), 3266–3302. https://doi.org/10.1257/aer.20171330
Eckel, C. C., & Wilson, R. K. (2004). Is trust a risky decision? Journal of Economic Behavior & Organization, 55(4), 447–465. https://doi.org/10.1016/j.jebo.2003.11.003
Fairley, K., Sanfey, A., Vyrastekova, J., & Weitzel, U. (2016). Trust and risk revisited. Journal of Economic Psychology, 57, 74–85. https://doi.org/10.1016/j.joep.2016.10.001
Falk, A., Becker, A., Dohmen, T., Enke, B., Huffman, D., & Sunde, U. (2018). Global evidence on economic preferences. Quarterly Journal of Economics, 133(4), 1645–1692. https://doi.org/10.1093/qje/qjy013
Fehr, E. (2009). On the economics and biology of trust. Journal of the European Economic Association, 7(2–3), 235–266. https://doi.org/10.1162/JEEA.2009.7.2-3.235
Fershtman, C., & Gneezy, U. (2001). Discrimination in a segmented society: An experimental approach. Quarterly Journal of Economics, 116(1), 351–377. https://doi.org/10.1162/003355301556338
Fréchette, G. R. (2012). Session-effects in the laboratory. Experimental Economics, 15(3), 485–498. https://doi.org/10.1007/s10683-011-9309-1
Glaeser, E. L., Laibson, D. I., Scheinkman, J. A., & Soutter, C. L. (2000). Measuring trust. Quarterly Journal of Economics, 115(3), 811–846. https://doi.org/10.1162/003355300554926 (NBER w7216: https://www.nber.org/papers/w7216)
Houser, D., Schunk, D., & Winter, J. (2010). Distinguishing trust from risk: An anatomy of the investment game. Journal of Economic Behavior & Organization, 74(1–2), 72–81. https://doi.org/10.1016/j.jebo.2010.01.002
Johnson, N. D., & Mislin, A. A. (2011). Trust games: A meta-analysis. Journal of Economic Psychology, 32(5), 865–889. https://doi.org/10.1016/j.joep.2011.05.007
Karlan, D. S. (2005). Using experimental economics to measure social capital and predict financial decisions. American Economic Review, 95(5), 1688–1699. https://doi.org/10.1257/000282805775014407
Karlan, D., Mobius, M., Rosenblat, T., & Szeidl, A. (2009). Trust and social collateral. Quarterly Journal of Economics, 124(3), 1307–1361. https://doi.org/10.1162/qjec.2009.124.3.1307
Lee, D. S. (2009). Training, wages, and sample selection: Estimating sharp bounds on treatment effects. Review of Economic Studies, 76(3), 1071–1102. https://doi.org/10.1111/j.1467-937X.2009.00536.x
Levitt, S. D., & List, J. A. (2007). What do laboratory experiments measuring social preferences reveal about the real world? Journal of Economic Perspectives, 21(2), 153–174. https://doi.org/10.1257/jep.21.2.153
Papke, L. E., & Wooldridge, J. M. (1996). Econometric methods for fractional response variables with an application to 401(k) plan participation rates. Journal of Applied Econometrics, 11(6), 619–632. https://doi.org/10.1002/(SICI)1099-1255(199611)11:6%3C619::AID-JAE418%3E3.0.CO;2-1