What it is
The Krupka-Weber method elicits injunctive social norms — what a community collectively believes people should do — by turning the rating task into a coordination game. Respondents are told they will be matched with another randomly selected member of their community and will earn a bonus if their rating of an action matches that person’s rating. This incentive pushes respondents away from personal opinion and toward the judgment they expect the typical community member to give. The logic is that of a pure coordination game: each respondent, wanting to match their partner, reports the rating they expect the plurality of others to choose. In the resulting equilibrium, responses converge on the community’s shared standard — not because respondents are told what others think, but because the incentive makes it rational to report it. The result is a norm index for each action that reflects shared community standards, not individual attitudes.
When to use it
Use this method when your research question involves social norms as an outcome or mediator — for example, does a sanitation programme change norms around open defecation? Does a women’s empowerment intervention shift norms around women’s mobility or work? Does a governance programme reduce the perceived acceptability of bribery?
The method is well-suited to:
- Baseline norm characterisation: understanding which behaviours are considered appropriate or inappropriate in the study context before treatment.
- Endline norm measurement: testing whether treatment shifted the community’s shared standards.
- Mechanism analysis: establishing whether norm change mediates observed behaviour change.
It is the right tool when you need to distinguish between three things that are often conflated:
| Construct | What it captures | Measurement approach |
|---|---|---|
| Social norm (injunctive) | What people think others should do | Krupka-Weber coordination game |
| Attitude | What the respondent personally believes or prefers | Standard Likert items |
| Behaviour | What people actually do | Observation, administrative data, behavioural tasks |
A high Krupka-Weber norm score for an action tells you the community thinks it is appropriate. It does not tell you that the respondent personally approves of it, and it does not tell you that people actually do it.
How it works
Step 1: Select the action set. Identify 8–15 specific actions from the domain of interest — for example, a woman working outside the home, a man refusing to pay a bribe, a household defecating in the open. Actions should vary in how normatively charged they are expected to be. Each action is described as something a named person does (“Suppose Rekha decides to…”), not as an abstract attitude statement.
Step 2: Construct the rating scale. Ask respondents to rate the social appropriateness of each action using the framing from Krupka and Weber (2013):
“How socially appropriate is this action — that is, how consistent is it with what most people in your community think is appropriate?”
Use a four-point response scale:
- Very socially inappropriate
- Somewhat socially inappropriate
- Somewhat socially appropriate
- Very socially appropriate
The phrase “what most people in your community think” is load-bearing — it directs respondents to report perceived community standards rather than personal views, which is the construct the coordination game elicits. Omitting this framing turns the elicitation into an attitude survey. This is the standard scale in the published literature. Avoid a midpoint — the four-point scale forces respondents to lean one way.
Step 3: Explain the coordination incentive. Before respondents begin rating, explain the payment rule clearly: “At the end of this survey, we will randomly select one of these actions. We will also randomly select another person from this study in your community. If your rating of that action matches their rating exactly, you will receive [amount].” Define the reference pool explicitly — respondents must understand they could be matched with any other participant in the study community, not just household members or people they know personally. The incentive must be real money. Hypothetical bonuses invalidate the elicitation.
Step 4: Collect ratings. Have respondents rate all actions individually. Collect responses in private where possible to reduce social desirability pressure — the incentive already does much of this work, but privacy helps.
Step 5: Compute the norm index. For each action, tabulate the full distribution of ratings across all respondents. The theoretical construct is the distribution itself: Krupka and Weber’s published norm scores report the proportion of respondents choosing each rating category, and the mode is the coordination focal point — the rating a strategic respondent chooses because they expect it to attract the largest share of others. In regression analysis it is common to use the mean rating (rescaled to 0–1 by subtracting 1 and dividing by 3) as a scalar norm index. Mean and mode are consistent when the distribution is unimodal — which is usually the case for clearly approved or disapproved actions. If the distribution is bimodal, the mean is uninformative; report both modes separately and treat this as evidence of genuine norm disagreement in the community.
Step 6: Compare across groups or rounds. Compare norm indices for the same actions across treatment and control arms, across sites, or across baseline and endline. A higher mean rating at endline in the treatment arm, relative to control, is evidence that the programme shifted community norms toward approving that action.
Key decisions
Vignette and scenario construction. This is the most consequential design decision. Scenarios must reference a specific named person performing a specific action, hold context constant across scenarios (only the action varies, not the setting or circumstances), be locally recognisable, and vary in their expected norm valence. If all actions are universally condemned or universally approved, the survey produces no useful variation. Pilot with a small community sample before finalising. Avoid abstract formulations like “Is it appropriate to be honest?” — these produce ceiling effects and are not anchored in local moral context.
Scale width. The 4-point scale is standard and recommended primarily for comparability with published work. A 6-point scale also reduces the probability of exact matching (the coordination payoff), which weakens the incentive at the margin — though the practical size of this effect depends on how concentrated responses turn out to be.
Coordination payoff amount. The payoff must be salient. A useful benchmark: the coordination bonus should be comparable to 30–60 minutes of local daily wage labour. Test comprehension of the payment rule in your pilot — if respondents cannot explain how they earn the bonus, revise the instructions.
Norm elicitation as outcome vs. mechanism. If norms are a primary outcome, pre-register the full action set and analysis plan before treatment assignment, and include the norm module in both baseline and endline surveys. If norms are a mechanism, the elicitation can be confined to endline.
Separating injunctive from descriptive norms. The Krupka-Weber method measures injunctive norms (what people should do). It does not measure descriptive norms (what people actually do). If your theory of change involves misperceptions about how common a behaviour is, you need a separate question format asking about prevalence (“What percentage of households in your village do X?”).
Field administration in low-literacy contexts. When respondents have low literacy or numeracy, administer the four-point scale verbally, with visual response cards showing the labels. Read the payment rule aloud and confirm comprehension before proceeding — ask respondents to explain in their own words how they would earn the bonus. The coordination logic is not difficult to grasp but respondents must understand the matching mechanism for the incentive to function as intended.
Caveats & common mistakes
Running the elicitation without real incentives. Without real money on the line, respondents have no reason to think strategically — they will simply report personal opinion. This turns the norm elicitation into an attitude survey, which could have been done more cheaply with standard Likert items. Budget for the coordination bonus from the start.
Confusing norm scores with personal approval or behaviour. A high norm index means the community collectively considers the action appropriate. It does not mean the respondent personally approves, and it does not mean people do it. These are three different constructs; conflating them in analysis is a common error.
Using abstract or poorly grounded vignettes. Scenarios must be specific enough to evoke a real social situation. Invest heavily in piloting scenario language with community members and local research staff before finalising the instrument.
Skipping the pilot on action selection. Actions that produce near-universal agreement have no analytic value. Pilot your action set with 20–30 community members and drop actions with little spread in responses. Target actions where the initial distribution spans at least two categories.
Small reference populations. The coordination game logic requires that matching is genuinely anonymous. If the study community has only 30 households, respondents may know personally who they might be matched with. The method works best when the reference community has 100+ respondents.
Not randomising action order. If a particularly salient action anchors early responses, ratings on subsequent actions can be contaminated. Randomise the order of actions across respondents where possible.
Analysis Guide
import pandas as pd
import numpy as np
import statsmodels.formula.api as smf
from scipy import stats
# norm_action1...norm_action10: respondent ratings (1-4)
# treatment: 1 = treatment arm, 0 = control
norm_vars = [f'norm_action{i}' for i in range(1, 11)]
# 1. Distribution of ratings per action — the mode is the coordination focal
# point; unimodal at 1-2 means inappropriate, at 3-4 means appropriate,
# bimodal signals genuine norm disagreement in the community
for v in norm_vars:
print(df[v].value_counts().sort_index())
# 2. Per-respondent composite norm index — rescaled to 0-1 so coefficients
# map onto category shifts on the original 4-point scale (0.10 is ~1/3 category)
df['norm_composite'] = (df[norm_vars].mean(axis=1) - 1) / 3
# 3. Compare norm distributions across arms per action — independent-sample
# t-tests on each action surface the items driving any composite-level effect
for v in norm_vars:
t, p = stats.ttest_ind(df.loc[df['treatment'] == 1, v].dropna(),
df.loc[df['treatment'] == 0, v].dropna())
print(v, t, p)
# 4. Regression of composite norm index on treatment with strata fixed effects
# and heteroskedasticity-robust SEs — main confirmatory specification
smf.ols('norm_composite ~ treatment + C(strata)', data=df).fit(cov_type='HC1').summary() Reading the output
- Each
tab norm_actionNshows the distribution of ratings (1–4) for that action. The mode is the theoretically relevant quantity — it is the coordination focal point that a strategic respondent targets. A unimodal distribution concentrated at 1–2 means the community considers the action inappropriate; concentrated at 3–4 means it is considered appropriate. - A bimodal distribution (mass split between two categories, e.g., 1–2 and 3–4) signals genuine norm disagreement within the community. The composite mean is uninformative in this case — report the full distribution and treat the action as one where norms are contested.
norm_compositeis rescaled to 0–1, where 0 = all ratings were “very socially inappropriate” and 1 = all ratings were “very socially appropriate.” Values near 0.5 can indicate either moderate norms or a bimodal split — always check the action-level distributions first.- A positive treatment coefficient in the regression means the treatment arm assigned higher norm scores to the action set on average. A coefficient of 0.10 on a 0–1 scale represents roughly a third of a category shift on the original 4-point scale.
References
Bicchieri, C. (2006). The grammar of society: The nature and dynamics of social norms. Cambridge University Press.
Bursztyn, L., González, A. L., & Yanagizawa-Drott, D. (2020). Misperceived social norms: Women working outside the home in Saudi Arabia. American Economic Review, 110(10), 2997–3029. https://doi.org/10.1257/aer.20180975
Krupka, E. L., & Weber, R. A. (2013). Identifying social norms using coordination games: Why does dictator game sharing vary? Journal of the European Economic Association, 11(3), 495–524. https://doi.org/10.1111/jeea.12006