What it is
Participatory Wealth Ranking (PWR) classifies households by relative economic status using knowledge held within the community rather than externally imposed asset indices or consumption measures. Trained community informants — typically three to five people who know the village well — are given a set of cards, each representing a household, and asked to sort them into piles according to wealth. The researcher records the sorting, elicits the criteria informants used, and constructs a ranking or categorical classification from the results. The method was first codified for development practice by Grandin (1988) and is one of the most widely used participatory targeting tools.
The core advantage is that community members observe indicators of economic status that are invisible to survey enumerators: livestock kept outside the village, seasonal migration, informal debts, off-farm income, and the quality of construction materials or food consumption that only close neighbours can assess (Adams, Evans, Mohammed & Farnsworth, 1997; Van Campenhout, 2007). PWR converts these locally observable dimensions into a usable within-community classification.
When to use it
PWR is most valuable in three settings. It supplements or replaces standard consumption and asset indices where those are unavailable or impractical. It targets the poorest households for a programme when administrative welfare data are absent or unreliable — the dominant field application, used at scale in the BRAC and CGAP ultra-poor graduation programmes (Banerjee et al., 2015; Hashemi & de Montesquiou, 2011) and tested experimentally against proxy-means tests by Alatas, Banerjee, Hanna, Olken & Tobias (2012) and Karlan & Thuysbaert (2019). And it captures a community-perception dimension of welfare that survey-based measures cannot — informative in its own right and useful as a complement to consumption rankings.
Chambers (1994) introduced participatory rural appraisal as the broader methodological frame within which PWR sits. Hargreaves et al. (2007) and Van Campenhout (2007) validated PWR ranks against consumption-based welfare measures and showed moderate-to-strong rank correlations across multiple field contexts. Vyas & Kumaranayake (2006) and McKenzie (2005) discuss asset-index alternatives to which PWR is most often compared.
PWR is less appropriate for cross-village or cross-regional comparisons — wealth categories are local and not calibrated to an external reference point. It is also unsuitable for programmes that require legally verifiable targeting criteria.
How it works
The standard implementation has four steps:
Step 1 — Household enumeration. Compile a complete list of households in the community from a recent census or rapid listing. Assign each household a unique card with the household head’s name, a stable household ID, and where useful a small photograph or sketch map.
Step 2 — Independent sorting. Recruit three to five informants who know the village well — selected for vantage variation (gender, age, hamlet, social position), not visibility. Ask each informant to sort the household cards into piles based on relative wealth using their own definition. Independent sorting first, joint reconciliation second, is the canonical sequence (Alatas et al., 2012; Hargreaves et al., 2007) — it preserves the inter-rater reliability check that joint sorting destroys.
Step 3 — Criteria elicitation. After sorting, ask informants what they were thinking when they placed households in each pile. Common criteria include housing quality, land and livestock holdings, schooling of children, seasonal food insecurity, and visible consumption. These criteria are the qualitative complement to the ranking.
Step 4 — Reliability check, then reconciliation. Compute an inter-rater reliability statistic across the independent rankings before averaging (see Analysis Guide). If reliability is adequate, average to a composite. If reliability is poor, return to the informants with the substantively contested households for discussion and re-ranking; do not average through disagreement.
Identification assumptions. A community ranking recovers the latent welfare ordering only under:
- Monotone observation. Each informant’s pile assignment is a monotone transformation of latent welfare plus mean-zero error. Without monotonicity the average pile across informants does not converge to the true rank.
- Independent errors across informants. Joint sorting violates this; one dominant informant drives the consensus and the rest contribute no independent information.
- No systematic social bias. Informants do not consistently up- or down-rank households based on caste, kinship, religion, or political alignment. Diverse informants mitigate but do not eliminate this.
- Comparable pile counts within an informant. Each informant uses the same pile structure across their entire ranking; the number of piles can vary across informants if rankings are normalised within informant before pooling (see code).
- Validation against an external welfare measure where feasible. Without validation, the ranking is interpretable as community perception but not as a welfare-error-free classification (Alatas et al., 2012; Hargreaves et al., 2007).
Key decisions
Number and selection of informants. Three to five is the field standard, justified by the inter-rater reliability literature (Koo & Li, 2016): with k informants, Kendall’s W or ICC(2,k) is computable; below k = 3 only pairwise Spearman is available and the joint reliability claim cannot be made. Recruit for vantage variation, not for community prominence — community-health workers, women’s-group leaders, long-standing residents of different hamlets. Avoid relying solely on the village head or the most politically connected residents.
Joint vs. independent sorting. Independent first, reconciled second, is the canonical sequence. Joint sorting produces a consensus faster but allows dominant informants to drive the result and forecloses the reliability check.
Number of wealth piles. Three to five piles emerge spontaneously in most settings. Forcing a specific number — say four, to match a programme targeting rule — is acceptable but should be flagged as researcher-imposed. Different informants may pick different pile counts; the analysis must normalise within informant before pooling (rank-based normalisation; see code) or the composite is meaningless.
Anchoring across communities. Each PWR is calibrated to a single community. Anchor households (a “clearly poorest” and “clearly richest” household per category) help with within-village reconciliation but do not solve the across-village comparability problem. For multi-village studies, the standard fix is to collect a common external welfare measure (asset index or short consumption recall) on the full sample and calibrate PWR ranks against it village-by-village before pooling.
Validation against a survey-based welfare measure. Default to yes, not optional. The targeting literature (Alatas et al., 2012; Karlan & Thuysbaert, 2019; Adams et al., 1997; Van Campenhout, 2007) treats validation as essential because PWR performance varies substantially across contexts. Collect an asset index, short consumption recall, or official poverty register category alongside PWR for a subsample. Use the convergent-validity correlation as a sanity check, not as proof that PWR is “right” or “wrong”; the targeting-error decomposition below is the more interpretable diagnostic.
Test-retest reliability. A separate dimension from inter-rater agreement: do the same informants reproduce their own rankings weeks later? Hargreaves et al. (2007) tested this in South African villages and found high stability for the bottom and top piles, lower for the middle. If the programme cares about the borderline category, plan a test-retest exercise on a subsample.
Sample size and power. For the convergent-validity test (Spearman or Kendall correlation between PWR composite and a survey welfare measure), the validation subsample power calculation is standard: n ≈ 60 detects ρ = 0.40 with 80% power at α = 0.05; n ≈ 30 detects ρ = 0.55. For the targeting accuracy decomposition (sensitivity / specificity), power is set by the smaller of the two cells in the 2×2 — n ≈ 200 households per village gives stable sensitivity estimates if the survey-poor share is ~20%.
Caveats & common mistakes
Elite capture and social bias. Informants embedded in social hierarchies may up-rank politically connected households and down-rank socially marginalised ones. Vantage diversity is the partial mitigation; explicit comparison of rankings across informant subgroups is the diagnostic.
Inclusion–exclusion error trade-off. Alatas et al. (2012) found that community targeting (close cousin of PWR) and proxy-means tests both produced inclusion errors of 30–40% in Indonesia — community methods better identified the community-perceived poor, PMT better identified the consumption-defined poor. PWR shifts the error composition, it does not eliminate error. Choose PWR over PMT when community perception is the relevant welfare construct or when PMT data are unavailable, not because PWR is more accurate against a survey-defined benchmark.
Confidentiality. PWR sessions name and discuss specific households in front of informants. In small communities, the identities of households placed in the “poorest” pile may become known. Ask informants explicitly not to discuss rankings outside the session, hold sessions in private locations, and store cards by household ID rather than name in the analysis dataset.
Non-comparability across communities. A household in the top pile in a poor village may be objectively poorer than the bottom-pile household in a wealthy village. PWR does not produce cross-community comparable rankings without an external welfare anchor.
Informant fatigue. Sorting 200 cards is cognitively demanding and reliability degrades through the session (Hargreaves et al., 2007). Split large communities into sub-areas of ≤100 households per session and use sub-area-specific informants; include at least one household in two sub-areas to allow a cross-section calibration.
Conflation of wealth with social desirability. Informants may hesitate to rank households as poor if the household head is present or if poverty carries strong stigma. Private session locations and framing the exercise in terms of economic resources rather than worthiness reduce this.
Analysis Guide
# 1. Normalise each informant's piles to within-informant fractional ranks BEFORE
# pooling — informants may have chosen different numbers of piles in Step 2,
# so averaging raw pile numbers across informants is not meaningful. Fractional
# ranks (1/N_i ... N_i/N_i, where N_i is informant i's pile count) place every
# informant's ranking on the same 0-1 scale.
# Input: long-format df with columns (household_id, informant_id, pile, village_id)
import pandas as pd
import numpy as np
import statsmodels.formula.api as smf
from scipy import stats
from pingouin import intraclass_corr
def fractional_rank(s):
return s.rank(method='average') / s.max() # 1/N ... N/N within informant
df['frank'] = df.groupby('informant_id')['pile'].transform(fractional_rank)
# 2. Compute inter-rater reliability across all informants BEFORE averaging.
# With k informants you want ICC(2,k) (two-way random-effects, average-rater
# agreement; Koo & Li 2016) — pairwise Spearman cannot summarise k > 2 raters.
# pingouin returns a tidy data frame; ICC2k is the row labelled "ICC2k".
wide = df.pivot(index='household_id', columns='informant_id', values='frank')
icc = intraclass_corr(
data=df, targets='household_id', raters='informant_id', ratings='frank')
print(icc) # ICC2k > 0.75 = good; 0.5-0.75 = moderate; < 0.5 = poor (Koo & Li 2016)
# 3. Average the within-informant fractional ranks into the composite. Require
# complete cases (every informant ranked the household) or denominator mixing
# silently averages over different numbers of informants for different
# households — the recurring failure mode in this library.
n_informants = df['informant_id'].nunique()
complete = wide.dropna(axis=0) # households ranked by all k informants
complete['pwr_composite'] = complete.mean(axis=1)
incomplete_n = len(wide) - len(complete)
print(f'Dropped {incomplete_n} households without complete informant rankings')
# 4. Convergent validity vs a survey-based welfare measure. Report the point
# estimate AND a bootstrap 95% CI — Spearman's rho on small validation samples
# has wide CIs and the point estimate alone is misleading.
merged = complete.reset_index().merge(survey, on='household_id')
rho, p = stats.spearmanr(merged['pwr_composite'], merged['asset_quintile'])
def boot_rho(x, y, B=2000, seed=20260606):
rng = np.random.default_rng(seed)
n = len(x); rhos = np.empty(B)
for b in range(B):
idx = rng.integers(0, n, n)
rhos[b] = stats.spearmanr(x.iloc[idx], y.iloc[idx])[0]
return np.quantile(rhos, [0.025, 0.975])
ci = boot_rho(merged['pwr_composite'], merged['asset_quintile'])
print(f'Spearman rho = {rho:.3f}, 95% CI [{ci[0]:.3f}, {ci[1]:.3f}], p = {p:.3f}')
# 5. Targeting threshold from the empirical distribution, NOT a hard-coded pile
# number. If the programme serves the bottom 20%, set the cutoff at the 0.20
# quantile of the composite. The 1.5 hard-coded threshold in earlier drafts
# of this guide was only valid for a specific 2-informant / 4-pile design.
target_share = 0.20 # programme serves bottom 20%
threshold = merged['pwr_composite'].quantile(target_share)
merged['pwr_poor'] = (merged['pwr_composite'] <= threshold).astype(int)
merged['survey_poor'] = (merged['asset_quintile'] <= 1).astype(int)
# 6. Targeting accuracy decomposition — sensitivity (true positive rate) and
# specificity (true negative rate) against the survey benchmark. Report the
# 2x2 table, both rates, and Cohen's kappa as the chance-corrected agreement.
tab = pd.crosstab(merged['pwr_poor'], merged['survey_poor'])
sens = tab.loc[1, 1] / tab[1].sum() # share of survey-poor caught by PWR
spec = tab.loc[0, 0] / tab[0].sum() # share of survey-non-poor excluded
print(tab); print(f'Sensitivity {sens:.2f} Specificity {spec:.2f}')
# 7. Does PWR carry targeting-relevant information BEYOND the survey measure?
# Regress programme enrolment on both; a positive coefficient on pwr_poor net
# of survey_poor means PWR captures information the survey misses. Cluster SEs
# at village level — within-village respondents are not independent.
fit = smf.ols('enrolled ~ pwr_poor + survey_poor + age + C(female)',
data=merged).fit(
cov_type='cluster', cov_kwds={'groups': merged['village_id']})
print(fit.summary()) SurveyCTO / XLSForm
PWR is primarily a face-to-face exercise; SurveyCTO is the digital recording layer rather than the elicitation tool. The recommended pattern is a session-management form with a repeat group over households and a parallel sub-form per informant:
| Item | Type | Notes |
|---|---|---|
village_id, session_date | calculate / date | Header for the session |
informant_id | text | Stable ID — same informant may appear in multiple sessions if sub-areas are split |
| household-repeat group | begin_repeat over roster.csv | One iteration per household card |
pile | select_one piles | Constrain pile ∈ [1, max_piles]; record max_piles per informant in the session header |
| inconsistency flag | calculate | Compare against prior informants in the session: count-selected(coalesce(...)) or post-session reconciliation script flags households with pile-variance above threshold for facilitator review |
criteria_text | text per pile | Open text where informants describe each pile’s criteria |
Pre-load roster.csv with stable census-time IDs (household_id, head_name, hamlet). Always export household_id and the per-informant pile, never the reconciled composite alone — the disaggregated data is needed for the inter-rater reliability check, and a composite-only export is irrecoverable.
For paper-based sessions (still the practical default in many field contexts), use a tally sheet with the same column structure and data-enter into the form afterwards. Match household IDs to the main survey dataset before analysis.
Reading the output
- ICC(2,k) (produced by
irr::icc(mat, model='twoway', type='agreement', unit='average')in R orpingouin.intraclass_corrin Python). Interpret on the Koo & Li (2016) scale: > 0.90 excellent; 0.75–0.90 good; 0.50–0.75 moderate; < 0.50 poor. Below 0.50 the ranking is uninformative — return to reconciliation. - Kendall’s W (produced by
irr::kendall(mat)in R). Landis & Koch (1977) interpretation: > 0.81 almost perfect; 0.61–0.80 substantial; 0.41–0.60 moderate; below 0.40 fair-to-poor. - Convergent-validity Spearman rho (
DescTools::SpearmanRho(..., conf.level = 0.95)in R or the bootstrap in Python) with its 95% CI. Field benchmarks: rho 0.40–0.60 is typical when PWR is compared with asset indices (Hargreaves et al., 2007; Van Campenhout, 2007). A low point estimate is not a failure if the CI is wide — that just means the validation sample is small. - Sensitivity (true-positive rate) and specificity (true-negative rate) against the survey benchmark, produced from the
pd.crosstab/table()2×2. Alatas et al. (2012) and Karlan & Thuysbaert (2019) report inclusion-error rates of 30–40% for community targeting under typical conditions — both rates landing in the 60–70% range is the realistic frontier, not a failure threshold. Below 50% on either, check the household enumeration and the informant selection before concluding the method failed. - Cluster-robust regression (
lm_robust(..., clusters = village_id, se_type = "CR2")in R;cov_type='cluster'in Python). A significant coefficient onpwr_poornet ofsurvey_pooris evidence that PWR contains targeting-relevant information the survey index misses. A null result is consistent with either redundancy or noise; the convergent-validity CI distinguishes the two.
References
Adams, A. M., Evans, T. G., Mohammed, R., & Farnsworth, J. (1997). Socioeconomic stratification by wealth ranking: Is it valid? World Development, 25(7), 1165–1172. https://doi.org/10.1016/S0305-750X(97)00024-7
Alatas, V., Banerjee, A., Hanna, R., Olken, B. A., & Tobias, J. (2012). Targeting the poor: Evidence from a field experiment in Indonesia. American Economic Review, 102(4), 1206–1240. https://doi.org/10.1257/aer.102.4.1206
Banerjee, A., Duflo, E., Goldberg, N., Karlan, D., Osei, R., Parienté, W., Shapiro, J., Thuysbaert, B., & Udry, C. (2015). A multifaceted program causes lasting progress for the very poor: Evidence from six countries. Science, 348(6236), 1260799. https://doi.org/10.1126/science.1260799
Chambers, R. (1994). The origins and practice of participatory rural appraisal. World Development, 22(7), 953–969. https://doi.org/10.1016/0305-750X(94)90141-4
Grandin, B. E. (1988). Wealth ranking in smallholder communities: A field manual. Intermediate Technology Publications.
Hargreaves, J. R., Morison, L. A., Gear, J. S. S., Makhubele, M. B., Porter, J. D. H., Busza, J., Watts, C., Kim, J. C., & Pronyk, P. M. (2007). “Hearing the voices of the poor”: Assigning poverty lines on the basis of local perceptions of poverty. A quantitative analysis of qualitative data from participatory wealth ranking in rural South Africa. World Development, 35(2), 212–229. https://doi.org/10.1016/j.worlddev.2006.10.011
Hashemi, S. M., & de Montesquiou, A. (2011). Reaching the poorest: Lessons from the Graduation Model (CGAP Focus Note 69). CGAP. https://www.cgap.org/research/publication/reaching-poorest-lessons-graduation-model
Howe, L. D., Hargreaves, J. R., Gabrysch, S., & Huttly, S. R. A. (2009). Is the wealth index a proxy for consumption expenditure? A systematic review. Journal of Epidemiology and Community Health, 63(11), 871–877. https://doi.org/10.1136/jech.2009.088021
Karlan, D., & Thuysbaert, B. (2019). Targeting ultra-poor households in Honduras and Peru. World Bank Economic Review, 33(1), 63–94. https://doi.org/10.1093/wber/lhw036
Koo, T. K., & Li, M. Y. (2016). A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15(2), 155–163. https://doi.org/10.1016/j.jcm.2016.02.012
Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174. https://doi.org/10.2307/2529310
McKenzie, D. J. (2005). Measuring inequality with asset indicators. Journal of Population Economics, 18(2), 229–260. https://doi.org/10.1007/s00148-005-0224-7
Van Campenhout, B. F. H. (2007). Locally adapted poverty indicators derived from participatory wealth rankings: A case of four villages in rural Tanzania. Journal of African Economies, 16(3), 406–438. https://doi.org/10.1093/jae/ejl039
Vyas, S., & Kumaranayake, L. (2006). Constructing socio-economic status indices: How to use principal components analysis. Health Policy and Planning, 21(6), 459–468. https://doi.org/10.1093/heapol/czl029