Status. This version draws only on already-public material: two commissioned public evaluations, the public workshop record, and the public post-workshop survey. A structured belief elicitation with participants is under respondent review; its numbers are deliberately excluded here and will be added in v2 (see What is still coming).
1. The decision this is about
A funder comparing a psychotherapy or mental-health programme against a bednet, a vaccine, or a cash transfer has to put both on one scale. In practice that means converting between DALYs averted and WELLBYs — one point on a 0–10 life-satisfaction scale, for one year — gained.
Almost every such conversion in current use passes through one step: treating one standard deviation of improvement on a depression or anxiety instrument as one standard deviation of improvement in life satisfaction. Founders Pledge's moral weights use it. Happier Lives Institute's cost-effectiveness work generates it. Anyone reading either inherits it.
The question funders ask us is not "is subjective wellbeing valid." It is: can I keep using this step, and how wrong might it make my ranking?
2. Bottom line
- The SD-to-SD step is defensible as a working assumption, and it is nobody's preferred method — including its originators'. Keep using it; report it as an assumption; run sensitivity analysis. Do not describe it as established.
- Its support is thinner than a near-1:1 point estimate implies, because conversion is a ratio operation. Scale-use correction leaves individual coefficients roughly intact but moves ratios of coefficients substantially — and ratios are exactly what a funder comparing two interventions is using.
- The likely error is directional, not symmetric. Compression and ceiling effects both push toward overstating the wellbeing gains from mental-health interventions relative to physical ones. A conversion factor that is wrong in a known direction should not be treated as noise.
- DALY disability weights and observed life-satisfaction losses do not line up across conditions. Disability weights top out around 0.3, while depression and anxiety show losses above a full life-satisfaction point. Any single global conversion factor therefore systematically mis-weights domains — the case for a domain-specific factor is stronger than the case for a better global one.
- The cheap fixes are known and mostly unused. Two vignette calibration questions is the minimum implementation of scale-use correction. Adding stated-preference tradeoff items to trials you fund would let you estimate the exchange rate rather than assume it.
3. What this is based on
| Component | Status |
|---|---|
| Commissioned public evaluations of Benjamin, Cooper, Heffetz, Kimball & Zhou, "Adjusting for Scale-Use Heterogeneity in Self-Reported Well-Being" (NBER WP 31728) — evaluators Caspar Kaiser and Alberto Prati | Public, with DOI: unjournal.pubpub.org/pub/evalsumheterogenity |
| Framing analyses: linear WELLBY reliability; DALY↔WELLBY conversion | Public: linear WELLBY, conversion |
| Workshop, 16 March 2026 — paper authors (Benjamin, Heffetz, Kimball), Happier Lives Institute (Plant, McGuire, Dupret), Kaiser, Julian Jamison, Dean Jamison, and practitioners from Founders Pledge (Lerner) and Coefficient Giving (Hickman) | Public transcript and video |
| Post-workshop participant survey | Public: survey results |
| Structured belief elicitation, 8 respondents (linear-WELLBY reliability, WELLBYs per DALY, adoption forecasts) | Under respondent review — not reported here |
Founders Pledge raised these questions with us through our Pivotal Questions programme; the work was scoped around a choice a funder was actually facing.
Quotations below are from the public, lightly edited workshop transcript (Zoom recording and automated captions, edited for readability). Check verbatim wording against the full transcript or recording before relying on it.
4. Why nobody defends the SD-to-SD step, including the people who use it
Founders Pledge arrived at its conversion by triangulation: WELLBY to income-doubling (from Joel McGuire's work), income-doubling to lives-saved, lives-saved to DALYs — then backing out the remaining side. Matt Lerner's own description:
a bit kludgy — we had two sides of a triangle and filled in the third for interconvertibility.
The reason for the SD step is data, not theory. Samuel Dupret (HLI) explained that HLI lacks rich 0–10 scale datasets for LMIC interventions, so it converts trial results to standard deviations, then to a 0–10 scale using typical Cantril Ladder SDs from the World Happiness Report — "This isn't because we love SD conversions — it's data constraints." Michael Plant: "We're being data omnivores." Lerner, agreeing: "If we had life satisfaction data for everything, we'd skip the SD conversion."
That consensus matters for how a funder should treat the number. It is a placeholder that survived, not a result.
5. Three problems that do not go away with a better point estimate
SDs inherit the population you measured them in
A standard deviation is a property of a sample, not of a person. Scaling a depression-trial effect from a rural LMIC sample by country-level Cantril SDs from the World Happiness Report assumes a transportability that has not been established. Mechanically, comparisons will tend to favour interventions run in more heterogeneous populations.
The instruments overlap without nesting
Depression and anxiety instruments load on negative affect and functioning; life satisfaction is a global cognitive evaluation. These are correlated constructs measuring different things, so a one-to-one mapping of changes is an additional assumption on top of the SD assumption.
Ratios move even when coefficients do not
This is the sharpest point in the evaluation package, and it comes from the material a funder is least likely to read. Caspar Kaiser's evaluation notes two readings of the Benjamin et al. results: individual coefficients are fairly robust to scale-use correction, but ratios of coefficients are not. Cross-intervention comparison is a ratio operation. Miles Kimball's related point: scale-use correction moves the income coefficient by roughly a factor of five in their unemployment conversion — first-order for anything routed through income-doubling equivalences.
And we know very little about scale use in intervention settings specifically. Kaiser's assessment is that the evidence here is close to absent, rich or poor country; the literature is dominated by population surveys.
6. What we recommend
Do now, at no cost
- State the SD conversion as an assumption wherever a comparison depends on it, and publish sensitivity bounds rather than a point estimate.
- Use directly measured life satisfaction wherever a trial reports it, in preference to a converted mental-health measure.
- Flag explicitly when a conclusion depends on the location of the neutral point.
- Do not compare a mental-health intervention against a physical-health intervention using a single global factor without saying so.
Do if you have budget or control an instrument
- Add two vignette calibration questions to any instrument you fund. This is the minimum viable implementation of scale-use correction; visual calibration gives partial correction where vignettes are infeasible.
- Collect calibration before and after treatment, and in controls, for any psychological intervention — otherwise you cannot separate a genuine welfare change from treatment-induced change in how people use the scale.
- Add stated-preference tradeoff items so the exchange rate can be estimated rather than assumed. Dan Benjamin's view in follow-up correspondence: absent extra data, comparing SD improvements is hard to beat, but a small number of added questions would let you estimate the rate directly.
Do not bother yet
- Full multi-dimensional personal wellbeing indices in LMIC field settings, absent dedicated funding.
- Chasing a single better global DALY↔WELLBY factor. The evidence points toward domain-specific factors, so precision on a global one buys little.
7. What would change this answer
Ranked by our assessment of value of information:
- Calibration plus response-shift diagnostics inside a mental-health intervention trial — separating welfare change from therapy-induced change in scale use. Raised independently by Kaiser and Kimball; appears to be a genuine literature gap.
- Calibration and value elicitation in a globally low-income population.
- Crosswalk harvest. Systematically collect trials that report both a depression or anxiety instrument and life satisfaction, and extract the distribution of the implied ratio rather than a point estimate. This is the tractable project that would answer the question asked here better than any single number, and it extends rather than duplicates HLI's conversion work.
- Direct validation of the linear WELLBY against better-specified alternatives — per Benjamin, not yet done.
We are looking for one funder or research group to take on (3). It is small, it uses existing published data, and it converts an assumption into a measured distribution.
8. What is still coming
Elicitation results (v2). Eight participants gave central estimates with 80% credible intervals on linear-WELLBY reliability, WELLBYs per DALY, and adoption forecasts. Respondents are reviewing a corrected analysis; we will publish with the correction round closed. Preview, without numbers: broad support for WELLBY-style data as useful, considerably less for a single linear life-satisfaction score as the endpoint, and genuine unresolved spread on the conversion factor. The spread is the finding, and a brief that reported a tidy central value would be misrepresenting the state of knowledge.
A documented elicitation limitation. Our form displayed submittable slider defaults, so some responses may be anchored or untouched. We diagnosed this after the fact, corrected the design for subsequent workshops, and will report sensitivity excluding exact-default responses. We would rather publish this than quietly drop it.
9. How to use and cite this
Written for funders and CEA practitioners who must choose between interventions measured in different units, now, without waiting for the literature to settle. It is not a verdict on whether wellbeing measures are valid.
Comments and disagreement are welcome, including via Hypothes.is annotation on this page. If you are making a grant decision that turns on the conversion factor and want to talk it through, write to contact@unjournal.org.