# Wellbeing Workshop Transcript

**Date:** March 16, 2026

**Duration:** ~4 hours 32 minutes

**Workshop:** The Unjournal Pivotal Questions — Wellbeing Measurement Workshop

---

## Speakers

- Caspar Kaiser
- Daniel Benjamin
- Julian Jamison
- Joel McGuire
- Matt Lerner
- Michael Plant
- Miles Kimball
- Ori Heffetz
- Peter Hickman (Coefficient Giving)
- Samuel Dupret
- David Reinstein (The Unjournal)
- Valentin Klotzbücher

---

## Transcript

### [00:00:05] David Reinstein (The Unjournal)

If I'm not there. Anyways, thank you so much for coming, it's really great to meet you all. And I'm really just gonna try to get right into it, because we're trying to keep to a strict schedule, so that people can jump into the sessions they're interested in, and obviously, I value your time, etc. So let me, let me get… get started here.

Okay, the AI companion is on if anyone needs to catch up on anything, but let me, yeah, let me get right into it. Okay, there are two breakout rooms you can join… oh, that's… let's not talk about that right now. Sorry, just doing a little setup. Does anyone know how to make, a second person a, co-manager?

I think I know how to do that, actually. Here we are. Okay, Valentin, you're now a co-host, so you can help out, thanks so much.

And yeah, if anyone's waiting to join… oh, just gotta let them in. Please, just to admit them all. Great, great to see you all, and once again, thanks for joining.

We're coming from all over the world. The time zones are a bit tricky. We have people from, I think, California to India, perhaps.

And… I wish there was a way to auto-admit people, I have never figured that out. Alright, let me just go ahead and share my screen. Sorry for the…

### [00:01:27] Valentin Klotzbücher

I'm admitting them, don't worry.

### [00:01:28] David Reinstein (The Unjournal)

Oh, thank you so much. Alright. Actually, my screen is shared.

Alright, well, let me get right into it. We're already 2 minutes behind. My fault, of course. Okay. I will give you a little bit of an introduction here, and I'll share some slides, but these are just You know, just to focus the discussion over here.

Everyone should be seeing my slides, but the same content is in the Google Doc. Okay, so… here we are. About me, not so important, I'm an economist, I'm the founder of The Own Journal. I'm not an expert in this area, so please don't let me dominate the discussion, although it's something I've been thinking about quite a bit. You can see links on what the own journal is, I won't go into detail right now.

I assume many of you are familiar with it. Here's the goals of the workshop. I want to bring… together researchers and practitioners in what I think's an unusual, particularly open, focused, and productive way. People who are funders, people who are doing technical research, etc. I want to sort of bring our knowledge base together.

or bring our needs together, help people understand what each other's highest value questions are, and bring practitioners up to date on what the research is and how we should interpret it. And you can see, Michael Plant had some sidebar comments on that that I think are particularly apropos. foster communication. what is it that we need to resolve between ourselves and can resolve, and what do we agree on? You know, let's not… what's the point of spending a lot of time, michael Plant just mentioned, what's the point of spending a lot of time arguing… talking about things we already agree on, or we agree don't have a high value?

Okay, there's a belief elicitation session. You know, ideally, and this is ambitious, we want to state and measure our beliefs openly within the group. Now and through the asynchronous follow-up. driving, hopefully, high-value Bayesian updating, and, you know, ultimately, better… better choices. and better decisions.

You know, obviously, that's our goal, and I think we're very well aligned here to try to actually drive, you know, better choices, particularly over interventions and funding in Low-income countries. Now let's talk about the focus. Because it's a bit different than the focus of, let's say, other well-being workshops you may have been to. So… The focus here is what are the best ways to measure and compare the relative benefits and cost-effectiveness of different interventions, particularly thinking about well-being, interventions involving different outcomes. How do the approaches we might use compare?

What are some practical guidance, and what are our next steps? Focusing on things like bed nets versus short cognitive behavioral therapy sessions in lower, middle-income… or low-income countries in… in… often in Africa, Sub-Saharan Africa, cash transfers, etc. Now, one, also, difference is… we're not trying to set a firm scientific precedent, necessarily. Policymakers need to make choices now and want to know what the best… their best options are, and want to have, you know, the most informed information about that. So we don't want to say, you know, we just don't know.

I mean, yes, we want to express how much uncertainty we have. And we're also not focusing on questions you might have seen in a lot of other well-being work, the sort of very interesting and perhaps deep questions about human behavior, as well as sort of comparisons. Now, here I'm getting a bit controversial. I'm hoping, and please jot this down in the doc… your comments down in the Google Doc. I'm hoping that we can agree on some Premises, or at least maintain them for now.

Well-being is an important goal. To the… but we might disagree about how measurable it is. Self-reports carry some information, they're probably not totally uninformative, therefore finding evidence that they correlate to some things is perhaps not enough on their own to recommend them.

But on the other side of the coin, you know. We probably can agree that these extreme Representative agent assumptions will be violated in some way, so showing a violation will not be highly informative to us or to policymakers. Now, I don't want to go into too much detail here, but, you know, there are… so what we're worried about is what measurement approaches lead to better decisions, and how much the violations matter.

And now here, you know, I don't have time to get into this, but we'll cover this throughout the workshop. You know, there are these different potential violations. I'm particularly worried about nonlinear scale use, and I'd like to see some links to the sort of gold standards That I haven't seen a lot, but heterogeneous scale use we'll also talk about. I guess what we could agree on here is that at least some of these are potentially important. Even though they may not be empirically important, but we also want to shift out which of these are, particularly important to the case that we're discussing here.

You know, some of them might be actually less important, like scale shifters might not affect comparisons between interventions. in LMICs in it using RCTs. Okay, so, just a thought, you know, type in the Google Doc, but we're going to move smoothly on to the next session. Let's try to work asynchronously.

Alright, how is this session working? We have a Google Doc. I'm trying to make it organized. Naturally, it's hard to make it organized, so I'm gonna bring some organization, hopefully, after the session as well. you know, synthesizing what we've written in that Google Doc.

You can put comments on the side. moving along Google Doc makes it hard to ask… you can take a look at it and see. It makes it hard to ask people… for people to add content if it's constantly moving, so be a little cognizant of that.

We put the questions for the speakers, etc, in sub-tabs below each section, and it's not meant to be just for this workshop. Hopefully, it could go on after the workshop, this conversation through this, and other asynchronous things we'll mention later. Trying to keep a strict timing, I'm already falling a little behind, so I'm gonna jump. Keep the Zoom chat mainly for procedural things, try to keep the discussion in the Google Doc if that's feasible. You can also go into breakout rooms with participants if you want to continue conversations.

I've set two up, which you should be able to access with the More button. And this is rather important. By default, this content might be shared, so if you're concerned about this being shared, you know, even outside this room at some point. please do let me know, but I will let you know before I do any sort of broad dissemination. Here's some discussion.

I think we can try to maintain good epistemic norms throughout this, and good focus, you know, not, trying to win the argument, but trying to achieve value here. And, yeah, you know, consider things like, what evidence could change my mind? Alright, well, let's, begin, and, thanks to everyone for coming. I will, pass the floor to our practitioner presenters, I believe, starting with, Peter Hickman. Thanks so much for joining.

Feel free to share your screen if you wish.

### [00:09:00] Peter Hickman (Coefficient Giving)

Hey, yeah, thanks for having me. Let me… let me share. And… Great.

And we go to… View as slideshow. Does that look good for everyone? Great. So yeah, I'm Peter, and I'm very happy to be joining this workshop and excited to learn from… I'm from Coefficient Giving, which, if you haven't heard, is the new name of Open Philanthropy.

And I wanted to set the table a little bit by talking about our current framework for valuing outcomes, and say a little bit about how self-reported well-being could fit into that. Give away the punchline, we don't currently have a value for a unit of well-being that is kind of canonical and we're using in our cost-effectiveness evaluations. But we're open to admitting that that's a blind spot and, you know, figuring out what is the best way for us to incorporate well-being going forward. to first just make sure everyone knows, Coefficient Giving is a philanthropic funder and advisor.

We have traditionally kind of been the outsourced Foundation staff for the Foundation Good Ventures, but we're increasingly working with other donors. And we, directed as much money as we have ever directed last year, reaching the $1 billion mark. Our mission is to help others as much as we can with the resources available to us.

And so, kind of part of the DNA is not picking certain cause areas initially, but being open to whatever helps others the most. And so, of course, making people Happier or more satisfied with their lives is potentially one way of doing that. Our framework is to focus on clauses, which are important, so affecting a lot of people, affecting them a lot, neglected, not just being covered by everyone else, intractable. So we can make a difference.

And potentially, improving mental health, improving well-being could fit into that. But about me, I'm an economist by training, finished my PhD in 2024. I did some work on development, econ with field work, and a little bit of behavioral as well. dabbled a little in understanding self-reported well-being and enjoyed reading some of the work of some of the people who are going to present today, but didn't get deep into that and am not an expert.

At CG, I'm on the cause prioritization team. which is sort of our research team, and I'm a research fellow, so I'm not making the grants for the most part, but kind of working more upstream on deciding with a team of people. What are we going to value and how? Trying to balance, kind of. being scientifically correct, with also, like, being something that we can actually use for the dozens and dozens of grants that we're making across many different areas.

So, yeah, putting numbers on very different outcomes, like lead exposure. Economic growth, housing. And just FYI, I don't work on catastrophic risk stuff, like AI and biosecurity, or on farm animal welfare, so don't ask me, questions about that.

But our current framework for valuing outcomes, okay. So our unit of impact's the coefficient giving dollar, which is defined as the value of giving $1 to someone with an income of $50,000 a year. So basically, you can think about that as the baseline, as we just give our donors money to people in high-income countries.

And everything else we would do is compare it against that. So our calculations of social return on investment, which we try to roughly calculate for all the larger grants that we're making, is the coefficient giving dollars created by the grant in expectation, divided by just the cost of the grant. And our bar is high.

We need an SROI of 2000X, in order to make the grant. And a lot of times, these calculations are rough, sensitive to assumptions, but that's sort of the benchmark that we're shooting for. The framework, is that all the grants cash out in either income or health.

And we don't currently use self-reported well-being, or have really any other canonical value that we're plugging in on a number of grants. So if we're valuing lead exposure reduction, the way that that's calculated is trying to estimate the downstream impacts of lead on IQ, which then translates to income, as well as Looking at the, health impacts as measured by DALIs. Just briefly, the way we're valuing income is using log utility, which was super familiar to me as an economist. So the formula would be that if you're increasing someone's income by Z% for Y years, and for W people, then the value would be this formula here. 50,000, coming from our definition of a CG dollar, being about people with $50,000 a year, times W, times the log of 1 plus Z%, times Y.

And so, a quick… an easy way to get, like, 100x return is just to give money to people in low-income countries who might be 100 times poorer than $50,000 a year, and you'd immediately have a 100X. And to get higher, you're gonna need to have some way of getting leverage on that. And we're valuing health, using DALYs, and defining a DALI to be worth $100,000… 100,000 CG dollars. So that is currently kind of our big… Assumption about comparing different outcomes that we are making. So I'll say a little about how that came about.

And it precedes me. I started at CG when it was open fill in 2024. But the big source for this is our, report on our website from Peter Favaloro, and Alexander Berger, who's now our CEO. So if you want more on this, you can Google technical updates, coefficient giving, and you'll… you'll find this.

So, in that work, they settled on this $100,000 per DALI value. By kind of triangulating between, primarily, studies of preferences that people have between income and health. Looking at value of statistical life literature.

Also. how each of them were affecting people's well-being, and then also GiveWell and other actors values, and they kind of settle on this Round number, that was… sort of seems to be in the middle of all of these different sources. But, you know, if I look at it, and if you look at it, you'll see there's a pretty big range of numbers you get out from the value of statistical life literature, and definitely fair to land in slightly different areas here.

But this was where, we've landed, and this is where we've been working, and kind of, you know, always open to revising this, but it's not, like, currently a high-priority project to tweak this number. So now, just to say a little bit about self-reported well-being. And the case for using it, in my view, would be, first, there's definitely more to life than income and health, and we don't want to just categorically say we're not going to value people feeling happier about their lives, or getting any other value in their life besides those two things.

And kind of, in my view, it's also potentially a unifying framework for comparing any outcomes. You just look at how these other outcomes affect self-reported well-being, and then you Are, are good to go. But the kind of questions that we would have at CG, and that I have individually. Kind of first is a practical one. Would a grant that focuses on self-reported well-being, that isn't cashing out in income or in Health.

plausibly be above our bar. So that's kind of… When we have that kind of case in hand, then it becomes more practically important to really figure out what we think about this institutionally. I don't think that's the focus of the workshop today, is thinking about what it is that we would want to fund, but, like, definitely eager to hear the latest thinking on this, and I'm sure the people from HLI will probably know where to point me.

And then, yeah, but then, in order to decide whether to make such grants, how to compare subjective well-being with DALIs and income, and, I think, yeah, the… relative to what David said earlier, the crux is… on my end, and I think others on my team would share, is a worry about experimenter demand, that if we get increases in self-reported well-being, it's because that's what, people thought that they should tell the surveyors and experimenters. And then just kind of general concerns about scale use. And looking forward to learning today, how people are dealing with these in the latest research, and kind of overall, I'm here with, a desire to learn, and I'm kind of looking forward to the conversations. Spending?

### [00:18:14] David Reinstein (The Unjournal)

Thank you, thank you so much. In the interest of time, let's move on quickly to Matt Lerner while, if you like, asking some questions within the Google Doc, either within the context or in that, in that table. And thank you so much, for, for, you know. Providing a picture on what your thinking is on this, and how it impacts your… decision making. Matt, please feel free to share the screen if you would like, but you don't have to if you'd rather just talk.

### [00:18:40] Matt Lerner

Yeah, well, I wouldn't rather just talk, but I don't have slides, so I will just… I will just talk, which those of the attendees who know me will be used to. I don't think I, will take the full 10 minutes, so I think we'll be in… we'll be in good shape. Thank you. What I'm gonna do, just in broad strokes, is talk through… I'll introduce myself.

And I will just talk through, sort of, our… our well-being measurement journey. So I… I work at Founders Pledge, I run the research team, I've been there for almost 5 years, and a key thing about… I'm a quantitative social scientist by training, and a key thing about our well-being journey as it is, is that When I started, the research team was, like, 3 people, including me, and now it's, like, 17 people. And so… our… Our way of approaching well-being. has maybe not kept pace with the needs of a growing team, and so that's part of the… part of the motivation for us wanting to get clear on this.

And also, to a large extent. the story of our… our well-being measurement methodology and our application of it has to do with the growth of our team, and sort of professionalization. So, where we started, where I started with this is back in 2021, Fighters Pledge was operating For those of you who don't know, you know, we're a membership organization, we function as a grantmaker and an advisor, we also operate DAFs.

But we were also functioning back then, to a significant extent, as something like a charity evaluator, so we had ongoing recommendations we would make to Founders Pledge members, of which there were then 1,700. And we had a very small number of interventions that were, in principle, well-being focused, and primarily they were mental health, depression, anxiety interventions. Like, we had, like, 2 or 3 of these.

And when I came on board, the way that these Cost-effectiveness analyses were working were just very, very straightforwardly Using disability weights to estimate the daily cost of major depression. And for those who know, which should be almost everybody here, that involves picking, between, mild, moderate, or severe depression from the disability weights. And so one thing that transpired when we started to do a few re-evaluations was just this… this issue that, like, the cost-effectiveness started to look different depending on, you know, how you decided to linearize, and whether you picked mild, moderate, or severe, and how you want to interpret those in the context of different measurement instruments, PHQ-9.

And that started to have implications for what we were recommending, and how we rated it. And we decided that it would be useful to be able to just measure things in the… in their native units. So the sort of, the principle there was, like, look.

There are already existing mental health and well-being measurement instruments. Is it possible to find some way to measure well-being interventions in well-being or mental health interventions in standard deviations of mental health, rather than try to make them interconvertible with having a toothache. And I use the toothache example because… We came to the view based on, you know, some but not very extensive data that Certain types of… Conditions are systematically over-underweighted, in the disability weights. So… Our initial attempt at this was based on HLIs and Joel McGuire's work On the convertibility of standard deviations of well-being in the, In general.

And so Joel published a paper in, I think, Nature with some co-authors. on this, not sure of his nature, I'm Tolkien correct, I guess, that I found fairly convincing. Like, the basic idea was, like, look, on different… Different measurements of mental health and subjective well-being. You can basically have a consistent valuation for, standard deviation in terms of income doublings.

So, around that time, my team embarked on a refresh of our moral weights, and this is… the interesting slash maybe funny part, you know, we ended up with a well-being to Dally conversion, but we didn't… calculate that. So what we did to formulate our new moral weights was this three-step process. We… I'm sorry. Hey, everyone being sick.

We started by doing a, well-being to income doubling conversion, which is coming from Joel's work. And then in step two, we did an income doubling to live-saved conversion. then from the live saved to DALIs, straightforward, live saved to Dally's conversion, and from that, we just backed out the DALI to Welby conversion. So we had, sort of, two sides of a triangle, and then we filled in the third to make it so that all of those different are interconvertible. So that's how we get the Welby Valley conversion we use now.

I think that's a little bit kludgy, but the basic goal was to just make all of our different… all the things we evaluate measurable in their, like, quote-unquote native units, which entailed sort of grafting on this interconvertibility. So it's, like, not a perfect interconvertibility, but by and large, it makes sense. We… Proceeded with that, and in cases where we felt there was strong reason to suspect that, for instance, disability weights are underestimating the effect. Understanding the effect, we use native units.

So, an oral pain is one example where if you look at it in terms of disability weights, it looks not so useful to work on, but if you look at it in terms of well-being, it looks a lot more promising. So the idea is, look. this is basically a quality of life issue, at least as far as we were interpreting it. You want to measure it in whatever instrument you use to measure quality of life issues.

And so that was the… that was the, sort of. motivating and initial use of measurement and wellbees. So we've been doing that for a couple years.

We… I guess, candidly, don't often actually measure things in Wellbees, like, sometimes we do. I would say it's been a valuable tool for doing prioritization, so, like, sometimes when we have something that It's just, like, difficult to measure in… sort of, like, more materialist terms. Like, a good example of this might be when you look at Issues of social equality. Like, you… if you really think that, like, this is the kind of thing you want to try and measure.

Then it probably seems like the most evidence-backed way to do it is to just look in terms of how happy people are in their lives. So it's been a useful cause prioritization tool, even though it hasn't cashed out in that much cost-effective stuff, but it makes us more confident in work on, people's subjective states. Through doing that, we have surfaced a few concerns. So… The primary one, as Dave gestured at, is this linearization. So I think, you know, if you… If you do linearize, which is what we do with a, you know, that's linear conversion.

Then it suggests weird things about the scales. And… mostly, I've been comfortable with this because… the actual movement across both scales is so small that the sort of conceptual issue with the linearization doesn't arise. Like, you don't end up with negative wellbees, or like… or negative dollies, or, you know, wellbees that are greater than 1, or something like that. That's one thing.

Another thing is this sort of, like, nagging uncertainty as to just how convertible all these different scales are, so, like. they, you know, I basically trust Joel's work on this, but the question of… Is one standard deviation on a depression instrument really the same as one standard deviation on, Cantrell's ladder is, is still nagging at me. Another… Similar question is, just… how much we want to trust the fairly limited data that we use to do these types of conversions.

So, you know, for converting to Wellbe SDs, we use SDs from the World Happiness Report. I sort of don't know how much to trust those SDs, at the country level. And I'm also not sure, don't have a lot of confidence, that STs estimate at the country level are going to be applicable for, like, a depression intervention implemented in, like, some rural village. Like, maybe there's a lot more variation or a lot less. So those are the kinds of open questions, and so, like, those sort of roll up into this big question, which is, like.

Is this actually better or worse than just using linearized disability weights, and is there something else we should be doing? So that's the… that's the origin of these… these uncertainties for me, and I think that the place I'd wrap up is sort of, like, this… this optimism that if we could get to a place of more confidence. And a more deployable methodology. then I do think there's a horizon of stuff that we could evaluate that we can't currently evaluate that effectively. You know, when we look at stuff.

Outside the sphere of depression and anxiety, So, like, schizophrenia. then I feel that we are somewhat… unsure how to look at stuff like that right now, and by default now we use disability weights or economic costs. It's not at all clear to me that that is not dramatically over or understating the cause. So coming to a point of, like, more certainty about when and how to use these instruments would be very valuable, and I think that is exactly time for me, so I did take the whole 10 minutes.

### [00:28:21] David Reinstein (The Unjournal)

Okay, So, thank you so much, Matt, and, you know, we're experimenting with this format, so sorry if there's a little bit of chaos here, but I hope that what, you know, you and, at, you know, what you are both able to convey is some sense of how you're using these measures, what the priorities are, what the concerns are, and I'm rather curious to hear what the academics and researchers, you know, have to say about that. So, in terms of the timing, we're still running a little bit behind, you know, my fault, I suppose, for getting started late. I just wanna… so I'll go very quickly. I don't have any slides for this, so that saves us a moment. Meanwhile, please do continue to add questions, and comments.

I'll… Either right here or in the tab. I'm coming around to thinking that maybe better to do it here than in the tab, because it's such a distraction to move… anyways. Alright, so I just wanted to mention, the… how this relates to our pivotal questions.

And, if you're familiar with our Pivotal Questions, please just have a look at, you know, the questions for Matt and maybe responses to that. So basically, the own journal, we evaluate research. publicly… we openly evaluate research, we commission people to do the evaluation of this research that we prioritize for impact.

We have this other trance where we're trying to focus more specifically on impact that's been mentioned by funders and stakeholders as high value, or questions that are identified as high value, and you can look at that at this link. One of the people that has engaged with us, one of the groups is Founders Pledge and Matt Lerner, and we've come up with a set of questions, or basically two questions, that seem to be particularly high value For being applied to research, have that research evaluated, have these pivotal questions, evaluators consider both the research and the questions, and weigh in with those beliefs. So, one thing to mention is we're hoping that this I'm hoping this workshop will catalyze the pivotal questions, and we're looking for people willing to sort of do that evaluation work. You can see what that involves at the link here.

So, I'm looking for nominations and self-nominations for people interested in taking on this project with, let's say, modest compensation. As you can see at that page, we built a little bit of a research base, we've operationalized the questions, And we have a bit of a plan, but I'm still interested in your feedback. on that plan.

Now, another thing that maybe you could take a quick look at now, or in the gaps, ideally, I'd love to get some statement of your beliefs at the beginning of this workshop, and then again at the end of this workshop. I know it's a big ask. Because this is very full-on.

But yeah, so you'll see, and we'll come back to this, there's an interface here, I think you can see it, where you can See how we're defining the questions, also give feedback on the questions in various ways, including annotate it with hypothesis. And also weigh in, in terms of your discussion, but also, You know, we're trying to give a sort of Quantification of some of these things. So I'm curious how that… this is a little bit early stage, but I'm curious whether I framed this in a useful way, and what your thoughts are on these questions. Which questions are we talking about?

Okay, so, you know, essentially, and I think you probably saw this in the material for the workshop. The major questions are. What measures should we be using to assess interventions in the context that, you know, that Matt and Peter talked about, to compare between interventions which affect multiple outcomes, including mental health, but also income, physical health, even mortality.

And how should we use well-being measures, including what we call the linear well-be measure, in that context? And then the second question is, if we're using these multiple measures, as Matt discussed, how do we… what's a reasonable way of converting between them? So, just to… just to give you a quick view on this, and then I'll jump right into the next, session, please have a look at the… oh, and please have a look at, just so you know what we're thinking about, and let me know, Matt, if this is inconsistent with what you had in mind. This was the idea of the focal case. You know, we're trying to compare among these interventions And we have different types of outcome data for each of these, and we need to decide, by some metric.

You know, social welfare, some wellbe… some metric of how much good is being done. We've had discussions of those. Which should we recommend? Or, you know, how much… which should we recommend under what conditions for funding and for interventions?

Okay, finally, just a little preview of some of the questions. You know, how should we use the metrics we have now To make these decisions. what… if we have a little bit more opportunity to collect more metrics, let's say through RCT in the pro… you know, in the context of RCTs, but also perhaps nationwide surveys that… or targeted surveys on the relevant populations that could be used for adjustment measures. What should we do? What should we be collecting?

You know, how reliable are these different measures? conversions, I mentioned that, calibration questions, etc. Okay, this is very full-on, so let's move right into the, next session, I don't think… We have time for, live questions. Oh, Dan is joining us, for live questions right now, at this very moment, although I think some of them will probably come up in the context of the next session, so we can kind of bring some of the questions we had on this session into the next session, which is. bring us into a, and perhaps less structured, discussion of the reliability of the, well-be measure.

And, you know, so here's the focal question, as I've just mentioned. I've also, as… for what it's worth, I've created a research background page. It might be more useful asynchronously, but just… some of this is AI, I've tried to vet it, it's bringing in… but it's bringing in things from my notes, our notes, and the readings. Note that we also have a Notebook LM page where you can, you can actually, sorry, I'm not seeing that page right now, the readings page just got… here it is… where you can actually ask questions of any of the readings, you know, that we sort of curated as being potentially important.

Okay, so, the well-be reliability discussion part of this session, we're… Pretty much on time. Actually, we're a little bit ahead of schedule, because I have scheduled in a break. So, why don't we take a moment to actually have a break, as well as, perhaps, if anyone has any questions before we begin the Wellbe Reliability session.

We'll just show what the… if anyone had added any questions here. You know, these are questions, basically, I've added. what is the native units consideration? You know, what… you mentioned the term native units.

I think also Michael Plant had a sort of catch-up question he wanted to see addressed. These were some preceded questions. They're not… I'm not super happy with those, but, Going back into the presentation here… Let's see, so Michael Plant mentioned, Michael Plant mentioned a question about Matt's concerns. Maybe, Matt, you want to take that as well as.

### [00:36:47] Michael Plant

Oh, no, I…

### [00:36:48] David Reinstein (The Unjournal)

Your desire to…

### [00:36:49] Michael Plant

I actually did use your AI tool, and they showed up, so, success.

### [00:36:55] David Reinstein (The Unjournal)

Okay, well, you know, we have time for anyone to take a break, as well as maybe one question at this point. If anyone wants to come on mic, that's also kind of nice. to ask any… I mean, maybe something for some of the academics. What were you surprised about, or what questions do you have about exactly how Founders Pledge, and, sorry, and coefficient giving, name change, and, you know, maybe other organizations, Are using these measures to make their decisions. Matt, would you like to say something about… I mean, maybe my question wasn't the best one, but something about the use of native measures, and why it's important, what you meant by that, and why it's important to use those native measures?

### [00:37:51] Matt Lerner

Sure. I mean, I'm not sure that I have, super, like, robust or principled way of describing this. I think the thought is, like. Something is going to be lost in conversion, no matter what, and so you want to do as few conversions as possible, that's sort of the intuition that I have, and so, like. fundamentally, to compare to GiveDirectly, we're gonna compete… we're gonna… So, so everything we do as with GiveWell, as, I guess, at some level with CG, Is, like, fundamentally being compared to cash transfers.

And so, there is the conversion to cash transfers happening no matter what. And… If we are using DALIs for mental health conditions, for instance, then we're doing two conversions instead of one.

### [00:38:42] David Reinstein (The Unjournal)

So that's.

### [00:38:43] Matt Lerner

Yeah.

### [00:38:44] David Reinstein (The Unjournal)

The motivation for the standard deviation-based approach, in some sense.

### [00:38:48] Matt Lerner

Yeah, I mean, I think that, like, if there was some direct measure of, like, whatever human experience that we could do on some scale, that would be the best thing. Obviously, it doesn't exist, but yeah, just getting as close as possible to the actual, experience is the motivating idea.

### [00:39:04] David Reinstein (The Unjournal)

Very cool, very cool. And, okay, I guess it's time to move on to the next session. I had one question for coefficient giving, but we can come back to it. By the way, GiveWell was mentioned There's no one in attendance from GiveWell here, as far as I know, but they are… they said they were interested in hearing the results of this conversation, so we will be, you know. after we discuss it, you know, reporting something to them.

I'm not sure everyone here is familiar with GiveWell, but I assume most people are. They are perhaps the leading, organization that, let's say, tries to… has for a long time been trying to rate And compare, and rank, Basically, global health and development charities in terms of their impact per dollar. Okay, just give me one sec. I got a private message from someone I just need to touch on before I start the next session. Yeah, okay.

Just a sec… oh, sorry, I hope you're not seeing my private message there on my screen, I hope that was hidden. I think it was. Okay, if Zoom's doing its job.

All right, so let me, just introduce this very briefly, and then I'll pass the, sort of, microphone to Casper for a little bit of, discussion, Casper Kaiser, and maybe fostering some back and forth, Here. So, I also, you know, as I said, I don't have slides on… on this, but the focal question is something like, is the linear Welby Now, what do I mean by the linear Wellbe? It's not exactly the st… the standard deviation approach seems a little different to me, but the linear Wellby, you know, is, I think, where others could… could define it better. It's defined better in the interactive research background page. Let me just jump over… Is basically… And tell me if I'm wrong, but it's basically… the… sorry, I had a better statement of it somewhere, it's a little bit informal.

It's basically… something drawn from asking people… nowadays, I think the main way is as a life satisfaction question, such as the Cantrell Ladder, asking people on a scale of 0 to 10, at least this is perhaps the canonical way we've posed it, as well as perhaps the modal way, asking people on a scale of 0 to 10, I wish I could find the… If anyone wants to paste it in the actual question, here we are. Please imagine a ladder with steps numbered from 0 to 10, the bottom at the top at the top. Top letter, the best possible life for you. I've seen some variations of this, and the bottom, the worst possible life for you. On which step would you feel you personally stand at the time?

Take everyone's measure of that, some loom, Average them, if you like. That's… the wellbe, basically. And we can talk about changes in the wellbe, and what we're really focused on here is perhaps, how does the wellbe change in response to an intervention, in response to bed nets. Or, psychotherapy, a cognitive behavioral therapy.

And of course, that could be gleaned from a randomized controlled trial between two groups, but then We're also comparing different interventions where those RCTs are run on different groups. If we're looking for… there's this… so I presented the, focal case. If you look at this, perhaps you can see the focal case as being embodied by the GiveWell assessment. Givewell was mentioned. GiveWell assessment here of… Sorry, that… that link died.

Sorry about that. GiveWell Assessment of Happy Hour… Happier Lives Institute. Sorry, I asked it to check for dead lengths. Let me just find this quickly. Give well.

here we are, of Happier Lives Institute's cost-effectiveness Analysis of Strong Minds. Which was, you know, focused on the… which used Welby's, And there's a comparison made here to these other… to two other interventions, in particular the Against Malaria Foundation. Let me just paste this into the chat so the dead link.

Okay, getting back to the focal question, is the linear Wellby the most useful feasible approach to cross-intervention comparison? Or, you know… sorry, I thought I pasted that into the chat. Does that play a role?

At least. I'll paste it into the chat. Is it… does it play a useful role?

Okay, so now here we operationalize this as a specific pivotal question. In a few different ways, and I think this could help crystallize our ideas, or our thinking on this. How reliable is it? Relative to other… Measures that may be able to pick up well-being. I can see I could potentially refine this question a bit.

Here's, again, the definition of the wellbe. We see some discussion of potential scale use heterogeneity, which may or may not matter in this context. And then, you know, here in this interface, I ask for your thoughts on this, particularly for the focal… context.

And, then I tried to crystallize this as an economist might do, if we think about, I'm generally using these measures to achieve a certain amount of value. You know, if I were going to use the simple linear Wellby measure as opposed to what I think is the best measure, or what is, let's say, what is the best measure, which may in fact be the linear Wellby, how much more would it cost me? Another way of thinking about it is, you know, given the available data, how should we be measuring the impact on well-being?

And perhaps people could give us, you know, a little more details on what we think is the available data in these contexts. I'd love you to, speak to that. we'll come back to this a bit later, and then, we'll get to the conversion question next.

Okay, so that's the focal question. we can think about the assumptions, but it's, as I said before, it's not… necessarily important to know that an assumption is violated in a, let's say, statistically significant way, in a way that we could say this is unlikely to be due to chance. What's important is to what extent is that going to matter for these comparative decisions in the context We care about.

All right, but I think you guys, you are all… everyone… all the people here are familiar with that, so I will now pass the mic to Casper to maybe have a discussion, or perhaps lead the, discussion here. He's raised these concerns previously. So, Casper, please feel free to share a screen if you would like. Otherwise, I'll just keep the screen share on the Google Doc.

### [00:46:21] Caspar Kaiser

Yeah, thank you, and thanks for having me. No slides for this section, though there will be slides when we get to, than Benjamin's and Ori's and others. paper. So the way I view this is, like, okay.

There's two ways of thinking about this. maybe pre… before all that, I should maybe briefly say who I am. I'm Kasper, I'm an assistant professor at Warwick Business School, have been working for most of my academic career on questions of well-being measurement. I'm also the chair of the board of the Happier Lives Institute, but all the views that I'm expressing here are strictly mine, and are with pretty large probability, divergent from those of HLI. Micah and I disagree on a great many things, though not are not things.

Okay, now… Do I think, how to structure this discussion? Well… I think there's two axes. One axis is, under what conditions will it be the case that the well-be will pick out the thing that we care about? Where the thing that we care about, roughly speaking, is… We are selecting that we are broadly viewing ourselves as. Consequentialists, so we care about the consequences of our actions, and we are, in particular, some version of a welfarist, so we believe that only the effects on the welfare slash credential value slash wellbeing of individuals is what matters in terms of determining the goodness of our consequences.

And then we say, well, the well-be is a measure of that. And then you can ask, well, under what conditions will it be true that we're picking out the right thing? And then there are these, kind of. I would call them statistical. assumptions that we need to take, and one of these statistical assumptions is comparability.

So that's the thing that Ben Benjamin's paper, that maybe all of us have read, and that has been at the start of this workshop, have focused on, and that is the question, well, you know, to maybe put it in terms of the To put it in terms of the… kind of framing of this workshop, suppose that I have this intervention, and the intervention causes, and I have causal evidence that people change the numbers that people say. Now, the question that I have to ask myself, well, is it the case that it changed the way people use the scale? So, you know, taking, for example, psychotherapy, is it the case that now people just think about these numbers in a different way? Or is it, in fact, the case that the latent underlying construct has changed?

And only in that case should we believe in the well-be, or at least that would be one way of weakening the well-be. Then this linearity assumption, which Dan Benjamin does not talk about, but I think that we should. I have a paper on this, others have a paper… have papers on that, I believe there's some in the reading list. Essentially just asked, well, you know, it might be the case that the difference between a 3 and a 4, in terms of the underlying well-being, is different from a move from a 7 to an 8.

And in that case, it matters strongly whether we're moving the 3s to the 4, or we are moving the 7ths to the eighths. And so, you know, you can fix this if you knew the particular form of nonlinearity. That you have. Two more things, and then I will stop talking.

A third assumption is this question of neutrality. So, especially in cases where you want to make comparisons between increasing the quantity of life to the quality of life. For example, what coefficient giving in some ways is doing by a straight-offs between income doublings and increases in life expectancy. I hope I didn't butcher that… that… that view, and if I did, Peter, please, please tell me. In the same way, the well-being needs to do this in some way.

And in order to be able to do this effectively, you need to know, roughly speaking, the position on these well-being scales, where we, as a impartial decision maker, would be indifferent between that particular person's life continuing or not. It's kind of a grim thing to be figuring out, but it seems, like, really, really key. And then the final thing, which we have… which surprisingly little talk is happening about, is, well. For the case of the wellbe, typically this is cashed out in terms of questions about a person's overall life satisfaction. this is making a substantive view on what makes a life go well.

It can be cashed out, I think, in terms of a philosophical view that might be called a global desire satisfaction view, where you say, like, my life is going well, to the extent that all the desires about my life and only my life are satisfied. But there are very many other different views about what might make a life go well. I mean, many of you, maybe in this workshop, including, I think, myself, might be hedonists.

And in that case, answers to a life satisfaction question might come apart from a question about hedonism. Now, these are all worries that you might have about the wellbe in particular, and how it might not capture the thing that ultimately matters, the master metric, the key metric. But of course. Every other metric that we have, like, for example, coefficient givings, income doublings, or the DALI, or the Quality Adjusted Life View, or inset, whichever other thing you currently care about, will have concerns that roughly map onto the structure that I've laid out onto the welding.

I think I should stop talking here. Sorry if this was too long.

### [00:52:09] David Reinstein (The Unjournal)

No, I think that was… that was a rather good layout. So let me just… I'm a little bit jumping. Maybe someone else wants to… label and summarize what you said, because I'm a little bit jumping between different windows, it's a little bit catastro… chaotic, so please, someone correct me if I'm… if I missed a point here, but I believe you addressed, you know, these… tell me… did I miss something?

But I believe you addressed these concerns, right, the ones that you see on my screen. Comparability. Do people use the measure? In the same way. Which I think is a… I see as a hard question to even conceptualize, because what do we mean in the same way?

You know, what is the gold standard here? Which is also… which is also… although, you know, we do have this vignette work, or vignette, and even more complicated than vignette work, which we'll get to in a moment. And… and then, of course, there's also the… The… the… and this is perhaps more… Relevant to the linearity question, or at least equally relevant Benchmarking against it, against what people could be seen to value in some way. So… Does a move from 3 to 4 mean the same as a 7 to 8? What we call linearity question, which I think people in this conversation are using the term cardinality and linearity fairly interchangeably.

I haven't… I've been trying to dig into that. Someone tell me if I'm Wrong, but cardinality, comparable cardinality, and linearity. But I'd like to see evidence I'm wading in too much, so please, others, jump on top of me. Benchmarked against things like the time trade-off. Or the person trade-off, where you say, okay, if you could move some people from 3 to 4 at the cost of moving others from 8 to 7, would you do it?

Or the standard gamble? And then you mentioned the neutral point, right? Which there's, I guess, we have, if you look at the notes, we have some measures of, but perhaps they're not we don't have a reliable measure of in this context, I'm not sure.

And then you talked about Are we measuring the life concepts? And I guess. two dimensions there. One could be, should we be asking life satisfaction questions versus the instantaneous happiness type questions?

And another, you know, the recall, how happy were you in the last X time. Another could be, should we be asking these overall questions? Or these more… There's, you know, a longer variety of questions about different aspects of your life.

And maybe one concern that I've heard mentioned is. Do… does asking, let's say, the cancel ladder question really reflect all of, let's say, the health and income measures we care about, or is it picking something else up? And obviously, there's a lot of discussion and evidence to be had on that.

And finally, I think you also mentioned the issue of the reporting function and whether interventions can affect the reporting function. In other words, and I think others have raised it already here, you know, how confident can we be, particularly when talking about a mental health intervention, that if I've been given cognitive behavioral therapy, that doesn't fundamentally affect the way I respond to, let's say, the cantral latter question. And there, if that leads to even a shift, that could potentially bias some of these estimates.

Anyways, this was… that was just my characterization of what you said. Is that… is that a fair… Breakdown or did I.

### [00:55:32] Caspar Kaiser

Yeah, absolutely. Yeah, no, I, I totally agree. agree.

So, I did the reporting function thing. Right, so in Dan's paper, and there's other papers, like the recent one from Alberto Pratti, I have a paper on this, a bunch of papers on this, is always put in terms of, like, oh, you know, associations with covariates in, like, these observational data sets. But the entire statistical machinery that's been developed In these papers, directly applies to you know, people doing RCTs.

And probably what we should just be doing is, like, asking questions about people's memories of past life satisfaction and vignette questions, and maybe calibration questions in these RCTs. Like, that would seem like the obvious thing to do.

### [00:56:21] David Reinstein (The Unjournal)

Although it's expensive, as we know, to do the longer vignettes, but I, you know, I raised the issue… again, I'm talking too much, other people should be talking, but could you do these vignette studies in a way that it's done… that at least if, like, maybe you do it on a representative population, and then can use it to adjust future metrics? I guess.

### [00:56:40] Caspar Kaiser

All you need is two vignette questions. The minimal thing to do for… to implement at least Dan Benjamin's method is two vineyard questions. That's the minimum you need.

### [00:56:51] David Reinstein (The Unjournal)

And that's easy.

### [00:56:52] Caspar Kaiser

too expensive.

### [00:56:53] David Reinstein (The Unjournal)

helps us get at the comparability question, but not at the linearity question, I would guess.

### [00:56:57] Caspar Kaiser

Correct.

### [00:56:58] David Reinstein (The Unjournal)

Well, just… and I'm really trying to stop talking, but just to go off mic a bit, let me just raise a few questions. One question is, which of these do we think are more important to the reliability of the linear well-be approach in this context, In other words, prioritize… and which are less, you know, prioritizing these interventions in… four countries. And maybe secondly, okay, what evidence am I missing, particularly on the linearity aspect?

And I see Miles is filling in some notes there. Maybe he even wants to come on mic. I'm gonna mute myself. To… as a good practice.

### [00:57:37] Miles Kimball

I'm happy to come on mic here. So… so I think, first of all, I think it's helpful to not get hung up on cardinality as if it's a technical question, because I think it's just as much an ethical question, and if we… If we had… if people understand the scale as it is, then you can ask them their attitudes about inequality on that. So this question of what's the appropriate curvature is really a question about inequality aversion for measures of well-being like this, and I think that would be a very useful thing to do. So anyway, I'd encourage us not to get hung up on it as a purely technical question, although there are technical issues in measuring inequality aversion on a scale like this.

The other thing I wanted to say is that scale use correction is highly relevant here, because if you do care about inequality in well-being as measured, then you're gonna count a corrected point as being worth more for somebody at a lower level, but you can't know if somebody has a lower level on the common scale without correcting for for scale use. So… so there's a variety… it's interesting, as we continue this discussion, I'm realizing a variety of reasons. Scale use correction matters that may be less obvious and haven't been in the documents.

The other one I'll just mention briefly is the income conversion. Scale use correction changes the income coefficients, so if you convert to income With, in the usual way, by dividing Whatever regression coefficient you have on the whatever effect size you have on the… on the income coefficient, you're gonna get a… you're gonna get a much bigger number. Like, when we converted unemployment, we got 5 times as big a number after scale use correction. Scale use correction has a particularly large effect on the income coefficient.

### [00:59:48] David Reinstein (The Unjournal)

Could you elaborate a little bit more on the income cost? Sorry, if someone else wanted to jump in, please, please do. I just wanted to elaborate a little bit more on the income.

### [00:59:54] Miles Kimball

Actually, let me have Ori elaborate on that. He's got the facts down better than I do.

### [01:00:01] Ori Heffetz

So, the… the question is how you convert from… between different things, but one practice has been to regress The so-called happiness regression, so you… Regress those answers to the question you showed before. You know, rate your… yourself on a letter from 0 to 10, or what's your life satisfaction from 0 to 10, or some other thing. And you, you regress that. On a bunch of features of the respondents, among which is their log income.

And other things. So, for example, employment status. And then you can find, by the ratio of the coefficients, under some assumptions, you can find, you can basically price, you can say, if you give people X dollars. Or give them a job, compare an unemployed to an employed person. there are indifference between the two in the sense of these two will have the same, increase, the same effect on their self-reported well-being.

And so this has been used to price all sorts of things that we can't price. Unemployment is one important example. And so we show in the paper that Dan will present later. and Casper will discuss, I think, etc, that, While scale use correction doesn't change the sign of things, so with or without scale use correction, unemployment is a bad thing. Pretty bad, you know, large negative coefficients, and income is a good thing, you know, large positive coefficients, so it doesn't change that, but it does dramatically change the ratios.

So, in one example there, we have the leading example with our data set that we collected, with, I think, around 10,000 respondents on the Understanding America panel. We show that you can change the ratio, so how you price unemployment in money Could change by a factor of 5. if you… a correct versus the raw data not correcting, and that's a big… right? That's a big deal. That's a pro… that's a project that, moves up.

Five-fold, when the funders here, consider which project to evaluate. There's a bunch of other, methodological issues with using this question that you showed before. that I'm not going to get into now, but I'm happy to discuss in the rest of the workshop. I'll just mention one. It actually matters which question you asked.

So sometimes little changes in the wording of the question Could affect… could make the people think, for example, about what time scale did you ask me about? So, one question, I would focus about the past few days or week, another question, I'll focus on the past week… year, or maybe my entire life. And this also makes dramatic differences, again, sometimes in order of magnitude. Because if I was answering about my entire life, you know, a letter, where I put myself on the ladder in the entire life, one intervention is not going to move me much. If I was focusing on this week, one intervention, even with short-run effects, would move me a lot.

Now, we priced it in money. But, you know, we priced in money what? This one week or the entire life, etc. So again, big issues.

But… but yeah, the… the… back to Miles' point, the… Even just a Scalius correction. Could move you, dramatically, in pricing these things.

### [01:03:59] David Reinstein (The Unjournal)

Just… If no one else wants to call Mike, just up. I guess my concern is, are we talking… I mean, this… so this, this, this, discussion we were just having, you know, this is one application of… and one… Context in which these well-being measures are used and considered. Is that relevant to the choice between interventions we're talking about now? Because, you know, the approach to sort of trying to see how much do people value income versus other things, and comparing the regression coefficients, is that kind of a different… anyways, you get the question I'm asking. Casper, please…

### [01:04:37] Caspar Kaiser

Yeah, two quick views on this. One is, Yes, because… so the worry would be that imagine you are either… you're doing one of two interventions, let's say you're either doing canonically, I guess, psychotherapy, or a direct cash transfer. And now, if we're only talking about the comparability question, let's leave aside the linearity stuff.

Then… What you might worry about is that if a person suddenly, like, you know, has received psychotherapy, they will interpret these questions in a different way, and you might now believe that I don't know, they… individuals now interpret these questions in a looser way, and therefore the… whereas when they receive a cash transfer, they interpret these questions in a stricter way. And as a consequence of that, the relative cost-effectiveness of psychotherapy would go down. But of course, this can also go the other way around. You could also believe that, like, as a consequence of psychotherapy, you now think of these terms as, like, much more weighty than you did before, and so therefore now you're changing your use of a life satisfaction scale in such a way that an underlying life well-being of a certain amount Suddenly now translates to a lower number. Currently, we don't, I think, know whether there are effects of cash transfers in particular, or psych therapy in particular, or lead interventions in particular, on the way people are using sketch.

It would surprise me in the case of lead, but I could imagine it for the other two. So, it's a source of uncertainty, but crucially, and I think this gets lost in this debate. many of these kinds of concerns are also occurring, variants of this, in other metrics. Like, maybe not the comparability stuff, but for example, the linearity thing. One might say that, you know, in QALYs and DALYs, linearity is baked in, but it's only baked in under extremely strong assumption about, like, how people reason about time trade-offs.

So, when we're attacking the wellbe with these… under these kinds of concerns, we should also do that for the other potential types of measures. Otherwise, I think we're attacking one thing. without attacking the other thing, while the other thing still has, you know, similar kinds of worries associated with it.

And we just don't worry about them because they are older.

### [01:07:16] Daniel Benjamin

I want to build on what Casper was just saying about Concerns with other approaches. I think it's useful to think about the, kind of, counterfactual to using wellbees? Like, what else would we do? Or license… or not Wellbees specifically, but… self-reported well-being measures.

I think the natural alternative that we would consider Is some kind of revealed preference approach. And… just as Casper was saying, I think… Some of the same issues come up. So, for example, if we use revealed preference, And… Think about… Costs and benefits. in dollar terms, There's an implicit assumption that a dollar… Is worth the same to everybody. You know, unless you do… And then there's attempts to deal with that with… Social marginal welfare weight, but… You know, the baseline cost-benefit approach.

Is making an assumption like a linearity assumption. There's also… Analogs for… Many of the other assumptions. And I think it's important for… two reasons. One, as Casper was emphasizing, It's, You know, we should be evaluating The self-reported well-being approach. relative to the next best alternative.

So the fact that it has these You know, that we may not have solutions to these problems doesn't mean we shouldn't do it, because we… The… the… what we'd be doing anyway would have the same or similar problems. And the second is… There are kind of… ad hoc solutions. That are, adopted in the reveal preference approach that, that could be carried over, and they'd be just as ad hoc, but similarly, you know, similarly so.

So, for example, the question of Where is the zero point? I mean, there's no… Definitive answer to that question. In the reveal preference approach, there's different ways of trying to estimate it, they're imperfect.

And I think you could use… You know, similar approaches in the self-reported well-being approach.

### [01:09:56] Matt Lerner

I wanted to… I wanted to bring two… two things to the conversation that are sort of related to this topic, but mostly I just wanted to introduce them as things to keep in mind. So… so one is this really, really painfully obvious point. So, sorry, these are both… both things coming from the sort of charity evaluation grant-making perspective, which I think are important to try and connect to the… academic conversation.

So, one is the painfully obvious point that, like, cash transfers are the benchmark via well-being, which I think is, like, probably obvious to everybody, but that seems like an important point, so… So, so Peter gestured at, other CG Peters, research on, on well-being and, like. So it's important, obviously, that, like. when we compare to cash transfers, we are directly comparing to a benchmark that's determined on the basis of well-being, well-being research about cash transfers, and not some other thing, like VSL. So that seems like just an important thing to bear in mind. So I guess, like, an alternative way we could do this would be to To just be thinking in, like.

VSL equivalents for cash transfers, and not talking about well-being at all. So, that just sort of seems like an important alternate reality to keep in mind, you know, it's an obvious point. The other thing which is probably more important to bring in the conversation is… My team's comfort, at least, with Kind of not having tackled some of these questions, and with stuff like linear conversions. Has to do with the fact that, like, ordinality is an overriding concern. For us, and so… to the extent… I think maybe Peter might have a different take for CG, I'm not sure how they do this sort of thing, but like… To the extent that, I forget what the textbook… math textbook phrase is, that these conversions are, monotonic under, Linear Athene transformation, I don't remember what it is.

To the extent that you can preserve the ordinality of different interventions using whatever… whatever conversion you use. it's basically been all the same to us, so we lose a little bit of fidelity, there. Like, we really do want to know, like, if stuff is, like, an order of magnitude or two orders of magnitude better than something else in well-being terms.

But a lot of the value is gained by having the ordinality preserved. And so, like, the idea that these different conversions preserve ordinality has less to be comfortable with, like, a less conceptually coherent… conversions, and so… That seems like an important practical, matter. I guess I… I do have this sort of unknown unknown as to whether, like, we are really missing a trick there, and it would still be better if we had more rigorous information. Like I said, we do… we do make some allocation decisions on, like, a cardinal basis, and so it would be good, but I do… I do want to flag that, like, the ordinality Of well-being conversions is, like, a… is a… is a principal concern, rather than… rather than getting the exact estimates right.

### [01:12:58] Peter Hickman (Coefficient Giving)

I think for us. you know, I mentioned this $100,000 value per Dally. If we doubled that to $200,000, or if we reduced it to $50,000, that would change how some of our program areas look relative to others. So it's important.

And I guess I'm… I'm looking forward to hearing more about the latest from the scale conversion research, but, like, if kind of correcting for scale use heterogeneity means that income is way less valuable relative to health than we previously thought, you know, that's an input that I think my team would want to consider. kind of… All that said, there's… kind of a… at CG, I guess we've got this recognition that comes up in that report from the other… from the other Peter, that different methods to kind of comparing income and health can give you pretty different answers, and… So we don't have, like, a formula that we're just gonna commit to and plug in, and then one tweak will then make us double it. But yeah, I do think that, like, how income affects well-being versus how physical health affects well-being is an important question for us.

### [01:14:08] Caspar Kaiser

I don't know if this is totally off-topic, but I would actually be quite interested in how historically it occurred to coefficient giving in particular, to look at DALYs in particular, and whether there's some story to be learned here for wellies. whether… Insofar as similar worries and, like, academic debates raging for decades. You know, remaining unresolved, and yet, you know, major organizations are adopting these measures.

So, I guess part of the reason why I ask is, I wonder, hey, how should I spend my days?

### [01:14:46] Samuel Dupret

Can I add something to Casper's question here, to make it simple? Is it… is it just because the global burden of disease has, like, 300 DALI values, and wellbees don't have that? state of… of, like, data richness? Like, is it just that there's loads of daily information, or is it something else?

### [01:15:11] Peter Hickman (Coefficient Giving)

I don't think I'll know, historically, what led to where we are. So I'll sort of give, I guess, my own… my own view. We talk about dallies.

A lot of it for us is just gonna come out in terms of lives saved, and then a multiplier for length of life, and we use some age adjustments that we just took from GiveWell. And so… kind you know, in a way, we're looking at how many people are dying from malaria, how many people are dying from the other different things, but we do pull from the GBD, the Global Burden of Disease, because then we're also factoring in the disability that's occurring, and it… yes, I agree that pulling from the GBD is this convenient fact that we can use, that, means that, you know, it's kind of a canonical source that we can pull from, and then have simple spreadsheets to do estimates of cost-effectiveness of different areas. Whereas we don't have the, like. GBD in terms of wellbees.

And as far as I know, actually, I'm kind of curious if there is, like, a source where you can get, like, what's the… affect what's the malaria… malaria's well-be burden worldwide versus other areas, but I think I know that then we have to think about the assumptions that's been brought up before about how bad is death, where you can get really different answers there.

### [01:16:31] Caspar Kaiser

Thank you. Yeah, that's really helpful.

### [01:16:36] David Reinstein (The Unjournal)

I know what?

### [01:16:36] Julian Jamison

We maybe need to move on soon, but this is interesting. I'm happy to keep talking about this, Julian. I also happen to have my father visiting, sitting next to me, which normally I wouldn't just bring my father onto the call, but it's Dean Jameson, and he… kind of helped originate DALIs and was the lead editor of GBD in this space, so if we want to know something about the history of Dallies, and if he's willing to say a few words, it might be useful. He's not a big fan of DALIs now, I think I can headline that, actually, but I'll let him talk.

Well… It was quite fortuitous. I've enjoyed the conversation so far. My name's, Dean Jamison, as Julian said. I happen to be a co-author of the first published, Global Burden of Disease Using DALYs. and… Have followed that literature.

Subsequently, on the disease burden measurement side, the disease burden measurement has ended up in a somewhat different place than the use of DALIs for Economic evaluation or cost-effectiveness evaluation. more narrowly. Mostly in… practice around economists using for economic evaluation, discounted DALIs so that a child death is worth 30 dollies, whereas the global burden of disease has a child death worth 70 or 80 dollies. So they're You can't talk about using cost-effectiveness as a way of judging how much you're reducing burden of disease, because the DALI, the burden of disease DALI, Has evolved very substantially, and in some ways, quantitatively, quite Importantly. I've just been embarked on an exercise… for the Lancet.

which has a Commission on Investing in Health, which I chair. And we're dealing… With… in this work, much more substantially than we have before with the question of non-fatal outcomes. So I have become interested again. in measurement of the YLD component.

The years lost due to disability. component. of the deli.

And just had one of my assistants look at what I thought were the 25 most important fatal outcomes that this Lancet Commission should consider. Blindness, deafness, major depressive disorders, manic depressive illness, dental carries, Pretty much anything that looked fairly big, but things like menstruation appear on that as a clear major use of health services. Menstruation is typical of about half Of the items on this list of Disabilities, we're trying to get a… Sense of the relative priorities of Dealing with the disability, but in particular, the relative priorities of interventions across alternative areas of disability, or… non-fatal, But half of them have no… easily ascertained. YLD measure of burden.

About half. Slightly fewer than half. And, in my view, many… of the ones that do have a YLD. version have the associated disability weights and the associated durations… Are prima facie simply not sensible? So even before you get to the question of comparison With mortality risk, and the this utility of… incurring more… mortality risk.

There's a serious incompleteness If you're thinking about healthcare, to the whole… LD structure. So my colleagues are not happy with me. about this conclusion, because the YLD and… than the Dally are. well-established As the measure both of burden and the cost effectiveness. variable.

Effectiveness variable. But it's very much on our minds, and as we get seriously into trying to look at All forms of non-fatal conditions addressed by healthcare systems, it's quite clear that the YLD is just not going to do that. So we're in the business of asking ourselves the question of. What do we do as economists trying to sort out the questions of priorities across dealing with this whole broad range from Alzheimer's on down to Deaths in childhood from malaria, too. pre-birth tests to stillbirths.

It's all got to be in the mix, but the focus And our current work is just on the… Non-fatal outcome side. Yeah, so my takeaway was, and I want to hear Michael's up, is even within health, a lot of this discussion is, like, well, bringing together health and income and other things, and well-bees can be good for that. because we've already got DALIs within health, but even within health, it seems like there's a lot of scope for something beyond the standard disability.

We could do better with the disability weights, but also even with those, there might be scope for well-based. I don't know if I'm in Toronto, or Michael, I'm curious to hear.

### [01:22:33] Michael Plant

Oh, hello, I, I… I don't know if I'm intriguing the next segment, or should I… should I.

### [01:22:39] Julian Jamison

Well, I guess that's up to David.

### [01:22:41] David Reinstein (The Unjournal)

Okay, yeah, sorry to be… here's a suggestion. Why don't you… maybe you could give a quick response, and then, you could… you're welcome to… we have two separate breakout rooms, although, I mean, I certainly would value your participation in the next segment also, but we should have two separate breakout rooms, so you could, I could suggest, if you want to… continue some specific conversations, you can join one of the breakout rooms, I think I labeled one of them… you see that as the three dots at the bottom? I think I labeled one of them. This seems like more of the theoretical, discussion breakout room, but whatever, that's just a fixed point.

Anyways, I'm talking over your time. Michael, do you want to give a brief response, and we'll move to the next, section?

### [01:23:18] Michael Plant

It's difficult to know what to say. I mean, I guess the first thing I'm… it's really a delight to have people discussing these issues. I recognize that HLI has played some role in getting them on the table, and it's really just wonderful to kind of hear people discussing them. I mean, I've got views on lots and lots of things, but I don't want to take up lots of, Lots of oxygen in this. I did, one thing I wanted to kind of draw attention to, or maybe two points, is, we've focused on this, linearity question, and, and this question about conversions.

So there's this other thing which, Dean, who I'm delighted to have here, I haven't followed up to the email from you and Julian, but I will do so later on. is, there's this lurking issue of kind of the madness of death, and so actually that's this kind of third separate thing. There's neutrality, there's also this age weighting, you could also complicate it more with population ethics and, you know, other sorts of things, so that's the kind of elephant in the room. On this table, on this kind of well-being weight, so actually in HLI, we're now working on this. So for a while, people have said to us, okay, what's the difference between a Dally and a Wellbe?

And we've said, look, we don't really know, one comes from people's expectations, the other comes from people's experiences, like. But we've realized that's kind of not really a good enough answer, so, we're actually going to try and put something together to say, here's kind of a standard set of weights you might get, and then here's what the wellby weights are from the data, so you can all get very excited about that appearing at some point. And hopefully it'll be useful for kind of seeing how this splits.

And that'll be, you know, that'll be the start of something, but it'll be kind of useful to have some more numbers for different things.

### [01:24:59] David Reinstein (The Unjournal)

Thank you very… thank you very much. So, we're gonna… I'm gonna move to the next session. We're about 10 minutes over, but it's okay, because we had scheduled some buffer and break. Just, like, in terms… in terms of following up, in this section… session, as you can see, there's… there are several… in the notes, there's several pivotal questions linked to this we'd like to get your feedback on.

And I would just say, To sort of… to sort of… consolidate. The main questions, perhaps, are. Or perhaps the most important questions are, which of these concerns are most important And, you know, that could motivate us encouraging people to do more research and collect more data. Which of these concerns are most important for the focal context of assessing these different interventions in LMICs?

And what is the evidence? You know, what do you feel is the best evidence that we have, either on these questions or on the reliability of these approaches, which, you know, involve wellbees? If you look… dig into the details, you'll see very often there's conversions from other psychological measures.

Anyways. So yeah, I'd love your continued feedback on this. This is chaotic, but hopefully productive.

As I mentioned, there's the breakout rooms. You're welcome to invite people to join. Okay, so now I will move to the next session, which, you know, is Certainly not unrelated to this material. Sorry, I was acting as if I was sharing my screen, which I wasn't, which is how… should Founders Pledge and others convert between dollies or QALYs, but it's generally dollies, and wellbees, if they are thinking of using each of these. What is the appropriate conversion.

You know, two ways you could consider to do it. One would be Standard deviation units, standardized relative to the variation in the populations you're dealing with. Another Could be, I… I forgot the other way, but I guess… I think the way that HLI does it is, I think that it's a matter of… of looking at the total value of the life and sort of maybe partitioning it equally into the different intervals of the… of the, of the response, or the well-be scale, rather than using standard deviations. Let me just jump to the workshop site on this, And I won't lead this conversation, but yes, so… This is the question. If we need to compare the two measures, if these are the only measures we have, maybe we could consider gathering a little more data to adjust or enrich one of those measures.

But mainly, thinking about the measures we do have, and I think the practitioners can discuss more what those measures look like, that they're collecting or able to collect in making these comparisons, particularly for that, core case I mentioned. What conversion, or what approach should they use to conversion these? You know, standard deviations, relative to what populations, you know, what normalized? Should we just use the scale of the life satisfaction question linearly. Should we use some other approach?

perhaps requiring some other instrument. What do we mean by best? You know, you can see there's some specific questions we've, Framed on this, which we'd like your feedback on, these pivotal questions well. Whatever you see as the welfare measure which leads to the decisions that lead to better… in expectation, let's say, better welfare outcomes. Should it vary by domain?

And how much does it matter? You know, maybe… maybe these yield approximately equivalent results. Okay, And, I will now… you can see there's a little bit of an AI-assisted summary of some of the readings and discussions we've had here, if you want to sort of bone up on that.

But I will now pass the mic for this discussion to Julian, and I know Sam suggested that he would want to discuss or present on this a little bit. So, Julian, would you want to go first, or would.

### [01:29:10] Julian Jamison

Yeah, I think I'll give the time to Sam. I'm happy to just keep the discussion going, although there doesn't seem to be a problem with this group, which is great, so I will… I will… if you maybe stop sharing, David, I think Sam has some.

### [01:29:21] David Reinstein (The Unjournal)

Yeah, no, you can just share right on top of me, that's okay. Okay. But… Thank you, Sam.

### [01:29:30] Samuel Dupret

Yeah, thank you, Julian. Yes, I… I am… Lots of interesting discussion here, I look forward to… hearing from Jameson's and others about this, so I thought… I don't know if I'll exactly answer the framing that, David, you've given, but I thought I'd give a bit of… a bit of mind framing about discussions of conversion between these two, and I'll give you a little… bit of an insight into what Michael was saying, like, what HLI has been looking at. So… Let's talk about wellbees. the… I just wanted to define well-bees again, it's like a one-point change on a 0 to 10 well-being scale across time.

So, the core thing here is that we are quantifying well-being over time, and as Michael mentioned, like, there's discussion about which which well-being scale might use? Is it life satisfaction, is it happiness, etc. And there's been a lot of mention of Wellbe SDs, and I just wanted to… take a moment to say what that is.

This is… I think, if I've understood correctly, this is because at the Happier Life Institute. when we do evaluations of cost-effectiveness of charities and interventions in low- and middle-income countries, we don't have these rich data sets using a 0 to 10 well-being scale. There's lots and lots of… RCTs on different scales.

And so we… we do what we typically do in a meta-analysis, is that we convert… we convert these things to standard deviations, and then… We could have stopped there and just been like, this is a standard deviation of on well-being scales, but we wanted to make it understandable and use the well-be framing, so we convert from standard deviations to what it would be like on a 0 to 10 scale. And we make some assumptions here using, the standard deviation typical on the Cantrell Ladder and the World Happiness Report, the Gallup World Poll. But basically, the core here is that we're quantifying well-being on self-reported scales, and then there's, like, different methodological things that we do with the constraints of the data we have. It's not like… We don't have a particular, like. love of doing this conversion through standard deviations.

The other thing we do is that we… there's so little data is that we convert effective mental health scales, we add them to these other well-being scales. Sometimes there's just not enough data, just, like, on life satisfaction or on happiness, so sometimes we add what we call effective mental health, so a depression scale will ask you, like, how low has your mood been over the last two weeks? And that's kind of a question that links to people's hedonia that links to well-being.

And we have a report about this where we look at the data we have and we compare, like. different interventions, what are the scores they give on these different scales? So that was just some clarification there.

And DALIs, there's two bits of the DALI. One is about life, years of life lost, so quantity of life, and there's years of life with disability, quality of life. and… I will make the case here that if we want to talk about conversions, like, we're… Mainly gonna talk about the differences in quality. Quantity can… you… you just take the… deaf and remaining life expectancy values, and you'd add the… you'd add the well-being level and questions of the neutral point.

But then, all the questions about philosophy that Michael has brought up in, The Elephant in the BetNet, they'd apply to the DALI as well. So, to say, like, I'm saying, like, when we're talking about converting dalies and wellbees. What we're really talking about comparing is what's the weight on health for the Dali, or what's the weight on well-being for the wellbe.

And if you're looking at quantity, then the philosophy about the badness of death affects both measures, and as Dean Jameson mentioned, like, the DALIs have actually changed their in the global burden of disease, there's actually changed their, implied philosophy about the badness of death, because they used to discount age, and now they don't. And also, this is not really a question about modeling, if you want to convert between these things. So, we add spillovers in our analyses. You could do spillovers with DALIs.

We… we don't typically do analyses where we say, this intervention reduced the cases of depression by X percent. We do, like. Oh, this intervention had this effect on the well-being scale or the well-being scales we have.

But what you're doing with Dally's… That you could just do with well-being, is that you say. I want to know how many people we've treated cancer for, and the well-being weight, or the daily weight for cancer is this, and this is how I'm doing my analysis. So, there's differences in how you do analyses that are kind of independent of the metric, although we tend to do things one way, it's not just because we use wellbees, it's just because we tend to do things some way.

The core difference, again, is, like, what's the weight that you give to conditions between these two things? What's the thing that you're quantifying over time? So we're quantifying well-being with the wellbees, and quantifying health with the dollies.

And so, one is doing well-being, one is doing health. The other difference is how these things are obtained. So, the daily weights are obtained by the general public, who doesn't necessarily have the condition, making a pairwise… making pairwise health judgments.

So, they ask, like, which… which of these is less healthy? Depression or cancer? And they're given lots of these questions, and then that gives you the weight. Whereas for well-being, we… For wellbees, we use self-reports, so it's the people with the condition receiving the intervention, etc, saying how satisfied they are with their lives, or how happy they are.

And so there's a 2x2 difference here. And these differences… what probably comes into these differences are problems of effective forecasting. So people making judgments about how hefty different conditions are, when they don't have these conditions, are likely to make mistakes, about How bad a different condition is to live with, and there's some literature about that.

So, Diane Melkaf, they did this with QALYs, but it's the kind of same idea. They find that people wait moderate mobility issues, just as bad as mental health issues, but if you ask people on their life satisfaction who have these issues, it's actually way worse to have a moderate mental health issue than a moderate mobility issue. Pine find that people who have experienced depression, they're way less likely to, Prefer the depressive states to other states.

And, there's different criticisms of the DALYs and how they incorporate the badness of mental health And then there's other intuitive things in the… unintuitive things in Dead Dally, so… In different versions of the daily weights, because these changed over years, the difference between having cancer with or without treatment is often very, very small. And that just seems very strange, as a health state. I haven't checked this one again, but… Apparently, blindness is not weighted very highly, because it's… it's… It's about well-being and not about health, so… Yeah, that's why we would… we would value well-being.

We'd value well-being, because it would take everything into account, not just health. Like, health is one determinant of well-being, but, if you're blind and that affects your relationships, it affects your wealth, it affects all sorts of different things, and not just your, your health, then, well-being would capture that. there seems to be some sort of, like… correlation between, DALIs and well-being at the, international level.

Although this then looks strange, but life lives with… lives… years lives with disability, that… This is suggesting that the more you have years lived with disability, the happier your country is, which is just strange, and this is probably this disappears when you control for GDP, it just seems to be that, there's other things here. So, like, rich countries live longer, they have welfare states, etc. They can deal with having years lived with disability. So if you were just looking at years lived with disability, you'd miss… be missing something.

Also, one is on 0 to 1, and 1 is changes on the 0 to 10 scale. I think that's rather trivial, you can kind of convert that for… Conversions, etc. So this is the differences between the scales. I'm going to talk about conversions, but Matt, I see you have your hand up. Do you want to…

### [01:38:56] Matt Lerner

Yeah, I wanted to address the… confusion or disconnect or something about the SD use, and possibly get a correction from you guys. So… So it's been a little bit of time since we did this, so I'm worried I'm gonna misrepresent what we're doing. But the basic thought was… sorry, so Samuel, you were saying, we just measure things directly in WellBees, why bother with SDs, right? That's basically your…

### [01:39:18] Samuel Dupret

No, no, I'm… I'm… I'm… I'm… I was just getting some… I was just saying, like, the SD thing is a… is, like, a methodological thing on top, that we do. We personally do that. But it's not that… it's not that this is, like, a core necessary function of using well-being, it's just a… it's kind of a consequence of having heterogeneous, limited data, so to add to all…

### [01:39:48] Michael Plant

That's what I'm saying. If we… if we had loads of, kind of, standard well-being measures, we would just, like, skip.

### [01:39:53] Matt Lerner

We would want. Yeah.

### [01:39:54] Michael Plant

Yeah, but the point is that we're being data omnivores, so we're using this stuff, and then because it's standard meta-analyses to kind of look at standard deviations, we thought, okay, what's, like, let's try and cross-compare them and make use of data which isn't.

### [01:40:07] Matt Lerner

Okay.

### [01:40:08] Michael Plant

an official subjective well-being.

### [01:40:09] Matt Lerner

Sorry, so, same position. I don't need to elaborate that unless anybody has any uncertainty about this, so let me know if there's any uncertainty, but it sounds like we're actually all in agreement on methodological questions here.

### [01:40:21] Samuel Dupret

Yeah, so that's… that's a… I feel like that's an extra methodological question. We have… a report in 2024 called Converting Effective Mental Health Scales, and we discussed this a little bit, but yeah, that's… It's not the ideal world of how wellbeing measurement should work, and I'm gonna focus on… on just the… the… the real… real true difference between delis and wellbees. There's been other conversions, maps mentioned how they did it with, like, this… these three steps. Stan Pinset tried to do something between as the years of depression and DALIs, and the Treasury has something where they go between… the difference between full health and death on the life satisfaction scale, and that's how they convert to DALIs.

All, Foundation methods and the UK Treasury methods, if I'm not incorrect, they involve, kind of, like. folding in the assumptions about the badness of death and philosophy in, and so I'm just gonna do something just about the disability weight, disability to well-being weight. We use data that the… Happiness Research Institute, which is not us, not the Happier Lives Institute.

They use the SHARE, which is a panel data of all the Europeans, and they have data on life satisfaction for different health conditions. And they compare this… they get coefficients for this on life subtraction, and they compare this to the disability weights. I'm gonna say disability weights for the rest of this. Two notes here.

They don't do exactly disability weights, they take the year's life years of life with disability, and they divide it, by the number of cases, because some people will have, like, moderate depression, and some people will have I can't remember the terms were, like, different levels of depression, and so they fudged this in some… a little bit in some way, instead of splitting them up. And also, these are disability weights. from 2015, I imagine, if it's in 2020, so someone could do this exercise again, with newer disability weights if they've changed for some of these conditions. I don't know if they have.

And so here, the 16 different conditions. You can see the disability weight doesn't go very high. The maximum would be 1, so this is not very high.

And then this is the loss in well-being. The… the condition that has the biggest loss in well-being is, more than one point on a 0 to 10 scale is depression. And you can see… There seems to be two of these that have more effect on well-being than the others, and these happen to be the mental health conditions. If you look at them in terms of, kind of, rankings within each, it looks like life satisfaction is much more affected by these mental health conditions than the physical ones, whereas the valley's, kind of, spread out. and then what I did is just a simple linear model with a moderation.

And, you can see if you take.

### [01:43:46] Michael Plant

Yup.

### [01:43:46] Samuel Dupret

Don't be too bad.

### [01:43:47] Michael Plant

That last one, that one last one is probably useful for the… For the viewers.

### [01:43:53] Samuel Dupret

Yeah, I can wait on this one for a bit.

### [01:44:01] Michael Plant

I mean, I guess the point here is… this I find kind of more intuitive than the previous one. The point is to share there's quite a lot of shifting around. If the… if the… if the weights between the well… between the disability weights and the well-be weights were the same, you just get the straight lines across, and actually here there's quite a lot of movement.

And, so we don't know exactly how far this generalizes, how much this is applied to, like, non-health things, and the story here seems to be about affective forecasting. So, you know, in part health, in part predictions. But this is why there's not a… there's not a kind of a neat one-to-one weight. If you're interested in… if you think what matters is well-being, and you think it's well measured by sort of active well-being scales, then you don't want to be like, oh, I just want to, you know, if you're trying to be… really kind of get to court matters, you won't want to say, oh, I just want a simple conversion ratio.

The point is that there are these differences, so if you're trying to capture well-being, you want to find the well-being weight. compare it to Dally weight, or the disability weight, and then go like, huh, how much difference does this make?

### [01:45:00] Matt Lerner

Can I… can I jump in quickly with another clarification for the academics? Because I feel like I might have glossed over this part as well. So the… I just posted in the chat that this is part of our rationale for using the native units thing. So the idea is, like, we need to have interconvertibility for the purposes of doing our work, which is comparing everything to cash transfers. So that's the point of having the conversion.

But I think that basically what Michael and Sarah are presenting is an indication that Possibly a fully coherent conversion is out of reach, and consequently, that's why we've done this native units thing. The idea is, like, okay, if you measure… if you conclude that depression and anxiety are sort of, like, out of step with disability weights, then we just sort of decide, okay, from now on, we use a different measurement for depression and anxiety, and we decide that we want interconvertibility For the purposes of ranking different interventions against each other across different cause areas, so… for E.G. for the purpose of ranking depression against something that operates on a life-saving margin. So that's just a clarification in general.

### [01:46:03] Samuel Dupret

Yeah, yeah, I guess that… that's… that's… That would be your approach, like, if we can't use this, we use this. We just… we just make a, like, we're just gonna use well-being, because that's what we do. And we have, like, other philosophical reasons why, so I'd add to that that I imagine that based on the other examples I gave of, strangeness with the DALI, I imagine that things like extreme pain might be strange, as well, like, with the cancer treated and not treated, so, like, palliative care, I would not use the DALI.

And, probably to what the, Jemison was mentioning, like, menstrual health or… or blindness, if… if… if these… if these things are not considered as health, and so they get low daily weights, but they still seem to affect well-being more largely, like, I imagine those would also be a case of something where you… you want to… Use a metric that will capture all of their facts. and we… I'll be very brief with this. We did some modeling, some linear modeling, I moderated between mental and physical. It's not statistically significant, but I think it's… It seems like it would be an… a real, like, difference in cluster if we had more data, you can do some fun Bayesian stuff, where you can be like, I don't have a lot of data, but I can put some priors on this to constrain things.

But seems like… Yeah, you'd… if you have a mental health DALI, it should produce more… you should adjust it so it produces more wellbees in your conversion. You could also force it to go through the origin with, like, a different type of modeling, So, if you will… worried about a linear model being too simple here, you could do these things. Happy to discuss this more, but I think the call here was giving some framing, saying how things change a lot, especially for mental health.

These are probably issues with… the delis, using the general population, making health… Comparisons, and so that's likely to be affected by affecting effective forecasting, so people don't necessarily know, really, how a condition they don't have is going to affect their well-being. And we'd be keen to do more in here, especially if… I personally think this could be interesting if it allowed a conversion between all of the global burden of disease things, and that helped people do more analysis in terms of well-being. But also, it seems like these conversion efforts are really limited by not a lot of data, and seemingly some conditions where the DALI just give something very different from well-being. So… Yeah, that's… That's for us.

### [01:49:13] Julian Jamison

Thanks, thanks very much, Sam, and happy… I'll follow your lead, David, on time, but I saw your comment, so people should definitely go out and take a break, you know, if they want, but a little bit of discussion might be useful. I… so if I understand correctly. The standard deviations is like, well, we're just getting… sometimes we have life satisfaction, sometimes we have anxiety and depression, sometimes we have other things. we want to look at well-being, so we convert them to standard deviations, and everybody's kind of on board with that.

But it sounded like maybe with what Matt was saying at the end, the native unit sources is not, that's a little bit more substantive of our, you know. would we ever use DALI's disability weights, or something else, because they're a more Native unit, or because we think that's right for consumption, and well-being is only best for mental health? I think I know HLI's view on this, and you said that, but… so that seems like potentially a point of… Yeah, a real point of comparison of, is it… is it just that it's hard to convert?

All right, well, that's a methodological thing, we don't know exactly what to do, but we try to do what we can. Or is it that we really think, well, wellbies are right for this purpose, something else is right for this other purpose, something else is right for a third purpose? Even if we had all the data. So even if we had life satisfaction data for everything, would we be using it for everything, or would we sometimes be using something else, whether it's DALIs or whatever?

### [01:50:44] Matt Lerner

I think I… That's an interesting question. So, like, if we had disability weights and… my satisfaction… for everything, what would we do? I am inclined to say that we would… I'm trying to say it depends on the details, I guess. Like, so I think, like, Nothing… So, disability rights are also, like, pretty, pretty imperfect, obviously, and, like, there is this funny thing where it's, like. I mean, no… no intervention is measured directly in disability weights, obviously, so… That… so, in that sense, the bar is really low, that if we had well-being measurement direct for everything, we would just use that, because that seems like it would be strictly better than having to… convert… whatever, like… Amputations averted into, like, something that was developed with a discrete choice model.

I guess the only hesitation I have is sort of, like, it's hard to know what that would look like. Like, I think that the level of reliability that I would want for something like that is, like, at the level of, like. Having, like, direct access to people's subjective experience, so it's like… Yo… if we were able to do, like, cancer's letter for every single intervention ever, I'm not sure that that would even be enough for me. I'm not… I'm not sure.

Although I can imagine getting to a point of comfort with that, like… like, I'm given a lot of… yeah, sorry, one thing that I didn't flag, I guess, about… About this whole thing is, like, The fact that… LSPs… Or Welby's, and… Effective… Well-being seem really, really highly correlated, gives me a lot of confidence in general. And that makes me feel a lot more comfortable about using wellbees. But I do still feel like there are… discontinuities that would make me feel uncomfortable with using them across the board, even if the data was available, and I can't totally put a name to what all those are.

### [01:52:57] Julian Jamison

other… Questions, reactions to… To the presentation, or in this… Comparability going back and forth between… Dolly's and Welby's.

### [01:53:09] Peter Hickman (Coefficient Giving)

Yeah, yeah, I have a question about… the… the nice chart that Sam showed, showing that, like, I took a little screenshot of it, so I'm still looking at it, like, strokes dally weight would be 0.17, but… the life satisfaction loss is just .05. And so these physical health conditions, as you said, seem to generally have a much lower effect on well-being than the disability weights imply. Right, that first, first, just quick question, is that that's the correct takeaway, right? Yeah, so then second, I guess a concern would be, is this because people… you know, so I guess the disability weights have this issue, and that people, you know, they're just guessing at what it would be like to be in a certain condition, but then if you're asking people who are actually experiencing the condition, maybe I'd worry that people have kind of gotten used to what life is like, and so they're using the scale differently. Do you think that's a concern here at all?

That, like… I guess, what's… Sam, mainly, what's your reaction to the idea that… These differences are small, because people are just getting used to saying, like, I'm 6 out of 10 on this scale again.

### [01:54:20] Samuel Dupret

I saw a lot of people, like, unmute themselves, so, But, there's two things here, it's like, if people are truly adapting, just like, it doesn't affect… it truly doesn't affect their well-being as much. then I think that's okay. That's just the nature of the thing. If it's that they're changing how they use their scale, and it's no longer representative of their latent well-being, then that's a problem, and that's the sort of thing that Casper and others have been discussing in the previous one.

Another thing is, like, this is… this is one panel data, one analysis, like, it's possible that, you know, the people that answer panel data and that have strokes, like, the people who had, like, the lighter… lightest possible strokes, or they've recovered from them, or… you know, I… that's also, like, I'm not, like, putting all the weight on this, but there seems to be an additional… an additional stone in the building saying, like, Dali's not very good at capturing mental health.

### [01:55:25] Miles Kimball

Well, I… I think that's… there's a big issue here. I mean, I just think neither… neither how good is your life overall, nor the dollies are capturing everything people care about. So, one is more focused on physical health, the other one is more focused on… especially life satisfaction is going to be more focused on mental health.

And so, I mean, there are a lot of other things you could worry about it, but I think the first order thing to worry about is that they're getting at two different things people care about. You need to get the conversion between how much people care about those two things, but the… I think it's a big, big leap to assume that any one measure is comprehensive. I mean, in our… in our work, we've found a lot of evidence that no single approach is comprehensive.

They… They might want to be comprehensive, and they might facially look like they're comprehensive, but they're not.

### [01:56:28] Michael Plant

So maybe I'll… I want to, pick up on that, and just kind of provide some… some context for people who aren't sort of deep in this world. So the… the kind of… the way that… from the thing that Sam showed us, and other things like this, will be, like, people are going about their lives, but they're in some sort of panel, and then you ask them questions about… you ask them about their well-being, and you ask them questions about their lives. So these people are not… these people are not being asked, okay, you're depressed, how bad is depression?

These people are just, you know, just the normal bits of their life is going on. Like, they report that, they also report their life satisfaction, and so this has… has this, and there are other measures of well-being can be. So it has this nice feature where people just sort of say how they're feeling about their life overall.

And then from the statistical analysis, you can work out, well, how important are these things as people are living their lives? And that's actually quite important, because it's not saying… if you ask people. okay, you have this thing, how bad is it? You're drawing their attention to it, and so… So by doing it this way, you're kind of capturing, well, how important does this person, you know, actually, as they're living their life in the wild, not when you sort of force them to focus on a particular thing. When you think about it in the panel, you could ask someone about all these different aspects of their life, you know, being asked about your, are you married, do you have income, you know.

Do you have a park near you? You could ask people to elicit all of those things, but that's, you know, you're going to get different answers for all these things, whereas you can sort of infer how much does this really make a difference to people's lives through the… through the well-being data. And, yeah, I think the… with the well-being stuff, I mean, I'm kind of… I'm sort of a known fan, but I think it has this feature that people can say what's important, you know, they can… they can say how their own lives are going on the basis of their… of their own judgment.

And, and if we're doing other things, we're sort of telling people how important we think their lives are. So that's the kind of… yeah, that's a bit more context on the hows and whys here.

### [01:58:18] Samuel Dupret

And to Miles' point, like, I think the thing here is, A question of, like, are we capturing the one thing that… Overarching thing that matters, are we capturing the true goods? And, I imagine it might be making a case that, like, life satisfaction next to health and other things is, like, one of the determinants of the true goods, and we might say at the Happier Alliance Institute, like, we think well-being is the true good, and… these measures like lifestyle transactional happiness, we think they directly measure that in some way. Like, there's a bit of a… of a debate here, a question here, but, .

### [01:59:06] Miles Kimball

I mean, it's a question of how people answer the survey questions, and basically, you know, we give people trade-offs between the one dimension of well-being and the other, and people will sacrifice any dimension of well-being for other things, so… And when we… when we do those kind of stated preference comparisons, the number one thing is actually the health of your children. It's not life satisfaction. Life satisfaction is actually way down the list.

Now, life satisfaction could be a very excellent proxy for a lot of things, which means in doing certain kinds of statistical analyses, it's going to be great. So if you're getting it in the wild, if you're getting the natural correlations that exist, I mean, basically, and… Honestly, it's a good proxy, partly because it has a… it has a very high loading. You can do the first principal component of thousands of aspects of well-being, which we did, and life satisfaction has a particularly high loading on the first principle component.

But actually, that high loading means, in some sense, it's a luxury. It's something that people who are well off care about especially much. But precisely because it goes up and down extra much. It's… it can be an excellent proxy in certain contexts, but if you ask people directly to say, would I rather have a little more life satisfaction or something else, life satisfaction itself is pretty far down the list.

And so, life satisfaction Might be related to other things in a way that it sometimes looks like a comprehensive measure, but it itself isn't a comprehensive measure of well-being.

### [02:00:59] Michael Plant

What? I just would… No, Casper had his hand up, sorry, Casper, go on.

### [02:01:03] Miles Kimball

It's a matter of how people interpret the survey question. I mean, you could wish that people would answer the survey question as if you'd asked them, what's your overall revealed preference utility in the economist's terms, and that'd be great if they did, but they don't. Because if they did, they wouldn't say they'd be willing to sacrifice it for other things, because that would make no sense.

### [02:01:27] Caspar Kaiser

let me just make one distinction here, and that is, it seems like Miles, and, you know, that's… that's one tradition, is to take what is valuable to be the thing that people, possibly idealized individuals, would be choosing, and that is the thing that is good for a person. But that's not the only view that one could have on what is good for a person. One could be a hardliner here and say, what is good for a person is what makes them satisfied with their life, independently of whether they would trade it off with some other thing.

Now, I don't believe that this is the correct view, but I just want to make the distinction between Well-being is the thing that possibly idealized individuals would prefer, and well-being is the thing that… philosophers, moral philosophers, would view as the thing that matters. The two might coincide, but they might not. And then, you know, it might turn out that life satisfaction coincides with the thing that philosophers correctly identify as the thing that matters. I don't think that will turn out to be the case, but that's another question.

### [02:02:34] Miles Kimball

Well, that's not…

### [02:02:35] Caspar Kaiser

How to make these distinctions.

### [02:02:36] Miles Kimball

That's not actually where the philosophers go, either. It is where many, many social scientists go, is to say that life satisfaction, per se, should be… should be, ethically, what we maximize. I don't think that's the view of folks working in philosophy departments, and it's true that I'm kind of assuming a view, a deferential view, of let's get people more of what they want. Let's make their lives better in whatever way they want. Let's make people's lives better in whatever way they want to make their lives better.

### [02:03:17] Julian Jamison

Let's… let's make sure, I'm sure we don't have a lot of time, but so Matt and Dan as well, please.

### [02:03:24] Matt Lerner

Absolutely. Thanks. Sorry, I could have sworn… same… I'm sorry, okay, sorry, I'll go. I… just as, like, a clarifying point about our use of this metric, like, one way that I tend to think about This question is just that well-being is a residual for a lot of our work.

So, like, I'm sort of, like… kind of had this conversation with a lot of you, but I'm, like. I'm on record a lot as being really against, institutionalizing well-being metrics at the current stage, and, like, when people… particularly in this context where people are like, we shouldn't measure GDP, we should just measure, like, national well-being, like, I really hate that. Sorry, everybody.

But the way that I think about it is, like, GDP is really good, and well-being is also pretty good, and, like, to the extent that there's stuff that GDP doesn't capture, I think that well-being measurement does a good job of capturing that residual, like, on all the… on all these GDP versus X charts that you… that you see floating around. And I think that that is… that intuition is basically carried through to the way we do conversions internally. So… so, like, right now, you know, one will be… We have at, you know, our conversions have this at half an income doubling, or a quarter of a DALI. equivalent.

And… I think what that basically cashes out to is… You have to be doing, like, a lot of aggregating of, like, well-being-specific Benefits to come out to… what Miles was gesturing at, which is, like, saving an infant's life, or something like that. So I do think that, like, the… The relative importance of… Things that are captured, like, more conventionally in disability weights, health that were alive, saved. at least as of right now, comes through in the conversions that fall out of, you know, HLI's research, for instance, on this. That is sort of priced in.

### [02:05:27] Daniel Benjamin

I'm… I, I, I, hope I… Matt, we can convince you that, Measuring national well-being is a worthy endeavor. Either at this conference or some other time. I was, I was just gonna react to what, for an, was saying about the life satisfaction as the goal, view, and philosophy, and… Or among… that one could take, philosophically, and just… Point that even if that were your view, that that's the ultimate good, that doesn't necessarily mean that the life satisfaction survey questions that we ask, capturing the concept philosophers would say is the ultimate good, and I think, in fact.

We have quite good evidence that it does not, and at the very least. There's a lot of heterogeneity across people in the way they interpret the question. which… Would seem to be inconsistent with it being, you know, for everyone, a measure of… of that con…

### [02:06:47] Ori Heffetz

Dan, you're muted.

### [02:06:50] Julian Jamison

He's gone on mute again.

### [02:06:51] Daniel Benjamin

I finished.

### [02:06:52] Julian Jamison

Okay, alright, you were, you were mostly there. Great. I'm… any… any last… thanks, everyone.

This is… has been really good. I am… I am cognizant of time, and David's being nice, but any last thoughts or reactions, or should we go to the next one? And thanks… thanks, Sam, especially for the presentation.

### [02:07:13] David Reinstein (The Unjournal)

Yeah, thank you. I suggest we take a little bit of an enforced Break, and sort of just go off mic, and then, because we've gone way over time, and the next session, Benjamin's presentation. is, is scheduled to begin at 1.35 PM. Obviously, you're welcome to continue any asynchronous, or, you know, any chat conversations that you like. Julian, did you want to have any sort of wrap-up on sort of where we've left things, or…

### [02:07:44] Julian Jamison

And the only thing I would say is. It's been great, and keep adding to the doc. On the survey that you sent around, David, and I did do my version earlier today, there is… one of the questions is, roughly, you know, if you had to come up with a number between dollies and wellbees, what would it be?

And we've heard some of them here, but I do think even if we… some of us disagree about the scales and which is optimal. It forces you to think think through that, and I would be curious to see the distribution of numbers from people here. So if you haven't done that, it'd be great to see that at some point.

### [02:08:20] David Reinstein (The Unjournal)

Just, just to… just to just… I'm sharing the, the screen right now, just what he's talking about is… You know, I… we've tried to… Operationalize these to try to get people to… force themselves to think about this and put themselves on paper in a way that might actually be useful, and as you see, there's Pivotal Question 2 right here. You know, what method should one use if you had to choose a linear method. What, you know, where would you put that at, And then there's some more detailed questions if you want to dig into them.

And, you know, I'd originally thought, let's not try not to anchor people, but I think it's important enough and useful enough for this conversation to share the results of this, so I will, you know, endeavor to share the results of this with everyone. participating, you know, so if you fill out the survey, or don't fill out the survey, but it'd be nice if you did. I'll… we will summarize the results, you can hear what other people thought about those questions.

Anyways, thanks for… thanks, Julian, and… and Sam. Sam, let me know if it's possible to make your slides shareable, they were very helpful, and I'll see y'all in about 15 minutes or so, 20 minutes. We're just gonna get started.

Now, I was a little unclear about whether we were starting at 25 past or 35 past. So, apologies for that. I don't know if anyone is here… hearing me, etc, but I'm just gonna start yapping.

So, I just want to give a little bit of a… If anyone's here, can you just, make me feel like I'm not speaking into the void?

### [02:24:47] Daniel Benjamin

I knew you.

### [02:24:47] Peter Hickman (Coefficient Giving)

Whoa.

### [02:24:48] Joel McGuire

You're not in the void. Okay, good.

### [02:24:51] Caspar Kaiser

Cheer.

### [02:24:52] David Reinstein (The Unjournal)

Good. Okay, so let me… I just… wanted to, just briefly, because I'm not sure everyone's, familiar with it, and there's a bit of a context to this, just give a very quick presentation on what the Young Journal is, and, how it connects to this Pivotal Questions project, and the evaluation of, Benjamin's paper. So I haven't… I have other slides, I haven't really prepared them here, but… Specifically for this.

But, basically, oh, there's nothing there, Basically, we did… alright, so what is the unjournal? The unjournal, I'll just… Jump into some quick slides here, the unjournal is… A nonprofit that evaluates… commissions the evaluation of research, both trying to bring more impact to research and more rigor to impactful research, trying to bring those two things together, just like we're doing here, trying to bring together the people making decisions based on the research and the people doing the research to try to get a sense of, you know. To, to try to, Orient it, better align it with what has value, and also to bring a better approach to research evaluation and the research conversation than the journal process, which you can see there's a lot of ineffici… you know, you can imagine, and I think we've all heard about a lot of inefficiencies in that.

So, we generally prioritize research based… that's either submitted, or someone on our team has suggested. Generally falling into certain categories, generally involving social science, economics, and measurement. I'm sharing the wrong screen right now. I thought I was sharing slides with you, sorry about that. You're just seeing a black screen, sorry about that.

Well, this is my typical spiel, I won't go into the detail. But yeah, so trying to address some of the problems with the academic research system. Career rewards, particularly in economics, or focused on economics, let's say, to try to make it more useful, efficient, waste less time, permit better formats of research by basically We simply… prioritize the research for impact. People can submit it.

We look… we have a manager who contacts evaluators who are experts in that area. And, then they write reports, they also give specific, evaluations, they also give specific ratings. It's all public.

They're compensated for their time. The author gets a chance to respond. We try to make it as visible as possible, and ultimately, the idea is that evaluating publicly hosted research has a lot of advantages and maybe is a more meaningful system than than the sort of system of journals that's evolved over, I don't know, a century or so.

And now, getting back to… so, we were doing that for a while, but we've also wanted to do things that… particularly aligned with impact and measurable impact, so that's why we put forward this Pivotal Questions project. And, Basically, the, the Pivotal Questions Project, involved us reaching out to organizations… here's a post about it, for instance, on the EA forum. To organizations that we thought were doing work focused on global impact, and asking them, what are your most important questions?

And trying to operationalize those questions. And as I've mentioned before, we've… we had some great back and forth with… with, with Matt Lerner, and… This is one of the three projects we're working on. The two others have a little bit more to do with animal welfare at the moment.

And, you know, we ask these questions to a group of organizations, and we focused on these three, and then we've operationalized them. The idea is we would then focus our evaluations and direct assessment… On research that informs these important questions, and also asks for direct assessment of those Of those target, questions. And if you see, I think I mentioned it very briefly before, what we're asking the, pivotal questions, evaluators… sorry, don't look at that. to do… It's a multi-step process.

We'll get to that perhaps later. But one of the papers that came out of this was Daniel Benjamin's paper. As… so if we look at the pivotal question, the well-being pivotal question here.

As we're posing it for the evaluators. this is a public pla- page. You have a different version of it on the webpage.

These were the ones that… and Julian and Martha Weiss Wisembart helped me… helped us come up and refine these questions. So we refined these in terms of what is the… goal-oriented question. What is the most important, or the two most important, sort of focal question.

And how can we define it in a way that would sort of be agreed upon, you know, verifiable? But… So, we posed those questions, and then Martha and Julian also helped me scope research having to do with that question, which includes of course, you know, much… some of the research that's been discussed here with people in attendance here, of course. And one of those papers that seemed particularly important in the discussion that we've had is the Benjamin paper.

And we've had two… alright, I still have four minutes, good. So we've had two, evaluations of the Benjamin paper, which is still, I think, you know, a work in progress. It's not… that's… that's… not speaking highly enough of it, but it's… it's still… there's work that's not available on the public version here, which I think Dan will be presenting about.

And you can see our evaluations. The two evaluations were asked to consider the paper as it's posted, as well as some of the material from Dan's, and his colleagues' presentation reporting on their newer work, and to provide feedback on it, and rating on it, rate it, allow them to be able to respond, and hopefully, when the new version comes out, an updated evaluation, again, focused a little bit on the pivotal questions, and a lot also on impact. So you can see the, and, Valentin helped me manage this, this evaluation. you can see what the work… what the output of our evaluation looks like here. I call this interim, since it's not the In some sense, not the final version of the paper, and we want to come back to it.

This is the document that that, governs the other documents. It has a DOI, etc. We try to make this as visible as possible. we allow the evaluators to be anonymous. In this case, they both chose to leave their names along with their evaluation, so, you know, you can see and re… you can read Casper's evaluation and his ratings here.

We also ask for claim elicitation, and the idea is that this is supposed to be of the quality of… and of the Of the quality and, of a top… Peer review, top… Referee report, but… Even better in some ways, because it's focused on improving the paper, it's focused on impact, we have these quality ratings, we're really… we're compensating people, we're giving them feedback and encouraging them to give us a little bit more, and, you know, interact anonymously or non-anonymously with the authors. And, you know, whereas a typical referee report might just be all about should the paper be accepted or rejected. So we're hoping this adds value, and after Dan's presentation, Casper is going to talk about his experience and some of the comments, and maybe we'll have some back and forth.

And I'm also… I also mentioned one point raised by the other evaluator who wanted to join, but he's on. maternity leave, I think some of the quest… one… the first question he raised is particularly relevant to this discussion. Okay, that's about it for now. Valentin, did you want to say anything before we introduce Dan to give his presentation?

### [02:33:32] Valentin Klotzbücher

Yeah, hey, you know, I'm not sure if I really have much to add. I think it's going really well. I'm, excited to… so, I don't know, when I was saying yes to be a valuation manager for this paper, I was, immediately excited because it's a, yeah. forward-looking, way of dealing with the questions here, and yeah, I'm all the more excited to hear now even more today.

### [02:33:56] David Reinstein (The Unjournal)

And I just really appreciate that the authors are really keen. I mean, I think we all appreciate that the authors are really eager to make this useful and make it used. So, anyways, without further ado, why don't I pass the mic over to Daniel, who's gonna, I think, share his screen… And yeah, I'll, I'll, I'll leave it, there. Daniel, are you here yet?

### [02:34:25] Daniel Benjamin

Yes, can you hear me?

### [02:34:27] David Reinstein (The Unjournal)

Excellent. Very well.

### [02:34:29] Daniel Benjamin

Great. Can you see my screen?

### [02:34:33] David Reinstein (The Unjournal)

I can't, but I… it might still be up, I'm just trying to… To check that. Can anyone see your screen? Yes.

Okay, it is, I just have to switch.

### [02:34:42] Daniel Benjamin

Okay, great. So, so first of all, thank you for, for having us. This has been a really interesting… I mean, I'm learning a lot from all of this discussion, and I'm eager, you know, to hear the rest of the conference as well.

And, we're very grateful to, to David, Valentin, the UnJournal, and, Casper and Alberti Prato, who were the, Alberto Pratti, who were the, evaluators for the paper. They had great comments in it, and great to hear the feedback. This paper's been referred to a few times as Daniel Benjamin's paper. I want to be clear, and this is joint work with Kristen Cooper, Ori Hepatz, Myles Kimball, and Jinan Zhao, the first three of whom are, I believe, also on this call, and… Could jump in to answer questions and reply to chats.

And and as David said, the paper is… undergoing a very substantial revision relative to the version we had previously made public. I'll try to highlight Some of the key ways in which that's true. With only 30 minutes, I can't go into a lot of the details of the paper, so I'm going to try to focus this presentation on the main questions that David asked me to address.

So, you know, the topic of the paper is heterogeneity and scale use, so that is, you know, if, say, it's a question about how satisfied you are with your life, maybe David says 70, and Valentin says 80, but maybe they you know, and so we would think Valencia is more satisfied with his life, but maybe those two numbers mean the same thing because they're using the scales differently. But that's the problem that we want to address. So… I'm going to show you the kinds of survey questions that, that we ask to study this question first, and one of the ways in which this new version of the paper is, I think, improved over the previous version is that… In the previous version, we had collected an online mTurk sample. In the new version, we've collected data from the Understanding America study, which is, study that aims to be nationally representative of the U.S. population.

It's an internet panel. About 10,000 people. So… The question… we actually study a number of well-being questions, but the one I'm going to focus on for this talk is how satisf… over the past year, on average, how would you have rated how satisfied you are with your life? I'm showing you the screenshot from the survey itself.

The scale is from 0 to 100. It's a slider. And we have the level, labels, lowest possible level, highest possible level. You can move the slider, or you can type in a number, and then you press next to move on to the next screen.

The… The fundamental problem that we face with scale use heterogeneity is that people differ from each other, both in what they're… True situation is. But also how they translate that situation to the response. And we can't disentangle those two things.

And so, we need an empirical tool to solve that identification problem, and that empirical tool is what we call a calibration question. So it's designed to measure that mapping from situation to response. So, one kind of calibration question that, that has been explored in one branch of the literature. and we examine it as well, is called a vignette.

But here's an example of a vignette for life satisfaction. There's a little story about a person. And then the question is. If this situation described your life during the past year, on average, how would you have rated how satisfied you are with your life?

We asked… three such calibration, three such vignettes, so that I'm just listing them here. I don't expect you to read them. But they are… One of them is designed to be a kind of low level of life satisfaction, one a medium level, and one a high level, so we get some… some variation. Hmm. as Casper said at some earlier point, what you actually need… the minimum you need for the method really is two such questions, but, but we think 3 or more is gonna be more, is gonna generally perform better.

And let you do more, more interesting kinds of analysis. We are also interested in exploring the possibility of Correcting for scale use heterogeneity using Calibration questions that are about a different dimension than the one you're interested in. So, you know, for example, you know, having vignettes about happiness, or having a totally different kind of calibration question, like this one. How dark is this circle? You rate that on a 0 to 100 scale.

We do, in fact, have 3 such questions like that. We call these visual calibration questions. And… You know, the reason why you might be interested in what we call cross-scale scale use correction… well, first of all, it may be that that's what you have in your data, you don't have calibration questions for the dimension you're interested in.

But, you know, there are also reasons why you might be… Worried about biases and the way people answer particular vignettes, and so… so… Cross-scale Scale used to mention could be worth exploring in that case. Oops. Although, our view is that Same dimension scale use correction, if you can… if you can rely on the assumptions. Is… is the best way to go, because you could be most sure that the scale… we're more… it's more plausible that the scale use is the same, or very similar, for the way you rate calibration questions in a dimension as to the way you rate yourself in that dimension.

Okay, so now what I want to do is show you some evidence, just some simple descriptive evidence that scale use heterogeneity is, in fact, present. So, what I'm showing you here is a plot where each point is an individual, so there's about 10,000 points. The x-axis is the mean response that a respondent gave to the three Life satisfaction vignettes.

And then the y-axis is their answer to the… Life satisfaction question. The red line is the… Best fitting regression line. And the point here is that the… The line is upward sloping.

The… so that… that means that people who are giving higher answers on average to the vignettes are also reporting higher life satisfaction. The correlation is 0.16, that may seem low, but you need to remember that, Both the x-axis and the y-axis variables are measured with error, so the correlation of the true underlying thing is definitely higher than that. So that suggests that CL use… that differences across people in the levels at which they rate things is mattering for their answering life satisfaction. What I'm showing you in this plot is… well, I should say, we actually re-surveied about 2,500 people a few months after the first survey on the Understanding America study.

And so, for that… that subsample, we can look at the… how their ratings of the vignettes changed, and that's what I'm showing on the x-axis, and on the y-axis, how their rating of their own life satisfaction changed. And those changes are also positively correlated with each other. And the reason why that's interesting is because, to the extent that economists deal with the issue of scale-use heterogeneity at all. It's by including… it's by studying panel data, which is data where you have repeated observations of individuals over time, and including a fixed effect for the individual, which is basically taking out, you know, the average Level that a given individual gives to a question over time.

But that only… works if the scale use of a person is fixed over time. And what this is showing is that actually the scale use is changing over time. For individuals, and those changes in scale use, or what it suggests is those changes in scale use are associated with changes in what they're reporting their life satisfaction to be.

There's actually another reason why fixed effects are not a complete solution to the scale use problem, and that is that it's… the issue is not just that people use different levels of the scale, different average numbers, but also that some people use more of the scale and some use less, and fixed effects don't do anything about that. That, okay, so then this, this next figure… is, is related to the issue of cross-dimension scale use correction. So here I'm showing on the x-axis the mean rating that's given to a number of these visual calibration questions, and on the y-axis, the answer to life satisfaction.

Again, there's a positive correlation. It's a bit smaller, or somewhat smaller here. So, that suggests that there is… Some transferability across dimensions to… to scale use.

Although, the lower correlation here is consistent with this idea that the scale use is more similar within a dimension. It becomes attenuated as you move to other dimensions. And then this last picture I want to show you is related to this point that people use different amounts of the scale. So… what I'm… what I'm plotting here on the x-axis is instead of the average response to a group of calibration questions, I'm giving the standard deviation of responses, so the spread in the answers. I can't calculate that, spread for a single question, like life satisfaction, but if I use a few different dimensions, like life satisfaction and happiness and, you know, not being anxious, then I can calculate a standard deviation of the Across those dimensions, and that's what's on the y-axis.

And here again, we see a positive correlation, and that again suggests that variation in how much of the scale people use is related to Their answers to the… Self-reported well-being questions. Now, a reason why all of these quest… all of these figures that I've showed you are just suggestive, they don't really nail the point that scale use matters, is because it could be That people… the differences in people's scale use is correlated with actual, true differences in their well-being. And we, you know, we can't rule that out on what I showed you so far, and we can't rule that out with purely subjective dimensions, like life satisfaction, but we can… Get at that if we look at some dimensions where we also have objective measures.

And so we… we studied height and weight. As two such dimensions. And with those dimensions, we asked people both to rate their height and their weight on a 0 to 100 scale.

We also asked them, in objective units, what their height and their weight is. And so what I'm showing you here are two regressions in the two columns. The first one is a regression for height.

And we have 3 vignettes for height. These are pictures where people are asked to rate the picture on a 0 to 100 scale. And then the control variables are a, a flexible function of the Your height in objective units.

And the numbers in the table are regression coefficients with standard errors. The point here is that two of these three vignettes have positive and, you know, clearly quite different from… statistically different from zero coefficients, indicating that… that controlling for your actual height the ratings that you give to the vignettes is predictive of how you rate your own height on the 0 to 100 scale. And we see a similar thing for weight, where all three of the vignettes we have about weight, positive.

But we think that's really, you know. Pretty powerful evidence that scale use heterogeneity really is driving the differences in the self-reports. Okay, so what I want to show you now is some evidence on the… On the scale use heterogeneity itself.

And… So, so I'm gonna show you a few figures, like this one. This figure… Each point in this figure is a calibration question, and there's a number of different dimensions, they're labeled, they're different colors. and I fitted a regression line, or we fitted a regression line through The points, there's actually 95% confidence intervals at every point, although they're pretty tight, so they're hard to see.

And the dotted line is the identity line. And… In this figure, the x-axis is the mean response to that calibration question. For people who have below median income in our sample, and the y-axis is the… mean response for people who have above median income. So we're basically comparing how people with lower income and higher income rate the same calibration questions. I want to make a few observations about this figure.

First is that all of the dimensions, you see positive relationships. What that means is that the… Lower income and higher-income people are ordering the vignettes in the same way. The second point is that, Most of these lines are… Pretty distinct from the identity line. Which suggests that there are scale use differences. If there were no scale use differences, then they should lie on the identity line.

The lower income and higher income people should be giving the same average responses. A third point is that there's some variation across dimensions, so it's not that the scale use is identical. That's an argument, again, in favor of using same dimension scale use correction, if you can.

And the fourth observation is that All of these, lines… all of the points, actually. Seems to be pretty well approximated by the lines. I'm… That's, So now, I'm showing you… I'm going to show you that those… all of those observations generalize. It's not just about… Lower and higher income. Here's a similar figure, but for white versus black respondents.

Okay, and actually, we have many such figures like that across many different kinds of splits, and it turns out it's not just you know, these observations hold not just in our data, not just in the Understanding America Study data, they held in our mTurk data, they also hold in other people's data. We've gone back in the literature and found other people who've studied vignettes and plotted their data, and they look similar. So, so one takeaway from… from this is that we're gonna… we're gonna assume… In what follows, that the… Translation function, the difference… the way one person's scale use translates to another person's scale use is linear.

And we don't… actually don't have to assume that all our methods can all be generalized to deal with nonlinear scale use, but we think that the linear… linearity assumption, I think what David called an affine translation function, in the notes to this conference is a good approximation, and it holds at the individual level as well. Okay, so… I want to just briefly tell you what the correction is that we're proposing you do. Like, if you have the calibration questions, what do you do with it?

And… The, the… The point here is gonna be that it's actually very simple. So… and transparent. So what we do, Is we first want to estimate each individual's A translation function across the, across the calibration questions. Dope.

So, the first question is translation brought to whose scale? We have to have somebody's scale as the, kind of. Common scale that we're gonna use for everybody.

And what we use as the common scale is the sample mean, but it doesn't matter. Turns out with a… if you have linear translation functions, you would get the same results in terms of things like ratios of regression coefficients, no matter who you picked. The sample mean is statistically convenient.

And so what we need to do then is estimate this equation, this equation RIC, this is respondent I's response to calibration question C, is some linear function of WC, where WC is the common common scale rating of that calibration question. So we need to estimate for each individual these two coefficients, an intercept A and a stretcher B, that govern how high or low their use of scales are and how Much of the scale they use, how spread out it is. Okay, so step one is we need to est… to estimate I'm… A and B for each person.

First, we need an estimate of the WC. But we do that in step one by calculating the average response to the calibration question in the sample. By definition, that is what we mean by the common scale rating for that calibration question, and that's going to be estimated very precisely. So I'm going to think of that as now a known quantity.

And then we can use that and run a regression of an individual… for each individual, we can regress Their response to the calibration questions. This would be a, you know, a regression, say, with 3 observations. If I have 3 life satisfaction calibration questions, I'm running a regression of my responses to those 3 questions. On the average, the sample mean response.

And that gives me an estimate of… of A and B. We get that for each person. Now, you might think that because we have an individual level measure of the scale use parameters, we could then correct for scale use at the individual level.

And you could, in fact. To make the common scale rating, then, of my life satisfaction by taking… there's this W star is gonna… is the notation I'm using for, like. my answer to the self-report question I care about, like the life satisfaction question, I take my answer that I gave, like a 80, that's the R star, I subtract my A estimate and divide all of that by my B estimate, and that would… that would be inverting the translation function to give me an estimate of my life satisfaction on the common scale. That would work if we had a large number of calibration questions.

But the problem is, with a small number of calibration questions, our estimates of A and B are noisy. And that means that if you try to do correction at the individual level, you would get… crazy numbers. It would be very, very unstable. You'd often be dividing by numbers close to zero.

So, we've instead developed some alternative Econometric procedures to adjust regressions. I'm going to tell you about one of them, which is the simplest. It's called the Method of Moments Estimator. It uses the estimates of A and B at the individual level, but it doesn't… but it uses them to adjust the regression coefficients you get from regression, rather than trying to get individual level scale use correction.

So, here's how it works. For each person. I take my estimates of A and B to get my estimated translation function, which is… which is this. It's a function of whatever W I put into it.

And then I evaluate that function, the sample mean for the self-report question. So, for life satisfaction, I calculate the sample mean life satisfaction, say it's… it's 85, and I calculate Oh, I would translate 85. And we call that the scale use only benchmark, because it is just capturing differences across people in their scale use.

We're entering the same 85 number, or the same R-bar for everybody. And, And the only thing that's causing this scale use only benchmark to differ across people is what their A-hats and their B-hats are, is about their scale use. So it's a measure of the scale use differences. and then I can run a regression. Of that scale use only benchmark, on… whatever covariates I have in the regression, so income, age, marital status, whatever, And… What we show is that that gives us an adjustment The vector of regression coefficients is the vector of adjustments that you make to a naive regression you would have run if you had just regressed life satisfaction on Income, age, marital status, and so on.

And… You need to take those coefficients and adjust them by The coefficients you get from this regression of the scale use only benchmark on X, and that gives you That gives you the right, you know, scale use adjusted, that's… Let me… let me illustrate that. What I'm showing you here are a set of regressions that are… Regressions of a scale-use-only benchmark on a set of covariates. Each column is a regression. So the… let me focus on the first column, which is the life satisfaction example we've been talking about.

But here, the variable on the left-hand side that… the… Independent… the dependent variable in the regression is this scale use only benchmark for life satisfaction. So it's a measure of how the individuals differ in how they use the scale when answering the life satisfaction vignettes. And then we're regressing that on… On each of these, covariates?

And the coefficients that we get Are telling us how much those covariates are… how they correlate with the scale use. So, de-meaned age negative 0.98 means that, that… For someone who is 10 years older, They give .98 Lower rating to the vignettes. Okay, and so… Now I'm showing you a second regression. This regression, well, actually, it's a number of regressions depending on the column. Let's focus on the first column.

unadjusted. This is the regression that People run all the time in academic papers, and that you might run if you were doing a policy evaluation with life satisfaction. Unadjusted, so it's not dealing with scale use, it's just regressing the raw reported life satisfaction on the covariates.

And… you know, you see… we're finding patterns that are pretty typical in the literature. So, for example, higher income is associated with higher life satisfaction, being unemployed is associated with lower life satisfaction, being married is associated with higher life satisfaction. I want to compare this first column to the fifth column. which is the same dimension, scale use corrected estimates using method of moments that I just have been talking about.

These are the coefficients from the regression after you do scale use correction. How do we get these coefficients? Well, let's look at this first row, D-meaned age.

The unadjusted coefficient is 2.49, We go back to this regression, the coefficient. In the first column on de-meaned age was negative .98. So what we need to do is subtract the negative .98 from 2.49, And that gives us 3.46. So it's very simple, so that's how we get each of the coefficients in each of the rows. You can see exactly where the correction is coming from, and how it's being driven by the relationship between the covariates and people's use of the scale.

Okay, so how do we know that, this scale use correction is giving us… Oh, yeah.

### [03:01:36] Ori Heffetz

just for one second, this is what I talked about before. If you look at the first column, the ratio between unemployed and low income is something like 1.6, if Eyeballing it, and then if you look in column 5, The ratio between unemployed and income. He jumps up to… Whatever it is. Several times that, so… These are the types of things that I was talking about before.

### [03:02:06] Daniel Benjamin

Yeah, Miles mentioned earlier that it matters a lot for income in our data, and that you can see that here, and so it's gonna… so, at least based on our data. This kind of scale use correction would matter a lot for any kind of pricing exercise that you wanted to do that involved dividing by the coefficient on income. But for… I mean, for other things, it may not matter, much, you know, I think that's going to be an empirical question. So how do we know that, that this, Chronic correction is giving us a better answer.

Well, to try to get some evidence on that, we come back to these dimensions where we have some objective benchmarks, the height and weight. And, and so… I'll explain a little bit of this table. The bottom row of this table, this .49 number under… in the height column, That's the… The… correlation. between… the height measured on an objective scale, feet and inches, with the 0 to 100 subjective rating of height, where we've standardized both of them. I mean, run a regression, essentially, where standardized measures gives you a correlation.

So it's .49. And I want to compare that with the number, directly above it, same dimension CQs. So this is what happens to that correlation. If instead of using the raw 0-100 response. to height and weight.

We instead use the responses that we get after scale use correction. It's slightly more complicated than that, because to do the standardization, we also need to use a method to correct the standard deviation of the calibration question responses, but we… for scale use differences, but we do that as well. And you see that in both cases, these correlations go up. So for height from 0.49 to 0.6, for weight from 0.23 to 0.51, I'm… So, you know, you may not… there may be reasons why they don't go all the way up to 1, even if you had perfect scale use correction, because maybe what you're rating on a subjective scale is actually something different than the objective reality, but the fact that they go up means that there's a closer correspondence. Between the subjective ratings and the objective measures.

After you do the scale use correction. Okay, oh, actually, I didn't mean to end with practical recommendations. I'm going to end… David asked me to end with limitations and open questions, so I want to say two things, two things about that. First is, what we're doing with our exercise, is converting Everyone's answers to the life satisfaction question to a common scale.

But, Even on a common scale, it's possible that… that if you could somehow measure the deep, true, underlying life satisfaction, you know, my, you know, David's common scale 80 is different than Valentin's common scale 80. I think that's… that's a fundamental philosophical limitation of any kind of self-report data. It's… it also applies equally to utilities. You know, there's… we don't have… Yeah, you have to make strong philosophical assumptions to compare utility levels across people as well. you know, we could talk about different approaches if you… to that problem, but I think one natural approach is to just treat the common scale level as the one that's normatively relevant for policy analysis.

I think that's a kind of Kind of solution that we use, in other contexts that kind of seems plausible, and in the absence of a better approach. could be pursued. I think the second limitation is that there are some assumptions that you have to make for scale use correction to make sense. I didn't spend time on them in this presentation, but, but we think that You know, it's more likely to hold in some dimensions than others, and in particular.

The assumptions are more plausible. For narrower dimensions. And that are more objective things, so things like, things related to particular aspects of your… of your health, or, you know, having enough Food to eat, or having shelter, you know, those things, we think… the assumptions are more likely to hold than if you ask about something like life satisfaction.

And one reason for that is, with life satisfaction. People could have different… weights that they put on the subdimensions that go into determining life satisfaction. And so there may not be an interpersonally comparable thing underneath, to be doing a kind of, scale use correction on.

And, you know, that's, I think, one reason One of several reasons why we would advocate for a well-being index approach, where you're measuring these more fundamental things and then aggregating them. Okay, so I will stop there.

### [03:07:41] David Reinstein (The Unjournal)

Thank you. I guess, you did say you had one slide. Did you want to just put it up on the screen with the, next steps? Was that, was that, was that something like that?

### [03:07:50] Daniel Benjamin

Oh, no, no, no, I was… I just… I forgot to change the outline.

### [03:07:53] David Reinstein (The Unjournal)

Oh, that's okay, that's okay.

### [03:07:54] Daniel Benjamin

Yeah, yeah, yeah.

### [03:07:55] David Reinstein (The Unjournal)

Terrific, terrific. So I think… I think the thing to do now is… Just, just to, so thank, thank you for, for coming, and thanks for your talk, and I just wanted to just, give an opportunity if people have clarification questions. There isn't much buffer right now, but if people have clarification questions, and just to sort of… Reorient, or move the discussion a little bit towards thinking about How could this be used, you know, in the focal, context? Because You know, it naturally… maybe… maybe sort of some discussion of… from the practitioners of… How they could think about using such an approach. what's feasible about it, what parts they maybe didn't quite get or didn't quite understand could… could be a little bit helpful.

And then we wanted to move on to, to, Casper to, yeah, to give his presentation, or to give his, sorry, his… Discussion, his response, again, focusing to some extent on the question of how relevant is this to this environment where We're running RCT trials in poor countries, you know, and we get these measures, some of which are on life satisfaction scales, and we're, you know, we observe these measures, sometimes there's a difference, you know, a fixed effect, sometimes it's, I think it's just one-off. But there's a randomized group that you're comparing the control and treatment, and then you're sort of comparing that difference to the difference in a… for a different intervention. could potentially use the same scale with another randomized treatment and control.

So, yeah, I just… I just want to, you know, think about how… to what extent your method is appropriate and practical. In that context. But yeah, has anyone written down some questions that they want to ask in terms of clarification, or… Or, you know, that they want to get out of the way now before we, have Casper, Yeah, respond and discuss, and maybe discuss a little bit about the evaluation process.

### [03:10:12] Joel McGuire

I have… I definitely have potential follow-up questions about the practical side of things, but I can hold off on those.

### [03:10:19] David Reinstein (The Unjournal)

Yeah. That makes sense. Okay, well… Let me see, let me… so let me introduce Casper.

So, as… mentioned Casper's… was commissioned by the UN Journal to evaluate this, this research. And it's an evolving research, so we're hoping to have sort of evolving response and evaluation to make it as useful as possible. And, if you look in the notes, you see, you can see a link to, to Casper is actually his full evaluation, and… which I think is rather comprehensive and helpful.

So, Casper, did you want to, share screen, or just come on camera at this point? I can't remember.

### [03:11:03] Caspar Kaiser

Yeah, I have some slides, let me see if I can…

### [03:11:05] David Reinstein (The Unjournal)

Yeah, thank you.

### [03:11:06] Caspar Kaiser

work. This is always a… Struggle. Can you see? Can you see me? I mean, they're obviously way too small, but…

### [03:11:18] David Reinstein (The Unjournal)

No, you can see them well. And by the way, just to say, if you want to gather your questions on either the presentation or Casper's discussion, we'll save some time to, you know, to have some back and forth on that.

### [03:11:29] Caspar Kaiser

Cool. So, people can see these slides, it's gonna be super quick. I assume people do.

So, comments on, adjusting for scarier heterogeneity and self-reported well-being, plus some wider thoughts, but I think many of the What are thoughts? have already been covered in our earlier discussion. So let me just… I mean, Dan has already said it, but here's my take on what Aishi, the… core contribution of this paper is… I think it's… it's twofold. One is that it proposes a kind of new framework for adjusting for scale use heterogeneity, and another one is that it provides some new data.

I think the first contribution is the more important than the second, given that others have Kind of introduced similar, kind of, Adjustment data as well. But in the methodological contribution, and I think this is underappreciated by some, at least, that I've talked to, there are really two core ideas, rather than one. So many people I talk to who are familiar with this paper really think that it's this first thing This idea of introducing calibration questions, understood as questions about kind of… Objective sense percepts, like brightness of a circle, numerosity of triangles. Introducing that idea as the novelty of that… of the paper. Which is true, that's one of them.

And then there's the other novel contribution, and that is a econometric approach to using data of that sort, which may be calibration questions. Or vignettes. I know that Dan has talked about vignettes as a subset of calibration questions, but most people would distinguish between them, like, sort of informal language.

And… Here, the key contribution is to show how one can estimate an or less regression, or fancier things that takes account not only shifts in people's scale use, but also stretches. Like, you know, that we can imagine that people, like, differing, okay, my 7 is kind of… you know, more stringent than Valentin's 7, so that would be a shifter, but we might also imagine that the jump for me between numbers is way larger than for Valentin, and then we would have a stretcher. And introducing this distinction is a new… new part of the methodology. I will not dwell for the moment on these assumptions, there's a bunch of assumptions. Broadly speaking, I am happy with all of these assumptions.

The crux of this entire literature, of whether you're using the old-school vignettes and a slightly different model that I will get to. or the new Benjamin et al. version, is… whether we believe essentially this. Whether we believe that Individuals, especially in the case of vignettes, have a common perception Of the thing that the vignette is being asked about.

So, in the case of a… of the brightness of a circle, You might easily believe that. Right? Because, I mean… There's little reason why, you know, my sense organs Should systematically differ from mild sense organs, or at least conditional and covariates should differ in some… you know, sensible way. So in that case, you know, that seems, like, easy to buy.

But, for such questions, the response consistency assumption, so the assumption that the way I'm using the scale to assess the brightness of a circle is kind of the same as the scale that I'm using to assess my own life satisfaction, now that assumption is kind of hard to argue over, like, is that really true that, you know, my scale use for this, like, brightness of a circle thing would be, you know, comparable to how I judge my life satisfaction? On the flip side, if somebody were to use a life satisfaction vignette, so a canonical example, I don't know if Dan and others are using this, but a canonical example might be something like, oh, here's John, John is 70 years old, he has back pains, but he has, like, a… really great marriage and two loving children. How satisfied do you think? John is.

Now, it seems very plausible that we have response consistency, that the way I'm, like, using the scale when answering this question about John is going to be the same scale as the one that I'm using. When answering questions about my… own life satisfaction. But now we have to worry that the common perception assumption might no longer be holding. For example. suppose that I am on… at a low income.

And now, nothing in living yet, let's suppose for that moment, does it exactly say what this income is and how a life with that income would be like. If I'm on a low income, I will now fill in these aspects, these economic aspects of John's life. And I would do this in a way that is different from a richer person, that will fill in these unmentioned aspects of Germ's life in a different way.

And so under that, in that case, the common perception… assumption would fail. So what I want to say here is. that it seems like there is this trade-off between the common perception assumption and the response consistency assumption.

I think there's ways of getting around that, but I want to flag that this is kind of the crux of the matter. This is not specific to Dan and others' work, it's general to the idea of using calibration questions and vignettes. Okay.

And I think we can get back to this. Let me just very quickly say something on the main empirical results. The version that I read, maybe this was now updated, has… one table where you can read this in two different ways. One way of reading this is to say, well, look, you know, you're doing the vignette and the calibration question adjustment in whichever way you're doing it, and it turns out That the sign and significance and, sort of. roughly qualitative magnitude of these coefficients doesn't change very much.

And you could read this as, like, really good news, at least for some people, maybe people like me, who, you know, spent their days running well-being regressions, because it would tell me, well, you know, it's roughly… roughly robust against worries about scale use differences. But… If you are specifically worried, as many people in this workshop, and parts of my life I'm also specifically worried about. specifically ratios of coefficients, allowing me to make trade-off between different kinds of variables, I guess EconTalk would be marginal rates of substitutions, then the ratios do turn out to vary quite a lot, as As pointed out by Ori.

One thing that I would here be worried about is that, of course, The ratio of two coefficients Can change much more than simply the magnitude of a single coefficient, just as a consequence of, like, weird idiosyncratic measurement error, and, like, sampling issues, and so on. And in general, you know, the expected value of the ratio of two random variables that is not even well-defined, and the ratio of two expected values is… it's not clear whether this is actually the thing we care about, and so there's… there's all these worries about ratios and the fragility of ratios in general that I just want to briefly… Wanna flag. That said, I really, really like the validation exercise, where you essentially show, like, hey, you know, once we are adjusting for these scale use differences. we're actually better at correlating, for example, objective and subjective height and weight. Like, this is… this is, to me, strong evidence that we're improving, in some way, our subjective measurement.

And I think we should, like. You know, make that finding more prominent in some way. I have some suggestions on things that I would want, Dan and colleagues, to do. One is, show… and you've done this now, you didn't do this in the, I think, in the original paper, like, oh, let's vary the calibration questions, do we get the same results?

There's, I think, a statistical test to be had here with a, you know… Things like, common understanding of the, of the, of the calibration questions. hold, but I think more importantly for the purposes of this workshop, is this sensitivity analysis. So… the data that you currently have is, you know, this particular US dataset, but… Increasingly, we're thinking of using well-being data in low- and middle-income countries, and we currently have no idea what's going on there in terms of, like, scale use differences and association of scale use differences with covariates, and specifically. what happens in response to a particular intervention, and I said that earlier on.

I think this is, like, the obvious frontier of where to push this type of work and related type of work. There's a point to be had on, like, I think this is trivial to extend it to the, panel data fixed effects case, but it would be useful to do that. And then, for the practitioners in the room, and more broadly, we need a Stata package, we need an R package, maybe for the nerds, we need a Julia package, and the other nerds, a Python package, that would just be really nice. Final thing to say from me.

As I said, there is this contribution of this paper that is associated with doing things on the stretcher and the shifter, and allowing things to be done on continuous variables using OLS regressions. But there's this older technique, the hierarchical order probate model. Which does the same kind of thing for dis- like, you know, discrete variables with 10 or less response categories. In principle, you could do it with more.

Apparently, things don't work when you do it with 100 response categories. But there's this… there's this other method, which has very similar assumptions. It also has, like, an assumption of, like, response consistency, so that the vignette questions are answered on the same scale as one's own life satisfaction. It also has this vignette equivalence assumption, so that my understanding of a vignette or a calibration question is the same as yours.

And then it has this other assumption of, okay, we must believe that the distribution of our dependent variable, I guess in our case here, life satisfaction, is conditionally normal. the… paper by Ben and others does not make that assumption, but it makes slightly other kind of technical assumptions. And so we are in this weird situation where there's these two techniques that do the same thing And they don't make nested assumptions, so it's not the case that Dan's and others' paper is, like, making strictly weaker assumptions. It's making slightly different assumptions, and that leaves me, at least, and maybe others, in this kind of uncomfortable situation where I don't know which assumptions I should believe, and how they relate to each other, and so on, and I would love if that could be clarified. I have a slide on other problems, but we talked about this in the previous session, so I will leave it here.

And maybe this can open discussion, maybe I should leave it… yeah, maybe this is a good idea to leave the slide open.

### [03:23:59] David Reinstein (The Unjournal)

Yeah, thank you, and we could take a quick look at your other slide, just to sort of recap, could be helpful. That last slide might be kind of interesting. So linearity, the neutral point, and the flow utility issues.

### [03:24:12] Caspar Kaiser

Should we be measuring life certification in the first place? Or whatever it is.

### [03:24:17] David Reinstein (The Unjournal)

Okay, well, let's open up for questions. I know that, Joel, you said you had some particular questions about the applicability of this, and I think related, I mean, Alberto Pratti's, evaluation raised a question which I think would also be relevant here, and I know… maybe you touched on this, you know, essentially, you know, can we use this in a small scale, with small scales, and how? So how well would a small… adding a small number of questions be likely to perform.

I think there was some response to that. You can see, Alberto's full comment as a footnote in the, in the Google Doc. He says, introducing two, Calibration questions can represent a large burden. It's not clear from the working paper how good it would perform with two to three, calibration questions.

But, you know, maybe, maybe, Joel, you want to add to that, and then… and then both speakers could… could respond and discuss a bit, perhaps?

### [03:25:22] Joel McGuire

Yeah, and, well, I guess first a comment on… on burden. Subjective well-being questions are presumably much less burdensome in terms of time than often… like, mental health questionnaires, like the PHQ-9 has 9 questions, whereas if you're asking life satisfaction, it's 1 question, so if you're adding, like, a few more additional calibration questions, so long as those calibration questions aren't themselves incredibly cognitive burdensome. That didn't strike me as something that's, like, necessarily, super, you know, going to be a deal-breaker there.

But in terms of the practicalities, I would, I would be very interested in trying to incorporate this more into our work and how we think of Evaluating different interventions. The… yeah, I guess there's a couple levels, is my… the thing that would probably be most useful, I imagine, is having, something on a general population, where we have… where we identify groups where we need to make this, you know, correction, where it's maybe the biggest correction, and then, you know, generalizing to this specific circumstances. Obviously, there's issues with the generalizability, and so… turning to the generalizability point, I imagine something else that would be useful, and I would… I would love feedback on… on this, is… To try and get… These calibration questions captured more, and… say, like, RCT settings, so, like. Ideally, we have it all in context, and, you know, maybe, maybe there's some way in which, at least We could trial these in pre-post settings, in, say, the actual charity contexts, because often they have varying degrees, but sometimes somewhat sophisticated, monitoring and evaluation processes, where, say, for depression charity, they'll… They'll do a screening exercise before and after. Sometimes they ask well-being questions, mostly they ask mental health questions, but yeah, it's like, you know, I would love to be like, here are the… here are the calibration questions you add, and then, you know, I could work on maybe, you know, piloting them in a specific charity that's… you know, actually, you know, running these things, and maybe, maybe that could help us with the generalizability.

So, I'll leave it… I'll leave it there, but those are kind of, like, some points in terms of what I'm thinking of, how I could actually use these and connect this to perhaps maybe making, Our evaluation's a bit more reliable.

### [03:28:04] Daniel Benjamin

Could I respond to a few, a few of the things?

### [03:28:07] David Reinstein (The Unjournal)

I was gonna say, you might want to respond to some of the direct questions, exactly.

### [03:28:10] Daniel Benjamin

Yeah, sure. So, first, and on the point Joel is raising, and Casper raised about, about generalizability in… in… Lower and middle-income countries. We do not have data on that, but, But calibration questions have been asked in such samples.

So, for example, the paper by Gary King and co-authors that we're building on. have data, report data from China and Mexico, where they have, vignettes And, We've analyzed their data, and the results, you know, things like the linearity of the translation function across groups seem to be… seem to hold in those data. But I, I mean, obviously, it'd be great to collect this data more broadly and continue to study generalizability. on… I wanted to address a couple of the points that Casper raised. One is on this trade-off.

Between response consistency, And… it was… I forgot quite how you… how you phrased it. It was response consistency versus…

### [03:29:26] Caspar Kaiser

So I call it vignette equivalence, but you have a different name for the assumption.

### [03:29:29] Daniel Benjamin

Yeah, yeah, yeah, okay. I can't, I… Common percent.

### [03:29:32] Caspar Kaiser

Yes, in response to a real concern.

### [03:29:34] Daniel Benjamin

Common perception, okay, good, thank you. Yes. So one of the things that we've done in the new paper is really, I think, made a lot of progress on the theoretical foundations of the scale use correction.

And… I think we understand much more deeply now You know, what need… what you need to assume. In order for it… it to work. And it turns out that actually, You… there's a lot of violations of common perception. That are perfectly fine. So it could be… you know, what we were calling fill-in-the-blank bias, which… which we were very concerned about, and shows, you know, you see it in the earlier draft.

We were worried about vignettes, people answering vignettes where you can't mention all the details, and people might fill in their own situation when there's unmentioned details. We thought that… we thought… we, and I think other people, have thought that that would invalidate the vignette approach, but what we've shown is that, in fact, it doesn't. As long as you don't have… complete overwriting, where you're ignoring the calibration question and just replacing it with your own situation.

Anything short of that. is actually accommodated. The one thing which… .

### [03:30:56] David Reinstein (The Unjournal)

We lose him?

### [03:30:58] Daniel Benjamin

Can you hear me? Can you hear me?

### [03:31:02] Valentin Klotzbücher

I hear you.

### [03:31:03] Daniel Benjamin

Okay. So I think the one… one major thing that we… You know, you might still worry about. is… is if people… you know, so we can also, I should say, we should also accommodate things like differences across people in… We can accommodate people, giving different answers to a vignette that they… than they would give to themselves.

So, for example, people might rate Pictures of height, of other people's height, objectively, but exaggerate their own height. It turns out that's also okay. As long as everybody who is at a given height exaggerates by the same amount. what it can't deal, but there's other things it can't deal with, like people of different heights, people of the same height exaggerating by different amounts.

But the main point is that, we're actually much less concerned about, violations of, Common perception than we were in the earlier draft, which makes us more favorable to To using vignettes. The other thing I wanted to respond to is about the HOPEIT model. And one of the things that… that, I think is important to highlight. You know, you said that the models are not nested, they make some assumptions, we make different assumptions.

But the additional assumptions that we make Are testable, and we have tested them, and to the extent we've tested them, they seem okay. The assumptions in the hope-it approach, like the assumption of normality for the latent variable, is fundamentally untestable. And in the presence of heteroscedasticity, which we think is gen… you know, usually present.

As you know, as related to the bond and line critique. You can get arbitrary results by making different assumptions than that. So, you know, I think there… there is a real sense in which, you know, our… we can be a lot more comfortable with the assumptions underlying our approach than… than the hope-it approach.

### [03:33:17] Caspar Kaiser

I'm glad to hear all of those things. I'm especially interested in seeing how it comes to turn out. that the, you know, these violations to a common understanding of the question, you know, turns out not to matter all that much.

I think that that would be a counterintuitive. And deeper result.

### [03:33:39] Daniel Benjamin

Yeah, we actually think that's a pretty major contribution of the revised paper.

### [03:33:43] Miles Kimball

Let me also add, I think that in terms of Method of Moments versus Chopit. The… I think the transparency of our method is really valuable, that in terms of CHOPIT, we've done a lot of work to list concerns about it. And CHOPIT, by the time you do a ProBit-like thing, you have enough opaqueness that… that you could have problems you don't realize. Like, for example, if you, not only do you have the Bond and Lang problem, but if you… If you transform the scale to force it to be normal, you also can introduce, interaction terms and quadratic terms that weren't there before.

And so there… there are a lot of, kind of, hidden pro… potential hidden problems within a more opaque procedure, and a probit-like procedure is inherently more opaque than an OLS-like procedure.

### [03:34:54] David Reinstein (The Unjournal)

If I could just jump in, in the interest of, sort of, timekeeping here, I think it seems we're having a very fruitful discussion about the paper, about the approaches. We should move to the Belief Elicitation methods section in just a moment, and I'm gonna need to take a 2-minute break before that, so… feel free to continue talking, and then if you want to have a more in-depth conversation about this, you know, hopefully I can still nudge you to do the belief elicitation exercise, but you may, you know, also find it fruitful to join the technical breakout room. I don't know if anyone's joined those rooms yet, but they're there.

The three dots, it could be an opportunity to continue the conversation. It seems so fruitful. So, I'll be back in two minutes.

### [03:35:49] Caspar Kaiser

Maybe I can just ask a quick question, then. in the new table, I don't know if it was in the old version, you showed that the… Cope it. Gives, like, a very different result from the… Methods of… a method of moments estimator. Do you have some intuition on why?

### [03:36:13] Miles Kimball

We know half. We know half of it, so… is due to the following. If you… you can easily modify CHOPIT to use only the calibration questions to identify differences in scale use, and that gets it half of the way towards our method of moments. So they're doing a weird thing in most of the literature where they're using the… the sort of self-ratings To also try to establish scale use.

And of course, that's mixed together with the true underlying levels. So, we wish we could do a full accounting, but half of the accounting is that. That very simple change of using calibration questions only to identify scale use gets it halfway towards our method of moments.

### [03:37:03] Daniel Benjamin

But also, I mean, given that you can get arbitrary results with Chopit, making different distributional assumptions. There's no reason why you should expect that we're gonna get the same results. You know, using our approach and their approach.

### [03:37:18] Miles Kimball

Yeah, our approach is invariant to heteroscedasticity. Theirs is very sensitive to heteroscedasticity. Heteroscedasticity doesn't matter any more for our approach than it does for OLS in general.

### [03:37:32] Caspar Kaiser

I know that now it becomes, like, completely arcane, what we're talking about, but I just wanna, like, make the footnote that it will turn out that In the special case, Where the distribution of, let's say. Reported life satisfaction is approximately normal. The ratios of coefficients that you can get from an ordered probate will equal those that you get from an or less regression. Like, this is, like, this old result from, van Track in, like, the 200…

### [03:38:08] Miles Kimball

Yeah, they don't look normal at all on the 0 to 100 scale.

### [03:38:12] Caspar Kaiser

So that's what's driving it, probably, then.

### [03:38:15] Miles Kimball

Well, I wonder if that result assumes no heteroscedasticity as well, because it's…

### [03:38:22] Caspar Kaiser

Yeah, well, you can extend the chill, but to account for dependence of the variance in the error term on covariates, but you get into all sorts of weirdness. I mean, I don't know if this is exactly Simply…

### [03:38:40] Miles Kimball

I think that's pretty…

### [03:38:41] Caspar Kaiser

Long problem, but yeah.

### [03:38:43] Miles Kimball

I think that's pretty sensitive and hard to be totally confident that you're getting the heteroscedasticity modeled right. So, ours is… ours is simply an estimator that doesn't get affected by heteroscedasticity at all.

### [03:38:57] Caspar Kaiser

Hmm. I think the more important point here is really, like. you know, one can quibble over or choke it or not, but I think the key thing is that we literally know nothing In the case of… interventions. whether this is in rich countries or in low- and middle-income countries. Like, we've understood, you know, with the big observational data sets, yeah, you can do these adjustments, and, you know, it makes a difference.

But at least for the people in this workshop, that does not matter, broadly speaking. What matters is, I'm going to give people a bunch of cash, or I'm going to give people, I don't know, bed nets, or I'm going to give people psychotherapy, will they change the way they're using their scales?

### [03:39:42] Miles Kimball

So, Casper, in the chat, I put this suggestion. I think the sort of thing we did with height and weight can be very valuable, because you worry about demand effects, so you worry about kind of changes Like, social desirability, people wanting to report things in a certain way afterwards. And so, if you can… if you can ask about… let's say your main interest is in a subjective outcome, but you would… ideally would also ask about an objective outcome, both objectively and subjectively, with calibration questions, and then you can identify plausibly, not only changes in scale use, but also changes in… in how… in misrepresentation based on social, you know, demand effects, or other kinds of social desirability concerns. So I think… I think… and that's something we're very interested in. That… that's a little bit beyond where we're at at the moment.

But I absolutely agree that in the context of an RCT or some other intervention, there are these… there's this additional issue of demand effects that scale use correction alone is not enough to deal with.

### [03:40:53] Michael Plant

I had a question on this, but are we… are we now… are we moving on to the next thing now, David?

### [03:40:57] David Reinstein (The Unjournal)

Yeah, I, I think so. Maybe, Michael, maybe just have the last question, and then we can move on to the next thing, and then if others want to continue this sort of econometric discussion, specifically about these techniques. You can hit the breakout room with those three dots at the bottom, then go to breakout rooms, and then choose the, technical one. So yeah, Michael, please.

### [03:41:16] Michael Plant

Yeah, so.

### [03:41:17] David Reinstein (The Unjournal)

Western or Westward.

### [03:41:18] Michael Plant

So, authors of this paper, I think I've heard you present this several times, and I confess, as a non-technical philosopher, I mean, still huge amounts of it go over my head. So I'm gonna maybe bring the tenor down to ask kind of a more philosophical question, which is, do you think this kind of approach should be treated differently for different measures of well-being, and do the results change? So, a kind of… a sort of view you might take about life satisfaction, remembering some of the earlier conversation, is like, look.

The point is to try and get how people rate their own lives, you know, like, you're the king of your own castle, you can say what you want, and so there seems something sort of philosophically unmotivated about saying, well, okay, you've said this, but like, if you were someone else, you would say this, it seems like, you know, you're sort of… the upper and lower bounds of your scale, like, that's for you to choose, and it's not really for anyone else to kind of mess around with. So there's kind of an argument there that this sort of calibration stuff doesn't make so much sense in terms of life satisfaction. It maybe seems a bit more plausible in the case of happiness, if we think, look, actually, there's, like, there is this real thing in the world of latent, sort of, sensations of happiness, and if people are using the scales differently. You know, you actually, you know, one person just has a different sense of what, you know, there's different actual amounts of it.

But then, how much difference does this make? between effective and, evaluative measures, because I can imagine people being, kind of, having more different, kind of, views on life satisfaction between each other than their, kind of, differences with respect to.

### [03:42:51] Miles Kimball

go.

### [03:42:52] Michael Plant

How much pleasure or pain you feel.

### [03:42:55] Miles Kimball

So we discussed that in the paper. We're kind of concerned about very broad measures. So life satisfaction has been attractive to people precisely because it's broad, and they hope that it brings in everything.

But that introduces the issue that maybe people really, our… our… are interpreting life satisfaction differently. You know, if you get to narrower measures, like how well do you hear, there's still different subdimensions, but it's probably a little bit more standard interpretation than life satisfaction. And so, we think there's a real trade-off between trying to have broad measures where you hope you can have just a few, and having narrower measures that are a little bit better defined.

So, so, technically, Our method requires that you'd be able… that it be philosophically possible to make interpersonal comparisons of whatever's being rated. So I think of that as you imagine you put people in a room with a moderator for a few days, they would come out being able to tell you which of the two people had higher life satisfaction. And if you think it's likely that after those days of discussing it with the moderator, they still couldn't agree, then then you've got a philosophical problem that also violate, you know, would violate some of the assumptions that we're making.

But you need this presumption of interpersonal comparability to have any of the exercises we're talking about with life satisfaction make sense.

### [03:44:32] David Reinstein (The Unjournal)

Okay, I'm gonna… I'm gonna wanna cut you off for thanks. Thank you so much for your, you know, involvement in the questions, and I think this was… seemed to be productive to me. Sorry, I'm not always so good with the words.

And by the way, I'll just mention that Michael Plant's paper, A Happy Possibility About Happiness Scales, An Exploration of the Cardinality Assumption, explores some of these questions, and it's linked In the workshop, and I even linked a Google Doc, which I think it could… I see some room for productive conversation about some of these issues, you know, between Michael's perspective and others. Obviously, he's coming from a philosophy background, so I see some fruitful conversations that could be had there. So, thank you for, Thank you for your presentations, and, let me just, As I said, you can jump into that breakout room, but I'd like to move on with the belief elicitation, so let me talk you through the idea of this exercise, alright, so I'm sharing my screen right now.

And everyone, let me just paste the interface in, just to make sure you all have access to it. I'm pasting it in right here. This is… Let me try to do a little bit of disambiguation and explanation of what this is about, and then… We've kind of cannibalized about half the session, so I might not have time to ask too… have too many discussion and questions. That said, That said, maybe there'll be some room for that, and of course, I'm always happy to jump on to answer questions asynchronously.

And I think, you know, we're not going to get all the questions answered in a 30-minute workshop. As I mentioned, Plant's… Michael Plant's paper, I'm just gonna paste it there, sorry. 5-hour, 3-hour workshop.

So, why are we doing this belief elicitation? And this is sort of a trial. We're working out the kinks, trying to see what will work in terms of Eliciting the beliefs of stakeholders and experts, and… What sorts of questions can we pose? How do we frame them?

So, the idea is, slash was, to capture your, we call them, what do you, if you like it, prior beliefs. Before hearing the full evidence presentations. And then at an intermediate stage. To sort of measure a sort of updating, if there's any.

And the low comparison. And then, you know, in principle, and this is a whole field, one wants to aggregate expert opinion to inform other people. So… Why do it, and why do it in this workshop? What are our goals? What are my goals?

you know, the direct and obvious goal. If we've framed the right question, and it's a high-value question. And we want to extract your knowledge, and And, you know, so the people making decisions can use that knowledge. Of course, this gets a bit meta, because obviously these are challenge… you know, framing questions are challenging things, scale use is challenging, so it kind of loops a bit in on itself, but that's the idea.

I think doing this helps focus our ideas. Oh, wait, that's what we're talking about. I didn't realize it until you were forced to frame a precise question.

This is also something we're trying to do more of, both for this question, for the other questions, and in the future, so I'd love your input, well, on the framing of the questions, and how we're doing this. And then, of course, we've been… We're an organization that we… Rely on grant funding, and grant funders are very interested in having a sense of what impact we are having, and this is one potential way of measuring impact. I mean, it's very hard to measure the impact of research work The meta-research work that we're doing, but, you know, this is the ambitious, approach. You know, to try to see, are we actually… before people read the thing versus after, does it actually shift their beliefs?

And if they can even identify how those beliefs enter into their policy choices and their value, well, then you've actually got a tip-to-tail, you know, purported measure of your impact. It's not easy, but one has to try. And you know, there's these meta-questions that are interesting, you know, this is just a preamble to that. Obviously, this is very… what we're doing right now is rather informal, but it's something that people do care about. You know, how much does this stuff matter to experts' beliefs, also to funders' beliefs?

Okay. more to funder's beliefs, perhaps. And there's some PhD students I'm talking to who are interested in making projects out of this stuff, as you can imagine. Obviously there's challenges in doing this well, you know, posing the questions correctly.

We all are experts at that, and it's not necessarily a much easier question just because we're posed… even if we're posing it in a group that, you know, we all have this expertise, we still understand questions and communication differently. And even for very quantitative academics and practitioners and researchers, I think. Stating one's beliefs in a quantified way, and saying. here's what… here are my 80% credible intervals is difficult.

And also, you know, this… if we're… people are trying to influence the agenda, that's also not incentive compatible, obviously. It only works among, quote, honest men. And aggregation, you know, again, we're getting very meta here, but, you know, should we do linear aggregation of beliefs?

Well, okay. So there's… you'll see three interfaces. And I admit it's confusing, so there's the CODA interface, which was the thing for our project for the people who are prioritizing the research, evaluating the research, and We're compensating them to, you know, evaluate these pivotal questions. This was the precise formulations for them. I then made it a little bit simpler in the conference website page, the one I just shared with you.

That's… may have bugs in it, though. And also, there's this I'm just curious, how many of you are familiar, before I mentioned it, with the site in Metaculous? Sorry, I don't think I have the full thing up, About maybe half of you?

Okay, I don't see everyone who's here necessarily, but it sounded like a half. So we're also putting it up there, maybe to get some more engagement, and because it's a very attractive interface, and I want to explore using that. Okay, so what are the key questions? Category 1, you could… you know, reduce to, is the linear well be reliable enough? Category 2, what's the best conversion factor?

Okay, so now… How do we deal with this? So, privacy notice, we're going to, you know. use this information. You can do it anonymously if you like.

We've discussed the focal case. State your best subjective probability. For open-ended questions, any input, you know, could be helpful. Reason… you're looking for your reasoning. I will try to synthesize this and share it with the group.

You can leave comments on this framing itself with the hypothesis tool. I'd greatly appreciate that, and I'll try to respond to all comments. The focal example I mentioned before.

Okay, well, here's… maybe it's a nicer statement of it. Suppose Founders Pledge, considering whether to donate recommend a donation of $100,000 to either Strong Minds. Or seasonal malaria chemo presentation… prevention.

They have substantial evidence, From a variety… a few different types of measures, including well-being surveys on these different interventions, which are in maybe similar but not identical context, maybe in different countries, affecting different groups. Definition of well-being, what do we mean by best? I'm a little bit vague, but, you know, think of a welfare metric. Think of the optimal decision maker. trying to optimize that metric.

And then, you know, what is calibration? What is uncertainty? I assume people here are fairly familiar with that, but, you know, maybe you haven't… haven't done it yourself. Okay. First question.

And please feel free to interrupt, or ask a question, or jump on top of me here, because that's the point of interacting with people. So… how much… I've mentioned this before, so I won't go over it again. How reliable is the linear well-be measure, stated in general terms, for the focal concept… context?

Again, giving you some more context here, if you need it. and then… yeah. can you… Oh gosh, did this get… did this question get changed? What share of value is obtained from using the linear Wellby as defined above for all interventions?

Okay, please just answer the question you see on the page, but I do see that there's some… this contrasts with what I may have said… had before, so sorry for being a little bit sloppy there. How much of the value is captured? Okay. Next question, given the available collected data. How should funders measure the impact on well-being.

What measure would you recommend? And then… Sorry. How would you convert between these measures? What would be the best mapping? What would be the best linear mapping?

And then finally, some, some, Broader, or sorry, some more, sort of. Simpler, but let's say, questions you could actually track. So, I'm asking about this hypothetical question. If we were to survey the relevant experts. What share would agree that the linear well B is a reasonably useful measure in this context?

And, you know, you could imagine the question stated exactly like that. Now, you're also encouraged to actually weigh in on the meticulous questions. Okay, then here's a question about GiveWell, and here's one that's very relevant to the discussion we just had. If calibration questions, etc, were added. Would this change meaningfully?

What do I mean by meaningfully? It says here. The cost-effectiveness ranking of the top 5 interventions recommended by Founders Pledge. Did you… sorry, does Founders Pledge recommend a top 5 interventions? Is that meant to say GiveWell?

### [03:55:23] Matt Lerner

Yeah, we don't… Ranked public intervention.

### [03:55:27] David Reinstein (The Unjournal)

Yeah, something got messed up there. That was meant to say, GiveWell. Okay, so… Let's… we can also see the metaculous questions, So, where is that again? Here's metaculous questions, and I think these, these are the ones that I… that I was a little bit more precise with. Okay.

So you could answer in either place if there's overlap, actually. This one is maybe a little bit more… thought out. Okay, yeah, this one I like a little bit better. Cost ratio to achieve the same welfare improvement with the linear well be relative to what you think is the best feasible measure.

Now, if you think the best feasible measure is the Wellbe. Well, then that cost ratio could potentially be 1, or close to 1, I guess I even allow above 1. Okay, I have to think a little bit more about the framing, and I'm like, love your feedback on that, to be honest.

### [03:56:22] Caspar Kaiser

Quick question on this.

### [03:56:24] David Reinstein (The Unjournal)

Yeah, please.

### [03:56:24] Caspar Kaiser

Maybe I missed this, because I was just gone for a second. It says… Relative to other available measures.

### [03:56:31] David Reinstein (The Unjournal)

Yep.

### [03:56:32] Caspar Kaiser

Which is different from other possible measures.

### [03:56:34] David Reinstein (The Unjournal)

Yeah, that's true. Which, so this one says available measures.

### [03:56:37] Caspar Kaiser

Hmm.

### [03:56:38] David Reinstein (The Unjournal)

The other one might… did I say possible measures in the other question?

### [03:56:43] Caspar Kaiser

Well, I think possible is the better question to ask.

### [03:56:47] David Reinstein (The Unjournal)

Yeah. Okay, no, that's good feedback. Let me… let me adjust it. You know, one thing you can do is, if you're answering the question, which I definitely appreciate, you can say, I'm answering the question, envisioning it as having asked this question. So I think… I think they're both… useful.

One is if I have to make a decision now, which should I use? The other is if I could collect other measures and took into the account the cost of collection. you know, which would be better.

And I get it that a little bit below, When I talk about… That's in the extended question, what other measures… yeah, here we are, what other measures should be collected? So, yeah, appreciate that. Any other feedback on the framing of these questions?

### [03:57:32] Michael Plant

Well, I would have comments on this, but I assume it's gonna… I mean, we're gonna answer what's in front of us, so it doesn't really matter, and I'm sure we'll work it out, and it'll become clear in the comments anyway.

### [03:57:42] David Reinstein (The Unjournal)

Okay, no, that's good. And then for… if there's a next iteration, or when I, you know, pose this to the… to the pivotal question evaluators, I can, make those, adjustments. Yeah, no, this one… this one, I… it's shared… talked about as a share of value, but the other statement of it… yeah, okay, I had as… How much would it cost relative to the other?

So, sorry about that. Any other questions just in navigating the page, or about the intention of this? Or about, you know, is something… what other things are unclear?

Alright, well, let me just give, we're… we're supposed to jump right into the next session, but I think it would be good to take a 5-minute break before we start the final session. So, yeah, let's just take a 5-minute break, and if you want to take that time to respond to this, I would appreciate it, and apologies if it's… You know, still a little rough, and still something that's being adjusted In response to feedback and information. So, to see you in 5 minutes.

All right, well, let's, move on to the final, session. I hope everyone's… again, if you're able to do the belief elicitation in at least some way, I think that that can add a lot of value to the process, as well as, you know, the objective, information that we'll be sharing. Okay, so let's, let's move on to the last session, where we bring back Matt and Peter, and Ori is going to be a discussant.

The thought was to… C, maybe… Report on what we think we learned, maybe unpack some questions we might still have. Maybe we can come to some ideas of what might be some practical recommendations, or at least Things to try in this context, as well as, sort of, next research directions. So, let's see, should we start? Ori, do you want to… do you want to sort of take the reins, and then we'll bring in Matt and Peter again? Would that make sense, or what should we do the other way around?

### [04:05:24] Ori Heffetz

Sure, so, yeah, so I, you know, I can say a few things, and then, Chair of the session. So, I guess a few thoughts. First of all, again, like everybody else said, you know, thanks, David, for putting this together, and beyond journal for inviting us. This has been… very inspiring, and… Gave me some new ideas to all of us. Let me say four things.

The first one came up earlier when my… in one of MISA's comments, but to me, it's the first question I asked myself. We started The day with the funders. And I kept asking myself, are they all utilitarian?

They never said anything about that. We talk about measures and metrics and stuff, and I thought, wait a second. Are they not prioritarian? Is it the same, you know, are all… $1… is $1 the same, regardless to who you give it to? No.

Is, health the same, regardless of who you give it to, etc. So, I think… We kind of jumped around that, but there are ethical questions and methodological questions that are quite distinct. And even if we could measure… Well-being on a great scale that is interpersonally comparable, and it means exactly everything to everyone, and everyone answers the same question, etc.

There are still questions about what to do with that information, which we didn't touch on. And to me, it's probably first order. We probably focus on second-order things, and we never discuss the first order thing.

So, I… I think, you know, we would love to see that discussion at some point. Today was on the methodology, so fine, but I expected from the funders, at least, to say something about that. That's the first point.

The second point is really, the practical advice that we give, it's clear that there are two types of practical advice. One type of practical advice is Given what you have now, and the metrics that we have, and all the data constraints, etc, and you have money in the field, and you want to direct it to the best users. What best advice we can give you?

And we can contrast that with advice as to, here is one or two or five more questions that if you add to your RCT, evaluation questionnaires. It would help you to do a much better job next time. And so I think a lot of the… Certainly, our paper is geared mostly towards new data collection, you know, you need at least two calibration questions, etc.

But we could also talk about, and we have talked about earlier today. What do you do, if you don't have that, could you… Could you use? You know, could you use some other data set to extrapolate to this data set? What do we do with the big problem was the fact that we only have some answers for some samples in rich countries, but we're actually interested in poor countries.

So, in this context, I want to say we have two, papers One of them is in the reading list. And one of them isn't. I don't know, actually, do I have a… Am I allowed to share a screen? Do I have permission to share a screen? Let me just try instead of asking.

### [04:09:20] David Reinstein (The Unjournal)

Please share your screen, please do.

### [04:09:22] Ori Heffetz

Here, it did say, wow, okay. So I'm just showing here… One table from one… one paper that we have, And we have that in different papers, right? So this is something that we came to like to put in our conclusions.

So, conclusions to our paper, you can see, are short. To view paragraphs. But then we stick a table.

And what the table does is. And these are all links to sections in the paper, right? So what the table does is it says, here's practical advice to practitioners, to you guys from the funding institutions. You don't have time for the paper, just look at the table.

And we split things along two dimensions of cost. So we are autonomous, we care about cost, right? So, one sort of cost is… Do we have something that is readily applicable, or do we need more research and development, and that takes time? So this column is the cheap stuff in terms of it's ready to use.

And then on this axis, we look at another metric of cost. Do you need additional survey design, implementation, effort, or duration cost? Yes, no.

And so this quadrant… is the staff advice that is readily applied, and we don't… you don't need more survey time. We don't tell you, add these questions, or repeat questions, or whatever. And so, this is a bunch of stuff that… and you can click and go to the section, and you get the advice. You have a little bit more funding, you have a little bit more survey time to spare, you have a little bit more interest or budget, go to this quadrant.

These do add costs, but they are ready to apply, certainly for your next survey, right? So this could be for the previous survey, this could be for your next survey. And maybe if you want to fund research into the thing, you can go to the other column. So that's the… That's about the research that we did, that's my second point. My third point that I wanted to say is, And again, this came up… there is a fundamental difference between QALY and DALY and, you know, all the [inaudible].

And the well-being. Which is, they don't rely on a question. They don't rely on a question.

We can calculate them with data, I mean, in some sense, with administrative data. And then we know that everybody's so-called answering the same question. But once you move to questions, and it doesn't matter if it's a life satisfaction question, or a letter, or a whole battery, you know, of 200 questions, and then you construct an index, you want to make sure that everybody answers the same question.

We also have work that shows that, no, the answer is no. People understand the question differently. The broader it is. the more… the more interpretation to the question you have. You know, you ask people, do you have leg pain?

Then probably everybody answers the same question. You ask people to put yourself on a leather. and think about your whole life, you get a whole lot of different interpretations.

We show that these interpretations are systematically associated with demographics. They could be systematically associated with the treatment, etc. So, we're in trouble.

And so now I'm not saying, okay, so I'm not a nihilist, right? So I'm not saying, so just do nothing. But we have to be aware of these things.

These are… some of these things are discussed exactly in these quadrants of the chief quadrants. You don't need to collect more data. You don't need to invest more money, you just need to think carefully about how you interpret the data you already have, and avoid mistakes.

Avoid mistakes in assuming that the data you have tells you something when it tells you something else, or, more likely, it tells you much less. than you think. And my last point actually summarizes everything. It also speaks directly to that point, which is you know, we do… we do what we can, we do the best we can with the best data we can. So again, I'm… I'm exactly the opposite of analyst.

We… we do well-based, it's better than nothing. We do other things that are better than nothing. Sometimes they are not great, but they are the best we have.

And here, I want to say, let's take maybe… let's follow the lead of politicians and of governments, right? So governments sometimes call us into the room. and listen to what we have to say, or read not quite the paper, but maybe the abstract or the, you know, the white paper.

But then they say thank you, and we leave the room, and somebody else comes and talks to them. And the somebody else maybe, you know, they don't know math, they don't know statistics, they don't have figures, they don't have numbers, but they highlight another angle of the problem. Sometimes they just listen to voters. Yes, those stupid voters who don't know anything about economics, they listen to them too, and that's also an input.

And I want to say to… To the… to the givers, you know, to the funders. It's a great, it's a great, you know, pie in the sky. It's, it's a great… M… aim to say, we're gonna… optimize the process.

We're gonna take, take out human judgment, and all the biases and all the… because these are multi-dimensional, difficult problems. We want to develop a script to do it right. And I think it's a great purpose, and you know, I devote my… I have devoted, and I'm still devoting my career to try to help you develop these metrics.

But these metrics are not there yet. We cannot fully delegate the decisions to these metrics, and eventually just compare a bunch of numbers, and you say, this one won't, because it came out at the top. We should definitely use all of these things and improve them, but it's just one input. It's just one input. Human judgment with all the biases and problems is another imperfect input, and a lot of others.

And so, yes, sometimes we will see that if we use health, you know, health-related measures, the mental health things come at the bottom, but if we use well-beat comes at the top. What to do? Well, what to do? To do what we've always been doing, to think.

And to think, and to think about other things, and to listen to other people, and to decide in that way. That's all I have to say. I am praying for the day where we will be able to automate… to, you know, to automate the system, which means we'll have really, really good metrics that we can trust. Hopefully, we're heading in that direction.

But for now, we should also maintain this, david called it, you know, the confidence interval, or whatever. It's also humility, or, you know… Okay. So now I'm switching to become the chair, so I guess Matt… The floor is yours.

### [04:17:02] Matt Lerner

Yeah, I'm… I think I'm gonna… I don't feel prepared to say anything very conclusive. I have some reflections. I guess… One reflection is… Yeah, I'll quote you, Ori, that we're in trouble, I guess. That would be… One reflection.

I think… I think we're in a different kind of trouble, so I think… One… one takeaway that I have is sort of that… getting to a point of higher confidence about wellbe measurements, is possible with the right attention, which I think already is part of what you were saying, like, sort of, like, being careful and attentive to the details of specific interventions and studies. with the right energy and time input can get you to a point of, like, feeling fairly… fairly okay about your conclusions, and I guess when I say we're trouble, I mean we… like, on my team or trouble, because I don't think we really have that time or energy. And so I do… I do worry about that, like, I'm… I guess, I guess maybe… Maybe part of the… part of the takeaway here might be, well, geez, like, maybe we're doing as well as we can, given the constraints, and that, like, doing better with respect to wellbees might require just putting in more time and energy than we have available. I feel a little bit more optimistic than that, because I feel like… Especially with, sort of, careful deployment of AI, I think we could do a better job, automating some of the conversion stuff. Like, I feel kind of good about that.

I feel like we've had some preliminary experiments where Okay, like, maybe it's simple enough to use wellbees, you just, like, or use DALIs, we match up the conditions and look up the disability weight, and you multiply out, and maybe our procedure for wellbees and subjective well-being just looks a little bit more complicated. Like, maybe even going to measure something in wellbees, we do need to do this kind of, like, intervention by intervention. Correction, which isn't out of the realm of possibility.

Okay, so the other reflection I have is, yeah, so… I think one underlying theme is the sort of, that I take away is the… Unavoidability of the core philosophical question of, like, what is the… what is the outcome, of interest, actually? And… I'm now… less confident than I was at the beginning that, for instance, types of things like The evident correlation between effective mood and subjective well-being are evidence that we're all just kind of getting at the same underlying, quantity of concern. And so, I guess, maybe perversely, like, my next step is, like, like, shit, I need to talk to Michael about moral philosophy some more, I guess, like, I don't know. Yeah, Michael once told me that under certain assumptions, like, dog shelters look really good, and now I'm like. Shit, maybe I need to explore those assumptions, along with other… other assumptions, about moral philosophy.

I think that's about it. Sorry, like, sorry I don't have, like, really, really confident conclusions, there's a lot to think about. I know, Ori, that you said, maybe the last point, Ori, I know you said that, like, in a lot of cases, collecting more data is not the answer, but I think one thing that I can't help thinking is… There is such a strong mythological basis for doing some of this work that I… Feel a lot more optimistic now about Just the research and funding agenda that looks more like Including subjective well-being measurement in other types of interventions, I f… I guess I have the intuition that that might actually just be the highest ROI. things you hear, just in terms of in-value information terms, but that's… That's all of my takes, yeah.

### [04:20:55] Ori Heffetz

Thank you, Matt. Peter?

### [04:20:57] Peter Hickman (Coefficient Giving)

Yeah, glad to have been here today. learning a lot about the state of the art here. So… I… like, I think my vague worries and ideas about Subjective well-being measurements have become more precise through seeing, some of the… Specific concerns and how they're being addressed. In… the Benjamin et al. paper.

I think this has… also highlighted what issues we didn't discuss today, namely the kind of badness of death point, and how that is probably The key crux for us in whether we'd want to Kind of, like, focus on Interventions where subjective well-being was the main outcome versus Focusing on life-saving interventions. I think… the kind of worries about scale use and experimenter demand have been addressed somewhat. It's good to see this new method that and correct for some of the scale use heterogeneity, which is really good.

I think… Would like to see some… Direct thinking about Like, whether people are just going to give you higher answers after an intervention. Because they want to please the research team. And that might inflate results.

So, that's kind of still an outstanding crux, but yeah, glad to have been a part of this today.

### [04:22:36] David Reinstein (The Unjournal)

Okay, thank you. So I just… we have a couple more, moments. Sorry, I've been trying to jump around a bit here to sort of show things in the screen share.

We have just a few more moments, and then, some of us will… maybe most of the people here will be going to a different sort of… sort of, well, a follow-up conversation, if you still have the energy. I just wanted to, Obviously, thank… sorry, I'm off video, sorry about that. Thank everyone. forecoming, and, you know, this is the first time we're doing this, and I… appreciate your participation, as well as your feedback on, you know, what this was able to accomplish, what we could do, how we could do it better. You know, is this… is this something worth doing?

It definitely seemed like we had some… useful debate, useful discussion, I should say, and it seemed like a very productive mode of discussion. If I just want to just mention a few things, which maybe we… we can give less takes on. Quickly, or just food for thought, I would, suggest, you know, I want to know… I'd like to get a sense of, you know, what… and I see, Peter, you just referred to it. What would change our beliefs? You know, what are the most important things That… that we want to hear or know more about.

What do we agree… what have we sort of come to agree on that we should communicate to others? What should be the next data collection, or the next… what should the next research do? What needs clarification and follow-up. Feel free to jot it down in this Google Doc, make sure to flag me.

As mentioned, what can we do right now? Now, some of these suggestions might be doable, certainly in some contexts. In other… some of them are not necessarily relevant to the, to the, to the context we're talking about, but some of them might be adaptable.

In terms of the next research, I guess my, my thought, and this was my, my take, was… Do we want an extension of a lot of the research that we've seen, you know, even the sort of… discussion and bringing theory to applied context research, but very much focused on this very… on this specific context of RCT, you know, interventions and RCTs. And in that context, would it be worth funding single… it doesn't seem like we have any available, choice Data, you know, of this nature in the low-income country context, which might provides some… Both a sense of how much do these findings carry over, but also potentially something that could be used to make at least a first pass adjustment. Yeah, okay, so does anyone want to have the… say anything else in terms of the last word, or responding to those? I know we're a bit pressed for time right now.

### [04:25:38] Michael Plant

Can I say something?

### [04:25:39] David Reinstein (The Unjournal)

Please, please.

### [04:25:40] Michael Plant

So it's, it's really fascinating hearing this, but I notice as I'm doing it, I'm having a slightly out-of-body experience with, with Matt and Peter being like, oh, this sounds like… I don't know how we would do this. How would we implement it? I'm like, we're doing it, like, so it's like, we, like, we built a car, we're driving a car, it's like, yeah, I don't know if you could possibly drive a car.

Now, it might be that the car is, like, imaginary, and in fact, it doesn't really go anywhere. But, yeah, I mean, you know, this is some kind of… this is stuff that we've been forging ahead with, and, I would be absolutely delighted to, to… to kind of talk to Founders Pledge, Coefficient, as you're now called, in greatest detail, to tell you how, like, how do we actually try and solve these problems in detail? Now, you might not agree with our answers, and that's, like, that's fine, and that's useful, but we probably have answers to some of these things, or at least have sort of thought about them.

And, yeah, I mean, this… some kind of… there's some bean bits about, like, kind of what's the philosophy going on here. I mean, we take a kind of… you know, maybe sort of simple approach, like, we think well-being is what matters, we think it can be reasonably well measured by subjective well-being, so I don't… we don't have, you know, that seems like a good… a good way to be making decisions. It is the case in doing this that we have to kind of break a huge amount of rocks, like. okay, what is, in fact, the well-being weight for this and this and this?

But we're hoping that those rocks have to be, you know, broken once, or you see, like, exploration, we can kind of work these things out. And really, part of the reason we exist is that, So, you know, much, much kind of bigger, more significant players in the space can make use of this. So I'm very happy if you, if you wanted to chat us further. I'd be even happier if I wanted to fund us to carry on more of this research. So that's my, that's my kind of pitch to the, to some of the funders out there.

And yeah, we really think… hope this can be… this… this can be useful. I do kind of want to add the limitation of it relating to kind of Peter's bit, that the wellbeing approach, it's a bit like a… a train, like, it's powerful if you're trying to do quality of life, if you want to do quality of life to saving lives, or animals, or the environment, like, you know, I don't have solutions for those, we're just kind of back into… into normal kind of difficult trade-offs, so it's useful in a kind of narrow… narrow domain. I'll stop there.

### [04:28:02] David Reinstein (The Unjournal)

Thank you. This… so I'll just… I'll say one more thing. I think we've… potentially, you know, certainly learned about these methods and potentially had some questions answered.

I think there were some questions that were raised that weren't sufficiently addressed that hopefully we can continue to raise and continue this discussion asynchronously. For instance, you know, the question about the Scale… how… how much the scale use changes in response to the sort… some of the sort of interventions we… we've talked about. my sense is that that's something that would be sort of crux-y, and we'd like… you know, I think presenting and sharing more evidence on that would be, would be high value. I also mentioned, you know, evidence relevant to the linear scale use against what you might argue to be sort of gold standards of how do people say that they would trade off, you know, between movements, between 7 and 8s and 1 and 2s, among other people, etc.

Anyways, anyone else want to maybe have time for one or two more responses before we sort of end and we have that little additional thing?

### [04:29:08] Matt Lerner

I sort of have yet another, well, I don't know, I hope Michael will agree. There's sort of a clarifying point for the academics, as we sort of bridge the academic versus donor world, which is… So when I say it's, like, hard, what I kind of mean, like, hard for us to learn to drive the car. what I sort of mean is, like.

I think… I think it's still true that, like, HLIs and our estimates of cost-effectiveness for these different interventions vary right now by a maximum of, like, a factor of, like, maybe 2 to 5x. And what I was sort of hoping… and I sort of… there… it's kind of a big deal, it's like a big deal in terms of allocation. But in the scope of, like, all philanthropy, it's kind of a narcissism, small differences thing, and sort of my… my hope… Was that we could, like. Get to… my hope remains that we can get to a world where, like. Even those differences are sort of washed out, and also that a team like mine without the expertise of Michael's team, can sort of derive, like, a heuristic… a heuristic approach here.

And I think kind of what I'm getting to is, like, man, maybe the heuristic approach is actually, like, not very well suited, and that this is, like, complicated. That's… that's mostly what I'm… what I'm coming out of this thinking.

### [04:30:27] Michael Plant

Yes, can I… can I maybe ask you… so, sorry, I've just gone very ghostly. So you're saying that… You think that… that… how you would estimate stuff, and how we would estimate stuff. looks quite similar, but you'd like… you'd like that difference to be… you'd like those differences to… to be reduced? I thought you were gonna say you hoped that there would be more kind of… like, there would be more radical, exciting differences from shifting from… from whatever the status quo is to wellbe usage.

### [04:30:57] Matt Lerner

Yeah, I mean, I think that's… that's a third thing, right? Like, that… I'm… my, like, the long-tail hope for me is that, well, the usage can also uncover things that we, like, aren't even really thinking about. I sort of think that that's sort of contingent upon getting, comfortable with the answers to the other two questions first.

### [04:31:20] Michael Plant

Right. I propose we chat more about it another time.

### [04:31:23] David Reinstein (The Unjournal)

Terrific, terrific. All right, well, I think we should… should wrap this up. I hope that we've at least sort of been able to understand what… a little bit more about what each other's… what we are doing, what our questions are, what our… you know, cruxes are so that can enable more communication and work that is more mutually valuable. I will, you know. Be around to continue to help facilitate future discussions and respond to questions.

And yeah, thanks to everyone for coming. I really do appreciate, your being part of this sort of experiment. Of sorts.

And yeah, hope you all have a great week, and I'll see some of you in a moment in the, in the other room. Thanks very much.

### [04:32:13] Caspar Kaiser

Thank you, that was great.

### [04:32:15] David Reinstein (The Unjournal)

I'm gonna end the call. Take care.
