Whose Gold Standard?
On the ethics of RCTs in development economics.
Around this time seven years ago, my first semester of college, during which I had begun studying economics, I remember seeing that Abhijit Banerjee and Esther Duflo—the husband-wife duo of development economists at MIT—had won the Nobel Prize for economics, along with Michael Kremer at Harvard, all for their work on randomized controlled trials (RCTs). Naturally, I did some digging, being curious and, in particular, thinking that their research might have something useful to say about India, a country I have spent an awful lot of time in for reasons too obvious to state.
But my enthusiasm faded pretty quickly, and I wound up feeling rather conflicted as I looked further and further into their entire RCT-oriented project: while (admittedly) rigorous, it struck me as essentially being experimentation on poor people in the developing world the likes of which would not stand in the developed world, all shrouded in the language of “gold standard” research. What I also found myself disagreeing with was the underlying idea that you could understand a poor country’s problems by simply running tests: they clearly believed social science could be as clear-cut as, say, medicine, in which RCTs are regularly used, and that instead of attempting to do big question research, economics as a profession should consider little questions instead and add up the results to understand the bigger picture.
I bring this all up because I got into a little back-and-forth on Twitter over the following paper today:
Reading the abstract alone, I felt my general antipathy toward development economics rise up for the first time in several years (although it was a little bit rusty). Here is a prize-winning paper based on (what was certainly expensive) research that involves playing with the professional and marital lives of women in Bihar, the poorest state in India.
Now, before I go any further, I want to say: One of the important things about research is that sometimes asking a question that seems entirely useless has rewards no one can envision. I do believe that some seemingly “worthless” questions are absolutely worth investing in.
But not all of them. And here, the question of “Do men with jealous tendencies prefer their wives work in single-sex spaces?” is, as I see it, not a particularly worthwhile question—and while I would shrug or be amused if someone studied it using some sort of methodology that didn’t require doing a randomized experiment on actual people, I’m more actively critical when it involves a randomized experiment that actually interferes, for better or worse, with the lives of real people who, I’ll add, are often very poor and often don’t really know what they’re signing up for, or what academic research even really is, leading to real issues regarding informed consent.
Let’s just look at this paper, as an example, which was only a two-week experiment to see the choices women and their husbands would make. The author found that women who reported more jealous husbands were also more likely to elect for a female-only workplace, but couldn’t rule out that that was just preference so he went further, I’ll just quote from his blog for the World Bank in which he summarized the results of this paper:
To measure preferences for female peers, I offered women and their husbands the option to forgo 20-35% of the salary to guarantee assignment to a female peer rather than risk having a (presumably male) standard peer. For ethical reasons, to ensure no potential future spousal conflict, when the peer support program was implemented, all women were matched with a female peer, even those that did not pay the fee, but importantly, they did not know this when making their choices. Over half (53%) of households paid the fee to guarantee a female peer when the discussions happened one-on-one (Figure 2). The fee was roughly equivalent to a full day of total income for the median household. This substantial willingness to pay reinforced the results from the first experiment that women have strong preferences for female co-workers that extend beyond workplace safety or social image concerns. However, it could still be that households preferred the female peer because they would make for a better mentor.
To test whether spousal jealousy drove these preferences, I introduced a treatment, called the husband program, where husbands could join the peer support discussions or receive recordings if they could not attend. Demand for female peers dropped substantially (by 19 percentage points) when husbands could monitor conversations (Figure 2). This decline was significantly larger in households with more jealous husbands.
One could still argue that perhaps female peers are valuable because they offer better advice but the value of this drops when the husband is present. To rule out peer quality completely, I introduce another treatment arm, called the video program, where women watched pre-recorded videos of their peers rather than interacting directly, and were told the script was identical regardless of the peer’s gender. If households cared about peer quality, preferences for female peers should disappear when the content is exactly the same. Yet one-third of households still paid for a female peer when simply watching videos, suggesting that peer quality is unlikely to explain the results and pointing to spousal jealousy being a constraint on women’s labor market choices. Qualitative interviews with a separate sample confirmed this interpretation: 84% of respondents said husbands would prefer their wives watch videos of a female peer due to jealousy.
OK, well, credit where credit is due: he and the IRB at least considered “potential future spousal conflict” here, but there are still things that are off to me. For instance, the inherent deception of having someone pay a fee—amounting to the median Indian daily salary—to guarantee a female peer when you know you’re going to assign a female peer regardless. Or the part where a husband is able to monitor his wife’s call, and asking the wife survey questions in front of her husband as to how she wants to be addressed at work.
Tell me: would that sort of deception have been allowed in, say, inner-city Baltimore? In rural Mississippi? In the poorest, roughest parts of the United States? As development economist and 2015 Nobel laureate Aengus Deaton wrote in a critique of the randomista direction in development economics:
Even in the US, nearly all RCTs on the welfare system are RCTs done by better-heeled, better-educated and paler people on lower income, less-educated and darker people. My reading of the literature is that a large majority of American experiments were not done in the interests of the poor people who were their subjects, but in the interests of rich people (or at least taxpayers or their representatives) who had accepted, sometimes reluctantly, an obligation to prevent the worst of poverty, and wanted to minimize the cost of doing so. That is bad enough, but at least the domestic poor get to vote, and are part of the society in which taxpayers live and welfare operates, so that there is a feedback from them to their benefactors. Not so in economic development, where those being aided have no influence over the donors. Some of the RCTs done by western economists on extremely poor people in India, and that were vetted by American institutional review boards, appear unethical, and likely could not have been done on American subjects.
Let’s take a look at some of the most unethical RCTs—not the minor cases, like this spousal jealousy paper, but the really bad ones. Like this one RCT that went viral for being almost cartoonishly bad in 2020, in which researchers randomized the “treatment” of…shutting off water to tenants in Nairobi slums to see whether or not their landlords would pay the utility bill. Another RCT, also in 2020, randomized exposure to a Protestant Evangelical theology program in poor Filipino households to examine what impact that would have on their economic outcomes. I mean: that is just bad on every conceivable level (see more criticism of the paper here). You do not get to mess with people’s beliefs and worldviews and families because you have an ivory tower chair, or because it might help your tenure case. You just don’t.
It’s also not clear to me that any of this is particularly good economics, in part because the entire randomista revolution in development studies treats economics as something that can be done completely technocratically, maybe even more so than other subfields within economics—there is a blindness to politics, to culture, to all the other little variables that might have implications for both the external and the internal validity of the RCT. As New School economist Sanjay Reddy writes in Foreign Policy:
The reliability critique contests the idea that RCTs provided a sure means—indeed the gold standard—for inferring, although in narrow terms, what worked in development. Those who have made the reliability critique, including eminent, statistically minded economists (a few of whom are also Nobel Prize winners), have argued that RCTs suffer from two problems of reliability. The first, external validity, concerns whether the estimate of the effect of a treatment from the place that the RCT is administered, even if it is accurate there, can be transferred elsewhere, given differences in the behaviors of different populations, as well as in the prevailing environmental, institutional, and social circumstances. For instance, public health information may influence behavior more where the government is trusted than where it is not. It may even have the opposite effect from that intended if the government is held in great suspicion. The second concern, internal validity, is about whether the results from a given context are really meaningful and accurate even there. RCTs are designed to measure the average effect of a treatment in a population and cannot generally tell us how it affects different parts of that population (in an extreme case, which is encountered frequently in medical trials, it may harm some people even as it creates a benefit on average). The effect of an intervention may moreover change over time even in a single place due to learning and behavioral responses. For these and other reasons, it is necessary to take care in interpreting what RCTs have actually measured and in employing their lessons, even in the very same place that they have been implemented.
I don’t want to be misconstrued as saying something I’m not. What I’m not saying is that RCTs are bad, or worthless, or should be thrown out from social science work in developing countries entirely. I’m not saying that contemporary empirical social science or development economics are bad or worthless either. What I am saying is that when you’re talking about something as expansive as development, the idea that you can piecemeal together RCTs to form a theory of a developing country without any understanding of politics or culture is suspect. More importantly, I’m saying that the ethics of experimental design are of utmost importance, especially when there is the sort of power differential between you and your subject that there inherently is in development economics.



Very well put, Ms. Neeraja. Social Science is not, and cannot ever be, similar to medical research. I remember reading the Angus Deaton critique (co-authored with his wife, if I remember it right) of RCTs many years back, but I didn't appreciate his concerns at that time. I suppose RCTs do have their place in the toolkit of economics, but it simply CANNOT replace the traditional methods of enquiry. The paper that you have chosen to highlight (to illustrate your points) should raise deep ethical concerns to whichever journal has published (or, is planning to publish) it.