One Unreported Reagent pH Buffer Skewed a Fairness Game Replication
In the mid-2010s, behavioral science labs faced an unsettling pattern. The ultimatum game—a simple bargaining task in which one player proposes a split of money and the other can accept or reject it—had long been a textbook example of human fairness. Early cross-cultural studies suggested that people everywhere reject low offers, even at a cost to themselves, as a way to punish unfairness. The finding was celebrated as evidence for evolved cooperation. But when labs around the world tried to replicate those results, many failed. Some even found responders accepting offers near zero. Rejection rates for offers of 20 percent or less dropped from the typical 50–60 percent to below 30 percent in some studies. The original results seemed to depend on unknown conditions that no one could identify.
The search for hidden moderators became a minor industry. Researchers tested variations in stakes, culture, anonymity, and even the gender of experimenters. Nothing fully explained the pattern. Then, in 2019, a team at the University of Toronto noticed a peculiar entry in their lab notebooks: the pH of the Tris buffer used in the experiment had drifted from 7.4 to 8.2 over several weeks. That single, unreported change, they suspected, might have been enough to skew the results.
A Fairness Game That Once Seemed Universal
The ultimatum game, first studied in the early 1980s by Werner Güth and colleagues, works as follows: one player, the proposer, receives a sum of money and offers a portion to the other player, the responder. If the responder accepts, both keep the agreed shares. If the responder rejects, neither gets anything. A purely self-interested responder should accept any positive offer, yet across dozens of studies, responders consistently rejected offers below about 20–30 percent of the total.
This rejection behavior was interpreted as a punishment for unfairness, a norm enforced even at a personal cost. Anthropologist Joseph Henrich and colleagues famously documented variations in ultimatum game responses across 15 small-scale societies, linking rejection rates to the degree of market integration and cooperation in everyday life. The study, published in 2005, became a landmark in behavioral economics, cited as evidence that fairness norms are both universal and culturally shaped.
By the early 2010s, the ultimatum game was a staple of introductory psychology and economics courses. Its rejection pattern was taught as a robust finding about human nature. Textbooks described it as a demonstration that people are not purely rational maximizers, but creatures with a sense of fairness. The game's simplicity made it a favorite for online experiments, lab studies, and even neuroimaging work. It seemed unassailable.
Yet as the replication crisis took hold in psychology after 2011, the ultimatum game came under scrutiny. Large-scale replication projects, such as the Many Labs initiative, attempted to reproduce classic findings. For the ultimatum game, results were mixed. Some sites found high rejection rates; others did not. Meta-analyses pointed to unknown moderator variables, but no one could identify them. The field was stuck.
The Replication That Wouldn't Replicate
Between 2012 and 2018, at least half a dozen preregistered replication attempts of the ultimatum game failed to match the original effect sizes. In some studies, rejection rates for offers of 20 percent or less dropped below 30 percent, far lower than the 50–60 percent seen in early work. Researchers tried a range of modifications: they varied the stake size from $1 to $100, changed the currency from dollars to euros or local equivalents, ran experiments in person versus online, altered the wording of instructions, and even tested different participant populations such as students versus community adults. Nothing consistently restored the original pattern.
A 2017 meta-analysis by David Rand and colleagues examined 44 studies and found that rejection rates varied widely, with no single moderator explaining more than a fraction of the variance. The authors concluded that the effect was real but fragile, dependent on unknown contextual factors. Some researchers wondered whether the original findings had been inflated by publication bias or small samples. Others suspected that the phenomenon itself might be culturally specific, fading in WEIRD (Western, Educated, Industrialized, Rich, Democratic) populations.
Then, in 2019, a team led by psychologist Jane R. at the University of Toronto decided to systematically test a hypothesis that had been lurking in lab chatter for years: could the buffer used to maintain the pH of the solution in which participants' physiological responses were measured be affecting their mood? In many ultimatum game studies, participants' skin conductance or heart rate was recorded using electrodes that required a conductive gel buffered with Tris. The buffer was prepared in bulk and stored at room temperature for weeks. The Toronto team noticed that their Tris buffer, initially at pH 7.4, slowly rose to pH 8.2 over time due to absorption of carbon dioxide from the air. They wondered whether this subtle alkalinity might alter participants' arousal or mood, making them more or less likely to reject unfair offers. It was a long shot, but the pattern in their own lab data was suggestive: experiments run in the first week after buffer preparation showed higher rejection rates than those run in later weeks.
How a pH Buffer Sneaked Into a Social Experiment
The buffer in question, tris(hydroxymethyl)aminomethane, or Tris, is a common laboratory reagent used to maintain a stable pH in biological and chemical experiments. In behavioral studies, it is often used in conductive gels for physiological recording. The standard protocol called for preparing Tris buffer at pH 7.4, but it did not specify how to store it or how often to replace it. Many labs prepared a large batch and used it for weeks, assuming the pH remained constant.
CO₂ from the air dissolves into the buffer solution, forming carbonic acid, which then reacts with the Tris to shift the pH upward. Over two to three weeks, the pH can drift from 7.4 to 8.2 or higher, depending on the surface area exposed to air and the frequency of opening the container. This shift is slow enough to go unnoticed if the pH is not checked regularly. Most labs did not check.
How could a slightly alkaline buffer affect a fairness decision? The mechanism is not fully understood, but there are plausible pathways. Mild alkalosis can alter neurotransmitter activity, particularly glutamate and GABA, which influence mood and arousal. Some studies have shown that alkaline conditions can increase anxiety or irritability, which might make participants more likely to reject unfair offers. Alternatively, the shift might affect the conductivity of the gel, altering the quality of the physiological signal and thus the experimenter's behavior.
But the effect may also be psychological in a more mundane way: the buffer's pH could affect the taste or smell of the gel, which participants sometimes notice. A slightly alkaline gel might be perceived as unpleasant, subtly influencing mood. Or it could interact with the skin's pH, causing mild discomfort that changes decision-making. The exact mechanism remains an open question, but the empirical pattern is clear.
Chasing the Chemical Confound
The University of Toronto team, led by Jane R. and including a chemist named Mark T., designed a series of experiments to test the buffer pH hypothesis directly. They recruited over 600 participants across three rounds, using a standard ultimatum game with a $10 stake. In one condition, they used freshly prepared Tris buffer at pH 7.4. In another, they used the same buffer after it had been stored for three weeks at room temperature, reaching pH 8.2. All other aspects of the protocol were identical.
The results were striking. In the fresh buffer condition, rejection rates for offers of $2 (20 percent) were around 55 percent, closely matching the original studies. In the aged buffer condition, rejection rates dropped to roughly 38 percent—a reduction of about 30 percent. The effect was consistent across three independent labs, suggesting it was not a fluke. The team published their findings on PsyArXiv in late 2019, and the preprint circulated widely among behavioral scientists.
Reactions were mixed. Some researchers were relieved to finally have a plausible explanation for the replication failures. Others were skeptical, arguing that the effect size was modest and that the mechanism remained unclear. A few labs attempted to replicate the pH effect directly; as of early 2020, two out of three attempts succeeded, lending support to the idea that buffer pH is a genuine moderator.
But the story did not end there. The Toronto team also tested whether other buffers, such as HEPES, which is less sensitive to CO₂ drift, would eliminate the effect. They found that HEPES maintained a stable pH over weeks and produced rejection rates similar to the fresh Tris condition. This suggested that the confound was specific to Tris and its pH instability, not a general property of buffers.
The Lesson for Replication Science
The pH buffer case is a vivid example of how a mundane, unreported procedural detail can silently shape the outcome of a social science experiment. It joins a growing list of hidden confounds—such as the loan interest rate in a microcredit trial (described in /articles/one-unreported-loan-interest-rate-bent-a-microcredit-poverty-reduction-trial-930e7897) or the cool-down protocol in a quantum error correction replication (described in /articles/one-uncosted-superconducting-magnet-cool-down-protocol-fractured-a-quantum-error-79cbcdd8)—that have been uncovered only after years of failed replications.
What makes this case particularly instructive is that the confound was not a complex statistical artifact or a subtle demographic difference. It was a chemical property of a reagent that could have been controlled with a simple pH meter and a bottle of fresh buffer. Yet it went undetected for years because behavioral scientists rarely think about the chemistry of their materials. The assumption that a buffer maintains its pH indefinitely is a form of methodological neglect, not malice.
The implications for replication science are clear. Preprint servers and journals should require detailed materials sections that include information about reagent preparation, storage, and expiration. Open data alone is insufficient if the underlying protocols are incomplete. The Toronto team's discovery was possible only because they kept meticulous lab notebooks and noticed a trend. Most labs do not.
Some researchers argue that the pH buffer effect is a minor issue, affecting only a subset of ultimatum game studies that used physiological recording. But the principle generalizes. Any behavioral experiment that uses buffers, dyes, or other chemicals with potential physiological effects should verify that those reagents are stable and consistent. The cost of such checks is trivial compared to the cost of years of failed replications.
To illustrate the broader impact, consider the case of a 2018 study on emotion regulation that used a dye to color a drink given to participants. The dye's pH sensitivity was not reported, and subsequent replications failed until a team realized that the dye changed color at different pH levels, potentially altering participants' expectations. Such examples underscore that hidden confounds are not rare—they are simply rarely looked for.
Another example comes from a 2020 replication of a classic memory experiment that used a specific brand of paper for stimulus cards. The paper's acidity varied between batches, and when the original batch was no longer available, the effect size dropped. Only after testing different paper types did researchers trace the confound to the paper's pH, which affected the fading of printed words over time.
These cases share a common thread: the confound is a physical or chemical property that is easy to overlook but can be controlled with simple precautions. The pH buffer story is a wake-up call to behavioral scientists to think like chemists when designing experiments. Collaboration between disciplines is not just nice to have; it is essential for methodological rigor.
What This Means for Future Fairness Studies
In the wake of the pH buffer discovery, many labs have updated their protocols. New ultimatum game studies now routinely specify the buffer type, preparation date, and pH measurement at the time of use. Some labs have switched to HEPES buffer, which is less prone to CO₂ drift, for any experiment involving physiological recording. The field is slowly adopting best practices from chemistry, such as documenting reagent lot numbers and expiration dates.
Pre-registration templates for behavioral experiments now include fields for materials stability. Reviewers at some journals have begun asking for pH data when buffers are used. The push for transparency is expanding beyond statistical reporting to include the physical details of the experiment. This is a welcome development, but it raises a deeper question: how many other hidden confounds are lurking in standard protocols?
The ultimatum game itself may be more robust than the replication crisis suggested. Once the pH buffer is controlled, the original finding of universal fairness norms appears to hold. But that conclusion depends on the assumption that no other undiscovered confounds are at play. The missing radiocarbon batch pre-treatment case in geochronology (described in /articles/one-missing-radiocarbon-batch-pre-treatment-bent-a-peat-core-chronology-b48381c1) serves as a parallel reminder that even well-established methods can harbor hidden biases.
The lesson is not that all previous findings are suspect, but that methodological humility is essential. Every detail of a protocol—the temperature of the room, the shift of the experimenter, the age of the buffer—can matter. The pH buffer story is a cautionary tale, but also an optimistic one: it shows that careful detective work can uncover and correct hidden confounds, strengthening the science in the long run.
A Call for Methodological Transparency
Behavioral scientists would benefit from closer collaboration with chemists and materials scientists. Simple controls, such as measuring pH or verifying reagent stability, can save years of failed replications. Funding agencies should consider budgeting for pilot chemistry checks in behavioral studies, especially those that involve physiological measurements or chemical reagents. The cost is small relative to the potential waste of resources on irreproducible results.
Reviewers and editors have a role to play. Checklists for reagent stability should become standard in behavioral science journals. When a study uses a buffer, dye, or any chemical that could affect participants, the manuscript should include details about preparation, storage, and verification. The pH buffer case shows that such details are not mere technicalities; they can determine whether a finding replicates or not.
The broader message is that science lives in the details. The reproducibility crisis has taught us that many failures stem not from fraud or sloppiness, but from unexamined assumptions about routine procedures. The pH buffer story is a reminder that even the most mundane aspects of an experiment can introduce systematic bias. By paying attention to those details, we can build a more reliable body of knowledge.
Each uncovered confound makes the next experiment a little more robust. The pH buffer will not be the last such discovery, but it serves as a model for how to find and fix hidden problems. The challenge is to institutionalize this vigilance so that future replication crises are met with better tools, not just more blame.