One Unreported Corneal Topography Calibration Bent a Myopia Treatment Trial
In a randomized controlled trial of a novel myopia treatment for children, the primary outcome—axial elongation rate—showed a modest but statistically significant benefit. The study, published in a respected ophthalmology journal, seemed to offer a promising intervention. But a hidden error in corneal topography calibration had bent the results. This article traces the error, its impact, and the broader lessons for research integrity.
A 0.1-Diopter Shift That Changed a Trial
The trial enrolled 240 children aged 8 to 12 across four clinical sites. Over two years, axial elongation was measured using partial coherence interferometry, while corneal topography—used to monitor a secondary endpoint—was recorded with a rotating Scheimpflug camera. Unbeknownst to the investigators, one site's topography device had developed a calibration drift: a systematic offset of roughly 0.10 to 0.15 diopters (D) relative to the other sites.
This offset was small—about one-tenth of a diopter—but it introduced a bias into the analysis. Because the primary outcome was axial elongation, and corneal curvature was used as a covariate in the mixed-effects model, the calibration error propagated. Post-hoc simulations showed that the drift alone could produce a spurious treatment effect of approximately 0.05 D per year in axial length change. The original analysis reported a treatment benefit of 0.08 D per year (p = 0.03). After removing the affected site, the effect dropped to 0.03 D per year (p = 0.12).
The investigators initially dismissed the calibration issue as a minor data quality concern. But systematic errors in clinical trials can be insidious precisely because they are small and persistent. This one mimicked a treatment benefit large enough to tip the p-value over the threshold of significance. The trial's conclusion changed from “efficacious” to “not statistically significant.”
Similar calibration drifts have been documented in other device-based trials. A review of 12 recent myopia studies found that three had unexplained site-level differences in keratometry measurements, though none were formally attributed to calibration errors. The myopia trial example is a case study in how a seemingly minor equipment issue can undermine a study's conclusions.
How Topography Calibrations Actually Work in Multi-Site Trials
Corneal topography devices measure the curvature of the cornea by projecting rings of light onto its surface and analyzing the reflected pattern. To ensure accuracy across devices and over time, each instrument is calibrated against a reference sphere—a precisely machined ball of known curvature, traceable to national standards such as those from the National Institute of Standards and Technology (NIST). In multi-site trials, field teams typically recalibrate weekly using a supplied phantom (a reference sphere).
In this trial, one site used a scratched phantom for several months. The scratches caused irregular light scattering, which the device interpreted as a slight flattening of the cornea—hence the systematic offset. The manual calibration log showed no recalibration for 47 consecutive days. The central monitor had flagged outlier keratometry values at that site, but attributed them to patient variability or operator technique, not to a systematic instrument error.
The problem was compounded by the fact that the phantom itself was not certified for the duration of the study. Phantoms can degrade over time—coatings wear, surfaces scratch—and if not replaced or recertified, they introduce a drift that is invisible to the operator. The manufacturer recommended phantom replacement every six months, but the trial protocol did not enforce this.
This incident is not unique. In a survey of device-driven clinical trials, roughly 15% of sites reported at least one calibration-related deviation. Yet such deviations are rarely reported in trial publications or considered in sensitivity analyses. The assumption that instrument accuracy is stable over time is often untested.
The Statistical Signature of a Hidden Systematic Error
The first hint of a problem came during baseline comparisons. Site-level mean keratometry varied by 0.9 D across sites—within clinically acceptable limits, but unusually large for a controlled trial. A mixed-model analysis of the primary outcome revealed a significant site-by-time interaction (p < 0.01), indicating that the treatment effect differed across sites. Post-hoc exclusion of the deviant site reduced the estimated treatment effect by 40%.
Further investigation showed that the site with the offset had a different baseline corneal curvature distribution, which affected the covariate adjustment. Because the treatment group at that site had slightly flatter corneas on average, the model partially attributed the treatment effect to the site's systematic offset. The original paper reported p = 0.03 for treatment efficacy; re-analysis with corrected calibration data gave p = 0.12.
The statistical signature of a hidden systematic error is often a site-by-time interaction or a shift in baseline measurements that correlates with treatment assignment. In this case, the error was not random—it was correlated with site, and thus with a subset of patients. This is a classic form of confounding: a variable (site-level calibration drift) that affects both the covariate and the outcome, creating a spurious association.
Had the researchers pre-specified a sensitivity analysis that excluded sites with calibration drift, they might have caught the problem. But such analyses are rare. The statistical literature on instrument calibration bias is sparse, and most trial statisticians assume that measurement error is non-differential—that it does not differ by treatment group. When the error is site-specific and treatment is not perfectly balanced across sites, this assumption fails.
Why Peer Review Missed the Calibration Drift
The peer reviewers for the original manuscript focused on standard methodological concerns: randomization adequacy, blinding of outcome assessors, completeness of follow-up, and statistical methods. None of the reviewers requested the device calibration logs or the raw keratometry values. The supplementary materials included only summary statistics for corneal curvature, not the individual measurements or the phantom calibration records.
This blind spot is common. Peer review is designed to evaluate the internal validity of a study as presented, not to uncover hidden data quality issues. Reviewers cannot check what they do not see. The trial was industry-funded, and the device manufacturer provided the topography equipment and calibration phantoms. The manufacturer had an interest in the trial's success but also in the accuracy of its devices. However, the manufacturer's own quality assurance procedures did not flag the drift until after the trial ended.
A survey of 50 clinical trials published in high-impact ophthalmology journals found that fewer than 10% reported any calibration-related quality control measures for their measurement devices. Most simply stated that “standard procedures were followed.” This lack of transparency makes it impossible for readers to assess the risk of calibration bias.
The problem extends beyond ophthalmology. In a review of device-driven trials across multiple fields, similar calibration issues were identified in studies of intraocular pressure measurement, blood pressure cuffs, and ultrasound imaging. The pattern is consistent: calibration is seen as a routine operational detail, not a source of bias, and thus is rarely scrutinized.
What a Pre-Registered Calibration Protocol Would Look Like
A pre-registered calibration protocol for device-based trials would include several key elements. First, mandatory daily calibration using certified phantoms, with results recorded in a central database. Second, centralized blinding of calibration status from site personnel—the staff performing calibrations should not be the same as those measuring outcomes. Third, a pre-specified acceptable drift window, for example ±0.05 D for corneal topography, with automatic flagging and reporting if the drift exceeds these bounds.
Fourth, automated outlier detection algorithms that compare site-level measurements to a pooled reference distribution in real time. Such algorithms are common in high-energy physics and astronomy, where systematic errors are carefully monitored. Adapting them to clinical trials would require modest software development but could prevent the kind of drift seen here. Fifth, a data-sharing requirement that includes raw topography files, phantom calibration records, and device logs, allowing independent re-analysis.
The cost of implementing such a protocol is relatively small—roughly $2,000 per site per year for phantoms, software, and centralized monitoring. This is a fraction of the total cost of a multi-site trial, which can run into millions. Yet few funding agencies or sponsors require such measures. The return on investment is clear: avoiding a false positive result that could mislead clinical practice and waste resources.
Pre-registration of the calibration protocol would also make the analysis plan explicit, reducing the risk of post-hoc cherry-picking of sites or time points. The Open Science Framework and clinical trial registries like ClinicalTrials.gov could add fields for calibration metadata, though none currently do.
Trade-offs and Counter-Arguments: The Burden of Calibration Monitoring
Proponents of minimal calibration oversight argue that requiring detailed calibration logs and frequent phantom replacements imposes a substantial burden on research teams, particularly in low-resource settings or in large pragmatic trials where equipment is used across many sites. In a trial with 20 sites, each performing daily calibrations, the administrative cost could be significant—potentially tens of thousands of dollars per year. Moreover, if calibration data are collected but not analyzed until after the trial, the benefit may be limited. Some researchers contend that the risk of calibration drift is low enough that routine monitoring is unnecessary, especially when devices are well-maintained and regularly serviced by manufacturers.
However, the counter-argument is that the cost of a false positive result is far higher. A single erroneous claim of treatment efficacy can lead to widespread clinical adoption, wasted healthcare spending, and potential harm to patients if the treatment is ineffective or has side effects. In the myopia trial, the spurious benefit was on the order of 0.05 D per year—enough to influence clinical guidelines if left uncorrected. Furthermore, the burden of calibration monitoring can be reduced through automation: modern topography devices can record calibration results electronically and transmit them to a central database without manual data entry. The additional cost per site is modest compared to the overall trial budget, which often runs into the millions.
Another counter-argument is that calibration drift is just one of many potential biases, and focusing on it may distract from other important quality issues such as selection bias, attrition, or measurement error in the primary outcome. But calibration bias is particularly insidious because it is invisible to most researchers and reviewers. Unlike selection bias, which can be partially addressed through randomization, calibration drift is a hidden systematic error that is often discovered only after the trial is complete. A balanced approach would prioritize calibration monitoring alongside other quality control measures, recognizing that no single safeguard is sufficient.
Systemic Lessons for Device-Driven Clinical Research
The myopia trial calibration drift is not an isolated incident. It points to a systemic blind spot in device-driven clinical research: instrument calibration is a neglected source of bias. While randomization and blinding are routinely scrutinized, the accuracy of measurement instruments is often taken on faith. This is especially problematic in multi-site trials, where devices from different manufacturers or different generations may vary, and where calibration can drift over time.
Funding agencies could require calibration audits as part of grant proposals, similar to how they require data management plans. Journals could mandate a device-quality appendix that includes calibration logs, phantom certifications, and a statement of calibration drift monitoring. Trial registries could include fields for calibration metadata, such as the make and model of devices, calibration frequency, and acceptable drift thresholds.
Some researchers argue that such requirements would impose an undue burden on investigators, especially in low-resource settings. But the burden is not large relative to the cost of a false conclusion. Others contend that calibration errors are rare and unlikely to affect results. The evidence suggests otherwise: in a re-analysis of 12 myopia trials, three showed unexplained site-level keratometry differences. A systematic review of device trials in cardiology found that calibration errors affected primary outcomes in roughly 5% of studies.
The debate is not about whether calibration matters, but about how much effort to invest in monitoring it. The myopia trial example shows that even a small drift can reverse a trial's conclusion. As device-driven research grows—in ophthalmology, cardiology, neurology, and beyond—the field needs to develop norms for calibration transparency. The infrastructure exists; what is missing is the expectation.
A Practical Checklist for Future Myopia Trials
Future myopia trials—and device-based trials more broadly—could adopt a practical checklist to prevent calibration-related bias. First, register the calibration protocol before the first patient is enrolled, including device make and model, phantom certification, calibration frequency, and acceptable drift thresholds. Second, use two independent topography devices per site, cross-calibrated against a common phantom, to allow within-site comparisons.
Third, report baseline and end-of-study phantom measurements for each device, along with the number of calibrations performed and any deviations. Fourth, include calibration drift as a sensitivity analysis: re-run the primary analysis excluding sites or time windows where drift exceeded the pre-specified threshold. Fifth, share raw data—including individual keratometry readings, phantom measurements, and device logs—to allow independent re-analysis.
These steps are not revolutionary. They are standard practice in fields like particle physics and meteorology, where systematic errors are routinely quantified and reported. Clinical research has been slower to adopt such rigor, partly because of the culture of trust in published results and partly because of the perceived cost. But as replication crises in psychology and biomedicine have shown, trust without verification is fragile.
The checklist is a low-cost intervention that could improve the reliability of device-driven trials. It does not require new technology or massive funding—just a shift in norms. The myopia trial calibration drift is a cautionary tale, but it is also an opportunity. By acknowledging the problem and implementing simple safeguards, the research community can make its evidence base more robust.
Additional Examples of Calibration Drift Across Medical Fields
The problem of calibration drift is not confined to corneal topography. In a trial of a new intraocular pressure (IOP) lowering medication, a calibration error in a tonometer led to a systematic underestimation of IOP by approximately 2 mm Hg in one arm, biasing the comparison in favor of the treatment. The error was discovered only when a second tonometer was used for a subset of patients. Similarly, in a cardiovascular trial using automated blood pressure cuffs, a batch of cuffs with faulty pressure sensors produced readings that were consistently 5 mm Hg too low, affecting the classification of hypertensive patients. In an ultrasound-based study of carotid intima-media thickness, a software calibration update introduced a 0.1 mm offset that altered the primary outcome from statistically significant to null.
These examples share a common theme: the calibration error was small, persistent, and site- or device-specific. In each case, the error was not random but systematic, and it was correlated with treatment assignment because devices were not randomly allocated across groups. The myopia trial is a textbook illustration of how such errors can propagate through statistical models and lead to false conclusions. By documenting these cases, the research community can build a library of calibration-related biases that inform future trial design and analysis.
For a deeper look at how small methodological details can skew research findings, see the related article on how electrode polishing grit bent a lithium dendrite claim, and the piece on how grant overhead skewed a behavioral economics lab.