Language Selection

Get healthy now with MedBeds!
Click here to book your session

Protect your whole family with Orgo-Life® Quantum MedBed Energy Technology® devices.

Advertising by Adpathway

         

 Advertising by Adpathway

Multi-institutional study benchmarks human variability in pediatric bone age AI

3 hours ago 7

PROTECT YOUR DNA WITH QUANTUM TECHNOLOGY

Orgo-Life the new way to the future

  Advertising by Adpathway

Bone age assessment, one of the most widely performed tasks in pediatric radiology, is far more variable than clinicians might assume — and in some respects, artificial intelligence already outperforms the humans doing it. That is the central message of a large new multi-institutional study published in Pediatric Radiology, which set out to quantify precisely how much radiologists disagree with one another when estimating skeletal maturity from hand and wrist X-rays, and to establish a rigorous human benchmark against which automated bone age algorithms can be judged.

The study, led by Jin Long of Stanford University School of Medicine together with David B. Larson, Curtis P. Langlotz, Sergios Gatidis and an international team of collaborators, analyzed 1,285 left-hand and wrist radiographs drawn from five US academic centers. Each image was independently interpreted by four different radiologists using the Greulich and Pyle atlas, the century-old gold-standard method in which a reader visually matches a child’s radiograph to standardized reference images of skeletal development. The interpretations came from a centrally administered pool of 22 radiologists recruited across nine institutions, creating one of the largest and most systematically controlled datasets ever assembled for this purpose.

The technical heart of the study lies in its statistical approach. Rather than simply reporting raw disagreement between readers, the researchers applied mixed-effects modeling to decompose variability into its constituent components: variability attributable to the image itself, to individual raters, and to the institutions from which readers and cases originated. This decomposition matters because raw disagreement conflates different sources of error. By adjusting for patient age and sex, the team could isolate how much of the observed spread reflected genuine differences in how radiologists interpret the same image, how much reflected systematic biases held by particular readers, and how much was merely an artifact of the case mix at different hospitals.

The results are striking. Among children younger than 12 years, the standard deviation of bone age estimates for the same image across four radiologists was 8.7 months — more than double the 4.2-month variability seen in older patients. This means that for a young child, two experienced radiologists looking at the identical X-ray can routinely disagree by well over a year in their estimates of skeletal maturity, a spread large enough to alter clinical decisions in fields ranging from endocrinology to orthopedic surgery. Variability was also consistently greater in male patients than in female patients, suggesting that the atlas-based method is inherently harder to apply in some populations than others.

Perhaps the most reassuring finding concerns institutions. Between-institution variability was initially measured at 8.9 months, a figure that might suggest systematic differences in practice across centers. But once the researchers adjusted for differences in patient age across institutions, that figure collapsed to just 0.3 months. In other words, hospitals were not really interpreting bone age differently from one another; their raw disagreement was almost entirely explained by the ages of the children they happened to scan. Inter-rater variability itself remained stable at 1.2 months after adjustment, indicating a consistent underlying level of human disagreement that persists regardless of context.

The mixed-effects models also uncovered something more troubling: several individual radiologists showed systematic biases related to patient age or sex. Rather than making random errors scattered evenly around the true value, these readers tended to overestimate or underestimate bone age in specific subgroups. Such systematic bias is clinically consequential because, unlike random noise, it does not average out when a single reading guides treatment — for example, in estimating remaining growth potential before limb-lengthening surgery or in dosing decisions for growth hormone therapy.

With this human variability map in hand, the team turned to the machines. Two automated methods were evaluated: BoneXpert, a commercially established software that models skeletal maturity from automatically segmented hand bones, and a deep-learning convolutional neural network algorithm trained on pediatric radiographs. Both were assessed in an interchangeability analysis against a three-rater consensus reference, with performance quantified using mean absolute error (MAE), root mean square error (RMSE), and the rate of substantial deviations, defined as estimates differing from consensus by more than 1.8 years.

The comparison produced a result that challenges the common assumption that AI must be held to superhuman standards before clinical adoption. BoneXpert achieved a mean absolute error of 4.8 months, notably lower than the 6.5-month average error of individual human readers, and it also exhibited lower variability. The deep-learning algorithm performed at a level essentially equivalent to human readers, with an MAE of 6.3 months. The most dramatic difference emerged in the tail of the distribution: substantial deviations of more than 1.8 years occurred in only 0.7% of BoneXpert estimates and 2.4% of deep-learning estimates, compared with 3.4% of single human readings. Even the less accurate AI system produced catastrophic outliers less often than an average radiologist working alone.

The authors are careful to frame these findings not as a victory lap for automation but as a methodological contribution. AI systems are typically evaluated against a single reference standard, which can make their errors look alarming in isolation. By quantifying the full structure of human variability — its dependence on age, sex, reader and institution — the study provides a quantitative benchmark for interpreting AI performance in context. An algorithm whose error rate falls within the range of disagreement between trained radiologists is, by definition, operating at a level of reliability that clinical medicine has long accepted from humans. The study’s interchangeability analysis, which asks whether an automated estimate could substitute for a human one without degrading the quality of care, offers a more meaningful test than simple accuracy metrics alone.

The clinical significance of bone age assessment extends well beyond radiology departments. Pediatric endocrinologists rely on skeletal maturity to diagnose and manage growth disorders, precocious and delayed puberty, and hormone therapy. Orthopedic surgeons use it to time interventions for leg-length discrepancies and scoliosis. Because the Greulich and Pyle atlas was developed from children of the mid-twentieth century and largely from populations of European descent, prior research has documented applicability problems across ethnic groups — a known limitation that compounds the reader variability documented in the new study.

The research also carries a pointed message for the rapidly growing field of medical AI evaluation. Headlines frequently celebrate algorithms that match or exceed specialist performance, but rigorous comparisons require knowing what human performance actually is — including its biases, its dependence on case difficulty, and its variability across readers and institutions. Studies like this one, built on multi-reader, multi-case designs borrowed from the methodology of reader studies in diagnostic imaging, supply the missing denominator. Without them, claims of AI parity or superiority rest on reference standards that are themselves noisy.

The study originated as a retrospective secondary analysis of data from a prospective multicenter randomized controlled trial, and its ethical approvals covered all participating institutions, including Stanford, Boston Children’s Hospital, Cincinnati Children’s Hospital, Children’s Healthcare of Atlanta, and Yale. One declared conflict of interest is worth noting: BoneXpert developer Visiana funded the study and provided technical support, with the company’s chief executive among the co-authors, although the investigators state that data analysis and interpretation were performed independently and that the company had no access to patient-level information.

For pediatric patients, the practical implications are immediate. The finding that variability is greatest in younger children — precisely the group in which growth disorders often present and in which early intervention matters most — suggests that single-reader atlas-based assessments deserve particular caution below age 12. The demonstration that automated systems can reduce both average error and the rate of large outliers supports the growing adoption of AI-assisted bone age reading, whether as a standalone estimate or as a second opinion alongside the radiologist. Previous randomized trial data from some of the same Stanford investigators had already shown that AI assistance improves radiologist performance in skeletal age assessment; the new study explains why, by revealing how much room for improvement human variability leaves.

As automated interpretation spreads through pediatric imaging, this work establishes something the field has lacked: a rigorous, multi-institutional measure of the human baseline. Against that baseline, both BoneXpert and deep-learning approaches now stand measured — and both are found to sit comfortably within, or beyond, the bounds of what experienced radiologists routinely deliver.

Subject of Research: Human reader variability in pediatric bone age assessment using the Greulich and Pyle atlas, and comparison of that variability with two automated bone age estimation methods as a benchmark for AI evaluation

Subject of Research: Cancer

Article Title: Quantifying human variability in pediatric bone age assessment: a multi-institutional multi-reader study and benchmark for AI evaluation

Article References: Long, J., Larson, D. B., Thodberg, H. H., Thrane, P. B., Phung, M., Silva, C. T., Lall, N. U., Towbin, A., Prabhu, S. P., Langlotz, C. P., & Gatidis, S. (2026). Quantifying human variability in pediatric bone age assessment: a multi-institutional multi-reader study and benchmark for AI evaluation. Pediatric Radiology. https://doi.org/10.1007/s00247-026-06758-0

Image Credits: AI Generated

DOI: 10.1007/s00247-026-06758-0

Keywords: Artificial intelligence, Bone age, Child, Machine learning, Observer variation, Radiography, Reproducibility of results, Greulich and Pyle atlas, Pediatric radiology, BoneXpert, Deep learning, Skeletal maturity

Cite Scienmag News
APA MLA Chicago

Harold Sullivan. (September 7, 2026). Multi-institutional study benchmarks human variability in pediatric bone age AI. Scienmag. https://scienmag.com/multi-institutional-study-benchmarks-human-variability-in-pediatric-bone-age-ai/

Copy citation Download RIS

Tags: AI outperforming human bone age estimationAI performance in pediatric radiologyautomated bone age algorithms benchmarkingclinical implications of AI in pediatric radiologyGreulich and Pyle atlas accuracyhuman vs AI bone age comparisoninter-rater agreement in pediatric imaginglarge dataset of pediatric radiographslarge-scale radiograph dataset analysismulti-center radiology research collaborationmulti-institutional radiology studypediatric bone age assessmentpediatric radiology accuracy metricspediatric wrist and hand X-ray analysisradiologist disagreement measurementradiologist variability in skeletal maturity estimationskeletal development assessment standardsskeletal maturity evaluationstandardized reference image matchingvariability in radiologist interpretations

Read Entire Article

         

        

Start the new Vibrations with a Medbed Franchise today!  

Protect your whole family with Quantum Orgo-Life® devices

  Advertising by Adpathway