Table of Contents
- Key Points
- Why This Research Matters
- Understanding the Ultrasound Models
- How the Research Was Conducted
- Key Findings: How Accurate Is Each Model?
- Subgroup Findings: Menopause and Cancer Prevalence Matter
- What This Means for Patients
- Study Limitations
- Recommendations for Patients
- Frequently Asked Questions
- Source Information
Key Points
- In a review of 99 studies and 42,496 ovarian tumors, expert subjective ultrasound assessment was the only model with both sensitivity and specificity above 90%.
- The ADNEX model at a 20% cut-off showed sensitivity of 86.7% and specificity of 87.9%, with performance similar to expert assessment in this analysis.
- The Risk of Malignancy Index had the lowest accuracy, identifying only about 70% of cancers, yet it remains recommended in several international guidelines.
- Postmenopausal women and those in higher-prevalence settings had lower specificity, meaning more false-positive results; sensitivity was lower in premenopausal women.
- These ultrasound models guide pre-surgical referral and treatment planning; definitive diagnosis requires tissue examination after surgery or biopsy.
Why This Research Matters
Ovarian tumors are a common finding in women of all ages. Each year, approximately 1 in 1,500 to 1 in 1,700 women undergo surgery for an ovarian tumor. Yet only a small minority of these tumors — between 10% and 17% — turn out to be malignant (cancerous). This means the vast majority of women who have surgery for ovarian masses have benign (non-cancerous) growths.
The challenge for doctors is figuring out before surgery which tumors are dangerous and which are harmless. This distinction is critical because the treatment path depends entirely on the tumor's nature. Women with ovarian cancer need to be referred to specialized oncology centers where a gynecological oncologist can perform surgical staging. In contrast, women with benign tumors can often be treated at a general hospital, sometimes with minimally invasive surgery or even conservative management. Avoiding unnecessary radical surgery reduces complications, shortens recovery time, and minimizes the impact on a woman's normal activities.
To assist in these decisions, doctors use various ultrasound-based models that evaluate the features of an ovarian tumor as seen on an ultrasound scan. This meta-analysis — a "study of studies" — evaluated the most widely used models to provide up-to-date, definitive evidence on which approaches work best.
Understanding the Ultrasound Models
The study evaluated five main approaches, each with its own method of assessing ovarian tumors:
- Subjective assessment (SA): An experienced ultrasound examiner looks at the images and uses their expert judgment to determine whether a tumor looks benign or malignant. This relies on a skill called "image recognition" that requires significant training and exposure to many cases.
- Risk of Malignancy Index (RMI) versions 1, 2, and 3: This model combines ultrasound features with a woman's menopausal status and a blood test for cancer antigen 125 (CA-125). Three different versions (RMI1, RMI2, RMI3) have been developed, typically using cut-off scores of 200 or 250. This was the preferred strategy in The Netherlands until 2021 and remains recommended in several international guidelines.
- Logistic Regression model 2 (LR2): Developed by the International Ovarian Tumor Analysis (IOTA) group, LR2 uses a mathematical formula incorporating clinical and ultrasound variables to calculate a patient's probability of malignancy. This study examined the commonly used 10% probability cut-off (LR2-10%).
- Simple ultrasound-based Rules (SR): Also from the IOTA group, SR uses five ultrasound features suggesting malignancy and five suggesting benignity. Tumors are classified as benign, malignant, or "inconclusive." When results are inconclusive, two approaches are used: considering all inconclusive cases as malignant (SR+Mal) or sending inconclusive cases for expert review (SR+SA).
- Assessment of Different NEoplasias in the adneXa (ADNEX) model: This IOTA model calculates the probability that an ovarian tumor is benign or malignant and can even distinguish between different types of tumors. It can be used with or without a CA-125 blood test, with various probability cut-offs (5%, 10%, 20%, 30%, and 40%) used in clinical practice.
The study excluded some other validated models — such as the Simple Rules Risk assessment (SRR), Risk of Ovarian Malignancy Algorithm (ROMA), and scoring systems by Ferrazzi et al., Lerner et al., and Sassone et al. — to maintain a clear and focused scope on the most commonly used and clinically implemented approaches.
How the Research Was Conducted
The researchers followed strict systematic review guidelines (PRISMA) and registered their protocol in advance on PROSPERO (registration number CRD42022282438), the international registry for systematic reviews.
The research team searched three major medical databases — Ovid/MEDLINE, EMBASE, and the Cochrane Library — from the start of each database until June 19, 2025. They also manually searched the reference lists of other systematic reviews and guidelines. Their search initially identified 3,310 potentially eligible articles: 1,451 from MEDLINE, 1,629 from EMBASE, 230 from Cochrane, and 27 from manual review of references.
After removing duplicates, 2,750 articles remained. Two reviewers independently screened these records, leading to 431 full-text articles being assessed for eligibility. Ultimately, 99 studies met the strict inclusion criteria and were included in the analysis.
To be eligible, studies had to evaluate at least one of the preselected models, collect model parameters prospectively (meaning the ultrasound data was gathered at the time of scanning, as would happen in real clinical practice), and provide enough data to construct a 2×2 table of test results. Studies that applied models retrospectively to stored images or videos were excluded because this does not reflect real-world clinical practice where models guide actual treatment decisions.
In total, the 99 included studies described 42,496 ovarian tumors, of which 31,371 (74%) were benign and 11,125 (26%) were malignant. Most studies (77%) were prospective in design, and 71% were conducted at a single center. The majority (71%) took place in oncology centers, while 22% were in mixed settings, 4% in non-oncology centers, and 3% had unclear settings. The prevalence of malignancy across studies ranged widely, from 3.0% to 63.9%.
The number of studies evaluating each model was substantial:
- RMI version 1: 32 studies
- RMI version 2: 18 studies
- RMI version 3: 10 studies
- LR2: 16 studies
- Simple Rules: 32 studies (28 using the SR+Mal strategy and 18 using the SR+SA strategy)
- ADNEX model: 22 studies
- Subjective assessment: 27 studies
For quality assessment, the researchers used the QUADAS-2 tool and its companion QUADAS-C extension for studies comparing multiple models. Two reviewers independently rated each study's risk of bias in four key areas: patient selection, index test, reference standard, and flow and timing. The results showed that 52% of studies had low risk of bias in patient selection, 82% in the index test, 94% in the reference standard, and 53% in flow and timing. However, only 41% of studies explicitly stated that patients were collected consecutively, and in 45% of studies there was concern about inappropriate patient exclusions.
Statistical analysis included calculating pooled sensitivity and specificity for each model, fitting bivariate models to generate summary receiver-operating-characteristic curves, and conducting meta-regression to determine statistically significant differences between models. The researchers also used funnel plots and Egger's regression test to assess publication bias — which was found to be present (P<0.001), meaning smaller studies tended to report higher accuracy than larger ones.
Key Findings: How Accurate Is Each Model?
The most important measure of a diagnostic test comes down to two numbers: sensitivity (how good the test is at catching cancer when it's truly present — high sensitivity means few cancers are missed) and specificity (how good the test is at correctly identifying benign tumors — high specificity means fewer false alarms and unnecessary surgeries).
Here is how each model performed in the primary analysis, based on 96 studies in which borderline tumors were classified as malignant:
Subjective Assessment (SA) — The Gold Standard
Subjective assessment by expert examiners was the only model in which both sensitivity and specificity exceeded 90%. It achieved a sensitivity of 90.2% (95% CI, 87.8–92.2%) and a specificity of 91.4% (95% CI, 89.3–93.2%). This means that among women with ovarian cancer, about 90 out of 100 were correctly identified by expert ultrasound examiners, and among women with benign tumors, about 91 out of 100 were correctly identified as benign.
Simple Rules plus Expert Review (SR+SA)
The strategy of applying the Simple Rules first and then having inconclusive cases reviewed by an expert demonstrated nearly identical performance to subjective assessment alone. Its sensitivity was 88.6% (95% CI, 85.7–91.0%) and specificity was 91.0% (95% CI, 89.0–92.7%). These differences were not statistically significant (P=0.397 for sensitivity, P=0.811 for specificity).
ADNEX Model
The ADNEX model's performance depended heavily on the cut-off value chosen. At the 10% cut-off, ADNEX had a sensitivity of 92.7% (95% CI, 90.8–94.2%) — statistically similar to subjective assessment (P=0.130) — but a lower specificity of 78.4% (95% CI, 71.7–83.8%), a difference that was highly significant (P<0.001).
At the 20% cut-off, ADNEX showed a more balanced profile: sensitivity of 86.7% (95% CI, 80.6–91.0%) and specificity of 87.9% (95% CI, 80.1–92.9%). Neither differed significantly from subjective assessment (P=0.095 for sensitivity, P=0.119 for specificity).
The study confirmed a predictable pattern: higher cut-offs decreased sensitivity (more cancers being missed) while lower cut-offs reduced specificity (more benign tumors being flagged as potentially malignant).
LR2 Model
At the 10% cut-off, the LR2 model achieved a sensitivity of 89.5% (95% CI, 85.8–92.4%) and a specificity of 82.3% (95% CI, 75.0–87.8%). This places its sensitivity close to that of expert assessment, though with somewhat lower specificity.
RMI — The Lowest Performer
The Risk of Malignancy Index had the lowest diagnostic accuracy of all models evaluated. RMI version 1 at a cut-off of 200 achieved only a sensitivity of 69.7% (95% CI, 67.0–72.2%), meaning it missed roughly 3 out of 10 ovarian cancers. Its specificity was 90.5% (95% CI, 88.3–92.4%). Similar limitations were seen across all three RMI versions.
Subgroup Findings: Menopause and Cancer Prevalence Matter
The researchers performed subgroup analyses to explore whether a woman's menopausal status or the underlying rate of cancer in the study population affected how well these models performed. The findings were striking: both menopausal status and the prevalence of malignancy significantly affected both sensitivity and specificity (P<0.01 for both measures).
Specifically, being postmenopausal and being in a population with higher cancer prevalence were both associated with lower specificity — meaning more false-positive results. Meanwhile, sensitivity was lower in premenopausal women.
This is an important clinical insight. Postmenopausal women with ovarian masses are at higher baseline risk of cancer, yet the models tended to produce more false alarms in this group. Similarly, in settings where many patients have cancer (such as specialized oncology referral centers), the models were more likely to flag benign tumors as suspicious.
What This Means for Patients
The study's conclusions offer reassuring news for women facing evaluation of an ovarian tumor. All of the approaches studied — with the notable exception of the RMI — performed well and can be used to differentiate between benign and malignant tumors.
However, there are important trade-offs to understand:
- Expert assessment remains the most accurate single approach. Subjective assessment achieved the best balance of sensitivity and specificity. Its major limitation, however, is that it requires a highly skilled ultrasound examiner, typically at the EFSUMB Level-3 expertise. In 16 of the 27 studies evaluating SA, all scans were performed by expert examiners.
- If an operator-independent strategy is preferred, ADNEX is the recommended choice. Because of its high sensitivity, the likelihood of missing a malignancy with ADNEX is low — a crucial safety consideration.
- In postmenopausal women, a higher ADNEX cut-off may be warranted. Because specificity is reduced in this group, the study authors say a higher cut-off may be worth considering, depending on how the impact of a false-positive result is weighed; however, higher ADNEX cut-offs lower sensitivity (more cancers missed), and few studies have tested higher cut-offs.
The RMI's relatively poor performance is particularly noteworthy given that it remains recommended in several international guidelines. For women who might previously have been evaluated using the RMI, this study's findings strongly suggest that newer models such as ADNEX or referral to an expert sonographer provide better diagnostic accuracy.
Study Limitations
While this meta-analysis is comprehensive, several limitations should be acknowledged:
- Publication bias: Statistical testing revealed significant funnel plot asymmetry (P<0.001), suggesting that smaller studies with favorable results were more likely to be published. This could mean the reported accuracy figures are somewhat optimistic.
- Operator variability: The skill level of ultrasound examiners varied across studies of subjective assessment, and some studies did not specify their examiners' experience level. This makes it difficult to generalize SA's excellent results to all clinical settings.
- Prospective data requirement: Studies that applied models retrospectively to stored images were excluded to reflect real-world practice. However, this strict criterion may have excluded some potentially informative data.
- Heterogeneity between studies: Studies varied in their healthcare settings, patient populations, ultrasound techniques, and definitions of borderline tumors. Five studies classified borderline tumors as benign or excluded them altogether.
- ADNEX versions combined: Both versions of the ADNEX model (with and without CA-125) were treated as a single strategy, as prior research suggests adding CA-125 doesn't meaningfully alter diagnostic accuracy.
- Settings skewed toward oncology centers: With 71% of studies conducted in oncology centers, the results may not fully reflect performance in general community hospitals or primary care settings where the cancer prevalence is lower.
Recommendations for Patients
If you are facing evaluation of an ovarian tumor, here are key takeaways from this research that you may wish to discuss with your healthcare provider:
- Ask which ultrasound model is being used. If your hospital uses the Risk of Malignancy Index as its primary tool, you might ask whether newer IOTA models such as ADNEX are available, as this study found RMI to be the least accurate approach.
- Request an expert ultrasound examiner if possible. Subjective assessment by an experienced sonographer remains the most accurate method. If your hospital has access to a specialist ultrasound department, this may be worth pursuing.
- Understand the meaning of your results. A "suspicious" finding does not necessarily mean you have cancer — the specificity of these models means that some benign tumors will be flagged as potentially malignant. Conversely, a "benign" result is generally reassuring, though no test is perfect.
- Discuss your menopausal status. This study found that menopausal status significantly affects test performance. Postmenopausal women may have different (and arguably more appropriate) cut-off values, particularly with the ADNEX model.
- Remember that this is about pre-surgical planning. These ultrasound models help guide referral and treatment decisions. The definitive diagnosis is ultimately made through histological examination (examination of tissue under a microscope after surgery or biopsy).
The bottom line: This large, carefully conducted analysis provides strong evidence that modern ultrasound-based models — particularly the ADNEX model and expert subjective assessment — are reliable tools for distinguishing benign from malignant ovarian tumors. The days of relying primarily on the RMI may be numbered, and patients can feel more confident that advances in ultrasound technology are helping doctors make more accurate, individualized treatment decisions.
Frequently Asked Questions
How accurate are ultrasound models at telling whether an ovarian tumor is benign or malignant?
In a review of 99 studies covering 42,496 ovarian tumors, expert subjective assessment was the only approach where both sensitivity and specificity exceeded 90%. The ADNEX model at a 20% cut-off showed sensitivity of 86.7% and specificity of 87.9%. The Risk of Malignancy Index had the lowest accuracy, identifying only about 70% of cancers.
What does it mean if my ultrasound result is called 'suspicious'?
A suspicious finding does not necessarily mean you have cancer. The models studied have specificities below 100%, so some benign tumors are flagged as potentially malignant. For example, ADNEX at a 10% cut-off had a specificity of 78.4%, meaning about 22 in 100 benign tumors were incorrectly flagged. The definitive diagnosis comes from tissue examination after surgery or biopsy.
What happens after my ovarian tumor is assessed with an ultrasound model?
These models help guide referral and treatment planning before surgery. If cancer is suspected, you may be referred to a specialized oncology center where a gynecological oncologist can perform surgical staging. If the tumor appears benign, treatment may be possible at a general hospital, sometimes with minimally invasive surgery or conservative management. The definitive diagnosis is made by examining tissue after surgery or biopsy.
Should I ask my doctor which ultrasound model is being used for my ovarian tumor?
Yes, it is reasonable to ask. This review found the Risk of Malignancy Index had the lowest accuracy, missing about 3 in 10 cancers, yet it remains recommended in several guidelines. You might ask whether newer IOTA models such as ADNEX are available at your hospital, as these showed better diagnostic performance in the studies analyzed.
If my hospital uses the Risk of Malignancy Index to evaluate my ovarian tumor, when should I seek a second opinion?
The Risk of Malignancy Index correctly identified only about 70% of ovarian cancers, missing roughly 3 in 10, while newer IOTA models such as ADNEX and expert subjective assessment performed better. If your hospital relies on RMI as its primary tool, a second opinion can review whether ADNEX or an expert ultrasound examiner is available, and whether your menopausal status warrants a different cut-off. Definitive diagnosis still comes from tissue examination. Diagnostic Detectives Network provides independent expert second opinions.
Source Information
Original article: "Diagnostic accuracy of ultrasound models for assessment of ovarian tumors: systematic review and meta-analysis" by E. Lems, A. H. Koch, E. J. L. G. Delvaux, J. C. Leemans, M. Y. Bongers, C. A. R. Lok, B. L. Ramaekers, and P. M. A. J. Geomini.
Journal: Ultrasound in Obstetrics & Gynecology, 2026; volume 67, pages 590–603. Published online December 5, 2025, in Wiley Online Library. DOI: 10.1002/uog.70135.
Funding: This is an open-access article distributed under the terms of the Creative Commons Attribution License.
This patient-friendly article is based on peer-reviewed research. The study was accepted for publication on September 24, 2025, and its protocol was registered in PROSPERO (CRD42022282438) in December 2022. Readers seeking the full technical details, including all supplementary tables and figures, are encouraged to consult the original publication.