Preview

Diagnostic radiology and radiotherapy

Advanced search

Diagnostic accuracy and agreement of AI services among themselves and with medical expert in evaluating mammography using the BI-RADS scale: a retrospective analytical study

https://doi.org/10.22328/2079-5343-2026-17-2-86-97

Abstract

INTRODUCTION: According to the literature, the diagnostic accuracy of mammographic AI services in determining the BIRADS category can vary quite widely. However, there are no studies that evaluate agreement between AI systems and radiologists in determining the BI-RADS category.

OBJECTIVE: Evaluation of the diagnostic accuracy and agreement of three AI systems with each other and with an expert reader in mammographic examinations through BI-RADS.

MATERIALS AND METHODS: A mixed study was performed, which included retrospective diagnostic and analytical components. The analysis included 99 anonymized mammographic studies. The studies were evaluated by an expert reader with the definition of BI-RADS 1–5 categories separately for the right and left breast; the scores obtained were used as a reference standard. The data set was processed by three commercial AI services, which also identified BI-RADS 1–5 categories.

Statistics: Statistical analysis was performed at the level of individual breasts (n=198). ROC AUC, sensitivity, specificity, and accuracy were evaluated with 95% confidence intervals for two binary BI-RADS scales: BI-RADS 1–3 vs. 4–5 and BI-RADS 12 vs. 3–5. Agreement was calculated using the Pearson and Cohen intraclass correlation method for similar scales and the full BI-RADS scale. The McNemar test was used to compare diagnostic parameters, and a bootstrap analysis with 5,000 repetitions was used to assess consistency differences.

RESULTS: Diagnostic accuracy estimates are presented as ranges of values obtained for individual AI services, and consistency estimates include values obtained for individual pairs of AI services or pairs of «AI service + reader». Diagnostic accuracy of AI services on the BI-RADS binary scale No. 1: ROC AUC — 0.644–0.876, accuracy — 0.788–0.909, specificity — 0.816–0.977, sensitivity — 0.333–0.883. According to the BI-RADS No. 2 binary scale: ROC AUC — 0.795–0.915, accuracy — 0.808–0.939, specificity — 0.818–0.961, sensitivity — 0.739–0.870. Agreement between pairs of AI services on the BI-RADS binary scale No. 1 ranged from 0.232 to 0.504, on the BI-RADS binary scale No. 2 — from 0.416 to 0.663, on the full BI-RADS scale — from 0.387 to 0.564. Agreement between AI services and an expert reader on the binary scale BI-RADS No. 1 ranged from 0.288 to 0.641, on the binary scale BI-RADS No. 2 — from 0.518 to 0.832, on the full BI-RADS scale — from 0.570 to 0.701. 

DISCUSSION: The diagnostic accuracy of AI services on the BI-RADS No. 1 binary scale and BI-RADS No. 2 binary scale corresponded to the ranges in similar studies. Compared to the results of our previous study, the agreement on the BI-RADS No. 1 binary scale, BI-RADS No. 2 binary scale, and the full scale was in the vast majority of cases was lower than the agreement between radiologists.

CONCLUSION: In most cases, there were no statistically significant differences between the diagnostic accuracy parameters of the AI services calculated for the BI-RADS binary scale No. 1 and No. 2. There were also no statistically significant differences in the agreement between pairs of AI services and between AI services and an expert doctor, depending on the type of BI-RADS scale. Considering the various BI-RADS scales, the agreement between AI services and an expert reader was in the vast majority of cases lower than the agreement between radiologists that we evaluated in a previous study.

About the Authors

A. S. Azaryan
Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies; Central State Medical Academy of Department of Presidential Affairs
Russian Federation

Avet S. Azaryan - postgraduate student; assistant Professor at the Department 

24 Petrovka St., Moscow, 127051; Timoshenko St., 19, p. 1A, Moscow, 121359



M. Yu. Khrustacheva
Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies
Russian Federation

Margarita S. Khrustacheva - postgraduate student

24 Petrovka St., Moscow, 127051



L. D. Pestrenin
Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies; MIREA Russian Technological University
Russian Federation

Lev D. Pestrenin - junior Research Fellow, Department of Medical Informatics, Radiomics and Radiogenomics

24 Petrovka St., Moscow, 127051



Yu. A. Vasilev
Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies
Russian Federation

Yuriy А. Vasilev - Dr. of Sci. (Med.), Medical Director

24 Petrovka St., Moscow, 127051



A. V. Vladzymyrskyy
Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies; I. M. Sechenov First Moscow State Medical University
Russian Federation

Anton V. Vladzymyrskyy - Dr. of Sci. (Med.), Deputy Director for R&D

24 Petrovka St., Moscow 127051



O. V. Omelyanskaya
Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies
Russian Federation

Olga V. Omelyanskaya - Chief Administrative Officer of R&D/CAO of R&D S

24 Petrovka St., Moscow 127051



A. S. Domozhirova
Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies
Russian Federation

Alla S. Domozhirova -Dr. of Sci. (Med.), Chair of the Scientific Problem Commission, Scientific Secretary

Moscow



K. M. Arzamasov
Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies; MIREA Russian Technological University; Samara State Medical University
Russian Federation

Kirill M. Arzamasov - Dr. of Sci. (Med.), Head of Department of Medical Informatics, Radiomics, and Radiogenomics

24 Petrovka St., Moscow, 127051



A. S. Gatsuk
Research and Practical Clinical Center for Diagnostics and Telemedicine Technologies
Russian Federation

Anderei S. Gatsuk - junior Research Fellow, Department of Medical Informatics, Radiomics and Radiogenomics

24 Petrovka St., Moscow, 127051



References

1. Bray F., Laversanne M., Sung H. et al. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries // CA: A Cancer Journal for Clinicians . 2024. Vol. 74, No. 3. Р. 229–263. doi: 10.3322/caac.21834.

2. Vladzymyrskyy A.V., Vasilev Yu.A., Shakhov A.V. et al. The impact of artificial intelligence technologies on active breast cancer detection: efficiency and scalability. The CIS Healthcare , 2026, Vol. 2, No. 1, pp. 11–19 (In Russ.)]. doi: 10.21045/3033-6341-2026-2-1-11-19.

3. Merabishvili V.M. The state of oncological care in Russia: breast cancer among the female population. Morbidity, mortality, reliability of accounting, detailed localization and histological structure. Population-based research at the federal district level. Issues of oncology , 2022, Vol. 68, No. 3, pp. 286293 (In Russ.)]. doi: 10.37469/0507-3758-2022-68-3-286-293.

4. Ren W., Chen M., Qiao Y., Zhao F. Global guidelines for breast cancer screening: A systematic review // Breast. Churchill Livingstone . 2022. Vol. 64. Р. 85–99. doi: 10.1016/j.breast.2022.04.003.

5. Spak D.A., Plaxco J.S., Santiago L., Dryden M.J., Dogan B.E. BI-RADS® fifth edition: A summary of changes // Diagnostic and Interventional Imaging . 2017. Vol. 98, No. 3. Р. 179–190. doi: 10.1016/j.diii.2017.01.001.

6. Vasiliev Yu.A., Son I.M., Pestrenin L.D., Vladzimirsky A.V. Effectiveness of preventive mammography in the Russian Federation: comparison of the results of the first stage of medical examination in 2019 and 2022. Healthcare Manager , 2024, Vol. 11, pp. 63–76 (In Russ.)]. doi: 10.21045/1811-0185-2024-11-63-76.

7. Cai H., Wang J., Dan T. et al. An Online Mammography Database with Biopsy Confirmed Types // Scientific Data . 2023. Vol. 10. Р. 123. doi: 10.1038/s41597-02302025-1.

8. Vasiliev Yu.A., Tyrov I.A., Vladzimirsky A.V. et al. Double viewing of mammography results using artificial intelligence technologies: a new model for organizing mass preventive research. Digital Diagnostics , 2023, Vol. 4, No. 2, pp. 93–104 (In Russ.)]. doi: 10.17816/DD321423.

9. Arzamasov K.M., Vasiliev Yu.A., Vladzimirsky A.V. et al. Application of computer vision for preventive examinations using mammography as an example. Preventive Medicine , 2023, Vol. 26, No. 6, pp. 117–123 (In Russ.)]. doi: 10.17116/profmed202326061117.

10. Dang L.A., Chazard E., Poncelet E. et al. Impact of artificial intelligence in breast cancer screening with mammography // Breast Cancer . 2022. Vol. 29, No. 6. Р. 967–977. doi: 10.1007/s12282-022-01375-9.

11. Liu J., Lei J., Ou Y. et al. Mammography diagnosis of breast cancer screening through machine learning: a systematic review and meta-analysis // Clin. Exp. Med 2022. Vol. 23, No. 6. Р. 2341–2356. doi: 10.1007/s10238-022-00895-0.

12. Le Boulc’h M., Bekhouche A., Kermarrec E. et al. Comparison of breast density assessment between human eye and automated software on digital and synthetic mammography: Impact on breast cancer risk // Diagn. Interv. Imaging . 2020. Vol. 101, No. 12. Р. 811–819. doi: 10.1016/j.diii.2020.07.004.

13. Ji H., Jang Mjin, Chang J.M. Variability in Breast Density Estimation and Its Impact on Breast Cancer Risk Assessment // J. Breast Cancer . 2024. Vol. 27, No. 5. Р. 334–342. doi: 10.4048/jbc.2024.0101.

14. Rigaud B., Weaver O.O., Dennison J.B. et al. Deep Learning Models for Automated Assessment of Breast Density Using Multiple Mammographic Image Types // Cancers (Basel) . 2022. Vol. 14, No. 20, Р. 5003. doi: 10.3390/cancers14205003.

15. Bobrovskaya T.M., Vasiliev Yu.A., Nikitina N.Yu. et al. Sample size for assessing the diagnostic accuracy of software based on artificial intelligence technologies in radiation diagnostics. Siberian Journal of Clinical and Experimental Medicine , 2024, Vol. 39, No. 3, pp. 188–198 (In Russ.)]. doi: 10.29001/2073-8552-2024-39-3-188-198.

16. Tan H., Wu Q., Wu Y. et al. Mammography-based artificial intelligence for breast cancer detection, diagnosis, and BI-RADS categorization using multi-view and multi-level convolutional neural networks // Insights Imaging . 2025. Vol. 16, No. 1. Р. 109. doi: 10.1186/s13244-025-01983-x.

17. Vasiliev Yu.A., Vladzimirsky A.V., Arzamasov K.M. et al. The first 10,000 mammographic examinations performed within the framework of the service «description and interpretation of mammographic examination data using artificial intelligence». Healthcare Manager , 2023, No. 8, pp. 54–67 (In Russ.)]. doi: 10.21045/18110185-2023-8-54-67.

18. Sasaki M., Tozaki M., Rodríguez-Ruiz A. et al. Artificial intelligence for breast cancer detection in mammography: experience of use of the ScreenPoint Medical Transpara system in 310 Japanese women // Breast Cancer . 2020. Vol. 27, No. 4. Р. 642–651. doi: 10.1007/S12282-020-01061-8.

19. Arzamasov K., Vasilev Y., Vladzymyrskyy A. et al. An International Non-Inferiority Study for the Benchmarking of AI for Routine Radiology Cases: Chest X-ray, Fluorography and Mammography // Healthcare . 2023. Vol. 11, No. 12. Р. 1684. doi: 10.3390/healthcare11121684.

20. Li H.Y., Lin Y., Hong Y.T., Chou C.P. Comparative Study of Artificial Intelligence-based System Alone in Synthesized Mammography Versus Radiologist’s Interpretation of Digital Breast Tomosynthesis in Screening Women // Journal of Radiological Science . 2024. Vol. 49. No. 1. Р. 59–65. doi: 10.4103/jradiolsci.JRADIOLSCI-D-23-00027.

21. Caldas F.A.A., Caldas H.C., Henrique T. et al. Evaluating the performance of artificial intelligence and radiologists accuracy in breast cancer detection in screening mammography across breast densities // European Journal of Radiology Artificial Intelligence . 2025. Vol. 2. Р. 100013. doi: 10.1016/j.ejrai.2025.100013.

22. Azaryan A.S., Pestrenin L.D., Vasiliev Yu.A., Akhmad E.S., Arzamasov K.M. Agreement between Moscow radiologists in the interpretation of mammographic studies using the BI-RADS scale. Almanac of Clinical Medicine , 2024, Vol. 52, No. 7, pp. 377–384 (In Russ.)]. doi: 10.18786/2072-0505-2024-52-035.

23. Taya M. Comparative Performance of Artificial Intelligence Algorithms for Screening Mammography. Radiol. Imaging Cancer . 2020. Vol. 2, No. 6. Р. e209034. doi: 10.1148/rycan.2020209034.

24. Renard X., Laugel T., Detyniecki M. Understanding Prediction Discrepancies in classification // Machine Learning . 2024. Vol. 113, No. 3. Р. 7997–8026. doi: 10.1007/s10994-024-06557-4.

25. Han T., Yun H., Sur Y.K., Park H. Optimizing Artificial Intelligence Thresholds for Mammographic Lesion Detection: A Retrospective Study on Diagnostic Performance and Radiologist-Artificial Intelligence Discordance // Diagnostics . 2025. Vol. 15, No. 11. Р. 1368. doi: 10.3390/diagnostics15111368

26. Chen Y., Partridge G.J.W., Vazirabad M. et al. Performance of Algorithms Submitted in the 2023 RSNA Screening Mammography Breast Cancer Detection AI Challenge // Radiology . 2025. Vol. 316, No. 2. Р. e241447. doi: 10.1148/radiol.241447.


Review

For citations:


Azaryan A.S., Khrustacheva M.Yu., Pestrenin L.D., Vasilev Yu.A., Vladzymyrskyy A.V., Omelyanskaya O.V., Domozhirova A.S., Arzamasov K.M., Gatsuk A.S. Diagnostic accuracy and agreement of AI services among themselves and with medical expert in evaluating mammography using the BI-RADS scale: a retrospective analytical study. Diagnostic radiology and radiotherapy. 2026;17(2):86-97. (In Russ.) https://doi.org/10.22328/2079-5343-2026-17-2-86-97

Views: 35

JATS XML


Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.


ISSN 2079-5343 (Print)