Preview

Herald of Dagestan State Technical University. Technical Sciences

Advanced search

Comparative analysis of big data processing tools on the Hadoop platform

https://doi.org/10.21822/2073-6185-2026-53-2-110-123

Abstract

Objective. This article provides a comprehensive comparative review of the leading tools used for processing big data in the Apache Hadoop ecosystem, with the aim of providing organizations with criteria for selecting the optimal solution for different types of workloads. The relevance of this research is driven by the fact that modern businesses require the analysis of data of all shapes and sizes, including structured and unstructured data, in order to gain a better understanding of their products, customers, and markets. The increasing volume, velocity, and variety of data (Big Data) has led to the need for specialized platforms to handle it. The Apache Hive instability issue is a key architectural challenge, as Hive is the most resource-intensive when performing identical tasks with Spark, as it starts processing input data without pre-analyzing it.
Method. A review of sources covering a wide range of tools and architectural solutions for processing big data in the Apache Hadoop ecosystem and beyond. The review included specialized scientific articles, technical guides, experimental research results, and industry analysis reports. The proposed methodology combines a comparative analysis of existing approaches and a comprehensive methodology for choosing optimal tools and architectural solutions for processing big data.
Result. The analyzed characteristics of Cloudera Impala and Apache Hive allow us to conclude that these SQL tools for analytical data processing, in the Hadoop ecosystem, complement each other.
Conclusion. It is premature to completely replace Apache Spark with Flink due to the potential need for additional infrastructure, but Flink is a viable solution for tasks that require minimal response time.

About the Authors

S. A. Reznichenko
Financial University under the Government of the Russian Federation
Russian Federation

Anna N. Alexeeva, Student, Department of Information Security

49, Leningradsky Ave., Moscow 125167



A. N. Alexeeva
Financial University under the Government of the Russian Federation
Russian Federation

Sergey A. Reznichenko, Cand.Sci. (Eng.), Assoc. Prof., Department of Information Security

49, Leningradsky Ave., Moscow 125167



References

1. Shaik A.S. Advancements In Real-Time Stream Processing: A Comparative Study Of Apache Flink, Spark Streaming, And Kafka Streams. International Journal Of Computer Engineering And Technology (IJCET). 2024;15(6). Accessed: 10.11.2025.

2. Lakshmipathy S. Apache Flink vs Apache Kafka Streams vs Apache Spark Structured Streaming — Comparing Stream Processing Engines [Электронный ресурс]. Onehouse Blog. – Cybersecurity Advisory AA23-059A. – 2025. – Режим доступа: https://www.onehouse.ai/blog/apache-spark-structured-streamingvs-apache-flink-vs-apache-kafka-streams-comparing-stream-processing-engines?, Accessed: 10.11.2025.

3. Kewalramani S. Impala vs Hive vs Spark SQL: Выбор правильного SQL движка для правильной работы в Cloudera Data Warehouse [Электронный ресурс] / Пер. с англ. mongohtotech // Хабр. 29 янв. 2020. Режим доступа: https://habr.com/ru/articles/486124/ Accessed: 10.11.2025.

4. ProjectPro. Impala vs Hive: Difference between Sql on Hadoop components [Электронный ресурс] // ProjectPro Blog. Обновлено 28 окт. 2024. Режим доступа: https://www.projectpro.io/article/impala-vs-hivedifference-between-sql-on-hadoop-components/180. Accessed: 11.11.2025.

5. Maayan G.D. Spark vs. Flink: Key Differences and How to Choose [Электронный ресурс] // Dataversity. 8 мая 2023. Режим доступа: https://www.dataversity.net/articles/spark-vs-flink-key-differences-and-howto-choose/?. Accessed: 11.11.2025.

6. Stream Processing with Apache Flink: Beginner's Guide 2025 [Электронный ресурс] // Ververica. [Б.г.]. Режим доступа: https://www.ververica.com/stream-processing-with-apache-flink-beginners-guide. Accessed: 11.11.2025.

7. Waehner K. Top Trends for Data Streaming with Apache Kafka and Flink in 2025 [Электронный ресурс] // Kai Waehner Blog. 2 дек. 2024. Режим доступа: https://www.kai-waehner.de/blog/2024/12/02/toptrends-for-data-streaming-with-apache-kafka-and-flink-in-2025/?. Accessed: 11.11.2025.

8. Ashyrova Y. Data management and analysis in BIG DATA. Symbol of Science. 2025. №1-1-2. URL: https://cyberleninka.ru/article/n/data-management-and-analysis-in-big-data. Accessed: 12.11.2025.

9. Belov V.A., Nikulchev E.V. Experimental Evaluation of the Time Efficiency of Processing Large Data in Specified Storage Formats. International Journal of Open Information Technologies. 2021; 9. URL: https://cyberleninka.ru/article/n/eksperimentalnaya-otsenka-vremennoy-effektivnosti-obrabotki-bolshihdannyh-v-zadannyh-formatah-hraneniya. – Accessed: 10.11.2025. (In Russ)

10. Altiev Aganazar, Bekdurdyev Gurbannazar, Ataev Allaberdi, Tachev Yusup Computer Technologies for Processing Big Data. Science and Worldview. 2025; 41. URL: https://cyberleninka.ru/article/n/kompyuternye-tehnologii-dlya-obrabotki-bolshih-dannyh. – Accessed: 12.11.2025.

11. Smilov N.K. Methods and Technologies of Data Quality Assessment and Management. Economics and Sociology. 2025;8(135). https://cyberleninka.ru/article/n/metody-i-tehnologii-otsenki-i-upravleniyakachestvom-dannyh. – Accessed: 12.11.2025.

12. Upaeva P.V. Optimization of ETL Processes: A Review of the Domestic Market. Bulletin of Science. 2024; 6 (75). URL: https://cyberleninka.ru/article/n/optimizatsiya-etl-protsessov-obzor-otechestvennogo-rynka. – Accessed: 11/12/2025.

13. S.G. Ermakov, M.M. Khalil, A.D. Khomonenko, V.A. Goncharenko, V.A. Khodakovsky, R. Abu Hassan The transition from the university data warehouse to the lake: models and methods of big data processing. IVD. 2024; 2(120). URL: https://cyberleninka.ru/article/n/perehod-ot-hranilischa-dannyh-universiteta-kozeru-modeli-i-metody-obrabotki-bolshih-dannyh. – Date of appeal: 12.11.2025.

14. Almeida, A., Brás, S., Sargento, S. et al. Time series big data: a survey on data stream frameworks, analysis and algorithms. J Big Data 10, 83. URL:https://doi.org/10.1186/s40537-023-00760-1. – Date of access: 12.11.2025.

15. Ovezdurdyeva Irina Kurbangeldievna, Myradov Maksat Tachmukhammedovich Technologies of working with big data “big data”: collection, storage and processing of big data // Science and worldview. 2024. No. 29. https://cyberleninka.ru/article/n/tehnologii-raboty-s-bolshimi-dannymi-big-data-sbor-hranenie-iobrabotka-bolshih-dannyh. – Accessed: 12.11.2025.

16. Kuzina E. Overview of Impala Architecture ADH Arenadata Docs [Electronic resource]. – Access mode: https://docs.arenadata.io/ru/ADH/current/concept/impala/impala.html. – Date of access: 12.11.2025. (In Russ)

17. Egorzev V. Working with Big Data: Introduction to Apache Hadoop and Spark [Electronic resource] // Tproger. [B.g.]. Access mode: https://tproger.ru/articles/rabota-s-bolwimi-dannymi--vvedenie-v-apache-hadoop-i-spark. – Accessed: 12.11.2025. (In Russ)


Review

For citations:


Reznichenko S.A., Alexeeva A.N. Comparative analysis of big data processing tools on the Hadoop platform. Herald of Dagestan State Technical University. Technical Sciences. 2026;53(2):110-123. (In Russ.) https://doi.org/10.21822/2073-6185-2026-53-2-110-123

Views: 20

JATS XML


Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.


ISSN 2073-6185 (Print)
ISSN 2542-095X (Online)