ROBUST APACHE SPARK-BASED DATA ENGINEERING
ARCHITECTURE FOR REAL-TIME ENTERPRISE
INTELLIGENCE AND DECISION SUPPORT

Abstract

Apache Spark has become a vital technology in enterprise data engineering, enabling the analysis and processing of large-scale, high-velocity, and diverse data sets in real-time to generate intelligence and support decision-making. The current research was conducted to evaluate how Apache Spark can be used within an enterprise to design a system which process data at scale, analyze streaming data, adopt cloud technologies, and enterprise intelligence. A qualitative study was performed using an inductive approach based on secondary research data. Only peer-reviewed articles, journals, and books, as well as reputable industry sources, were included in the current analysis. The review identified several ways in which Apache Spark can be leveraged to improve processing capabilities, support distributed systems, and allow the building of real-time analytics systems within an enterprise environment. Six performance areas were revealed by the analysis: scalable distributed processing, real-time data analytics, enterprise intelligence, cloud technologies, fault tolerance and processing efficiency, and opportunities for artificial intelligence. According to the reviewed articles, Apache Spark can process petabytes of data, analyze streaming data in real-time with low-latency queries, provide up to 100x faster processing than traditional enterprise data platforms, and offer linear scaling of processing power with additional servers. In addition, the use of cloud technologies allows for better management of resources and creates opportunities for implementing AI capabilities for operational analytics and decision-making. Thus, the study confirmed the potential of Apache Spark as a technology which can be used to build large-scale and reliable data analytics and processing systems. Moreover, the current research has provided a thorough understanding of how Apache Spark can be used in enterprises and what opportunities it creates for modern business.

Citation details of the article



Journal: International Journal of Applied Mathematics
Journal ISSN (Print): ISSN 1311-1728
Journal ISSN (Electronic): ISSN 1314-8060
Volume: 36
Issue: 2
Year: 2023

Download Section



Download the full text of article from here.

You will need Adobe Acrobat reader. For more information and free download of the reader, please follow this link.

References

  1. [1] AIJCST (2021) Cloud Computing and Distributed System Reliability. Available at: http://aijcst.org/index.php/aijcst/article/view/34 (Accessed: 4 August 2026).
  2. [2] Collier, K., Brand, M. and Pramod, N. (2019) 'Models of enterprise intelligence', Thoughtworks. Available at: https://www.thoughtworks.com/en-in/insights/articles/intelligent-enterprise-series-models-enterprise-intelligence
  3. [3] Fu, Y. and Soman, C., 2021, June. Real-time data infrastructure at uber. In Proceedings of the 2021 International Conference on Management of Data (pp. 2503-2516). https://dl.acm.org/doi/abs/10.1145/3448016.3457552
  4. [4] IEEE (2019) Conference Paper. Available at: https://ieeexplore.ieee.org/document/8600738/
  5. [5] IEEE (2020) Conference Paper. Available at: https://ieeexplore.ieee.org/document/9174921/
  6. [6] IEEE (2021) Conference Paper. Available at: https://ieeexplore.ieee.org/document/9428549/
  7. [7] IEEE (2022) Conference Paper. Available at: https://ieeexplore.ieee.org/document/9987472/
  8. [8] Khan, M.N.I., 2022. A systematic review of legal technology adoption in contract management, data governance, and compliance monitoring. American Journal of Interdisciplinary Studies, 3(01), pp.01-30. https://ajisresearch.com/index.php/ajis/article/view/27
  9. [9] Lekkala, C., 2022. Integration of Real-Time Data Streaming Technologies in Hybrid Cloud Environments: Kafka, Spark, and Kubernetes. European Journal of Advances in Engineering and Technology, 9(10), pp.38-43. https://www.academia.edu/download/117095652/EJAET_9_10_38_43_3_.pdf
  10. [10] Mao, Y., Fu, Y., Gu, S., Mu, W., Cheng, L. and Liu, Q., 2020. Resource management schemes for cloud-native platforms with computing containers of docker and kubernetes. arXiv preprint arXiv:2010.10350. https://arxiv.org/abs/2010.10350
  11. [11] MarketsandMarkets (2022) Big Data Market with COVID-19 Impact Analysis, by Component, Deployment Mode, Organization Size, Business Function, Industry Vertical and Region – Global Forecast to 2026. Available at: https://www.marketsandmarkets.com/Market-Reports/big-data-market-1068.html
  12. [12] McKinsey & Company (2022) The Top Trends in Tech 2022. Available at: https://www.mckinsey.com/~/media/mckinsey/business%20functions/mckinsey%20digital/our%20insights/the%20top%20trends%20in%20tech%202022/mckinsey-tech-trends-outlook-2022-full-report.pdf
  13. [13] Rathore, S. and Soni, M. (2021) Cloud Computing Reliability and Fault Tolerance. Available at: https://www.academia.edu/download/68311421/83_1570665530_24597_EM_13jan21_16dec20_9aug20_Ry.pdf (Accessed: 4 August 2026).
  14. [14] Shah, S.D.A., Gregory, M.A. and Li, S., 2021. Cloud-native network slicing using software defined networking based multi-access edge computing: A survey. IEEE Access, 9, pp.10903-10924. https://ieeexplore.ieee.org/abstract/document/9317860/
  15. [15] Talluri, S., Overweel, L., Versluis, L., Trivedi, A. and Iosup, A. (2022) 'Empirical Characterization of User Reports about Cloud Failures', IEEE. Available at: https://research.vu.nl/en/publications/empirical-characterization-of-user-reports-about-cloud-failures/