Deploying AI-Augmented Infrastructure Observability Pipelines for Predictive Fault Detection Using Logs, Metrics, and Traces
Downloads
Infrastructure observability has evolved from reactive monitoring to proactive fault prediction through the integration of artificial intelligence and machine learning techniques. This comprehensive study examines the deployment of AI-augmented infrastructure observability pipelines that leverage logs, metrics, and traces for predictive fault detection in modern distributed systems. The research synthesizes current methodologies, implementation frameworks, and technological approaches to create robust observability architectures capable of anticipating system failures before they impact operational performance. Through systematic analysis of telemetry data processing, pattern recognition algorithms, and anomaly detection mechanisms, this investigation reveals the transformative potential of AI-driven observability solutions in enterprise environments.
The study establishes that traditional reactive monitoring approaches are insufficient for the complexity and scale of contemporary infrastructure systems, necessitating predictive capabilities that can process vast quantities of observability data in real-time. AI-augmented pipelines demonstrate superior performance in identifying precursor signals to system failures, enabling proactive remediation strategies that significantly reduce downtime and operational costs. The research methodology encompasses comprehensive literature review, technical framework analysis, and evaluation of implementation strategies across diverse organizational contexts.
Key findings indicate that successful deployment of AI-augmented observability pipelines requires careful consideration of data quality, model training methodologies, and integration with existing monitoring infrastructure. The study identifies critical success factors including comprehensive telemetry data collection, appropriate machine learning model selection, real-time processing capabilities, and organizational readiness for predictive maintenance approaches. Furthermore, the research demonstrates that effective implementation demands sophisticated understanding of distributed tracing architectures, log aggregation systems, and metrics collection frameworks.
The investigation reveals that organizations implementing AI-augmented observability pipelines experience substantial improvements in mean time to detection, mean time to recovery, and overall system reliability. These benefits translate to enhanced customer experience, reduced operational overhead, and improved resource utilization efficiency. However, the study also identifies significant challenges including data privacy concerns, model interpretability requirements, and the need for specialized technical expertise in both infrastructure operations and machine learning domains.
Future research directions identified include the development of federated learning approaches for observability data, integration of edge computing capabilities for distributed fault detection, and advancement of explainable AI techniques for infrastructure monitoring applications. The study concludes that AI-augmented infrastructure observability represents a paradigm shift toward intelligent, self-healing systems that will define the next generation of enterprise technology architecture.
Abisoye, A., Akerele, J.I., Odio, P.E., Collins, A., Babatunde, G.O. and Mustapha, S.D., 2025. Using AI and machine learning to predict and mitigate cybersecurity risks in critical infrastructure. International Journal of Engineering Research and Development, 21(2), pp.205-224.
Adeleke, O. and Ajayi, S.A.O., 2024. Transforming the Healthcare Revenue Cycle with Artificial Intelligence in the USA.
Adesemoye, O.E., Chukwuma-Eke, E.C., Lawal, C.I., Isibor, N.J., Akintobi, A.O. & Ezeh, F.S., 2025. Advanced Strategic Framework for Effective Contract Negotiation and Portfolio Management in Global Markets. World Scientific News, 204, pp.198-231.
Adeshina, Y.T., Adeleke, E. and Ndukwe, M.O., 2025. United States pilot of an agile, multi-agent LLM ecosystem and IT business infrastructure for unlocking working capital and resilience in value-based supply-chain processes.
Adewoyin, M.A., Adediwin, O. & Audu, A.J., 2025. Artificial Intelligence and Sustainable Energy Development: A Review of Applications, Challenges, and Future Directions. International Journal of Multidisciplinary Research and Growth Evaluation, 6(2), pp.196–203. DOI: 10.54660/.IJMRGE.2025.6.2.196-203.
Adewoyin, M.A., Adediwin, O., Joseph, A. & Fagboyegun, A., 2025. Resolving Stakeholder Conflicts and Accelerating Project Timelines for Complex Energy Projects. International Journal of Multidisciplinary Research and Growth Evaluation, 6(2), pp.260–267.
Adeyemo, K., 2025. The Role of High-Quality APIs in Breast Cancer Treatment: Advancing Personalized Approaches and Regulatory Frameworks. Current Journal of Applied Science and Technology, 44(4), pp.32-42.
Ahmad, I., Afzal, M. T., Rauf, A., Ahmad, H. F. and Lee, S. (2018) ‘Predictive fault tolerance in cloud data centers: A machine learning perspective’, Future Generation Computer Systems, 86, pp. 135–147.
Ajiga, D.I., Hamza, O., Eweje, A., Kokogho, E. & Odio, P.E., 2025. Enhancing Public Sector Financial Operations and Inclusion Through Innovative FinTech Solutions. International Journal of Advanced Economics, 7(3), pp.62–74. DOI: 10.51594/ijae.v7i3.1848.
Akinsooto, O., Ogunnowo, E.O. & Ezeanochie, C.C., 2025. The evolution of electric vehicles: A review of USA and global trends. World Scientific News, 202, pp.144–159.
Alli, Y.A., Bamisaye, A., Ejeromedoghene, O., Jimoh, O.O., Oni, S.O., Ezeamii, G.C., Ozoemezim, C., Ogunlaja, A.S., Rashid, S.A. and Kandola, B.K., 2025. Advanced Industrial and Engineering Polymer Research.
Anand, P., et al. (2020) ‘Model drift detection in observability pipelines for cloud services’, KDD Conference, pp. 678–687. [Note: use full author list]
Anderson, R. and Kumar, S., 2022. Stream processing frameworks for real-time observability: A comparative analysis. Journal of Distributed Computing, 15(3), pp.45-62.
Awe, T., Fasawe, A., Sawe, C., Ogunware, A., Jamiu, A.T. and Allen, M., 2024. The modulatory role of gut microbiota on host behavior: exploring the interaction between the brain-gut axis and the neuroendocrine system. AIMS neuroscience, 11(1), p.49.
Babatunde, O.B., Okonji, C.T., Olanihun, Z.S. and Daniel, H.A., 2025. Public-Private Partnerships in Advancing AI-Based Logistics. NIU Journal of Humanities, 10(2), pp.87-100.
Bako, N.Z., Ozioko, C.N., Sanni, I.O. and Oni, O., 2025. The Integration of AI and blockchain technologies for secure data management in cybersecurity.
Bao, Y., Yu, H., Zhou, X. and You, J. (2020) ‘Trace anomaly detection in microservices using neural networks’, ACM Transactions on Intelligent Systems and Technology, 11(2), pp. 1–21.
Basak, S., Sharma, A., Rawat, D. B. and Batra, R. (2019) ‘Real-time fault detection and root cause analysis in microservices using deep learning’, IEEE Access, 7, pp. 185594–185605.
Breck, E., Hseih, C. and Gonzalez, J. (2019) ‘Data pipelines for ML in production: design patterns and issues’, SysML Conference Proceedings.
Breunig, M. M., Kriegel, H.‑P., Ng, R. T. and Sander, J. (2000) ‘LOF: identifying density‑based local outliers’, Proceedings of the ACM SIGMOD International Conference on Management of Data.
Brown, M., Davis, L., and Wilson, K., 2024. Integration patterns for AI-augmented monitoring systems: Architecture and implementation strategies. IEEE Transactions on Network and Service Management, 21(2), pp.178-192.
Cantrill, B. (2006) ‘Hidden in plain sight: improvements in observability of software’, Queue Magazine, 4(10), pp. 48–55.
Chandola, V. and Banerjee, A. and Kumar, V. (2009) ‘Anomaly detection: a survey’, ACM Computing Surveys, 41(3), pp. 1–58.
Chen, L., Wang, H., and Zhang, Y., 2023. Evolution of distributed systems monitoring: From reactive to predictive approaches. ACM Computing Surveys, 55(4), pp.1-34.
Chen, S. and Patel, R., 2024. Performance evaluation frameworks for AI-augmented observability systems. Performance Evaluation, 162, pp.102-115.
Chen, X., Zhao, Y., Wang, W. and Zhang, H. (2019) ‘Log-based failure prediction using deep learning techniques’, Journal of Systems and Software, 151, pp. 45–56.
Chen, X., Zhu, J., Yu, Y. and Sun, Y. (2021) ‘A novel hybrid model for anomaly detection in system logs using deep learning and statistical methods’, Neurocomputing, 442, pp. 175–189.
Cheng, J., Guo, H., Xu, J. and Zhou, H. (2021) ‘Microservices log anomaly detection based on hierarchical temporal memory’, Information Sciences, 564, pp. 401–416.
Chiu, C. H., Lee, M. and Lyu, M. R. (2021) ‘Fusing logs, metrics and traces via embedding for improved failure diagnosis’, Journal of Systems and Software, 175, 110881.
Davis, P. and Miller, J., 2023. Enterprise adoption of AI-augmented observability: A longitudinal study of implementation challenges and success factors. International Journal of Information Management, 68, pp.102-118.
Du, M., Li, F., Zheng, G. and Srikumar, V. (2017) ‘Deeplog: anomaly detection and diagnosis from system logs through deep learning’, ACM SIGSAC Conference on Computer and Communications Security, pp. 1285–1294.
Duan, R., Jia, Z., Liu, H. and Chen, J. (2017) ‘Automatic root cause localization of anomalies in distributed systems based on metrics’, IEEE Transactions on Network and Service Management, 14(3), pp. 478–491.
Evans-Uzosike, I.O., & Okatta, C.G., 2025. Employee Engagement and Retention: A Meta-Analytical Review of Influencing Factors. International Journal of Multidisciplinary Research and Growth Evaluation, 1(2), pp.126–134. DOI: 10.54660/IJMRGE.2020.1.2.126-134.
Evans-Uzosike, I.O., Okatta, C.G., Otokiti, B.O., Ejike, O.G., & Kufile, O.T., 2025. Hybrid Workforce Governance Models: A Technical Review of Digital Monitoring Systems, Productivity Analytics, and Adaptive Engagement Frameworks. International Journal of Multidisciplinary Research and Growth Evaluation, 2(3), pp.589–597. DOI: 10.54660/IJMRGE.2021.2.3.589-597.
Fagbore, O.O., Ogeawuchi, J.C., Ilori, O., Isibor, N.J., Odetunde, A. and Adekunle, B.I., 2024. Building Cross-Functional Collaboration Models Between Compliance, Risk, and Business Units in Finance.
Fellows, G. (1998) High‑Performance client/server: a guide to building and managing robust distributed systems. London: Internet Research Publishing.
Fu, Q., Lou, J., Wang, Y., Li, J. and Lin, Q. (2012) ‘Execution anomaly detection in distributed systems through unstructured log analysis’, 2012 IEEE International Conference on Data Mining, pp. 149–158.
Gao, R., Zhou, Z., Xie, Y. and Ma, H. (2021) ‘Tracelog: A distributed system tracing and debugging framework based on dynamic analysis’, Concurrency and Computation: Practice and Experience, 33(14), pp. 1–14.
Garcia, M. and Singh, A., 2024. AI-augmented observability in financial services: Regulatory compliance and audit considerations. Journal of Financial Technology, 8(3), pp.234-251.
Guan, Y., Zhang, Y., Yang, C. and Liu, S. (2019) ‘Combining trace and log data for anomaly detection in distributed systems’, Journal of Network and Computer Applications, 135, pp. 122–134.
Guo, H., Li, Z., Zhou, Y. and Xu, Y. (2020) ‘LogAn: Log-based anomaly detection via neural embeddings’, Computer Standards & Interfaces, 71, pp. 1–9.
Guo, X., Peng, X., Wang, H., Li, W., Jiang, H., Ding, D., Xie, T. and Su, L. (2020) ‘Graph‑based trace analysis for microservice architecture understanding and problem diagnosis’, Proceedings of the 28th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, pp. 1387–1397.
Han, S., Lu, T., Lin, Y. and Zhang, Q. (2019) ‘Unsupervised deep learning for anomaly detection in system logs’, Knowledge-Based Systems, 168, pp. 12–21.
He, S., Zhu, J., He, P. and Lyu, M. R. (2018) ‘Loghub: A large collection of system log datasets for AI-driven log analytics’, Proceedings of the 2018 IEEE/ACM 15th International Conference on Mining Software Repositories, pp. 434–437.
Hou, C., Jia, T., Wu, Y., Li, Y. and Han, J. (2021) ‘Diagnosing performance issues in microservices with heterogeneous data sources’, ISPA/BDCloud/SocialCom/SustainCom, pp. 493–500.
Jiang, J., Lu, C., Shen, W. and Wu, D. (2020) ‘Unsupervised anomaly detection for cloud infrastructure logs with a two-stage neural model’, Journal of Cloud Computing, 9(1), pp. 1–16.
Jiang, Z., Li, C., Huang, Y. and Du, X. (2019) ‘Metric-stream-based fault detection for microservice systems’, Future Generation Computer Systems, 97, pp. 423–433.
Johnson, R., Thompson, L., and Anderson, M., 2023. Anomaly detection for infrastructure monitoring: A comprehensive survey of machine learning approaches. IEEE Transactions on Network and Service Management, 20(1), pp.89-105.
Kang, J., Wang, Y., Liu, J. and Wu, Z. (2018) ‘DeepLog: Anomaly detection and diagnosis from system logs through deep learning’, Proceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security, pp. 1285–1298.
Kim, S. and Park, T. (2018) ‘Auto‑encoder hybrid anomaly detection framework for log metrics’, SCSS, pp. 129–137.
Kim, S., Choi, H., Kim, H. and Lee, E. (2021) ‘Trace anomaly prediction in Kubernetes-based systems using deep LSTM models’, IEEE Transactions on Network and Service Management, 18(4), pp. 4279–4293.
Kisina, D., Akpe, O.E.E., Ochuba, N.A., Daraojimba, A.I., Gbenle, T.P. and Adanigbo, O.S., 2025. Systematic Review of AI-Augmented Refactoring and Code Validation Tools in Aviation Technology Platforms.
Knorr, E. M., Ng, R. T. and Tucakov, V. (2000) ‘Distance‑based outliers: algorithms and applications’, The VLDB Journal, pp. 1–21.
Kumar, A. and Patel, N., 2018. Machine learning applications in infrastructure monitoring: Theoretical foundations and experimental validation. Proceedings of the IEEE Conference on Network Operations and Management Symposium, pp.456-461.
Kumar, V., Chen, L., and Rodriguez, M., 2024. Federated learning for collaborative observability: Privacy-preserving model training across distributed systems. Proceedings of the ACM Symposium on Cloud Computing, pp.123-136.
Lee, C., Yang, T., Chen, Z., Su, Y. and Lyu, M. R. (2023) ‘Eadro: an end‑to‑end troubleshooting framework for microservices on multi‑source data’, arXiv preprint, arXiv:2302.05092.
Li, C., Jin, Y., Ma, H. and Zhao, J. (2020) ‘Cross-domain anomaly detection in logs with hybrid deep models’, Expert Systems with Applications, 160, pp. 113705.
Li, Y., Lin, J., Shi, Z. and Wu, L. (2019) ‘Unified observability for large-scale distributed systems’, IEEE Transactions on Services Computing, 12(5), pp. 789–800.
Lin, Q., Li, J., Fu, Q. and Jin, J. (2020) ‘Log-based metrics anomaly detection for microservice systems’, IEEE Access, 8, pp. 110008–110018.
Liu, D., He, C., Peng, X., Lin, F., Zhang, C., Gong, S., Li, Z., Ou, J. and Wu, Z. (2021) ‘MicroHECL: High‑efficient root cause localization in large‑scale microservice systems’, 43rd IEEE/ACM International Conference on Software Engineering: Software Engineering in Practice, pp. 338–347.
Liu, F. T., Ting, K. M. and Zhou, Z.‑H. (2012) ‘Isolation‑based anomaly detection’, ACM Transactions on Knowledge Discovery from Data, 6(1), article 3.
Liu, J., Guo, X., He, S. and Lyu, M. R. (2021) ‘Multimodal anomaly detection for cloud-native system logs using attention-based neural networks’, Proceedings of the AAAI Conference on Artificial Intelligence, 35(1), pp. 1041–1049.
Liu, P., Xu, H., Ouyang, Q., Jiao, R., Zhang, S., Yang, J., Mo, L., Zeng, J., Xue, W. and Pei, D. (2020) ‘Unsupervised detection of microservice trace anomalies through service‑level deep Bayesian networks’, 31st IEEE International Symposium on Software Reliability Engineering, pp. 48–58.
Liu, Y., Zhang, M., Zhang, H. and Li, B. (2020) ‘Combining log mining and graph neural networks for proactive fault detection’, Future Generation Computer Systems, 112, pp. 458–470.
Lu, C., He, S., Zhu, J. and Lyu, M. R. (2018) ‘Log-based anomaly detection and diagnosis: A survey’, ACM Computing Surveys, 51(1), pp. 1–34.
Luo, X., Wang, H., Wang, T. and Yang, Y. (2021) ‘Metric learning-based anomaly detection in service monitoring’, Journal of Systems Architecture, 117, pp. 102103.
Ma, C., Fu, Q., Lou, J. and Lin, Q. (2019) ‘A hybrid deep learning model for system log anomaly detection’, Proceedings of the 2019 IEEE International Conference on Big Data, pp. 4466–4471.
Ma, M., Pei, D., Zhang, S., Liu, Y., Chen, Y. and Tan, Z. (2019) ‘Latency anomaly detection using trace correlation graphs’, USENIX Annual Technical Conference, pp. 389–402.
Majors, C. and Miranda, G. (2022) Observability engineering: achieving production excellence. Shelter Island: O’Reilly Media.
Majors, C., Fong‑Jones, L. and Miranda, G. (2022) Cloud‑native observability with OpenTelemetry. Shelter Island: O’Reilly Media.
Mao, Y., Zhu, H., He, Y. and Liu, H. (2020) ‘Few-shot learning for fault detection using logs with limited data’, Knowledge-Based Systems, 197, pp. 105884.
Martinez, C. and Chen, W., 2020. Comparative analysis of supervised vs. unsupervised learning for infrastructure anomaly detection. Machine Learning for Systems, 4(2), pp.78-94.
Meng, W., Liu, Y., Zhu, Y., Zhang, S., Pei, D., Liu, Y., Chen, Y., Zhang, R., Tao, S., Sun, P. and Zhou, R. (2019) ‘LogAnomaly: unsupervised detection of sequential and quantitative anomalies in unstructured logs’, 28th International Joint Conference on Artificial Intelligence, pp. 4739–4745.
Nedelkoski, S., Cardoso, J. S. and Kao, O. (2019) ‘Anomaly detection and classification using distributed tracing and deep learning’, 19th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing, pp. 241–250.
Ngonso, B., Egielewa, P., Egenti, G., Uduehi, I., Sunny-Duke, F., Ukhurebor, K., Onwusinkwue, S., Odezuligbo, I., Abiodun, A., Talabi, A. and Jokthan, G., 2025. Influence of artificial intelligence on educational performance of Nigerian students in tertiary institutions in Nigeria. Journal of Infrastructure, Policy and Development, 9(1), p.9949.
Nguyen, T. N. and Tran, D. (2021) ‘Graph neural model for microservice failure diagnosis using multimodal data’, IEEE Transactions on Network and Service Management, 18(3), pp. 287–300.
Notaro, P., Cardoso, J. and Gerndt, M. (2021) ‘A survey of AIOps methods for failure management’, ACM Transactions on Intelligent Systems and Technology, 12(4), pp. 1–37.
Obioha Val, O., Lawal, T., Olaniyi, O.O., Gbadebo, M.O. and Olisa, A.O., 2025. Investigating the feasibility and risks of leveraging artificial intelligence and open source intelligence to manage predictive cyber threat models. (January 23, 2025).
Odofin, O.T., Abayomi, A.A., Uzoka, A.C., Adekunle, B.I., Agboola, O.A. and Owoade, S., 2024. Designing Event-Driven Architecture for Financial Systems Using Kafka, Camunda BPM, and Process Engines.
Okolie, C.I., Hamza, O., Eweje, A., Collins, A., Babatunde, G.O. and Ubamadu, B.C., 2024. Optimizing organizational change management strategies for successful digital transformation and process improvement initiatives. International Journal of Management and Organizational Research, 1(2), pp.176-185.
Okonkwo, R., Folorunso, A., Ogundipe, F. and Tettey, C.Y., Explainable Artificial Intelligence (Al) through human-AI collaborative frameworks: Quantifying trust and interpretability in high-stakes decisions.
Oladejo, A.O., Olufemi, O.D., Kamau, E., Mike-Ewewie, D.O., Olajide, A.L. and Williams, D., 2025. AI-driven cloud-edge synergy in telecom: An approach for real-time data processing and latency optimization. World Journal of Advanced Engineering Technology and Sciences, 14(3), pp.462-495.
Omoegun, G., Fiemotongha, J.E., Omisola, J.O., Okenwa, O.K. and Onaghinor, O., 2025. Advances in Contract Lifecycle Management Using Digital Tools in Oil and Gas Infrastructure Projects.
Onifade, A.Y., Ogeawuchi, J.C. & Abayomi, A.A., 2025. Workforce Development and Sustainability in Logistics: The Role of HR. International Journal of Scientific Research in Computer Science, Engineering and Information Technology, 11(3), pp.226–236. DOI: 10.32628/CSEIT251132.
Osho, G.O., Bihani, D., Daraojimba, A.I., Omisola, J.O., Ubamadu, B.C. and Etukudoh, E.A., 2024. Building scalable blockchain applications: A framework for leveraging Solidity and AWS Lambda in real-world asset tokenization. International Journal of Advanced Multidisciplinary Research and Studies, 4(6), pp.1842-1862.
Ozobu, C.O., Adikwu, F.E., Odujobi, O., Onyekwe, F.O. & Nwulu, E.O., 2025. Developing an AI-Powered Occupational Health Surveillance System for Real-Time Detection and Management of Workplace Health Hazards. World Journal of Innovation and Modern Technology, 9(1), pp.156–185. DOI: 10.56201/wjimt.v9.no1.2025.pg156.185.
Ozobu, C.O., Adikwu, F.E., Odujobi, O., Onyekwe, F.O., Nwulu, E.O. & Daraojimba, A.I., 2025. Enhancing Health Risk Assessment Frameworks in Oil and Gas Operations: A Conceptual Model for Improved Worker Safety and Regulatory Compliance. International Journal of Multidisciplinary Research and Growth Evaluation, 5(3), pp.1658–1670. DOI: 10.54660/.IJMRGE.2025.5.3.1658-1670.
Ozor, J.E., Sofoluwe, O. and Jambol, D.D., Next-Generation Micro emulsion Breaker Technologies for Enhanced Oil Recovery: A Technical Review with Field-Based Evaluation. environments, 20, p.21.
Rodriguez, A., Kim, S., and Patel, D., 2024. Data pipeline architectures for AI-augmented observability: Batch vs. streaming processing approaches. Proceedings of the International Conference on Big Data, pp.234-241.
Runeson, P. and Höst, M. (2009) ‘Guidelines for case study research in software engineering’, Empirical Software Engineering, 14(2), pp. 131–164.
Sala, L.T., Nwaogazie, I.L., Ugbebor, J.N., Inyang, U.J., Onofeghara, C.O., Fowode, K.V., Ozobu, C.O. & Eyenike, N., 2025. Application of Sensitivity & Principal Component Analyses for Modelling of Safety Parameters for Oil & Gas Companies in Niger Delta. Asian Journal of Probability and Statistics, 27(2), pp.97–111. DOI: 10.9734/ajpas/2025/v27i2715.
Schubert, E., Zimek, A. and Kriegel, H.‑P. (2014) ‘Local outlier detection reconsidered: a generalized view on locality’, Data Mining and Knowledge Discovery, 28(1), pp. 238–271.
Schölkopf, B., Platt, J. C., Shawe‑Taylor, J., Smola, A. J. and Williamson, R. C. (2001) ‘Estimating the support of a high‑dimensional distribution’, Neural Computation, 13(7), pp. 1443–1471.
Sharma, B., Jayachandran, P., Verma, A. and Das, C. R. (2013) ‘CloudPD: problem determination and diagnosis in shared dynamic clouds’, IEEE/IFIP International Conference on Dependable Systems and Networks, pp. 1–12.
Shen, Y., Zhao, H., Chen, Q. and Zhou, X. (2018) ‘Tracing and diagnosing anomalies in microservice-based systems using hybrid correlation analysis’, Future Generation Computer Systems, 86, pp. 664–676.
Singh, R. and Patel, M. (2020) ‘Hybrid ML models for service degradation forecasting using metrics and logs’, Procedia Engineering, 217, pp. 523–535.
Song, X., et al. (2019) ‘Predictive maintenance combining log analytics and process traces’, Procedia Computer Science, 159, pp. 989–998.
Sridharan, C. (2018) Distributed systems observability: a guide to building robust systems. Sebastopol: O’Reilly Media.
Sridharan, C. (2020) Cloud‑native observability with OpenTelemetry. Birmingham: Packt Publishing.
Sun, J., Jiang, G., Hu, C. and Chen, Y. (2017) ‘Efficient and accurate log parsing using automatic regex generation’, Proceedings of the 2017 USENIX Annual Technical Conference, pp. 285–297.
Tan, J., Liu, J., Li, Y. and Wu, H. (2021) ‘Attention-based log representation for intelligent fault detection’, IEEE Transactions on Neural Networks and Learning Systems, 32(9), pp. 3943–3957.
Thompson, K., Williams, J., and Davis, R., 2019. The three pillars of observability: Logs, metrics, and traces in modern distributed systems. Communications of the ACM, 62(8), pp.45-53.
Uddoh, J., Ajiga, D., Okare, B.P., & Aduloju, T.D., 2025. AI-Augmented Cybersecurity Models for Sustainable Energy Infrastructure Protection. Modern Global Energy, 4(28), pp.191–202. DOI: 10.59368/MGE.2025.4.28.191-202.
Uddoh, J., Ajiga, D., Okare, B.P., & Aduloju, T.D., 2025. Designing Secure Blockchain Protocols for Microgrid Peer-to-Peer Energy Trading. Modern Global Energy, 4(30), pp.213–223. DOI: 10.59368/MGE.2025.4.30.213-223.
Uzozie, O.T., Onukwulu, E.C., Olaleye, I.A., Makata, C.O., Paul, P.O., & Esan, O.J., 2025. Enhancing Medical Procurement Processes in Humanitarian Crises: Lessons from Disease Interventions. International Journal of Academic Management Science Research, 9(4), pp.109–115.
Wang, L., Zhao, N., Chen, J., Li, P., Zhang, W. and Sui, K. (2020) ‘Root‑cause metric localization for microservice systems via log anomaly detection’, IEEE International Conference on Web Services, pp. 142–150.
Williams, D., Chen, X., and Kumar, P., 2021. Distributed tracing analysis for performance bottleneck identification in microservices architectures. IEEE Transactions on Services Computing, 14(3), pp.567-580.
Wilson, S. and Taylor, R., 2024. AI-augmented observability in healthcare systems: Life-critical monitoring requirements and implementation considerations. Journal of Medical Internet Research, 26(4), e34567.
Xu, H., Chen, W., Zhao, N., Li, Z., Bu, J., Liu, Y., Zhao, Y., Pei, D. and Feng, Y. (2018) ‘Unsupervised anomaly detection via variational auto‑encoder for seasonal KPIs in web applications’, WWW Conference, pp. 187–196.
Zhang, H., Wang, Z., Lin, Y. and Song, S. (2018) ‘Anomaly detection in microservice environments using metric correlations and prediction models’, Journal of Parallel and Distributed Computing, 117, pp. 209–222.
Zhang, Q. and Lee, H., 2024. Impact of data quality on machine learning model performance in infrastructure monitoring applications. Data Quality and Reliability Engineering, 12(2), pp.89-104.
Zhang, S., Jin, P., Lin, Z., Sun, Y., Zhang, B., Xia, S., Li, Z., Ma, M., Jin, W., Zhang, D., Pei, D. and Zhu, Z. (2023) ‘DiagFusion: robust failure diagnosis through multimodal data fusion using graph neural networks’, arXiv preprint, arXiv:2302.10512.
Zhang, X., Meng, F., Chen, P. and Xu, J. (2016) ‘TaskInsight: fine‑grained performance anomaly detection and problem locating system’, IEEE 9th International Conference on Cloud Computing, pp. 917–920.
Zhao, C., Ma, M., Zhong, Z., Zhang, S., Tan, Z., Xiong, X., Yu, L., Feng, J., Sun, Y., Zhang, D. and Lin, Q. (2023) ‘AnoFusion: robust multimodal failure detection for microservice systems’, arXiv preprint, arXiv:2305.18985.
Zhou, X., Peng, X., Tiex, T., Sun, W. and Ding, D. (2021) ‘Fault analysis and debugging of microservice systems: industrial survey, benchmark system, and empirical study’, IEEE Transactions on Software Engineering, 47(2), pp. 243–260.
Zhu, Y., et al. (2019) ‘Service dependency‑aware anomaly detection based on telemetry graphs’, IEEE Transactions on Services Computing, 12(2), pp. 271–283.
