Research topic

Microservices Performance and Reliability Research

This topic gathers work on understanding multi-service applications through distributed traces, service-invocation graphs, profiling metrics, and performance models. The papers address latency variation, service interactions, anomaly detection, trace reduction, and root-cause localization in microservice environments.

Related search terms: microservice systems · microservice performance · distributed service analysis

12 papers in this topic, ordered newest first. The detailed paper records contain the evidence-grounded methods, tools, datasets, findings, and citation guidance.

Selected papers

2026 · IEEE Transactions on Software Engineering

CARE: Context Aware Root Cause Identification Using Distributed Traces and Profiling Metrics

Mahsa Panahandeh, Naser Ezzati-Jivan, Abdelwahab Hamou-Lhadj, James Miller

CARE combines distributed traces and profiling metrics with graph- and spectrum-based analysis to localize performance root causes in microservices.

Keywords: distributed traces · profiling metrics · context-aware RCA · microservice diagnosis · TrainTicket

Read the detailed paper record · · Authoritative source

2026 · Journal of Systems and Software

DTraComp: Comparing distributed execution traces for understanding intermittent latency sources

Maryam Ekhlasi, Fatemeh Faraji Daneshgar, Michel Dagenais, Maxime Lamothe, Naser Ezzati-Jivan, Matthew Khouzam

DTraComp is an open-source Eclipse Trace Compass framework that compares groups of distributed requests and attributes span time to user-space, kernel, thread-state, and system-call evidence.

Keywords: DTraComp · distributed trace comparison · OpenTracing · LTTng · LTTng-UST

Read the detailed paper record · · Authoritative source

2026 · SIGSOFT FSE Companion

Rethinking Performance Debugging: From Optimization to Collaborative Reasoning

Mahsa Panahandeh, Naser Ezzati-Jivan, Abdelwahab Hamou-Lhadj

The paper reframes performance debugging as collaborative reasoning over multiple evidence-grounded hypotheses rather than optimization for one supposedly best explanation.

Keywords: performance debugging · collaborative reasoning · AgentDebug · Reasoning Surface · hypothesis generation

Read the detailed paper record · · Authoritative source

2025 · IEEE International Conference on Software Maintenance and Evolution (ICSME)

HybridRCA: Lightweight Critical-Path-Aware Hybrid Tracing for Root-Cause Analysis in Production Microservices

Maryam Ekhlasi, Arnaud Fiorini, Michel R. Dagenais, Naser Ezzati-Jivan, Maxime Lamothe

HybridRCA combines critical-path-aware span analysis with targeted kernel metrics to reduce production trace volume while preserving root-cause localization evidence.

Keywords: critical path · hybrid tracing · production microservices · LTTng · OpenTracing

Read the detailed paper record · · Authoritative source

2025 · ACM/SPEC International Conference on Performance Engineering (ICPE)

Utilizing Graph Neural Networks for Effective Link Prediction in Microservice Architectures

Ghazal Khodabandeh, Alireza Ezaz, Majid Babaei, Naser Ezzati-Jivan

The paper applies graph attention networks to predict future interactions in microservice call graphs, supporting proactive monitoring.

Keywords: microservice call graphs · link prediction · graph attention networks · temporal segmentation · negative sampling

Read the detailed paper record · · Authoritative source

2024 · ACM/SPEC ICPE Companion

Analyzing Performance Variability in Alibaba's Microservice Architecture: A Critical-Path-Based Perspective

Alireza Ezaz, Ghazal Khodabandeh, Naser Ezzati-Jivan

The paper identifies response-time variability in Alibaba microservice traces through critical-path extraction and variability analysis of service interactions.

Keywords: Alibaba microservice architecture · critical path · distributed traces · response-time variability · critical interactions

Read the detailed paper record · · Authoritative source

2024 · The 37th Canadian Conference on Artificial Intelligence

Automatic Reduction of Execution Trace Data Volume Using Gradient Boosting in Large-Scale Microservice Systems

Amir Haghshenas, Naser Ezzati-Jivan, Michel Dagenais

The paper uses gradient boosting and feature importance to reduce the amount of trace data needed for microservice performance modeling.

Keywords: trace data volume · feature importance · CPU demand · memory demand · Alibaba microservices

Read the detailed paper record · · Authoritative source

2024 · ACM/SPEC ICPE Companion

Context-aware Root Cause Localization in Distributed Traces Using Social Network Analysis (Work In Progress paper)

Mahsa Panahandeh, Naser Ezzati-Jivan, Abdelwahab Hamou-Lhadj, James Miller

The work-in-progress paper combines service-call graph context, social-network analysis, and spectrum-based fault localization to rank distributed-trace root causes.

Keywords: context-aware RCA · service-call graph · distributed traces · service communities · Louvain

Read the detailed paper record · · Authoritative source

2024 · ACM/SPEC International Conference on Performance Engineering (ICPE) Companion

Efficient Unsupervised Latency Culprit Ranking in Distributed Traces with GNN and Critical Path Analysis

Mahsa Panahandeh, Naser Ezzati-Jivan, Abdelwahab Hamou-Lhadj, James Miller

The paper combines an unsupervised GraphSAGE model with critical-path-specific latency profiles to detect anomalous requests and rank likely microservice culprits without labelled training data.

Keywords: latency culprit ranking · distributed traces · GraphSAGE · graph neural networks · critical path

Read the detailed paper record · · Authoritative source

2024 · ACM/SPEC International Conference on Performance Engineering (ICPE) Companion

Network Analysis of Microservices: A Case Study on Alibaba Production Clusters

Ghazal Khodabandeh, Alireza Ezaz, Naser Ezzati-Jivan

The paper applies graph community detection and service-graph clustering to expose recurring microservice communication structures in an Alibaba production-cluster snapshot.

Keywords: microservice networks · Alibaba production clusters · service call graphs · community detection · Louvain

Read the detailed paper record · · Authoritative source

2021 · Wireless Communications and Mobile Computing

Debugging of Performance Degradation in Distributed Requests Handling Using Multilevel Trace Analysis

Naser Ezzati-Jivan, Houssem Daoud, Michel R. Dagenais

The paper correlates LTTng traces from user space through kernel, storage, network, and multiple hosts in a disk-backed state model, enabling top-down diagnosis of distributed request latency.

Keywords: distributed requests · multilevel trace analysis · LTTng · Apache · PHP

Read the detailed paper record · · Authoritative source

Related topics