| International Journal of Computer Applications |
| Foundation of Computer Science (FCS), NY, USA |
| Volume 187 - Number 127 |
| Year of Publication: 2026 |
| Authors: Sneha Gullapalli |
10.5120/ijca0959b18e3a0a
|
Sneha Gullapalli . Policy-Governed Self-Healing Loops for Observability Pipelines in Distributed Cloud Systems. International Journal of Computer Applications. 187, 127 ( Jul 2026), 25-32. DOI=10.5120/ijca0959b18e3a0a
Observability pipelines have become a critical component of distributed cloud systems, collecting and processing metrics, logs, traces, and events that support monitoring, diagnosis, and automated operations. Failures within these pipelines, including queue saturation, exporter outages, schema drift, timestamp delays, memory pressure, and backend unavailability, can significantly reduce system visibility and affect operational decision-making. This paper presents SHIELD-OP, a policy-governed self-healing framework designed to improve the resilience of observability pipelines. The framework extends the Monitor–Analyze–Plan–Execute (MAPE) loop by incorporating policy constraints, verification mechanisms, rollback procedures, cooldown controls, and risk-aware remediation actions. A reproducible simulation study involving 50 random seeds and 120 failure episodes per seed compares SHIELD-OP with static alerting, retry-based recovery, ungoverned automation, and MAPE-K without verification. Results indicate that SHIELD-OP reduces recovery time and telemetry loss while maintaining a lower rate of unsafe actions, demonstrating the value of governance-aware self-healing for cloud observability infrastructure.