Fa: A System for Automating Failure Diagnosis
Failures of Internet services and enterprise systems lead to user dissatisfaction and considerable loss of revenue. Since manual diagnosis is often laborious and slow, there is considerable interest in tools that can diagnose the cause of failures quickly and automatically from system-monitoring data. This paper identifies two key data-mining problems arising in a platform for automated diagnosis called Fa. Fa uses monitoring data to construct a database of failure signatures against which data from undiagnosed failures can be matched.