California Fault Lines: Understanding the Causes and Impact of Network Failures
Of the major factors affecting end-to-end service availability, network component failure is perhaps the least well understood. How often do failures occur, how long do they last, what are their causes, and how do they impact customers? Traditionally, answering questions such as these has required dedicated (And often expensive) instrumentation broadly deployed across a network. The authors propose an alternative approach: opportunistically mining "Low-quality" data sources that are already available in modern network environments. They describe a methodology for recreating a succinct history of failure events in an IP network using a combination of structured data (router configurations and syslogs) and semi-structured data (email logs).