How do you effectively lower your Mean Time to Detect (MTTD) when managing highly complex, distributed microservices architectures? Furthermore, this critical site reliability engineering metric measures the average time it takes for a team to identify an operational issue from the exact moment it occurs. Why does failing to automate your alerting pipeline directly undermine your overall system availability?