Ask Debugging Failure Large Interconnected Back
1 signals · 1 source
Evidence
Ask HN: Debugging failure in large interconnected back end systems — I’m trying to understand how teams actually debug production issues in systems made up of multiple services and external integrations (e.g. Stripe, Twilio, internal microservices, queues, webhooks, etc.).<p>In practice, when something breaks, it seems like the workflow is usually:<p>an alert fires (Datadog/Sentry/CloudWatch/etc.)<p>or a customer complains<p>engineers then start checking logs, traces, dashboards across multiple systems<p>and eventually manually reconstruct what happened across services<p>What I’m curious about:<p>How do you actually trace a single failed request or transaction across multiple services today?<p>What tools do you rely on most in practice (not in theory)?<p>Where does it usually break down — logs, tracing, instrumentation, or just missing context?<p>How long does it typically take to go from “something is wrong” → “we know exactly why it broke”?<p>What part of this is still mostly manual stitching together of information?<p>Trying to understand what the real pain points are in practice, especially in systems with lots of external integrations and async flows.