Our integration broke 3 times before we fixed it.
We integrate with a third-party accounting API. Between March and September, the integration failed in production three separate times. Here is what happened each time:
Failure 1 — March 14. API returned a 200 with an empty body instead of an error code. Our code assumed 200 meant success. We wrote null values into 47 customer records. Took 9 days to notice because no alert fired. Fix: validate response body, not just status code.
Failure 2 — June 2. The vendor changed their rate limit from 100 requests per minute to 60 without updating their changelog. We hit the limit during a bulk sync, got throttled, and queued 2,300 jobs that never retried. Took 4 days to notice. Fix: exponential backoff and a dead-letter queue with a daily digest.
Failure 3 — August 21. Their OAuth token expiry changed from 90 days to 30 days. We cached tokens and refreshed every 60 days. Every customer's connection silently broke on day 31. Took 6 hours to notice because a customer emailed us.
Total customer impact across all three: 63 accounts affected, 12 churned, roughly $4,800 in lost MRR.
The common thread was that we trusted the vendor's documentation and status page instead of monitoring the integration from our side. We now run a synthetic transaction every 15 minutes against every customer connection.
What third-party dependency has bitten you the hardest, and did you keep using it afterward?
Replies