PulseBoard — AI incident investigator that tells you why your site broke

by

Most monitoring tools tell you when your site is down. PulseBoard tells you why.

I built PulseBoard because I was tired of getting alerts at 3 AM and spending hours digging through logs, CloudWatch, and GitHub commits just to figure out what went wrong.

Here's how it works:

🔍 Monitors your websites and APIs
🧠 Vigil AI investigates incidents automatically
📦 Correlates failures with GitHub commits and deployments
☁️ Pulls AWS telemetry (EC2, RDS, ALB, Lambda, CloudWatch) — and now Lightsail!
📊 Scores commits by time, file relevance, deployment status, and infrastructure correlation
✅ Filters out noise below a 40/100 threshold to save AI tokens
🎯 Shows you the evidence, not just a guess

Recent improvements:
• Lightsail support (requested by a real user!)
• AI commit summaries cached and reused
• Provider failover (Groq → Mistral → Cerebras)
• Circuit breakers and anti-flapping logic
• "Cause unknown from available data" — Vigil admits when it doesn't know

One thing I've learned: a confidently wrong AI answer during an outage is worse than admitting uncertainty. Vigil is honest about its confidence level.

I'm currently refining the organization/workspace schema to support teams and agencies. If you've ever spent hours debugging an outage, I'd love to hear how you handle it today.

🔗

Feedback welcome — good, bad, or ugly. 🙏

8 views

Add a comment

Replies

Be the first to comment