AI SRE for K8s: what’s actually working vs. what’s hype?
There's a lot of noise right now around AI-driven SRE for Kubernetes (auto-remediation, anomaly detection, incident triage, all of it). Vendor demos look great, homegrown Claude-built tools can be slick - but I want a real read on what it's like running this stuff in production.
If you've actually adopted this seriously, not just a POC, I'd like to hear:
- What is it actually catching or fixing that you couldn't before?
- Any false-positive or bad-remediation stories?
- How much trust have you given it? Read-only suggestions, or does it actually take action on your clusters?
- Anything you wish you'd known before rolling it out?
Not looking for product recommendations, more interested in the operational reality (good, bad, or ugly) from real-world stories.