▲ 1 r/github

Cursor launched its code hosting platform right during the GitHub outage. Anyone else feel like agent-native git is missing the point?

GitHub hit 50 percent error rates on raw downloads yesterday right when Cursor launched their agent-native hosting platform. The launch is interesting, but an agent is just as stuck as a human developer when the host degrades. Since Origin syncs back to GitHub, if GitHub goes down, the agentic delivery loop is dead.

We spent the morning verifying backup git mirrors. Running a simple Bash cron job to mirror repositories to a private instance is thankless, unglamorous work, but it saved our deploy window.
What is everyone else doing for repository redundancy during these outages?

reddit.com
u/Holiday-Record7341 — 2 days ago
▲ 1 r/devopsindia+1 crossposts

Here's what changed in AI SRE vendor land this past week (6-13 Aug)

1) incident.io shipped Investigations and named the platform underneath it Nexus. Nexus is free on every plan. Investigations itself is gated to Pro and Enterprise, sold as an add-on to their Response product. The launch post has the most candid competitor admission I've seen this year: they had a version of this 18 months ago, called it "AI SRE," shipped it to design partners, and it was "confidently wrong" often enough that they pulled it back and rebuilt. That's the whole reason "AI SRE" quietly became "Investigations" on their site.

2) Cleric put real prices on the table for the first time. $1 per credit, an investigation costs 10 credits, plans start at 100 credits a month. First hard public number in this category: roughly $10 per investigation as a starting point to compare against. Same update added a second billable unit, change verification, where Cleric follows a deploy into prod for up to 14 days and flags regressions while the change context is still fresh.

3) Datadog beat on Q2 earnings, revenue up 36% to $1.12B, big customers up 23%, and the stock dropped anyway, reportedly on a usage cut from their biggest AI customer. Consumption-based observability revenue just got punished for being concentrated. Worth remembering next time someone pitches usage-based pricing as the safe default.

  1. Dynatrance is in this category properly now: an Autonomous SRE Agent plus a no-code agent builder, coordinating remediation across AWS, Azure and GCP. An incumbent with the telemetry and the enterprise contracts already in hand is a bigger deal here than another funded startup launching.

5) Resolve AI published head to head comparison pages naming incident.io and PagerDuty's SRE agent directly, the same week incident.io launched Investigations. They also shipped a benchmark arguing Sonnet at medium effort gets close to Opus on incident investigation for a fraction of the cost, which lines up with a separate Anyshift post finding smaller models plus deterministic execution erased most of the accuracy gap between model tiers on bounded infra questions.

6) Traversal quietly deleted a sentence naming American Express, Capital One, PepsiCo, DigitalOcean and Kraken as Fortune 100 customers, and pulled the PepsiCo case study video. Second time they've retracted named customer proof. Their newer customer story goes with "a leading global crypto exchange" instead of naming Kraken outright, which tracks.

  1. Rough week for autonomy claims generally. OpenAI disclosed agents coordinating across sandboxes to reach Hugging Face during testing, the UK AI Security Institute reported agents creating fake identities to get around access rules, and OpenAI paused a system after it started finding zero days on its own. Every enterprise buyer read some version of this, probably why incident.io leading with an honest failure story landed the way it did.

  2. AWS had its fourth reliability incident in four months, another us-west-2 failure on close to the same network path as the July 24 outage. Same-path repeat failures are exactly what a change-aware agent should flag, and what a human paged at 3am usually doesn't connect.

Smaller stuff: PagerDuty's sitemap has an unlinked "AI Startups Trial" page sitting there, Rootly quietly dropped the explicit gpt-3.5-turbo mention from its privacy policy for a generic subprocessors list, and Neubird published 22 new glossary pages in one shot plus a "top 25 autonomous ops platforms" comparison.

Sources below.

reddit.com
u/Holiday-Record7341 — 5 days ago

Here's what changed in AI SRE vendor land this past week (25-31 July)

  1. Resolve AI now charges for results instead of tokens. They say if their agent wastes tokens, that's their cost to eat, not yours.
  2. Datadog has made Agent Observability free for up to 40,000 LLM spans a month. They didn't announce it explicitly, they just changed the line at the bottom of their AI blog posts.
  3. Opsgenie will be shutting down in April 2027. PagerDuty wrote a post to win those customers, and put incident.io in a comparison table. incident.io hit back point by point. As said by incident.io "calling our AI "limited" is "like describing PagerDuty as a pager."
  4. NeuBird put up a page comparing itself to Datadog Bits AI. It's the best public summary of where Bits AI is right now. Bits AI SRE is now called Bits Investigation. It's billed in AI Credits sold in 500-credit bundles, and unused credits expire each month. Their remediation and detection features were shown at DASH 2026 but are still in Preview.
  5. Traversal says most AI SRE tools pull data too late. Their point is that the other tools query your observability APIs during the incident and get stuck behind rate limits. They stream the data ahead of time instead.
  6. PagerDuty has a new CEO as John DiLullo takes over, Jennifer Tejada moves to Executive Chair. They also shared four straight quarters of GAAP profit and a $100 million buyback.
  7. Everyone is now selling prevention, not just faster fixes. NeuBird wrote a buyer's checklist around it. Their test question: "Show me an incident you prevented that never generated an alert."
  8. Sherlocks AI shipped an automation builder. Pick a trigger, describe the job in plain English, choose where the result lands. Their pitch: "Describe the job. It runs." The examples lean preventive rather than reactive, watching canary deploys, flagging infrastructure changes in PRs, tracking pod memory drift.
  9. Resolve AI is pushing background agents. Agents that keep running between incidents, watching deploys and doing checks. A customer quote sums it up: "The alerts are already investigated. The deployment summaries are already written."
  10. Traversal is renaming things. "Chat with Prod" is becoming "Production Support." "AI-Native Compressor" is becoming "Causal Indexer." Both old and new names are live on their site right now.

Sources linked below in comments.

reddit.com
u/Holiday-Record7341 — 19 days ago

Here's what changed in AI SRE vendor land this past week (25-31 July)

  1. Resolve AI now charges for results instead of tokens. They say if their agent wastes tokens, that's their cost to eat, not yours.
  2. Datadog has made Agent Observability free for up to 40,000 LLM spans a month. They didn't announce it explicitly, they just changed the line at the bottom of their AI blog posts.
  3. Opsgenie will be shutting down in April 2027. PagerDuty wrote a post to win those customers, and put incident.io in a comparison table. incident.io hit back point by point. As said by incident.io "calling our AI "limited" is "like describing PagerDuty as a pager."
  4. NeuBird put up a page comparing itself to Datadog Bits AI. It's the best public summary of where Bits AI is right now. Bits AI SRE is now called Bits Investigation. It's billed in AI Credits sold in 500-credit bundles, and unused credits expire each month. Their remediation and detection features were shown at DASH 2026 but are still in Preview.
  5. Traversal says most AI SRE tools pull data too late. Their point is that the other tools query your observability APIs during the incident and get stuck behind rate limits. They stream the data ahead of time instead.
  6. PagerDuty has a new CEO as John DiLullo takes over, Jennifer Tejada moves to Executive Chair. They also shared four straight quarters of GAAP profit and a $100 million buyback.
  7. Everyone is now selling prevention, not just faster fixes. NeuBird wrote a buyer's checklist around it. Their test question: "Show me an incident you prevented that never generated an alert."
  8. Sherlocks AI shipped an automation builder. Pick a trigger, describe the job in plain English, choose where the result lands. Their pitch: "Describe the job. It runs." The examples lean preventive rather than reactive, watching canary deploys, flagging infrastructure changes in PRs, tracking pod memory drift.
  9. Resolve AI is pushing background agents. Agents that keep running between incidents, watching deploys and doing checks. A customer quote sums it up: "The alerts are already investigated. The deployment summaries are already written."
  10. Traversal is renaming things. "Chat with Prod" is becoming "Production Support." "AI-Native Compressor" is becoming "Causal Indexer." Both old and new names are live on their site right now.

Sources linked below in comments.

reddit.com
u/Holiday-Record7341 — 19 days ago
▲ 65 r/devopsindia+1 crossposts

Here's what changed in AI SRE vendor land this past week (25-31 July)

  1. Resolve AI now charges for results instead of tokens. They say if their agent wastes tokens, that's their cost to eat, not yours.
  2. Datadog has made Agent Observability free for up to 40,000 LLM spans a month. They didn't announce it explicitly, they just changed the line at the bottom of their AI blog posts.
  3. Opsgenie will be shutting down in April 2027. PagerDuty wrote a post to win those customers, and put incident.io in a comparison table. incident.io hit back point by point. As said by incident.io "calling our AI "limited" is "like describing PagerDuty as a pager."
  4. NeuBird put up a page comparing itself to Datadog Bits AI. It's the best public summary of where Bits AI is right now. Bits AI SRE is now called Bits Investigation. It's billed in AI Credits sold in 500-credit bundles, and unused credits expire each month. Their remediation and detection features were shown at DASH 2026 but are still in Preview.
  5. Traversal says most AI SRE tools pull data too late. Their point is that the other tools query your observability APIs during the incident and get stuck behind rate limits. They stream the data ahead of time instead.
  6. PagerDuty has a new CEO as John DiLullo takes over, Jennifer Tejada moves to Executive Chair. They also shared four straight quarters of GAAP profit and a $100 million buyback.
  7. Everyone is now selling prevention, not just faster fixes. NeuBird wrote a buyer's checklist around it. Their test question: "Show me an incident you prevented that never generated an alert."
  8. Sherlocks AI shipped an automation builder. Pick a trigger, describe the job in plain English, choose where the result lands. Their pitch: "Describe the job. It runs." The examples lean preventive rather than reactive, watching canary deploys, flagging infrastructure changes in PRs, tracking pod memory drift.
  9. Resolve AI is pushing background agents. Agents that keep running between incidents, watching deploys and doing checks. A customer quote sums it up: "The alerts are already investigated. The deployment summaries are already written."
  10. Traversal is renaming things. "Chat with Prod" is becoming "Production Support." "AI-Native Compressor" is becoming "Causal Indexer." Both old and new names are live on their site right now.

Sources linked below in comments.

reddit.com
u/Holiday-Record7341 — 19 days ago
▲ 0 r/sre

Our agent nailed the correlation in an incident but Then it suggested restarting the wrong pod.

I have had this happen twice to us now. The agent pulls dashboards, logs, and deploy history together and lands on a hypothesis fast, way faster than a person doing it manually at 2am.

Last time it flagged a memory metric that lined up almost perfectly with the incident window. It turned out to be correlated, not causal and it still suggested restarting the pod.

Someone caught it before we shipped in the end but If the person on call had been more tired or newer to the system, I'm not sure they would have.

So the agent owns the correlation and the first hypothesis now, and a person signs off on anything that actually touches production which is a perfect split that has held up so well for us so far

Still struggling with the correlation part, any help?

reddit.com
u/Holiday-Record7341 — 24 days ago
▲ 0 r/sre

GPT-5.6 Sol deleted prod databases after its own warning system flagged the risk over 6x. Anyone else nervous about how agents handle their own alerts?

Saw this Sunday and I'm still thinking about it a lot, the latest OpenAI agent, GPT-5.6 Sol, deleted production files and a database, all while an internal warning system had already flagged the risk at 6.3x before it happened. TechTimes has the writeup if you haven't seen it (linked in comments).

As usual, my first reaction was "okay, another agent-goes-rogue story," which, fine, we've all read a few of those this year and mostly skimmed past them. What actually stopped me was the 6.3x number. This shows that the system yelled and got ignored which is never a good sign.

We've all had a human do the ignoring version of this, Alert fires for the fifth time in a week, looks like noise, on-call swipes it away, and 9 times out of 10 nothing happens. That's a known failure model which we've built runbooks and escalation paths around. What I haven't seen anyone actually solve is what happens when the thing doing the ignoring is the same system that's about to take the destructive action. There's no separate person in that loop to go "hang on, why does this keep flagging."

We run agentic stuff in a couple of low-blast-radius places, config rollout checks mostly, and I caught myself last month trusting the agent's own confidence score more than I should have, purely because it's a number and numbers feel objective. Took me a minute to realize a warning generated by the same policy that's deciding to act isn't really an independent check. It's the same brain grading its own homework.

How are you guys actually gating this in practice if at all?

reddit.com
u/Holiday-Record7341 — 29 days ago
▲ 0 r/sre

AWS CloudFront outage last week, 59 minutes just to confirm it was real. Why is this gap still existent?

AWS's CloudFront outage on July 16 ran from 07:45 to 11:18 UTC, about three and a half hours. The Root cause was a fleet limit inside the subsystem that routes edge traffic to private VPC origins.

As usual the part that got less coverage was that it took 59 minutes just to confirm the outage was real and scoped, before any fix work started. Most public post-mortems bury that detection-to-confirmation gap inside one "identified" timestamp instead of breaking it out.

Once AWS called it, resolution took another 2 hours 34 minutes based on the public timeline. Most of which could have just been rollback, AWS did not clarify anything.

For anyone who's run infra at that scale has lived some version of that first hour, a wave of alerts from unrelated services, and someone has to decide fast whether it's one upstream problem or five separate ones. Get that call wrong and it burns the rest of the window which looks like the case of the AWS issue.

reddit.com
u/Holiday-Record7341 — 1 month ago
▲ 61 r/sre

Google SRE's new AI ops whitepaper, the separate execution control plane is the part I haven't wrapped my head around yet.

We're working through how to add AI-assisted mitigation to our on-call workflow, while referencing Google SRE's white-paper from May. I noticed it's more concrete and more complicated at the same time.

The architecture has three pieces, AI Operator for autonomous mitigation, Actus as an execution control plane, and IRM Analyzer for continuous readiness evaluation against historical incidents. The Actus piece is just confusing, The mitigation agent can't exceed what Actus allows, even when the agent's own reasoning suggests otherwise. Actus is an architectural constraint, baked into the control plane which is very different from a permission model or a flag you configure per environment.

The IRM Analyzer evaluates readiness nightly against past incidents, so there's an actual record of where the agent failed. This help earn trust through measurement.

The honest question here is what a non-Google version of Actus looks like. We don't have dedicated infrastructure for a separate execution control plane. The constraint we have today is just the on-call engineer reviewing before anything runs. That works until the volume doesn't let it.

Whitepaper: sre.google/resources/practices-and-processes/ai-engineering-reliable-operations/

reddit.com
u/Holiday-Record7341 — 1 month ago
▲ 80 r/sre

Meta runs 50,000 automated root cause analyses per day. Has anyone else built a separate Investigation Layer?

Stumbled upon the Meta DrP paper from last December. The headline - 50,000 automated root cause analyses (RCA) per day across 300+ teams, five years in production, with MTTR improvements between 20% and 80%.

The design choice that surprised me is that they built investigation as a completely separate platform from their observability stack. The playbooks, the pattern matching, and the hypothesis loop all live in a different system from the dashboards and traces.

We've been treating "better runbooks" as the answer to slow investigations for the past year. Runbooks help when the incident is a shape you've seen before, but they're inert artifacts someone still has to read and follow at 2am. What DrP encodes is the senior engineer's diagnostic sequence as something that runs at alert time automatically.

The part I genuinely don't understand is whether the separation is worth the engineering cost at smaller scale. 300 teams gives you enough incident volume that pattern learning is actually viable. We're running maybe 5 serious incidents a month. I'm not sure if that's enough signal for automated RCA to get smarter over time, or if it just stays static.

Has anyone tried building investigation as a separate concern from the observability layer?

reddit.com
u/Holiday-Record7341 — 1 month ago
▲ 1 r/sre

OpenAI API was degraded for 90 minutes on July 5. Anyone else spend the first 20 minutes looking at their own stack?

Hit elevated latency and rising error rates on our end around peak US hours on Saturday. First assumption was something we must have deployed. We checked recent changes, looked at our own services, and found nothing obviously wrong.

Took us about 20 minutes before someone checked status.openai.com. Degradation on GPT-4o and Assistants endpoints, started roughly when our numbers went sideways too.

The failure presented entirely inside our system before the upstream status page updated. Our error budgets were moving, our latency was spiking, and the root cause lived completely outside our stack. The SRE playbook assumes you're investigating your own infrastructure. It doesn't have a great answer for this pattern.

Anyone else hit this on Saturday? And how are people handling third-party API dependencies in SLOs? We treat OpenAI as a hard runtime dependency now and I'm not sure our incident process actually reflects that yet.

reddit.com
u/Holiday-Record7341 — 1 month ago

The Tata DC fire knocked out Google Cloud India for hours. Anyone running production in India think through what happens to your on-call when the monitoring stack burns down with the data center?

A fire at Tata Communications' Delhi facility on June 24 took down Google Cloud India connectivity. Reuters reported one firm lost 20 years of operational data.

Most of the coverage is on the connectivity outage. The part I can't stop thinking about, when the physical facility fails, all your observability infrastructure in that region fails with it. Logs, metrics, traces, dashboards. All dark at the same moment as the services you're trying to investigate.

Every incident response flow I've seen assumes the monitoring layer survives. You get paged, you open Grafana or whatever, you start correlating. A fire removes that entirely. Your on-call is staring at alerts with nothing behind them.

Anyone actually running India-facing infra with a real plan for this, or is it mostly "hope the secondary region has enough context"?

reddit.com
u/Holiday-Record7341 — 2 months ago
▲ 104 r/sre

AWS DynamoDB was down for hours on June 28 while the status page said "operating normally." Cost us 3 hours of assuming it was our fault.

DynamoDB us-east-1 was having a bad day on June 28 and we lost about 3 hours assuming it was our fault.

Errors started climbing, we went straight to our own code. Questioned a deploy from earlier that morning, pulled in two people who weren't on call, spent time we didn't have going through changes that turned out to be fine. The AWS status page was green the whole time, so we kept looking inward.

Eventually someone just tried writing to DynamoDB directly from their laptop and it was clearly broken on AWS's end. That's when we checked Twitter and found a bunch of other people hitting the same thing.

The status page didn't update for another hour after that. What stung was that this was a solvable problem. A simple check on our own write success rate, with our own threshold, would have told us within minutes that the failure wasn't in our code. We've since set that up for every external dependency we use. Obvious in hindsight, annoying that it took this to get there.

reddit.com
u/Holiday-Record7341 — 2 months ago
▲ 3 r/sre

DORA has tracked MTTR for years. For most teams it hasn't moved. What actually moved it for you?

We've been grinding on incident response time for the past year. The DORA (DevOps Research and Assessment) 2023 report shows the elite cohort at under an hour for MTTR (mean time to recovery); the bottom 60% still sitting at 1 to 24 hours, same as 2019.

The frustrating part is we added observability tooling over that period, more dashboards, better alerting, structured logs, and none of it moved the number.

What we eventually noticed is that the actual wall-clock time in most incidents goes to the hypothesis loop, you think you know the cause, you check 3 tools, you're wrong, you form another theory. The fix itself is usually fast, sometimes anticlimactic, once you find the root cause.

Is this a universal pattern or just something very specific to our stack. If you and your team actually moved the number, help a fellow redditor?

reddit.com
u/Holiday-Record7341 — 2 months ago
▲ 0 r/sre

Discord's 2.5-hour RCA was a correlation problem, not a data problem. Anyone solved this?

On June 19, Discord had message delivery lag up to 4 minutes for a subset of users. Root cause was Redis keyspace eviction from memory pressure caused by an unrelated deploy. The post-mortem line that stuck: the team had Redis metrics, latency dashboards, and error rates visible throughout, and the causal chain still took 2.5 hours to reconstruct by hand.

Every link was instrumented. The sequence only became obvious after someone stitched together timestamps across systems, matching numbers that were off by a few seconds because two services logged in different time zones.

I've done that stitching. It's the same unglamorous 90 minutes whether you're 2 years into SRE or 10.

What I often find myself coming back to is whether this is a tooling problem or a mental model problem. Discord isn't understaffed or undertooled. Even if a system had flagged the correlation automatically, someone still has to decide whether to trust that read or keep looking. That validation step has a floor.

Has anyone found a setup that actually shortens the reconstruction phase, not the detection phase?

reddit.com
u/Holiday-Record7341 — 2 months ago
▲ 0 r/sre

A HN thread past weekend, "why does on-call still feel broken after years of investment?" got over 300 upvotes.

The complaints aren't about the page volume. People were complaining about the same 4 alerts, 2 hours of manual cross-referencing, one root cause that the alerts were pointing at the whole time.
That pattern caught my attention because the routing problem did get better, definitely. Smarter grouping, better noise suppression, more granular escalation policies. On-call noise came down for a lot of teams over the last few years. Unfortunately the burnout didn't follow it down. The comments are describing is the correlation step. Holding context across Datadog, PagerDuty, Kubernetes events, and your database at 3 AM while building a coherent timeline.

Honestly a HN thread is not at all a good sample to judge on but it is a very common problem i see people face every other day.

reddit.com
u/Holiday-Record7341 — 2 months ago
▲ 84 r/sre+1 crossposts

Anthropic's own safety team is now documenting failure modes that SRE tooling has no coverage for

The Claude 4 system card has a section on agentic deployment risks that I keep coming back to. "Long tool-call chains with irreversible side effects" is how they categorize one of the primary risk categories. That's a real production concern now, not a hypothetical.
The problem is that every existing observability primitive is built around metrics, logs, and traces. None of those tell you why an agent took a sequence of actions. You can see that a tool was called. You can't reconstruct whether the decision chain leading to it was coherent or had drifted somewhere upstream. Mean time to detect something in this category is probably not great. Mean time to understand it is going to be a lot worse.

Anyone running Claude 4 agents in production right now: how are you handling the investigation side when something goes sideways? Curious whether teams are building anything specific for this or just falling back to log correlation.

reddit.com
u/Holiday-Record7341 — 2 months ago

Anyone hosting a side event at KubeCon this year? Drop it here.

KubeCon India is June 18-19 in Mumbai.

If you're hosting a side event, meetup, dinner, or any gathering around the conference dates, drop a comment with the details: event name, date, time, location, and a link if you have one.

Trying to keep a running list so people have one place to check.

reddit.com
u/Holiday-Record7341 — 2 months ago

Sessions I'm looking forward to at KubeCon India 2026 (June 18-19, Mumbai)

AI on Kubernetes:
"Beyond VLLM: Distributed LLM Inferencing With llm-d on Kubernetes" — Ravindra Patil (Red Hat). GPU management and model routing at scale. Practically useful if you're running AI workloads in production.

Observability:
"Who Watches the Watchers? From Closed Observability to Open Control at Scale" — Aditi Gupta (JioHotstar), Madhu Patel (Adobe), Sandeep Kanabar (Gen). Three practitioners, one stage. Telemetry at real production scale.

Security:
"SPIFFE & OpenFGA Based Identity/Authz for Agentic AI" — Rahul Jadhav (AccuKnox). Zero trust for AI agents is a problem most teams haven't solved yet. Curious what the proposed architecture looks like.

Operations:
"The Leapfrog Upgrade Playbook: Upgrading When You're Years Behind" — Yug Gupta (Walmart Global Tech). Every team has this problem. Nobody talks about it openly.

Full schedule: https://www.cncf.io/announcements/2026/03/10/cncf-unveils-kubecon-cloudnativecon-india-2026-schedule/

Which ones are you planning to attend?

u/Holiday-Record7341 — 2 months ago

OpenAI’s June 4 outage traced to a K8s config change that degraded traffic routing across regions. How do you encode the blast-radius pattern for config rollouts?

OpenAI's status page on June 4 attributed a multi-hour ChatGPT and API outage to a Kubernetes
configuration deployment that degraded traffic routing across regions. Hours of impact, not minutes.
Config-change-induced routing failures have a recognizable fingerprint if you've seen them before:
latency spike first, then partial 5xx, then regional skew starts appearing in the distribution. A senior
SRE who's debugged one of these before gets to the right hypothesis fast. Someone without that
pattern in their head takes much longer, because every symptom is consistent with 4 other failure
modes too.
The question I keep coming back to: how do teams actually transfer that "I've seen this before"
knowledge? Runbooks capture resolution steps, not the diagnostic reasoning that led there.
Postmortems capture what happened, not the hypothesis path the on-call ran.
We've tried annotating our own runbooks with "if you see X + Y together, this is the failure class to
check first." Kinda works. Doesn't survive topology changes well.
Curious how others handle this. Specifically for config-change blast radius: is there a format you've
found that actually helps a junior on-call reach the right hypothesis faster, or is it mostly pairing and
osmosis?

reddit.com
u/Holiday-Record7341 — 2 months ago