r/devopsindia

▲ 4 r/devopsindia+1 crossposts

Resigning without an offer

Hi guys , has any Devops engineer in the recent past resigned without and offer and still managed to get a job with good hike ?

#toxicjob #resign

reddit.com
u/Predator1110 — 1 day ago
▲ 13 r/devopsindia+1 crossposts

Hiring - Sr DevOps Engr , Remote in India

Job Title: Senior DevOps Engineer - India

Experience: 10+ Years

Location: Remote
Max budget 40 L

Job Summary:

We are looking for a highly skilled Senior DevOps Engineer with 10+ years of experience and strong hands-on expertise in Oracle Cloud Infrastructure (OCI) and AWS (both are mandatory). The ideal candidate should have experience in CI/CD, Kubernetes, Docker, Terraform, Linux, and Infrastructure as Code (IaC).

Key Skills:

- Oracle Cloud Infrastructure (OCI) – Mandatory

- AWS – Mandatory

- Kubernetes & Docker

- Terraform / Ansible

- CI/CD (Jenkins, GitLab CI, GitHub Actions)

- Linux Administration

- Bash/Python Scripting

- Monitoring & Logging Tools

Preferred: OCI/AWS certifications and excellent troubleshooting skills.

sandeep@xociate.com

 

reddit.com
u/Maximum-Ad6056 — 4 days ago
▲ 1 r/devopsindia+1 crossposts

Here's what changed in AI SRE vendor land this past week (6-13 Aug)

1) incident.io shipped Investigations and named the platform underneath it Nexus. Nexus is free on every plan. Investigations itself is gated to Pro and Enterprise, sold as an add-on to their Response product. The launch post has the most candid competitor admission I've seen this year: they had a version of this 18 months ago, called it "AI SRE," shipped it to design partners, and it was "confidently wrong" often enough that they pulled it back and rebuilt. That's the whole reason "AI SRE" quietly became "Investigations" on their site.

2) Cleric put real prices on the table for the first time. $1 per credit, an investigation costs 10 credits, plans start at 100 credits a month. First hard public number in this category: roughly $10 per investigation as a starting point to compare against. Same update added a second billable unit, change verification, where Cleric follows a deploy into prod for up to 14 days and flags regressions while the change context is still fresh.

3) Datadog beat on Q2 earnings, revenue up 36% to $1.12B, big customers up 23%, and the stock dropped anyway, reportedly on a usage cut from their biggest AI customer. Consumption-based observability revenue just got punished for being concentrated. Worth remembering next time someone pitches usage-based pricing as the safe default.

  1. Dynatrance is in this category properly now: an Autonomous SRE Agent plus a no-code agent builder, coordinating remediation across AWS, Azure and GCP. An incumbent with the telemetry and the enterprise contracts already in hand is a bigger deal here than another funded startup launching.

5) Resolve AI published head to head comparison pages naming incident.io and PagerDuty's SRE agent directly, the same week incident.io launched Investigations. They also shipped a benchmark arguing Sonnet at medium effort gets close to Opus on incident investigation for a fraction of the cost, which lines up with a separate Anyshift post finding smaller models plus deterministic execution erased most of the accuracy gap between model tiers on bounded infra questions.

6) Traversal quietly deleted a sentence naming American Express, Capital One, PepsiCo, DigitalOcean and Kraken as Fortune 100 customers, and pulled the PepsiCo case study video. Second time they've retracted named customer proof. Their newer customer story goes with "a leading global crypto exchange" instead of naming Kraken outright, which tracks.

  1. Rough week for autonomy claims generally. OpenAI disclosed agents coordinating across sandboxes to reach Hugging Face during testing, the UK AI Security Institute reported agents creating fake identities to get around access rules, and OpenAI paused a system after it started finding zero days on its own. Every enterprise buyer read some version of this, probably why incident.io leading with an honest failure story landed the way it did.

  2. AWS had its fourth reliability incident in four months, another us-west-2 failure on close to the same network path as the July 24 outage. Same-path repeat failures are exactly what a change-aware agent should flag, and what a human paged at 3am usually doesn't connect.

Smaller stuff: PagerDuty's sitemap has an unlinked "AI Startups Trial" page sitting there, Rootly quietly dropped the explicit gpt-3.5-turbo mention from its privacy policy for a generic subprocessors list, and Neubird published 22 new glossary pages in one shot plus a "top 25 autonomous ops platforms" comparison.

Sources below.

reddit.com
u/Holiday-Record7341 — 5 days ago
▲ 12 r/devopsindia+2 crossposts

Top 7 Claude skills for frontend, by install count. Two repos hold five of the seven.

  1. vercel-react-best-practices, 631,348 (vercel-labs/agent-skills)
  2. vercel-composition-patterns, 288,538 (vercel-labs/agent-skills)
  3. shadcn, 271,674 (shadcn/ui)
  4. minimalist-ui, 247,075 (leonxlnx/taste-skill)
  5. develop-userscripts, 224,031 (xixu-me/skills)
  6. image-to-code, 210,463 (leonxlnx/taste-skill)
  7. vercel-react-native-skills, 187,251 (vercel-labs/agent-skills)

Method first, because a ranking is only worth as much as it: every skill filed under Frontend in the Skillselion catalog, sorted by installs, top seven. Nothing hand-picked.

Vercel takes three of the seven. leonxlnx/taste-skill takes two. Five from two authors.

The fair question is whether that is five decisions or two. The gaps answer it. Vercel's three run 631,348, then 288,538, then 187,251, a 70 percent spread inside one repo. If people were installing the repo wholesale those numbers would sit almost on top of each other. They do not, so this is individual choice, and first place really did beat its own sibling by 343,000.

shadcn is the only name here that people were already using before agent skills existed. Everything else was written for agents from the start.

Installs measure adoption, not quality, and nobody publishes uninstall data, ours included. If a number looks off, say so and I will pull it again.

The full list: Claude skills for frontend, ranked by installs.

u/skillselion — 6 days ago
▲ 6 r/devopsindia+3 crossposts

Researching production log analysis & RCA — looking for engineer feedback

Researching production log analysis & RCA — looking for engineer feedback

Hi everyone,

I'm doing some early market research around production troubleshooting and log analysis and would really appreciate feedback from people who actually deal with production systems.

I'm exploring a tool where engineers could interact with their logs using natural language, investigate production incidents, correlate events across different systems, perform RCA, and generate custom analysis/visualizations from historical log data.

Before building this further, I want to understand the actual problems engineers face today — how they investigate incidents, where existing tools fall short, how much time RCA takes, and whether this is a problem worth solving.

I've made a short 2–3 minute anonymous survey:

https://docs.google.com/forms/d/e/1FAIpQLSfbJt7moEOhZR9Xr8HXbKu5Y0F2Ep0Yv5xPFQlohKbH63r6Dg/viewform?usp=publish-editor

If you work with SRE, DevOps, infrastructure, backend, Kubernetes, networking, observability, or production operations, your experience would be especially valuable.

I'm looking for honest feedback, including negative feedback. I'm trying to validate the problem, not just validate my idea.

Thanks!

u/SmartGHST056 — 7 days ago
▲ 1 r/devopsindia+1 crossposts

I want to build an AI-powered enterprise investigation system and I want your advice on the best approach before I start building it.

Want to build an AI-powered enterprise investigation platform where users can ask natural-language questions across enterprise systems such as Jira, New Relic, Azure, AWS, etc. For now I'm starting with Jira + New Relic, with all integrations through MCP servers.

The goal is not just searching tickets; users should be able to ask arbitrary operational questions such as "Why did this incident happen?", "When did it happen?", "What caused it?", "Show related incidents", or "Give me correlated tickets including questions that require multiple systems and multiple steps of investigation.

I'm looking for advice from people who have built production-grade agentic systems. If you were starting this project from scratch, what approach would you choose for the agent/orchestration framework, MCP/tool selection, multi-step investigation and reasoning, evidence collection, and handling complex cross-system questions? What are the biggest challenges or failure modes I should expect, and what architectural decisions would you make differently to keep the system reliable, scalable, and maintainable as I add more enterprise systems? I'm deliberately not specifying my preferred framework or architecture because I want unbiased recommendations before I continue building.

reddit.com
u/Bluu_moon_ — 8 days ago
▲ 5 r/devopsindia+1 crossposts

Doubt !!!CONFUSIONS!!!

i am really new to devops and i have been studying networking as of now but i get to see like there is no entry level jobs in it neither intern roles as a devops engineer requires experience so what should i do?? as i am actually interested in the subject. if there are no jobs for freshers or even as an intern...

reddit.com
u/Opening_Dog533 — 13 days ago