r/devops

▲ 30 r/devops

I actually enjoy my job, I just hate the way deadlines are set

I genuinely love my work, but the way deadlines are given is slowly killing my interest in it

My manager will be like, "This should take 2 hours.Its just matter of 1-2 days not more than it, Use Claude and get it done by 5 PM.”

Meanwhile, I needed to do development, testing for quite a good time. I am just getting way too distracted from my field and not learning much because of claude.

reddit.com
u/throwaway-well — 20 hours ago
▲ 1 r/devops+1 crossposts

AWS DevOps Professionals: Seeking Real-World Experience & Interview Insights

Hi Everyone,

I am moving from DEV to DevOps.

I’m preparing for AWS DevOps interviews and would love to hear from people currently working in AWS DevOps.

Could someone please share briefly:

- Your current role & responsibilities

- Your DevOps/AWS experience

- Your current project architecture

- Tools/services you use in your day-to-day work

Real-world examples would be really helpful for interviews and for others in the community too.

Thanks in advance! 🙏

reddit.com
u/sandy_24998 — 16 hours ago
▲ 60 r/devops

Officially a KubeAstronaut Now

Just passed the KCSA exam and officially became a Kubestronauts!

If anyone is preparing for KCSA or any other exam and has questions, feel free to ask.

​edit: Reddit titles are permanent, spelling mistakes are forever 🫠

reddit.com
u/root0ps — 1 day ago
▲ 31 r/devops

How aggressively should non-prod AWS environments be shut down?

I've been looking into non-prod AWS costs lately, and I'm wondering how far people actually go with shutting these environments down.

Scheduling dev/test environments to shut down overnight or over the weekend seems like an easy win. The tricky part seems to be taking them all the way to zero, especially when someone suddenly needs the environment and has to wait for it to come back.

What's the practical approach here? Do people just accept the startup delay, keep a minimum capacity running, or is there a better way to handle it?

reddit.com
u/whispered_word12 — 2 days ago
▲ 2 r/devops

Removing a key from a file doesn't remove it from your repo. How are you handling history scanning?

I always assumed deleting a key from a file was enough but from what I’ve learned that’s not how git works

Pre-commit hooks and CI only look at the diff. So if a key gets committed and you delete it in the next commit, every check turns out OK from then on, but the value is still sitting in history and still valid. The recommendation is to scan full history on a schedule and treat anything you find as exposed, and rotate it, even though it's long gone from the current files.

So the scan is really just telling you a leak already happened, and rotation is the part that actually contains it.

I have side projects from years ago where I don't remember what was committed before I knew better. Some of those keys are probably still valid.

Do you run scheduled history scans, or just pre-commit and CI? And when something surfaces from years back, do you rotate it or make a call based on whether the repo was ever public?

reddit.com
u/Chris__Codes — 1 day ago
▲ 64 r/devops

Cloud sovereignty is starting to become an architecture problem

There’s a lot more to sovereignty than choosing an EU cloud region.

With the EU Data Act, NIS2 and DORA, teams also need to think about where control and state live, where observability data goes, who can access the environment, and how dependent the workload is on a particular provider.

This recent CNCF article looks at those questions from the cloud-native architecture side and shows how separating different platform responsibilities can help.

Thought this was worth sharing given how much the sovereignty discussion is growing in Europe.

https://www.cncf.io/blog/2026/08/18/cloud-native-platform-sovereignty-through-multi-plane-architecture/

u/No-Income-2235 — 2 days ago
▲ 114 r/devops

How do you deal with multitasking throught the day?

I start the day with vscode, firefox and iterm opened

I finish the day with 5 vscode windows, 10 terminals and infinite tabs - and somehow my task backlog has growed...

How do you guys deal with this?

reddit.com
u/GBT55 — 2 days ago
▲ 2 r/devops

Building an Observability Pane

Hi Observability & DevOps Experts,

I'm looking for guidance from teams that have successfully scaled observability across large enterprise environments.

We operate a large-scale estate spanning AWS, Azure, and on-premises environments and have been using Datadog for several years. Over time, a significant amount of technical debt has accumulated around our observability implementation.

Current challenges include:

  • Datadog Agents managed differently across teams and platforms.
  • Custom log collection configurations distributed across hosts and applications.
  • APM, RUM instrumentation owned by individual application teams.
  • Inconsistent tagging standards and monitor configurations.
  • Outdated agents and instrumentation libraries.
  • Heavy dependency on multiple teams for upgrades and configuration changes.
  • A large portion of Datadog provisioning and onboarding is still handled manually.

As a result, maintaining and evolving observability at scale has become increasingly difficult.

We are considering building a centralized "Observability Foundation" or "Observability Platform" that teams would consume as part of their standard deployment process.

Our goal is to provide reusable Terraform-based observability components that application and infrastructure teams can adopt during provisioning and releases.

Examples of what we would like to standardize:

  • Datadog Agent deployment and upgrades
  • Custom log collection configurations
  • Standard tags and metadata
  • Monitors and alert templates
  • Dashboards
  • OpenTelemetry / APM instrumentation standards
  • Synthetic monitoring configurations
  • Cloud integrations
  • Security and governance controls

Questions:

  1. Has anyone implemented a similar centralized observability platform or observability-as-code model at enterprise scale?
  2. What worked well and what were the biggest challenges?
  3. What observability components can realistically be centralized through Terraform modules, deployment pipelines, or platform services?
  4. What components typically must remain application-owned or infrastructure-owned and cannot easily be centralized?
  5. How do you handle APM instrumentation ownership, versioning, and upgrades across hundreds of services?
  6. What governance model have you found most effective:
  • Central observability team ownership
  • Platform engineering ownership
  • Federated ownership with standards enforcement
  • Something else
  1. How do you prevent observability drift over time, especially around:
  • Agent versions
  • APM libraries
  • Log configurations
  • Tags
  • Dashboards
  • Monitors
  1. If starting again today, would you build around:
  • Datadog native tooling
  • OpenTelemetry
  • An internal observability platform
  • A combination of the above
  1. What are the biggest architectural mistakes or anti-patterns we should avoid when designing this platform?

Our provisioning and infrastructure management are heavily Terraform-based, so we're especially interested in Terraform-centric implementation patterns and real-world lessons learned.

Looking forward to hearing how other organizations have approached observability standardization at scale and what you would recommend before we begin designing this solution.

P.S. - One of our key design goals is to avoid vendor lock-in. While Datadog is our current observability platform, we want the architecture to remain flexible enough that a future migration to another observability stack (e.g., Grafana, New Relic, Dynatrace, Elastic, Azure Monitor, or an OpenTelemetry-native platform) would require minimal changes to application teams and infrastructure code.

reddit.com
u/JayDee2306 — 1 day ago
▲ 3 r/devops

When should a small Node service stop accepting work during shutdown?

For a small service with HTTP requests, background jobs, and maybe WebSockets, I’m trying to keep shutdown behavior simple.

My current order is to stop accepting new work, let active requests finish, flush bounded notifications, then close the queue and database connections. I’m less sure how long to wait before forcing the process to exit.

What shutdown steps have actually mattered in production, and which ones turned out to be unnecessary complexity?

reddit.com
u/UkrMalt — 2 days ago
▲ 0 r/devops

Can there be only 1 SRE Engineer in a company?

Is there a possibility of being the only sre engineer in a company? Before choosing it I want to clarify it because if there could be then hes life could be problematic because he has to stay on call everyday. Also if it is not there then will be on call rotation right? Because I don't want to stay on call everytime that could be problematic. How many SREs are there in your team?

reddit.com
u/AcanthaceaeUnlucky18 — 2 days ago
▲ 9 r/devops

How do you manage and organize big tasks?

I work at a bank and we provide a PaaS for our internal clients.

recently we scaled the architecture a lot and so our monitoring system demanded a change.

this change was basically a complete refactor from the ground, a task that I singlehandedly worked out.

so the state of this feature is currently at first stage, every piece is bootstrapped, the architecture is deployed and working and now we are at the stage of deprecating our old monitoring system into the new one. This means migrating alerts, products, dashboards, bridges, etc...

the thing is, I find myself stumbling across the different problems and errors of the day to day operation without a real traced path of what I should be doing. I know what the end goal is but I dont quite know how to get there.

seems like the task is unfinishable, there's always something that lights an alarm after being deployed for 2 weeks, something I didn't take account for, or something that I realise I dont want it that way because another piece demands something different from it etc..

I've probably did 10 or 20 deployment iterations of some or all pieces of the system because of this...
Now I am at a point where my task backlog is so overwhelming that I dont know how to face it.

I dont know if its better to deploy the architecture in layers, deploy everything at once in one environment, deploy everything and update...

I think the main problem is my task management capacity...

tips, courses, anything on this topic is appreciated, thanks!

reddit.com
u/GBT55 — 2 days ago
▲ 42 r/devops

CKS Exam Experience 2026: What Helped, What Didn’t, and the Mistake That Cost Me Time

https://preview.redd.it/xwwdz2ozw9kh1.png?width=1246&format=png&auto=webp&s=f19f1cfadf6e4e48d2f56a0135f3fd5e49011bd1

Today, I passed CKS with 75% (not a good score) and wrote a detailed blog about my exam experience, preparation approach, the resources I used, and the Kubernetes security topics that helped me the most.

DMs are open if you are preparing for the exam. I can help with whatever is still fresh in my memory.

My biggest takeaway: CKS is noticeably harder than CKAD and CKA. It is not only about knowing Kubernetes commands. You need to understand why a configuration is insecure, how to fix it, and how to verify that your change actually worked.

The biggest mistake I made was spending around 10–15 minutes too long on one question because I felt I was close to solving it. That created unnecessary pressure towards the end and probably led to a couple of avoidable mistakes.

So my strongest advice is: if you are stuck and don’t see a clear path after a few minutes, mark the question and move on.

A few things that helped me:

Don’t memorise solutions. Understand the security reasoning behind them. If a NetworkPolicy, API server flag, securityContext, audit policy, or admission control changes slightly, memorised YAML will not help much.

Always verify your work. Security changes can easily break workloads or cluster components. Check Pods, control-plane components, logs, services, NetworkPolicy connectivity, admission behaviour, audit logs, node readiness, and systemd services wherever required.

Be comfortable with Linux as well as Kubernetes. CKS can require you to work with configuration files, systemd services, container runtimes, permissions, certificates, and node-level settings.

Use documentation whenever required instead of trying to remember every flag or custom resource.

Topics I would strongly recommend practicing:

  • Kubelet and etcd hardening
  • kube-apiserver authentication and authorization
  • Admission controls and ImagePolicyWebhook
  • Secure Dockerfiles and non-root containers
  • Container immutability and securityContext
  • Audit policies and API server logging
  • NetworkPolicy
  • HTTPS Ingress and TLS
  • ServiceAccount token security
  • Worker node administration and upgrades
  • SBOM and software supply-chain security
  • Restricted Pod Security Standard
  • Docker/container runtime hardening
  • Istio STRICT mTLS
  • Cilium network security
  • CIS benchmarks and kube-bench remediation

Resources I used:

  • KodeKloud CKS course
  • KodeKloud Ultimate Mock Exam Series
  • iximiuz Labs
  • KillerKoda
  • Killer.sh CKS simulator
  • ChatGPT/Claude for topics that needed a simpler explanation or extra practice scenarios

Between the KodeKloud course mocks and Ultimate Mock Exam Series, I had around six mock exams. I found them very useful and reasonably close to the level of difficulty you should prepare for.

Killer.sh felt a little off-track compared with the actual exam in some areas, but I would still recommend doing it. It is useful for practicing under time pressure, discovering knowledge gaps, and improving troubleshooting skills.

I also used ChatGPT and Claude quite a lot during preparation. CKS has many small security topics, and sometimes a course or lab explanation may not immediately click. In those cases, asking AI to explain the concept differently, compare configurations, or generate a small practice scenario was very useful.

The simplest advice I can give is: practice a lot, understand the security reasoning behind what you are doing, verify every change, and don’t let one difficult question consume your exam time.

I also wrote a full blog with more details on my preparation strategy, resources, task areas, mistakes, and lessons from the exam.

Blog: https://blog.prateekjain.dev/cks-exam-experience-2026-preparation-strategy-and-lessons-learned-1fad785a430b?sk=52b58a9c6d812444bc1340f15eda7dd6

reddit.com
u/root0ps — 3 days ago
▲ 0 r/devops

New to DevOps/AWS CodePipeline - Could Someone Help Me Understand Our Setup?

Hi everyone,
I recently joined a new company as a DevOps engineer, and I’m still trying to understand the existing infrastructure and CI/CD setup.
Our pipelines are running on AWS CodePipeline, and there are quite a few things I’m struggling to understand how the pipeline is structured, how the different stages work, deployments, and how everything is connected.
I’m currently learning, so I’d really appreciate it if someone experienced with AWS CodePipeline/DevOps could guide me through the basics and help me understand how to approach an existing setup.
If someone is willing to spend some time explaining it to me over a remote call/screen-sharing session, I’d be extremely grateful. I’m not asking anyone to access my company systems or credentials — I just want to learn and understand the concepts and workflow.

reddit.com
u/No-Heat-1958 — 3 days ago
▲ 11 r/devops

What are the best Wiz alternatives for mid sized company?

We've been evaluating Wiz over the past few weeks and I can definitely see why larger organisations like it. The visibility is impressive, but for a company of our size (around 200 people), it's difficult to justify the cost.

We're mostly running AWS, Kubernetes and containerised applications, with quite a small infrastructure team. We don't have the time to babysit another platform or sift through thousands of findings every week. We're looking for something that helps reduce risk without adding more operational overhead.

So far we've looked at Orca, Rapidfort and a few others, but it's hard to tell how they compare once you get past the sales material.

If you own or work at a small or mid sized organisation and you found the same challenge, what did you end up choosing instead of Wiz, and has it worked out?

reddit.com
u/Successful-Bat9218 — 3 days ago
▲ 0 r/devops

Is learning arch really worth it for devops

Hi everyone. I have a simple question is learning Arch Linux really worth it?

Most of what I've heard is that it helps with troubleshooting and gives you a better understanding of Linux. I'm already comfortable with Ubuntu, though, so I don't want to switch to Arch if the benefits are only marginal.

Would learning Arch actually give me a significant advantage, especially for someone interested in DevOps?

reddit.com
u/Strong-Interview45 — 3 days ago
▲ 98 r/devops

Did GitHub Just Gaslight Our Monitoring System?

Did anyone else notice that the GitHub status page reported an incident with GitHub Actions, only to deny it 47 minutes later?

Our monitors captured it, paged our on-call team, and then GitHub denied that any incident had occurred.

https://www.githubstatus.com/incidents/gx7js8bd0jpz

u/pod_army — 4 days ago
▲ 4 r/devops

QA Engineer (9 YOE) Looking to Transition into DevOps — Need Guidance on Roadmap & Resources

Hi everyone,
I’m a QA Engineer in the gaming domain with around 9 years of experience, and I’m seriously considering transitioning into DevOps. I’d really appreciate some guidance from people who have made a similar transition or are currently working in DevOps.
Here’s where I currently stand:
I have 9 YOE in QA/testing, primarily in the gaming domain.
I have a good understanding of SDLC and STLC B and how software moves through different stages from development to production.
I’ve been involved in the complete feature lifecycle — from initial specification/discussions, through development and testing, to production release.
I’ve used Jenkins for build creation and server deployments.
I use Git mainly for creating/raising PRs, but I haven’t worked extensively with Git commands and workflows such as push, pull, branching, rebasing, etc.
I’ve used Grafana for tracing application logs and investigating issues.
We use AWS SSM to log into different server boxes, tail server logs, modify server-side files/configs, etc.
I’ve recently started learning the basics of Python and Java.
I’m also fortunate to have a good relationship with our internal DevOps team and manager. I’m considering approaching them for an internal transition when I feel I’m ready.
I’m aware that my current skill set is far from what would typically be expected from a DevOps engineer, and I don’t want to underestimate the amount of learning required.
I’m 34 now, so I do sometimes feel like I’m starting this transition quite late. However, I genuinely want to make the move, and I’m willing to put in the time and effort.

  1. What I’m struggling with is where exactly to start and what order to learn things in.
    For someone coming from a QA background like mine:
  2. What would be a realistic DevOps learning roadmap?
  3. Which skills should I prioritize first — Linux, networking, Git, Docker, Kubernetes, CI/CD, Terraform, AWS, etc.?
  4. Are there any beginner-friendly courses/resources you would strongly recommend?
  5. How much programming/scripting should I learn, and should I focus on Python or Bash first?
  6. Given my existing experience with Jenkins, AWS SSM, deployments, logs, and the software release lifecycle, are there areas where I can leverage my QA experience?
  7. Would an internal transition into a DevOps team be a reasonable approach, even if I don’t yet meet all the requirements of a typical DevOps job?
    I’m not looking for shortcuts. I’d just really appreciate some practical guidance from people who have been through this journey.
    If you were in my position, what would you learn over the next 6–12 months, and in what order?
    Any roadmap, resources, project ideas, or personal experiences would be hugely appreciated.
    Thanks in advance! 🙏
    Rephrased with GPT
reddit.com
u/Full_Acanthisitta873 — 3 days ago
▲ 38 r/devops

What Kubernetes misconfigurations have caused you real production problems?

After working with Kubernetes in production, I've noticed that some of the most annoying incidents aren't caused by obvious failures. They're often caused by small configuration decisions that look perfectly reasonable during review.

Things like:

  • missing resource requests/limits
  • incorrect probes
  • overly permissive RBAC
  • missing PodDisruptionBudgets
  • unsafe container configuration
  • incorrect readiness behaviour
  • services without appropriate timeouts
  • configuration drift between environments

I'm curious what the DevOps community has actually encountered in production.

What's one Kubernetes configuration mistake that caused you a real incident?

I'd especially like to hear about the less obvious ones that aren't caught by the usual linters.

reddit.com
u/nerd3n — 4 days ago