
Day 7 of the daily scenario-based interview prep series
DevOps Interview Prep Day 7: Secret Sprawl, Failed K8s Rolling Updates, and Prometheus OOM [Daily Series]
Hey r/devops,
Day 7
Scenario 1: The Secret Sprawl
You discover database passwords are hardcoded in 12 different repositories. Some in docker-compose files, some in Kubernetes manifests, some in shell scripts.
Question: What's your approach to centralize and secure these secrets? Name 2 tools/methods you'd use.
Hint: External secret stores, K8s native secrets vs external operators, CI/CD secret injection.
Scenario 2: The Rolling Update Gone Wrong
You triggered a Kubernetes rolling update for a new image version. Now half your pods run v1, half run v2, and v2 pods keep crashing. Traffic is partially broken.
Question:
- What's your immediate action to restore stability?
- What setting could have prevented this mess?
Hint: Rollback commands, maxUnavailable, readiness probes.
Scenario 3: The Prometheus Memory Explosion
Your Prometheus server keeps getting OOMKilled. Memory usage grows from 2GB to 16GB within hours. You have 50 microservices being scraped.
Question: What are 2 likely causes and how would you investigate?
Hint: High cardinality labels, retention period, number of active time series.
Drop your answers below. Solutions tomorrow.
Week 1 Recap:
- Day 1: Container restarts, Registry auth, Pending pods
- Day 2: Zombie processes, Pipeline timeouts, Volume mounts
- Day 3: SSH lockouts, Disk space alerts, Grafana gaps
- Day 4: Git credential leaks, Docker networking, Nginx 502s
- Day 5: Slow Docker builds, CrashLoopBackOff debugging, Merge conflicts
- Day 6: Env variable issues, ALB health checks, Terraform state lock
Week 1 done! What should Week 2 focus on? More K8s? Jenkins pipelines? Linux troubleshooting? Let me know in the comments.
YT Link: https://www.youtube.com/watch?v=ul9z-SBo53w&list=PLqOrZmpwbWUKRQTrFpqAKhChaTq0l5bIw