r/openstack

▲ 0 r/openstack+1 crossposts

[Help] How does OpenStack manage and integrate dozens of physical servers?

I am a beginner who has just started working in the data center industry. As a beginner, I have many doubts and questions about the technical aspects. Regarding the OpenStack platform, I am not sure how more than ten or even more servers are managed to enter the platform, including computing nodes, storage nodes, and GPU computing power servers. What technologies can be used to be recognized and managed by OpenStack, or what recommended YouTube tutorials can help me understand these technologies and learn more about related knowledge? Feel free to leave a comment. Thank you very much.

reddit.com
u/Chamo_Link_3309 — 2 days ago

What is the AWS/Azure/GCP big cloud equivalent of Zun+Heat+Gnocchi+Aodh?

Like I want autoscaling, and I want maybe 1 or 2 gb for the containers. They can autodelete the docker logs after 1 day or 2 days .

I have a vm seperately for db and redis.

Why is everythign so costly??

I want zun containers to just autospin and scale depedning on the traffic the minimum being 1.

Our country does not have a strong openstack public clouds (self-managed) :( . So want to know if there are any equivalents that are extremely cheap and provides autoscaling without costing a bomb.

Here I mean, containers, that directly run on baremetal so as to provide maximum power output like zun.

reddit.com
u/Rare_Purpose8099 — 3 days ago

Am I too niche, targeting the wrong roles, or just in the wrong market?

I’ve spent almost 4 years at the same company since university. Am I too niche, or am I just in the wrong market?

I'd really appreciate some honest perspective from people working in cloud infrastructure, platform engineering, private cloud, telco cloud, networking, or infrastructure engineering.

I graduated from university and joined my current company shortly afterwards. I've now been there for about 4 years, and I've basically built my entire professional career in the same environment.

That's actually one of the reasons I'm finding myself a little stuck.

I've learned a huge amount and had a lot of freedom to build things, but I haven't really experienced working at other companies, especially larger engineering organizations where I could work at a bigger scale and see how these environments operate.

My background is quite infrastructure-heavy. I've worked across ISP infrastructure, Linux, networking, Kubernetes, virtualization, distributed storage, cloud, DevOps and security.

Some of the technologies I've worked with include:

  • Kubernetes
  • Talos Linux
  • KubeVirt
  • Rook/Ceph
  • Cilium/eBPF
  • BGP, VLANs and LACP
  • RADIUS/AAA
  • AWS
  • Terraform
  • ArgoCD/GitOps
  • GitHub Actions
  • Prometheus/Grafana
  • Infrastructure automation and security tooling

I'm also currently working hands-on with OpenStack because I'm trying to deepen my understanding of traditional private cloud and telco/NFV infrastructure, and understand how it compares with the Kubernetes-native infrastructure I've been working with.

A lot of my experience has come from actually building things.

For example, I've worked on a Kubernetes-based private cloud running on bare metal, combining Kubernetes, KubeVirt, Ceph, Cilium networking, tenant isolation, BGP routing and public/private VM connectivity.

I've also worked on ISP infrastructure, including subscriber authentication, RADIUS/AAA, billing integration, MikroTik routing and network service delivery.

But there's an important caveat to all of this.

I don't work for a huge cloud company, hyperscaler, or major technology company. I work for a relatively small ISP in Nigeria.

I've actually been very lucky in that environment.

My original job description didn't say that I needed to build a private cloud, learn distributed storage, work on Kubernetes virtualization, or learn all of these different areas.

I had to find opportunities to do those things and deliberately build those skills.

Whenever I came across a problem or something interesting that could improve the infrastructure, I'd learn what I needed, experiment with it and try to make it useful to the company.

The company gave me the freedom to do that, and I'm genuinely grateful for it. A lot of the experience on my CV probably wouldn't exist if I hadn't been given that freedom.

But now I'm starting to feel limited by the environment I'm in.

Not necessarily because the company is bad, but because I don't think I can continue building the kind of experience I want at the pace I want.

I've been there since university, so I also don't know what I'm missing.

I don't know what infrastructure engineering looks like at a larger company.

I don't know how much of what I've been doing would normally be handled by dedicated platform, networking, storage, SRE, cloud or infrastructure teams.

I don't know whether my experience is actually unusual or whether I'm simply getting a distorted view because I've had to wear so many hats.

And that's one of the biggest reasons I want to leave.

The other reason is compensation.

I've grown considerably beyond the scope of the role I originally started in, but my compensation hasn't really caught up with the level of responsibility, complexity and expertise I've accumulated.

I've tried communicating the value and complexity of the work I've been doing, but it hasn't really changed the situation.

So, I want to move on.

Not because I hate my current company. In fact, I'm grateful for what it has allowed me to do.

I want to move because I want to experience a different engineering environment, work with people who are operating infrastructure at a different scale, learn how other organizations solve these problems, and hopefully be somewhere that values this kind of work more appropriately.

The problem is that getting another job has been much harder than I expected.

And because I've only really known one company, I'm honestly not sure whether the problem is my profile, my positioning, the market, or all three.

Am I too niche?

Is the combination of Kubernetes + virtualization + distributed storage + networking + ISP/telco infrastructure something that has relatively little demand outside certain companies?

Am I searching for the wrong roles?

Or is this simply a case of being in a country where there isn't enough demand for this kind of infrastructure engineering?

At the moment I'm considering roles like: Cloud Infrastructure Engineer, Infrastructure Engineer, Platform Engineer, Kubernetes/Platform Engineer, Private Cloud Engineer, Cloud Infrastructure Architect, Telco Cloud/NFV Engineer, SRE, Infrastructure/Cloud Networking Engineer, OpenStack Infrastructure Engineer

But I'm honestly not sure which direction makes the most sense.

If you saw this background on a CV, what kind of engineer would you consider me to be?

What roles would you actually search for if you had this experience?

And perhaps more importantly:

What am I missing by having spent so much of my career at one company?

If you've moved from a small company into a larger engineering organization, what surprised you about the difference?

If you were in my position, would you:

  1. Go deeper into private cloud/telco cloud/OpenStack?
  2. Focus heavily on Kubernetes/platform engineering?
  3. Broaden into general cloud infrastructure?
  4. Move toward infrastructure/cloud networking?
  5. Target architecture-oriented roles?
  6. Or deliberately look for a role where I can be exposed to larger-scale infrastructure and learn how mature engineering organizations operate?

I'm not looking for reassurance. I'm genuinely trying to understand where I fit and what I should be doing next.

If geography wasn't a constraint, what kinds of companies, teams and roles would you target with this background?

And for anyone who has been in a similar position spending most of their career at one company and then trying to make that first big move, I'd really appreciate hearing what you wish you'd known before making the jump.

reddit.com
u/According-Capital522 — 7 days ago
▲ 4 r/openstack+1 crossposts

At a point in my career where I’m not sure whether to start over or just play along..?

I’ve been thinking about this for a while, and maybe writing it out will help me make sense of it.
I’ve been working in cloud infrastructure for around 4–5 years now, and I’m at a weird point in my career where I don’t know whether I want to start again somewhere else or just continue playing along with the way things are.
I recently had an injury that has put me on medical restrictions and, among other things, made me realize how dependent my current job is on being physically present. With a no-WFH policy, I don’t really know how long the company can accommodate me, especially when I don’t even know how long recovery is going to take.
That has pushed something else to the front of my mind.
I’ve been thinking about making a career move for a while now. But honestly, I’m at a stage where I’m questioning whether I should keep pushing toward something better or simply accept that this is how the industry works and learn to play along.
I’m not really the typical “yes sir, whatever needs to be done” kind of person.
I don’t enjoy being another piece of furniture that clocks in for nine hours because that’s what company policy says.
I like troubleshooting.
I like building things.
I like getting my hands dirty, learning something I don’t know, figuring out why something isn’t working, trying something unconventional, and then finding a way to implement it properly within the constraints of the environment.
Basically, think outside the box — then figure out how to make it work inside the box.
And when I pick up a problem, I tend to get restless until it’s either solved or I’m confident that it’s under control.
That’s always been the part of engineering I’ve enjoyed.
But lately I’ve started wondering whether that mindset is actually valued anymore.
Maybe I’m looking for something that the industry doesn’t really need from someone at my level.
Maybe I’m overthinking it.
Or maybe I’ve just been in the wrong environments.
I’m at a stage where I want to look at the bigger picture without losing the eye for detail. I want to keep evolving instead of simply getting comfortable doing the same thing because that’s how it has always been done.
And now, because of the injury, there’s also a practical clock ticking in the background.
I need to start looking seriously at remote opportunities because I genuinely don’t know how long my current situation is going to remain workable.
So I guess I’m asking people who are further along in their careers:
Does this mindset still have a place in the industry?
Is there still room for engineers who genuinely enjoy troubleshooting, building, experimenting and taking ownership — or does career progression eventually become more about managing expectations, following processes and playing the corporate game?
And if you’ve reached a similar point in your career, how did you decide whether to start over somewhere new or just learn to play along?

reddit.com
u/xpokeraxx — 7 days ago

Openstack Job Openings

Hey all, been looking for openstack openings for a while but currently finding no luck. I've great experience with bootstrapping & designing openstack & CEPH based clouds. I've integrated several openstack projects like Octavia, rancher, barbican etc. I've majorly worked on kolla ansible, cephadm, Netapp, HPE 3par and am capable on working and deploying these services on my own. Worked on many backup and migration projects as well using tools like commvault or hystax. Have a good understanding on both private and public cloud and good knowledge on the metal side of the stack as well. Have worked on several L3 troubleshooting (reviving dead rabbits💀) as well as worked on whole monitoring stack for openstack using prometheus, grafana, zabbix, dynatrace and currently working on a plan to upgrade Zed Openstack to epoxy.

All in all I've good experience under the belt on openstack but roles seem to be shying away from me. Thought I'd try the OpenStack linkedin to see if any redditors have any leads.

reddit.com
u/Gold_Fix_5309 — 9 days ago
▲ 29 r/openstack+3 crossposts

[Tool] I built an Oh My Zsh plugin to manage multiple OpenStack clouds, auto-venv, and fuzzy-find SSH / VNC consoles

If you work with more than one OpenStack cloud/project day to day, you know the drill: source ~/clouds/prod-openrc.sh, remember which venv has the right client version, openstack server list, ssh into whatever floating IP you copy-pasted from the output... I got tired of it and wrote a Zsh plugin to automate the parts I do 20x a day.

What it does:

- openv [cloud] — pick a cloud from clouds.yaml via fzf (with a live preview card showing auth URL / project / region, no secrets shown), activates the matching Python venv, and exports OS_CLOUD

- ops-ssh [user] [-i keyfile] [--insecure] — fuzzy-pick a server and SSH straight into its floating IP (prioritizes public IPs over private ones automatically); --insecure skips host-key checks, handy since floating IPs get recycled between instances constantly

- ops-console — fuzzy-pick a server and pop its Horizon VNC console URL open in your browser

- opcheck — quick openstack token issue sanity check so you find out your token expired before you're 3 commands deep into something

- opwho / opls — status of what's active / table of everything in clouds.yaml

- ophelp (or openv --help) — cheatsheet of everything below, printed in your terminal

- ~25 short aliases for the commands I run constantly (ops, opnet, opvol, opsec, opfl, oplb, etc. — full list in the README)

It's a normal Oh My Zsh plugin, MIT licensed, no telemetry, no dependencies beyond python3 + PyYAML + fzf (the plugin checks for these on load and warns if something's missing).▎

Install:

git clone https://github.com/whoami96/openstack-zsh-plugin.git ${ZSH_CUSTOM:-~/.oh-my-zsh/custom}/plugins/openstack

# add "openstack" to plugins=(...) in ~/.zshrc, then reload

Repo: https://github.com/whoami96/openstack-zsh-plugin

It's a fairly small, personal-scale tool — built it for my own workflow managing a handful of clouds — but figured it might save someone else the same repetitive typing. Happy to hear feedback or take PRs if it's missing something obvious for your workflow.

u/__Who_Am_I_____ — 11 days ago

Hey, we built an eBPF thing that tells you which OpenStack tenant used which bytes

We've been building a little thing called Lachesis and figured this crowd

might find it interesting. It's the network telemetry bit behind CubeCOS, our

private cloud on top of OpenStack.

Basically: OpenStack can tell you a tenant moved 5 TB, but not that 3 TB of it

went to the internet, 1 TB to another tenant, and 1 TB stayed put. For billing

that's the whole thing.

So Lachesis is a small agent that sits on each compute node, hooks eBPF onto

every VM's tap, and sorts traffic in the kernel: internet-out, internet-in,

same-tenant, cross-tenant — and figures out who the other end actually belongs

to. It untangles Octavia so the bytes land on the real tenant, and just spits

out Prometheus counters you can feed into CloudKitty or whatever you bill with.

Still very much a work in progress — core stuff works and we've run it on real

OVN clusters, but there's plenty left (hardening, IPv6, docs). OVN only for now.

Repo's here if you wanna poke at it: https://github.com/bigstack-oss/lachesis

Would love any feedback, especially if you've tried to bill OpenStack networking

and hit the same wall. Roast away 🙂

u/arashi87_ — 12 days ago

Cinder backed by LVM is making my Instances go read only

I'm having a problem where suddenly all my instances (all run some flavor of Ubuntu) have their filesystems go Read Only. It happens randomly and at least once it happened with nothing really running on the VMs.

Looking at one of the Compute/Storage nodes, I noticed a broken iSCSI connection. I run "dmesg -T" and got something like:

[Fri Aug 7 09:57:36 2026] connection8:0: detected conn error (1019)
[Fri Aug 7 09:57:38 2026] connection8:0: detected conn error (1019)
[Fri Aug 7 09:57:39 2026] sd 15:0:0:1: [sdc] Synchronizing SCSI cache
[Fri Aug 7 09:57:39 2026] sd 15:0:0:1: [sdc] Synchronize Cache(10) failed: Result: hostbyte=DID_TRANSPORT_FAILFAST driverbyte=DRIVER_OK

Restarting a bunch of Docker containers, followed by restarting the VM instances fixed the problem (specifically I restarted iscsid, tgtd, cinder_volume and nova_compute on all my storage and compute nodes).

Of course this is a bad fix if I have to do it every week.

Now, Gemini is telling me this is a consequence of using Cinder with LVM which, according to it "LVM + iSCSI is notoriously brittle for production OpenStack" and I should move to Ceph.

Is this true, or should a Cinder/LVM setup be a bit more resilient?

Context/extra info: my deployment is a Kolla-Ansible one (2025.1) and Ceph is no longer deployed by this version. I would need to deploy it separately.

reddit.com
u/Darkblood18 — 13 days ago