Reddit DevOps
287 subscribers
81 photos
32.5K links
Reddit DevOps. #devops
Thanks @reddit2telegram and @r_channels
Download Telegram
DevOps embedded in a non-technical process team, no technical sparring partner

Looking for some outside perspective on how DevOps fits into org structures at your companies.

This is my third or fourth DevOps job, this time in embedded. In previous companies there were always some organizational or cultural hiccups, but stuff got worked out as a team, including on the DevOps side - people weren't DevOps themselves, but they were technical, so they got it.

Now, a few months into a new company, things feel different. The embedded department has already gone through two managers, each with a different vision for how DevOps should work. The current one seems pretty caught up in the AI hype - wants to "bring in AI" because everyone else has it, without really knowing what that means in practice. There was no DevOps in the department before me, so a lot of groundwork needed cleaning up first - repo structure, templates, CI/CD, that kind of thing. Garbage in, garbage out: you can't just bolt AI onto a messy foundation and expect it to work. He also doesn't seem to register that I've already been automating things with AI where it makes sense since I joined, once the basics were actually in place.

I also ended up placed in a non-technical/semi-technical team. The original plan was to hire a second DevOps so I'd have a sparring partner, but now it looks like that role might turn into some mixed "does a bit of everything" position, or even a "senior AI agent specialist" (lol). My current team is process/methodology people, Jira and workflow folks, and a release manager.

They don't really understand what I do technically, and they're very focused on "why." During planning, they want to rewrite everything in my Jira summaries into long descriptions based on their own interpretation, which usually doesn't match what the epic/task is actually about. The idea is that management should be able to glance at a board and understand what we're doing just from the title. In my previous companies, Jira was a tool for us, not something management used to check on our work.

My team manager (not the department head, and also non-technical) said he didn't understand what one of the tasks was about. When I asked if he'd actually read the epic's description, he said no - and explained that it's hard to tell what your people are doing from just a title. But that's the thing: I'd already written a detailed description for this project, which is a fairly big one made up of related chunks, at the epic level. He never opened it. From what I can tell, part of the disconnect is also that he doesn't have a technical background (never worked with AWS, for example), so even a well-written description might not land the way he expects.

From an organizational standpoint, how do you see this - is it a good idea for a DevOps role to sit inside a team like this, especially when there's already a separate cloud team in the department? I don't feel good there. There's a lack of understanding, they think in very corporate/process terms while I think technically, and since I'm the only DevOps, I genuinely don't have time for some of this overhead.

What worries me is burning out on micromanagement and feeling cut off from technical people in a few weeks - I also just work better under a technical manager.

Has anyone been in a similar spot? Do I wait and see how it develops, or should I propose alternatives to the embedded department head?

https://redd.it/1wqke4a
@r_devops
Anyone else had a 340% AI cost spike overnight? Here's what caused ours

At 02:17 UTC, a finance dashboard flashed red: LLM budget jumped from an expected $1,200 to $4,080 in a single day. A 340% spike, no warning.

Nobody on the team caused it on purpose. That's the annoying part about AI spend, it doesn't announce itself until the invoice does.

Walked through it step by step afterward: checked for model or config changes first (none), checked usage volume (flat versus the prior week), then pulled per-request cost and found one endpoint, a background retry job, was the outlier. Average cost per completed task had tripled even though cost per individual call hadn't moved. The retries were getting logged as separate normal-cost calls instead of one task that cost 3x.

Fix was attaching a stable task ID and rolling cost up by task instead of by raw call.

I want to know if others have chased a spike like this, and what your actual debugging path looked like, model change, retry storm, something else entirely?

https://redd.it/1wqm1lt
@r_devops
How can I deploy this microservices project for free without a credit card?

Hey everyone,

I'm building a graduation project with this architecture:

Frontend (Next.js)
↓
Traefik Gateway
├── Auth Gateway Service (nestjs)
├── Issues Service (nestjs)
├── Ingestion Service (nestjs)
├── User Service (nestjs)
└── SmartQA Vector Service (fastApi)

I also have the following infrastructure services:

Kafka for event-driven communication
MongoDB for issues and vector data
PostgreSQL for users
MySQL + Keycloak for authentication
Consul for service discovery
Traefik as the reverse proxy/API gateway

The backend services are deployed as separate Docker containers, and I'm currently looking for a way to deploy the whole application for free, if possible.

My main requirements are:

Free tier
No credit card required
Very low traffic — I'm essentially the only user, with at most 1–2 users
Support for Docker containers/multiple services
Ideally, persistent storage for MongoDB/PostgreSQL
I don't mind cold starts since this is mainly a graduation/portfolio project

What platforms or deployment approaches would you recommend for this kind of setup?

Thanks in advance!

https://redd.it/1wqpwys
@r_devops
Indian devs who self-host tools: what actually becomes painful after 6–12 months?

I’m curious about the long-term side of self-hosting from developers in India who actually run things in production.

Setting up a VPS + Docker Compose is usually the easy part.

What I’m more interested in is what starts becoming annoying after a few months:

backups and restore testing
upgrades breaking something
monitoring/alerts
disk usage
SSL/DNS issues
database maintenance
security patches
managing multiple servers
documenting setups so someone else can understand them
getting paged when something dies

Especially for small startups/agencies where there isn’t a dedicated DevOps/SRE person.

At what scale did self-hosting stop feeling “cheap” for you?

And which tools do you still prefer to self-host versus paying for a managed version?

Would be useful to hear actual production experiences rather than just setup tutorials.

https://redd.it/1wr1vsk
@r_devops
How are people verifying unattended/CI agent runs actually did what they claimed?

Running into something I'm curious how others handle: when a coding agent runs unattended (CI, scheduled job, autonomous pipeline), the only account of what happened is the agent's own closing summary. Nobody is watching in real time, and if the summary is wrong or incomplete, there is no independent way to know without going back and manually digging through logs after the fact.

Concrete case that got me thinking about this: had a subagent's tests genuinely fail during a run, and the main agent's final summary still said all tests pass, without mention of the failure. Would have gone completely unnoticed in an unattended context, nobody was reading the transcript live, and the summary gave no reason to suspect anything was wrong.

I am curious what people are actually doing here such as deterministic post-run checks, separate log validation, something else? Are you treating the agent's own report as trustworthy at all in these contexts, or building around the assumption that it might not be?

(For context: this is what pushed me to build a small open-source tool called Rashomon that independently records what actually happened during a Claude Code session, separate from the agent's own account. Happy to share details if it's relevant to what people are describing, but mainly want to hear how others are approaching this.)

https://redd.it/1wr2lng
@r_devops
How to maintain helm charts of multiple micro-services on k8s at scale ?

hey everyone,

how do you guys manage helm charts for multiple micro services.

we have multiple microservices and our CI-CD flow is when the docker image is built and pushed to a registry then the values file will be updated for the image tag and this change is detected by the gitops tools like argo or flux to deploy.

\- My question is we are having these different templates i.e a helm chart like below for each and every service and is there a way to handle this at scale with something like DRY ?

\- Is it possible to maintain a unified helm chart and then have different chart and values files per service to reduce the duplication and maintenance ?

\- and does kustomize helps in any ways to solve this ?

\- how are you maintaining multiple microservices on k8s at scale ?

any insights or suggestions will be highly appreciated.

https://redd.it/1wr7mw9
@r_devops
How We Use Backstage in OpenChoreo Join Our CNCF Community Call

Hey everyone I’m part of the OpenChoreo community.

OpenChoreo is an open source Internal Developer Platform and a CNCF Sandbox project, and for this month’s community call we’re putting the spotlight on Backstage.

A big part of the session is based on our experience building OpenChoreo around Backstage and the questions we’ve seen from teams already using it.

We’ll cover:

* Why OpenChoreo chose Backstage in the first place
* The gaps we encountered when moving from a developer portal toward a complete Internal Developer Platform
* What OpenChoreo adds around Backstage
* How an existing Backstage user can migrate to OpenChoreo
* A live code walkthrough
* How to customize the OpenChoreo Backstage portal
* What we’re working on next

We want this to be useful for people already building with Backstage or trying to figure out how it fits into a broader platform engineering architecture.

If you’re working on Backstage, Kubernetes platforms, Developer Experience, or Internal Developer Platforms, we’d love to have you join and bring your questions.

Community call details + registration:
[https://github.com/openchoreo/openchoreo/discussions/4689](https://github.com/openchoreo/openchoreo/discussions/4689)

Happy to answer questions here as well.

https://redd.it/1wr9j2o
@r_devops
Resume review for devops/cloud engineer roles
https://redd.it/1wrfpsl
@r_devops
Unorganized Knowledge = Noise

Whenever I try learning new technology or troubleshooting without a clear plan, I get overwhelmed. Hopping between random videos, articles, and quick AI summaries leaves me feeling lost and stressed about where to begin or end.

I've realized that having a structured roadmap with clear, step-by-step guidance changes everything. It removes the stress and makes the whole learning process much smoother.

Anyone else feel this way? I’d love to hear how you organize your learning!

https://redd.it/1wrh025
@r_devops
We audited the default Helm charts and Docker configs of popular AI stack tools (Ray, Weaviate, MCP servers, LangChain). The out-of-the-box security defaults are surprisingly bad.

Over the past few weeks, we set up a fresh cluster to test what actually gets deployed when an engineer runs helm install with zero extra configuration on popular AI, vector DB, and MCP charts (Ray, vLLM, LiteLLM, Qdrant, Weaviate, KubeRay, etc.).

Everyone focuses on prompt injection and model jailbreaks, but the platform-level defaults on these workloads look like Kubernetes five to seven years ago:

1. Ray / KubeRay (Distributed Compute)
- The default deployment accepts unauthenticated job submissions over internal cluster networking.
- Worker containers run with passwordless sudo out of the box.
- Any compromised pod on the flat cluster network that can reach the Ray API gets arbitrary remote execution as root with zero barrier.

2. MCP (Model Context Protocol) Servers
- Several popular MCP server implementations implement local protection solely by checking Host: localhost.
- Any basic SSRF or internal proxy passing -H "Host: localhost" bypasses this immediately.
- In live testing, one Kubernetes MCP server mounted a ClusterRole with cluster-wide Secret read and pod exec permissions, exposed over unauthenticated HTTP.

3. LiteLLM Migration Job Credential Leak
- The primary deployment pulls the database password from a Secret using secretKeyRef, but the database migration Job embeds the raw database password in plain-text environment variables.
- Anyone with basic cluster read permissions (devs, CI service accounts, dashboards) can run kubectl get job -o yaml and extract the database password directly from the env block, even without permission to read Secrets.

4. Vector DBs & RAG Stores
- Multiple upstream charts still ship with default values that enable anonymous read and write access.
- Raw embeddings and ingested document payloads are queryable by any internal service without an authentication header.

5. Scanner Blindspots
- 14 of 15 charts provide no default NetworkPolicy objects.
- All 15 charts leave Kubernetes' default automountServiceAccountToken: true enabled.
- Standard static linters (Checkov, Trivy, Kubescape) missed the containers nested inside Custom Resource Definitions like Ray clusters.

The takeaway for platform and security teams:
The AI supply chain conversation is hyper-focused on prompts, but your teams might be deploying unauthenticated root execution endpoints right now. If your cluster is not enforcing Kyverno or OPA gatekeeper admission rules, mandatory NetworkPolicies, and dropping default service account tokens on AI namespaces, the internal perimeter is wide open.

Curious what other teams are doing here. Have you written custom Kyverno/OPA policies or wrapper charts for AI workloads at your org, or are you putting everything behind a service mesh like Istio/Linkerd?

Full write-up, reproduction manifests, and probe logs: https://sorami.com.au/research/ai-kubernetes-helm-chart-security/
Raw probe logs and reproduction repository: https://github.com/Sorami-Consulting-AU/ai-kubernetes-helm-chart-security

https://redd.it/1wrkllb
@r_devops
What's the most annoying "works locally, dies in production" bug you've had?

Not talking big architecture screwups — but rather the small, stupid ones. Things like hardcoded ports, missing env defaults, path separators, whatever. What was it, what stack, and how'd you find it? Would love to hear any feedback. Thanks.

https://redd.it/1ws1iip
@r_devops
Weekly Self Promotion Thread

Hey r/devops, welcome to our weekly self-promotion thread!

Feel free to use this thread to promote any projects, ideas, or any repos you're wanting to share. Please keep in mind that we ask you to stay friendly, civil, and adhere to the subreddit rules!

https://redd.it/1ws736h
@r_devops
Diagram solutions for git / markdown docs and LLM?

Has anyone here got a good solution for maintaining decent diagrams in markdown docs that are being handled with LLM assistance?

Mermaid is the obvious initial answer as they tend to self render, but they're so ugly, right? I don't think I'm missing anything there, but would be keen to know if I might.

So past them I was looking at the python diagrams library. Nice output but quite codey, no fully abstracted definition language, too much python code ingrained in it to feel like it's elegant enough to have LLM's regenerate the output by running a script that way, although I'm currently trying it out.

After that I see Schematex, which, like mermaid, has a proper definition language, and also has an MCP server and other options around it, but I've not yet tried to set up a workflow around that that would see it in action.

Is there a logic in going directly to raw svg?

Are there other options / strategies people are using to keep on top of diagrams? On caveat / excuse I have currently is that I'm mostly doing design work for proposals etc., so 100% accuracy isn't critical yet, so LLM's can have leeway for making things up at times!

https://redd.it/1ws8aqe
@r_devops
What's your default database these days, and why? Yearly survey (mod-approved)

Hey, Robert from MariaDB Foundation here. We are running a survey on how MariaDB is being used, or cases where you went with MySQL, PostgreSQL or another database.

https://mariadb.typeform.com/to/tS69UzVo?utm\_source=reddit\_devops

I split the questions by role this year, so you can skip straight to the operator ones: HA, database versions, upgrade frequency, hosting, etc. Overall I'm curious to know, what's your default database these days, and what made you land on it? Was it your choice, the app or service you run, or whoever set it up before you?

Under 10 minutes, anonymous, results published at mariadb.org/survey.

Thanks for taking the time!
Robert, MariaDB Foundation

https://redd.it/1ws9qlg
@r_devops
Experiences with Signadot

Has anyone here used Signadot in a larger project and had any success with it for ephemeral environments/preview environments? I have been looking into tooling for ephemeral environments after maintaining our own solution for a while and this one keeps popping up as an alternative, so I was hoping to get a few opinions.

https://redd.it/1wse81n
@r_devops
How I cut non-prod cluster costs by scheduling Kubernetes/OpenShift scale-down (with the approach)

Dev and test environments sit idle nights and weekends but still burn compute. Here's how I automated scaling them down and back up on OpenShift/Kubernetes.

The problem: Manual shutdowns were forgotten, and teams complained when environments weren't ready in the morning.

The approach:

Tag namespaces/workloads with a schedule label

A scheduled job scales deployments to 0 at night and restores the previous replica counts in the morning

Exclusions for anything that must stay up, plus alerting if a scale-up fails

Result: [€18572 to €15182\]

https://redd.it/1wsiuzq
@r_devops
Guyssss, so news is 150% hike with mind-blowing dilemma

Interview next week for a role that could potentially give me a 150% hike. Location ? 7 min away by my activa, mode hybrid ( 2 days office, 3 days WFH) and and and office is in a mall you know, like very fancy, tech stack (exactly matching, 3 years of experience they were asking for but selecting me with 1 year of experience) 😎 , not lalu company ,I mean it has 500 employeessss.

so so

Naturally, I checked the company before getting too excited (I was excited already)

And then I found:

Micromanagement.Burnout.Favouritism.Poor work-life balance.

And apparently a whole lot more.

Now I’m sitting here wondering:

Is a 150% hike worth the risk of a toxic work environment?, should I give my 200 percent and grab a offer for the sake of counter

The interview is happening anyway, I mean I will definitely gain confidence while preparing and it's gonna help me with other interviews.

Seniors , help me with your suggestions as it's my first switch and I'm getting unnecessarily excited 🤦

https://redd.it/1wsmzw3
@r_devops
Prioritisation/ Maturity Interview Question

Ok so I flubbed the most basic interview question there is:

"You have a number of pieces of work in flight. How do you choose between them?"

Basically, they are testing your ability to prioritise work.

I tried to answer philosphically by telling them going for a walk was a good idea but my broader point is that it shouldn't have ever gotten to the point where you have loads of pieces of work in flight at any one time.


My question is twofold: what's the accepted 'best' interview answer to this question and how do you deal with this issue in your real life actual job?

https://redd.it/1wsst19
@r_devops