Reddit DevOps
288 subscribers
82 photos
32.6K links
Reddit DevOps. #devops
Thanks @reddit2telegram and @r_channels
Download Telegram
How did u find ur first job as a DevOps engineer?

Hi, first of all sorry for my English I'm not a native speaker. I have probably a very common question but I didn't find any info here or anywhere else that gives me any light.

Context: I'm a full stack developer and I have been working for almost five years developing different kind of apps (web apps, desktop apps, libraries, APIs, installers, etc). Two years ago I had a chance in my previous job to build pipelines with gitlab and github, and since then I haven't stop learning. I even deploy a local Jenkins server and build fully functional pipelines that build, test and deploy the back and the front of an app I'm currently developing. I'm now learning a couple things of Azure DevOps.

Now I want to know how did u find DevOps offers? I'm sending at least 1 CV per day. I'm apllying as a junior, also as a semi senior and senior. I don't know why but I didn't even receive a reject mail. I mean there has to be a reason, maybe is the way I built my CV, which is in ATS format. How did u guys (and girls) find DevOps jobs.

Note: I almost forgot this but in my current job I tried to get involved in DevOps stuff, since they are starting to migrate to this practices but they always push me aside, that's why I want to move on to another job.

https://redd.it/1wodqzw
@r_devops
When did infrastructure engineering become the art of telling developers ‘no’?

This a semi-serious semi-affectionate rant from an old man. No need to pop a gasket.

These fucking AI infrastructure people are killing me.

And to be clear: this shit did not start with AI. We have been heading down this road for at least ten years. Cloud infrastructure, enterprise security, identity, zero trust, compliance, service meshes, gateways, tokens, policies, orchestration layers — every year another layer gets inserted between the engineer and the fucking computer.

AI deployment has just taken the existing insanity and turned the knob to eleven.

Compare this shit to developing for the Internet in the 90s. You wanted to build something? Great. Here’s TCP/IP. Here’s a socket. Here’s HTTP. Go fucking build it. Every year you got more capability, more bandwidth, better machines, better tools. The infrastructure was there to let you do things.

Now I want to move twelve fucking bytes from one machine to another and suddenly I need an identity provider, three OAuth flows, a gateway, a service account, a token exchange, a policy exception, a certificate rotation strategy, some fucking YAML, and a meeting with six people whose entire job appears to be telling me why the computer is not allowed to talk to the other computer.

And AI makes every part of this worse.

Now there are model gateways. Model access policies. Prompt policies. Data policies. Regional restrictions. Safety filters. Rate limits. Per-model permissions. Per-project permissions. Per-user permissions. API keys that aren't really API keys because now everything needs an identity token generated from some other identity which itself needs permission to request the token that gives you permission to make the fucking request.

We already spent ten years building an infrastructure culture where the default answer to everything is NO.

AI has arrived and instead of saying, holy shit, here is this incredible new capability, let's make it easy for engineers to use, the first instinct is apparently to surround it with seventeen concentric rings of authorization machinery.

And then the best part: IT’S ALWAYS BROKEN.

The token expired. The gateway ate the request. The policy changed. Your role disappeared. The certificate rotated. The secret is in the wrong secret store. Your service principal can authenticate but isn’t authorized, except it was authorized yesterday, and nobody knows why because the authorization policy is generated by another policy which lives behind another fucking gateway.

We have constructed this gigantic paranoid security cathedral where every component assumes every other component is actively trying to murder it.

And yes, I understand why security exists. I’m not proposing we FTP the fucking model weights onto an anonymous server and put the password in README.txt.

But at some point you have to notice that the security apparatus is consuming the actual product.

That is the part that drives me fucking crazy.

The 90s Internet engineering mentality was: here are some new powers. Try not to fuck it up.

The last ten years have been: you probably shouldn't be allowed to do that.

And the AI deployment revolution has somehow become: you absolutely cannot do that until seventeen unrelated systems agree that you are the correct human, running the correct service, from the correct machine, in the correct region, for the correct model, for the correct purpose, using a token that expires in nine fucking minutes.

Then everybody stands around wondering why nothing ships.

https://redd.it/1woo526
@r_devops
KubeCraft or TechWorld w Nana - Thoughts?

Newbie, learning from the bottom up.
KubeCraft 6000
TWN 1800

Pros and Cons would be helpful. Thanks!

https://redd.it/1woq05c
@r_devops
What is your approach to identifying areas for improvement to make impact at your company?

What’s steps, processes, or simply just things do you do to identify areas for improvement at your company to make an impact?

Do you learn about trending tools in the industry and look for ways to use them in your org or do you find a use cases and look for tools that can help improve a product or process?

I feel like this is the biggest thing I am lacking in my career that is keeping me back from moving up and can use some advice.

https://redd.it/1wp28x3
@r_devops
Any of you use alternatives to GitLab when self-hosting is a requirement? (and general toolchain question)

Considering jumping to Gitlab because we're looking into ditching Atlassian ASAP. We need to stay on-prem, but tbh it's pretty difficult for me to understand what different products and platforms don't support.

I'd love to stay open source but there's also a lot of ease-of-use to be taken into consideration. I looked into things like Forgejo, but the actions seem quite lackluster if you want to integrate test reports, publish findings from opengrep in a PR - or am I just missing something obvious?

What does your on-prem solution look like in terms of SCA, SAST etc?

https://redd.it/1wp8eyh
@r_devops
OC gha-oidc-claimsim (RC): offline CLI — GHA OIDC claims vs IAM trust JSON

Built a small Apache-2.0 CLI for platform/security/CI folks wiring GitHub Actions → AWS OIDC.

It takes workflow event context JSON + IAM trust policy JSON, predicts default OIDC sub/aud, and prints ALLOW or DENY with reasons — no STS, no AWS credentials, no network.

Main wedge: branch-pinned trust (repo:ORG/REPO:ref:refs/heads/main) often DENYs pullrequest jobs whose default sub is repo:ORG/REPO:pullrequest.

Honesty:
- Release candidate 0.1.0-rc.2 — not stable 0.1.0
- Not on PyPI — install from git / GitHub Release
- Simplified IAM Condition model (not bit-identical to AWS)
- Optional hcl-dup is convenience only; prefer tflint for HCL duplicate keys (not unique IP)
- Not a Checkov missing-sub linter and not a live-account scanner

Repo: https://github.com/mapleleaflatte03/gha-oidc-claimsim
Release: https://github.com/mapleleaflatte03/gha-oidc-claimsim/releases/tag/v0.1.0-rc.2

Happy to take questions on the claim grammar and limits.

https://redd.it/1wp7aex
@r_devops
Question about KernelModuleConfig

Hi all 👋

I have a question. I am installing Thalos. I have not applied my configuration yet. If I apply this file

apiVersion: v1alpha1
kind: KernelModuleConfig
name: cdc_ncm


and then apply my controlplane.yaml, and then reboot, will the cdc_ncm module be loaded?

https://redd.it/1wpdazb
@r_devops
Beste Embedded-Messe mit DevOps-Fokus im DACH-Raum?

Ich bin DevOps Engineer und betreue vorrangig ein Embedded-Entwicklerteam. Ich suche eine Messe oder Konferenz, bevorzugt im DACH-Raum, gerne aber auch anderswo in Europa, die sich klar auf Embedded-Software konzentriert und DevOps-Themen wie CI/CD für Firmware, HIL-Tests und OTA-Updates fundiert behandelt. Neben der technischen Praxis interessieren mich auch strategische Fragen, etwa wie man DevOps im Embedded-Umfeld langfristig aufstellt und skaliert.

Auf dem Schirm habe ich embedded world und ESE Kongress. Welche Veranstaltung hat euch wirklich weitergebracht? Gibt es Geheimtipps?

Danke!

https://redd.it/1wpbkcl
@r_devops
A release claim that does not match git

A release is a claim. A fix version, a change ticket, or release notes you typed: these are the changes you say are in this ship. A Jira fix version is the same claim.

Git does not know about that claim. Generating release notes from the tag always matches, so it cannot catch a claim that is already wrong. A commit can cite a ticket the version never listed. A ticket on the version can have no commit. If the previous tag is not an ancestor of this one, the range itself is not proven.

Shipledger checks the claim you wrote against a local git range. It does not fetch Jira, it does not generate the notes, and it does not deploy.

The public demo uses GitHub releases only because those notes are visible:

v0.2.0 lists #1 and #2. Both are in the tag. The check exits 0.

https://github.com/kacxx/shipledger-demo/releases/tag/v0.2.0

v0.3.0 lists only #3. The tag also contains #4. The check exits 1.

https://github.com/kacxx/shipledger-demo/releases/tag/v0.3.0

README has the commands. Not on npm yet.

https://github.com/kacxx/shipledger

https://github.com/kacxx/shipledger-demo

https://redd.it/1wpjbxh
@r_devops
How do you check a release list against what is actually in the tag?

Between "we're ready to cut" and "it's deployed," most teams I see do some mix of:

\- pull the ticket list from Jira, Linear, or a GitHub project

\- compare it to the tag

\- write or generate the notes

\- fill in a change request

\- have someone look at the diff and say it looks right

Generating the notes from the tag always matches, because the tag is the input. That doesn't check a list a person already maintains.

Things I've seen go wrong:

\- a hotfix is cherry-picked onto the release branch and never added to the tracker

\- a PR merges just before the tag and the notes are already written

\- a ticket is Done, and the PR was reverted

\- two repos ship together and the notes only cover one

That check is still "a senior engineer looks at it," or we find out afterward.

If you ship from a ticket list rather than from generated notes, how do you catch the ones that don't match git? Where is that still a person, and where have you made it mechanical?

https://redd.it/1wpmoyl
@r_devops
How do you learn DevOps

I’m a software developer, and in my own projects I already handle some basic operational work like Docker, VPS deployment, Nginx, DNS, Cloudflare, Git, etc. I also had some little experience in network administration.

What I’m trying to understand is how people approach learning DevOps and system architecture properly.
One thing that has always confused me is entry-level DevOps roles. A lot of DevOps work seems to exist specifically to support the software development lifecycle, things like CI/CD, GitHub Actions, deployments, environments, infrastructure, monitoring… I find it hard to imagine understanding these things deeply without first experiencing the development workflow itself.

At the same time, I don’t necessarily want to switch away from software development. My goal is to become the kind of engineer who understands the whole system: application development, system design/architecture, networking, cloud infrastructure, deployment, observability, scalability, and production operations.

Basically, I want to be able to build a system and also understand how to design, deploy, scale, and operate it properly at a much larger level.

For people who have gone down this path, how would you approach learning these areas while continuing to work primarily as a software developer, as of course I dont have the knowledge for a switch? And in your opinion am I on the right path to stay employable in this uncertain future?

https://redd.it/1wpr4q3
@r_devops
Mobile App DevOps

What do people normally use for CI/CD of mobile apps (both IOS and Android)?

https://redd.it/1wps92p
@r_devops
Stuck in a slow platform support role with a DevOps title. How do I make sure I don't get pigeonholed?

Hey everyone,

Just want to get some advise from you all. I’m early in my career working in an enterprise DevOps role. but my role is mostly focus on company's platform but I do not write the platform only solve the issue kinda like support for platform. No code writing and my team is slow and sometimes there's not much things to do.

Usually, what I do is learning outside tech by doing small project outside work hour and during work hour if there's nothing to do then I try to understand the flow of the platform written by others like Ansible and etc.. and this is my current strategy.

My goal is to become senior or mid level in 3 years but with my current team tech and work I don't think I can become one in the future and I will be stuck in one team becuase I don't know other stuffs and I'm seeing it currently in my team. I am trying to move to other team after a year at my current team because I wanna learn more and to meet my goals to be marketable.

(I'm also very grateful for my current job given by current team which give me my first year experience but....)

Can you all give me some adivise that

>Appreciate any feedback or reality checks.

#

Can you please give me some advise on what I'm doing is fine or please let me know any advise?

https://redd.it/1wpt1t3
@r_devops
How do I test if my website can handle a traffic spike?

My company is getting ready to launch a huge marketing campaign and I’m not sure if

our website is ready for the increased traffic once it goes live. What would a realistic

pre-launch test look like? Should we have it simulate traffic levels we’ve seen before or

should we push it beyond our typical peak? What’s the best way for us to even test it?

My company is pretty big and we expect a lot from this campaign so I want to make

sure everything runs smoothly come launch day.

https://redd.it/1wpufyt
@r_devops
How do you know which engineering issues need attention before they become expensive?

I'm doing a bit research on how DevOps/SRE teams handle this.

Let's suppose you have

* a Datadog alert
* a Sentry issue
* a Jira/Linear ticket

It gets noticed, but doesn’t get resolved for hours or days. Eventually it affects customers, causes downtime, delays something important, or starts costing the business money.

# How do you know which issues actually need immediate attention before that happens?

Do you already have something that connects technical issues to their business/customer impact, or is prioritization mostly based on severity, alerts, and engineering judgment?

Trying to understand whether this is a real problem before building anything around it.

Would especially love to hear from DevOps, SRE, platform, or engineering ops people who deal with this regularly.

https://redd.it/1wq6da2
@r_devops
How do you know which engineering issues need attention before they become expensive?

Curious how DevOps/SRE teams handle this.
Say you have:
\- a Datadog alert
\- a Sentry issue
\- a Jira/Linear ticket

It gets noticed, but doesn’t get resolved for hours or days. Eventually it affects customers, causes downtime, delays something important, or starts costing the business money.

So…
1. How do you know which issues actually need immediate attention before that happens?
2. Do you already have something that connects technical issues to their business/customer impact, or is prioritization mostly based on severity, alerts, and engineering judgment?

Trying to understand whether this is a real problem before building anything around it.
Would especially love to hear from DevOps, SRE, platform, or engineering ops people who deal with this regularly.

https://redd.it/1wq5qd2
@r_devops
Kubernetes scaling is much more interesting when you look at what happens internally.

Two things I found very interesting:

1. Only the API Server writes to etcd

Controllers don’t directly modify etcd.
If HPA decides that the application needs 5 replicas instead of 3, it doesn’t go directly to etcd and change the value.
It talks to the API Server.
The API Server handles the request, applies authentication, authorization and admission controls, and then persists the desired state in etcd.
This gives Kubernetes a single controlled entry point for modifying cluster state.

It also means the other components don’t need to understand how etcd works or deal with its consistency and access directly.
Single responsibility principle acting beautifully

2. All other components use watch instead of constantly polling

This is another really nice design decision.
A controller doesn’t need to keep asking:
“Did something change?”
“Did something change?”
“Did something change?”

Components don't call each other. The HPA doesn't call the scheduler. The scheduler doesn't call the kubelet. Each one opens a long lived watch on the API server for the one kind of object it cares about, and reacts when that object changes.

Here's the whole scaling chain:

• HPA controller reads metrics and updates replicas on the Deployment. That's all it does.
• Deployment controller watches Deployments, sees the new count, and updates the ReplicaSet.
• ReplicaSet controller watches ReplicaSets, sees it has 3 pods but wants 5, and creates 2 Pod objects. They have no node yet.
• Scheduler watches for Pods with no node, picks the best node, and writes a binding.
• Kubelet on that node watches for Pods assigned to it, starts the containers, and reports status back.

Every component does exactly one job, writes its result to the API server, and walks away. The next component picks it up through its own watch. Nobody knows who comes next, and nobody needs to.

That's why the system is so resilient. If the scheduler restarts, it relists, sees the pending pods and carries on. Controllers compare desired state with actual state, so a missed event doesn't break anything. They just reconcile again.

The part that surprised me most

After all that machinery, the only real thing that changed is the number of Pods. The Deployment and ReplicaSet are just records in etcd with a different number. The HPA, controllers and scheduler never run your app. At the end of the chain, it's only pods that get scaled.

If you want to see all of this graphically I have explained this in detail below

https://youtu.be/bwiHEG2NLpE?si=hPEEnlOR17qXydAQ

https://redd.it/1wq952x
@r_devops
How do DevOps team notice issues before they become a big problem?

Last week I’ve spent it researching about this specific problem before I go deeper and build something for this so PLEASE reply if you suffer from it.

So issues talking about Jira or Linear tickets, Sentry issues, Datadog alerts or whatever that sit for a LOTTT of time before getting resolved which annoys both the customer and the dev after he notices, so the thing is this means “Revenue Leakage”.

So, I might be wrong but DevOps teams use dashboards like any team in the world, but the thing with dashboards is they NEVER EVER show you that a Jira/Linear ticket is indeed of resolving it, or a fix the the issue that caused a Datadog alert or fix the bug(Sentry) or whatever I’m not so deeply involved with dev operations.

What the dashboards do is they don’t immediately tell you that “HEY BRO A JIRA TICKET IS SITTING AND REVENUE IS GONNA DROP SO HURRY”, they just show analytics after a month of the revenue being down 7%.

So I’m wondering:
1. How quickly does your team spot these types of issues?
2. How much revenue what you say has been leaked from your org (don’t necessarily answer that but would help TON)
3. How much of these types of issues are you facing per month?(Jira tickets, Datadog alerts, etc.)

I would be MORE than happy to get a reply from an experienced professional and tell me if this is a problem you want resolved in your org now.

https://redd.it/1wq53pm
@r_devops
Why doesnt DevOps have an ACLS?

Edit: Tl;Dr at bottom. I write too much.

Bit of an odd thought I had after reading another thread here about gaining confidence during high-severity incidents. Thank you to that poster, BTW, for inspiring my brain.

I have a weird fascination with ER and critical-care videos. Not really because of the medicine itself, but because I’m a massive compulsive systemizer and emergency medicine is basically human systems engineering under some of the worst possible conditions. Someone is actively dying, information is incomplete, the situation is changing by the second, and multiple people have to coordinate without getting in each other’s way. Oh, and if you don't solve the issue soon, the patient will be dead in 5 minutes, and you'll have to tell the family.

What fascinates me the most though is that, when a resuscitation is being run properly, it usually doesn’t look like ten people frantically improvising. There’s command, defined roles, structured communication, task ownership, checklists, algorithms, simulation, debriefing, and a whole lot of deliberate training behind the scenes.

This is basically crack cocaine for my dopamine-addicted brain.

Obviously software going down is not equivalent to someone dying. But from a systems perspective, we’re solving a strangely similar problem: how do you get a bunch of fallible humans to reason and coordinate effectively while something important is actively failing?

This is also something I’ve had to think about in practice.

My last employer had essentially no incident-management practice in place when I joined. Outages were fairly common, roughly once a month, and it wasn’t unusual for them to last several hours or even a whole day.

So I volunteered to build one.

I wrote the process, defined how incidents would be declared and coordinated, laid out responsibilities, escalation, communications, and what we were supposed to be doing while production was on fire. I had it submitted to change committee, got board sign-off, and felt confident in the process.

Two months later, we got to stress-test it for real.

And it failed horribly, because I was the only person that saw the problem and opted for preparation rather than eternal deferral of The Boring Stuff (tm). I even got written up for telling a nontech manager who kept DM'ing techs to "try this fix they found on Google" to please follow the board-approved procedure and message their designated PoC rather than interrupting responders who were already sweating buckets of bullets while elbow-deep in code hell**.** But that's a story for another day. Sigh.

There’s something very different about designing an incident process on paper and then watching actual humans try to execute it while systems are failing, information is incomplete, everyone wants an answer immediately and fragile egos are at stake. Frankly though, that experience probably has a lot to do with why this subject fascinates me.

And it’s also why I think our industry is kind of shit at training people for this.

A lot of on-call training basically amounts to, “Here’s Kubernetes, Grafana, AWS, some architecture diagrams, and a wiki page Dave wrote eighteen months ago. You’re on call next week. Good luck and don't fuck it up, you're still on probation.”

Then eventually there’s a massive Sev1, thirty random people you've never heard of pile into a call, Slack explodes, somebody begins restarting pods and unwittingly knocks over three adjacent services, three people start making production changes at the same time, a C-level joins the call and says its costing them 100 billionty dollars every 7 seconds, somebody adds 64 replicas to RDS and forgets about it, half the engineering team find out their access doesn't work... and some poor sod who was voluntold that they'd be IC during an on-call incident channels their freaking deer-in-headlights-energy and freezes.

And then we conclude they’re “not good under pressure.”

Well... did we ever actually train them to be?

That got me
thinking about what the equivalent of ACLS for production engineering would actually look like.

I don’t mean another multiple-choice SRE certification, a webinar, or a tabletop where everyone knows the database is going to “fail” at 2 PM. I mean actual practical training where the environment itself is part of the course.

Imagine you show up Monday morning with seven other engineers. You’re given a cheap company laptop and told you now work for a fictional company. Behind a VPN there’s a real training environment with APIs, databases, Redis, queues, monitoring, deployments, source control, chat, ticketing, and all the usual operational cruft. Even "real" loads, facilitated by magical load testing fairies.

And the fictional company is deliberately badly run.

The docs were written by some guy named "George," and last updated 3 years ago... and are deliberately incorrect. Some credentials don’t work. Only a third of the class discovers their SSH access actually works. A dev DB server is named "prod," and a prod DB server is named "dev." Nobody is quite sure who is supposed to declare an incident. The escalation documentation is so out of date it may as well've been written on stone tablets. Maybe one person has production database access for reasons nobody can explain.

Then, at some completely unannounced point Monday morning, the pager goes off.

Checkout is fucked.

Go.

The first incident would probably be an absolute dumpster fire, and that would be intentional.

Three people make conflicting changes. Nobody establishes command. Half the team can’t access the systems they need. Nobody knows exactly what customers are experiencing. Somebody spends twenty minutes debugging something that could have been rolled back in two. Somebody follows the documented steps to restore service, and end up taking out marketing. Somebody didn't thought to set DEBUG_MODE to true when testing something, and now 320,000 "customers" receive emails about their accounts being terminated for nonpayment.

Then you stop and debrief.

Why did that happen? What made the response harder than it needed to be? What did we assume somebody else was doing? What information did we not have? What process simply didn’t exist?

Now the lesson actually means something because the students just felt the consequences of not having it.

You teach the relevant concepts, then a few hours later, while everyone is doing something completely unrelated, the pager goes off again.

Two surprise incidents every day for five days.

Everybody rotates through Incident Commander, technical lead, communications, scribe, responder, and the other supporting roles. The outspoken engineer doesn’t get to be IC all week. The brilliant debugger has to run communications. The quiet mousey person eventually has to take command.

But then... something amazing happens.

As the week progresses, the fictional company improves because the students improve it.

They fix the access process because the access process screwed them. They improve the runbooks because the runbooks screwed them. They establish incident roles because Monday was chaos. They improve monitoring because a dashboard sent them down the wrong rabbit hole. They throw the inaccurate documentation in the trash, and silently curse whoever "George" was.

Later incidents get nastier.

Maybe the telemetry is misleading. Maybe a senior engineer confidently insists on the wrong diagnosis (and they were set them up for this failure by deliberately leaking convincing logs to them). Maybe an executive barges in and starts asking responders for ETAs directly. Maybe a downstream dependency causes cascading failures that make five unrelated systems look broken. Maybe the team restores service, brings everything back too quickly, and knocks production over again.

The point isn’t just to teach technical troubleshooting.

It’s to teach people how to operate as a team while they’re uncertain, overloaded, and under pressure.

By Friday, ideally, the pager goes off and the exact same group that looked