cannot silently enter a new high-cost category, fan out across regions, and generate material spend before humans can react?**
My current takeaway is that mature DevOps / platform governance for these categories probably needs a layered model:
* Organizations/SCP default-deny for high-cost or unused service categories;
* explicit approval workflow for Bedrock, Marketplace, and similar services;
* IAM permission boundaries for human, workload, and CI/CD identities;
* no long-lived credentials where avoidable;
* SSO / IAM Identity Center for humans;
* short-lived role assumption for automation;
* region restrictions by default;
* service quotas reduced where meaningful;
* first-time service/category usage detection;
* emergency containment automation for accounts with narrow expected usage;
* forensic readiness through CloudTrail and billing correlation.
I’d appreciate feedback from people who have implemented similar controls in real cloud environments.
Which controls actually reduce financial blast radius in practice, which ones look good only on paper, and what would you prioritize first to optimize for time-to-stop, not just time-to-detect?
https://redd.it/1uptjrx
@r_devops
My current takeaway is that mature DevOps / platform governance for these categories probably needs a layered model:
* Organizations/SCP default-deny for high-cost or unused service categories;
* explicit approval workflow for Bedrock, Marketplace, and similar services;
* IAM permission boundaries for human, workload, and CI/CD identities;
* no long-lived credentials where avoidable;
* SSO / IAM Identity Center for humans;
* short-lived role assumption for automation;
* region restrictions by default;
* service quotas reduced where meaningful;
* first-time service/category usage detection;
* emergency containment automation for accounts with narrow expected usage;
* forensic readiness through CloudTrail and billing correlation.
I’d appreciate feedback from people who have implemented similar controls in real cloud environments.
Which controls actually reduce financial blast radius in practice, which ones look good only on paper, and what would you prioritize first to optimize for time-to-stop, not just time-to-detect?
https://redd.it/1uptjrx
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Follow-up: How does your team actually move important Slack discussions into documentation?
https://www.reddit.com/r/devops/comments/1unvlkj/engineering_managers_how_do_you_prevent_valuable/
https://redd.it/1upqr3q
@r_devops
https://www.reddit.com/r/devops/comments/1unvlkj/engineering_managers_how_do_you_prevent_valuable/
https://redd.it/1upqr3q
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Looking at tooling and curious how your team actually using AI in ITops?
Curious how teams are actually using AI in day-to-day ITops, not the vendor pitch version, but what's genuinely saving time , improves results vs. what's been overhyped.
https://redd.it/1upvgyr
@r_devops
Curious how teams are actually using AI in day-to-day ITops, not the vendor pitch version, but what's genuinely saving time , improves results vs. what's been overhyped.
https://redd.it/1upvgyr
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Security you can't justify is a vicious cycle
https://bogomolov.work/blog/posts/security-you-cant-justify-is-a-vicious-cycle/
https://redd.it/1upxytm
@r_devops
https://bogomolov.work/blog/posts/security-you-cant-justify-is-a-vicious-cycle/
https://redd.it/1upxytm
@r_devops
bogomolov.work
Security you can't justify is a vicious cycle
Dependabot is on fire, the model yells, the meeting says add TLS. Fine. Is this stopping an attacker, or feeding a checklist?
Security you can't justify is a vicious cycle
https://bogomolov.work/blog/posts/security-you-cant-justify-is-a-vicious-cycle/
https://redd.it/1upzq3o
@r_devops
https://bogomolov.work/blog/posts/security-you-cant-justify-is-a-vicious-cycle/
https://redd.it/1upzq3o
@r_devops
bogomolov.work
Security you can't justify is a vicious cycle
Dependabot is on fire, the model yells, the meeting says add TLS. Fine. Is this stopping an attacker, or feeding a checklist?
DevOps is one of the most vaguely defined roles in tech, and I think that's exactly the point
​
DevOps spans a huge range: CI/CD, security, SRE, DevEx, each with its own sub functions. At big companies, engineers often specialize in one. At startups, you juggle all of them at once, usually by necessity.
Here's the distinction I keep coming back to: DevOps isn't about mastering one function in isolation. It's about holding all of them in your head at once, even when you're only actively working on one. Building CI/CD? You're also thinking observability, security, and DevEx, because a pipeline that ships fast but can't be debugged, or isn't secure, isn't actually done.
Curious how this plays out elsewhere. Does your team specialize by function, or does everyone think across all of them regardless of company size?
https://redd.it/1uq3g3d
@r_devops
​
DevOps spans a huge range: CI/CD, security, SRE, DevEx, each with its own sub functions. At big companies, engineers often specialize in one. At startups, you juggle all of them at once, usually by necessity.
Here's the distinction I keep coming back to: DevOps isn't about mastering one function in isolation. It's about holding all of them in your head at once, even when you're only actively working on one. Building CI/CD? You're also thinking observability, security, and DevEx, because a pipeline that ships fast but can't be debugged, or isn't secure, isn't actually done.
Curious how this plays out elsewhere. Does your team specialize by function, or does everyone think across all of them regardless of company size?
https://redd.it/1uq3g3d
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Anyone else spend half their week just chasing down which tag blew up your cardinality?
we had a service push a UUID out as a label last month and our metrics query latency went sideways for two days before anyone traced it back to the actual deploy. nobody noticed until dashboards started timing out.
now i'm paranoid about every new label that goes into prod, and "just use fewer labels" isn't really an answer when you don't know which label is the problem until it's already in your bill.
please share how does your team actually catch this before it happens?
https://redd.it/1uqfflz
@r_devops
we had a service push a UUID out as a label last month and our metrics query latency went sideways for two days before anyone traced it back to the actual deploy. nobody noticed until dashboards started timing out.
now i'm paranoid about every new label that goes into prod, and "just use fewer labels" isn't really an answer when you don't know which label is the problem until it's already in your bill.
please share how does your team actually catch this before it happens?
https://redd.it/1uqfflz
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Need an Azure learning path for an internal transfer to DevOps (2-3 month timeline)
https://redd.it/1uqjb9j
@r_devops
https://redd.it/1uqjb9j
@r_devops
Optimising workspace for multiple laptops?
What setup have you put in place on your desk to manage being on multiple laptops?
https://redd.it/1uqm1f2
@r_devops
What setup have you put in place on your desk to manage being on multiple laptops?
https://redd.it/1uqm1f2
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Why are intermittent production bugs so hard to reproduce?
I had one of those bugs last month that made the whole team question reality for about four days. A checkout service would randomly throw 500s, maybe 3-4 times a day, always on the same endpoint, never on a predictable schedule sometimes it happened during peak traffic, sometimes at 3am with almost nobody hitting the system. Logs showed the error but the stack trace pointed to a null reference in a place that should have been impossible given our validation logic. We tried the obvious stuff first: replaying the same request payload from the failed logs, same user, same cart contents, down to the byte worked fine every single time locally and in staging couldn't reproduce it on demand no matter how hard we tried. Turned out three things had to line up at once to trigger it, a specific cache eviction timing on one node, a race condition between two async calls that only mattered under a certain concurrency level and a stale feature flag value that got cached slightly differently depending on which pod served the request none of that shows up in a single request replay because the bug wasn't really about the request, it was about the state of the system at that exact moment. This is why intermittent bugs are so brutal to chase down the failure depends on system state not just input, things like cache state, connection pool exhaustion, background job timing or which pod handled the request all shift constantly and aren't captured by a normal log line. Concurrency and timing windows are nearly impossible to force, since race conditions might only trigger under load patterns that don't exist in staging or in a single manual test and by the time you are looking at it, the state that caused it is already gone. logs tell you what happened not the exact conditions leading up to it, so you are reconstructing a crime scene from a few photos instead of watching it happen. We eventually solved it by adding much more granular tracing around the cache and flag evaluation logic then waiting for it to happen again with all that extra visibility on took another two days after that just to catch it in the act. How do other people approach this class of bug, better tracing and waiting to catch it live or actual techniques for forcing race conditions to surface faster before they become customer-facing?
https://redd.it/1uqm2hr
@r_devops
I had one of those bugs last month that made the whole team question reality for about four days. A checkout service would randomly throw 500s, maybe 3-4 times a day, always on the same endpoint, never on a predictable schedule sometimes it happened during peak traffic, sometimes at 3am with almost nobody hitting the system. Logs showed the error but the stack trace pointed to a null reference in a place that should have been impossible given our validation logic. We tried the obvious stuff first: replaying the same request payload from the failed logs, same user, same cart contents, down to the byte worked fine every single time locally and in staging couldn't reproduce it on demand no matter how hard we tried. Turned out three things had to line up at once to trigger it, a specific cache eviction timing on one node, a race condition between two async calls that only mattered under a certain concurrency level and a stale feature flag value that got cached slightly differently depending on which pod served the request none of that shows up in a single request replay because the bug wasn't really about the request, it was about the state of the system at that exact moment. This is why intermittent bugs are so brutal to chase down the failure depends on system state not just input, things like cache state, connection pool exhaustion, background job timing or which pod handled the request all shift constantly and aren't captured by a normal log line. Concurrency and timing windows are nearly impossible to force, since race conditions might only trigger under load patterns that don't exist in staging or in a single manual test and by the time you are looking at it, the state that caused it is already gone. logs tell you what happened not the exact conditions leading up to it, so you are reconstructing a crime scene from a few photos instead of watching it happen. We eventually solved it by adding much more granular tracing around the cache and flag evaluation logic then waiting for it to happen again with all that extra visibility on took another two days after that just to catch it in the act. How do other people approach this class of bug, better tracing and waiting to catch it live or actual techniques for forcing race conditions to surface faster before they become customer-facing?
https://redd.it/1uqm2hr
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
What are you guys using for KYC API for identity verification and onboarding?
Getting straight to the problem: our current KYC API is hitting its limits. Slower verification times, too many false rejections on international documents and support goes quiet the moment something breaks in production.Every vendor sounds identical on paper so we fail to understand what does AI powered verification actually mean when an edge case hits at 2am(lol) and thousands of verifications are queued up? Looking for something API first with good docs, reliable webhooks, global document support, and AML screening that is baked in. Anyone who has run more than one of these in production and can give an honest take would be really helpful.
https://redd.it/1uqookv
@r_devops
Getting straight to the problem: our current KYC API is hitting its limits. Slower verification times, too many false rejections on international documents and support goes quiet the moment something breaks in production.Every vendor sounds identical on paper so we fail to understand what does AI powered verification actually mean when an edge case hits at 2am(lol) and thousands of verifications are queued up? Looking for something API first with good docs, reliable webhooks, global document support, and AML screening that is baked in. Anyone who has run more than one of these in production and can give an honest take would be really helpful.
https://redd.it/1uqookv
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Do people actually use their GPUs as much as they expected?
When I first started looking at upgrading my setup, I was convinced I'd be using a GPU every day.
Reality turned out to be pretty different.
Some weeks I barely run anything. Then I'll have two or three days where I'm constantly launching jobs.
Now I'm wondering if most people actually use their hardware that much, or if it mostly sits there waiting for those busy days.
https://redd.it/1uqpme2
@r_devops
When I first started looking at upgrading my setup, I was convinced I'd be using a GPU every day.
Reality turned out to be pretty different.
Some weeks I barely run anything. Then I'll have two or three days where I'm constantly launching jobs.
Now I'm wondering if most people actually use their hardware that much, or if it mostly sits there waiting for those busy days.
https://redd.it/1uqpme2
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Could you please share a typical daily routine for a DevOps professional and outline the types of tools used for application hosting?
The objective here is to investigate the DevOps tools currently employed by various companies within the market, thereby enabling adequate preparation for a potential career transition.
Also share the year of experience do you have?
https://redd.it/1uqqjho
@r_devops
The objective here is to investigate the DevOps tools currently employed by various companies within the market, thereby enabling adequate preparation for a potential career transition.
Also share the year of experience do you have?
https://redd.it/1uqqjho
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
How do you use agents and LLMs in your work
I was wondering how you guys use LLMs and agents in your day to day work, do you use them to write yaml/boilerplate or beyond that?
https://redd.it/1uqsg66
@r_devops
I was wondering how you guys use LLMs and agents in your day to day work, do you use them to write yaml/boilerplate or beyond that?
https://redd.it/1uqsg66
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Docs as Code implementation for Infrastructure
Hi there
Recently I was tasked to write documentation for our infrastructure in "doc as code" way but I have not very well grasped what it is
The only requirement my team leads has is that the documents should be enough for any new person to understand our infra setup and tools we are using.
They also mentioned that any changes in the documents should have a PR and only after reviewing and approving any changes should be visible.
What I understand till now is that we would have a central repository in confluence or version control with documentation files.
There should be a way to navigate to different documents
All .md files are similar in structure, how they are written
Architecture diagrams to show infrastructure
I had a look at kubernetes documentation as I get what it is everything is in markdown it is being rendered to the website and has different documents for different versions.
But I still have no idea how to start on this.
Can I know what are some common points to note down or industry standard for these kind of documentation. And how to implement it
https://redd.it/1uqwjuf
@r_devops
Hi there
Recently I was tasked to write documentation for our infrastructure in "doc as code" way but I have not very well grasped what it is
The only requirement my team leads has is that the documents should be enough for any new person to understand our infra setup and tools we are using.
They also mentioned that any changes in the documents should have a PR and only after reviewing and approving any changes should be visible.
What I understand till now is that we would have a central repository in confluence or version control with documentation files.
There should be a way to navigate to different documents
All .md files are similar in structure, how they are written
Architecture diagrams to show infrastructure
I had a look at kubernetes documentation as I get what it is everything is in markdown it is being rendered to the website and has different documents for different versions.
But I still have no idea how to start on this.
Can I know what are some common points to note down or industry standard for these kind of documentation. And how to implement it
https://redd.it/1uqwjuf
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Looking for advice on transitioning from Sysadmin to DevOps
I'm looking to apply to devops/sre positions to change my current job and have a profesional glow up, i no longer feel challenged project from my job and i am stating to just doing app maintaenance and helpdesk tasks.
I am a sysadmin with hands on production enviroments and automatation background (scripting and low code) but where i learn and enjoy the most is in my homelab, I have around 2.5 years of professional experience.
I'd like to learn technologies such as Terraform (mainly because I see it requested in many LinkedIn job postings) and Ansible, as well as deepen my knowledge of CI/CD pipelines. I've already worked with GitHub Actions.
I've also used AI to help me create a learning roadmap and prioritize milestones. One of the strongest recommendations was to document everything on GitHub.
Beyond following a roadmap, I'd like to hear what you think are the most important things to focus on when transitioning into DevOps. If you've seen what helps people land their first DevOps role, or if you have any advice on common mistakes, skills to prioritize, or portfolio ideas, I'd really appreciate your perspective.
https://redd.it/1ur3a7j
@r_devops
I'm looking to apply to devops/sre positions to change my current job and have a profesional glow up, i no longer feel challenged project from my job and i am stating to just doing app maintaenance and helpdesk tasks.
I am a sysadmin with hands on production enviroments and automatation background (scripting and low code) but where i learn and enjoy the most is in my homelab, I have around 2.5 years of professional experience.
I'd like to learn technologies such as Terraform (mainly because I see it requested in many LinkedIn job postings) and Ansible, as well as deepen my knowledge of CI/CD pipelines. I've already worked with GitHub Actions.
I've also used AI to help me create a learning roadmap and prioritize milestones. One of the strongest recommendations was to document everything on GitHub.
Beyond following a roadmap, I'd like to hear what you think are the most important things to focus on when transitioning into DevOps. If you've seen what helps people land their first DevOps role, or if you have any advice on common mistakes, skills to prioritize, or portfolio ideas, I'd really appreciate your perspective.
https://redd.it/1ur3a7j
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Additional burden of hosing AI apps
With AI, business and product teams are creating apps left and right. They dont understand what the code is doing, no clue about security or how to host it.
This burden falls on DevOps/Engineering to now maintain it, fix it. Authors are still considered the owners of these apps. I wanted to know how are you guys handling this situation?
\- Is Engineering/DevOps the defacto owners of such apps in your company?
\- How are you deploying these - in your prod AWS or some hosted env?
TIA
https://redd.it/1ur4nwc
@r_devops
With AI, business and product teams are creating apps left and right. They dont understand what the code is doing, no clue about security or how to host it.
This burden falls on DevOps/Engineering to now maintain it, fix it. Authors are still considered the owners of these apps. I wanted to know how are you guys handling this situation?
\- Is Engineering/DevOps the defacto owners of such apps in your company?
\- How are you deploying these - in your prod AWS or some hosted env?
TIA
https://redd.it/1ur4nwc
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
How to use my old laptop
Anyone help me what I can do useful with my old laptop related to server or something and I don't have router.
https://redd.it/1urbjdi
@r_devops
Anyone help me what I can do useful with my old laptop related to server or something and I don't have router.
https://redd.it/1urbjdi
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
What is your optimal infrastructure
we're all striving to optimize our infra, but what would you say is the PERFECT infra configuration in your mind?
I'll go first
# Product:
micro services
each service has its own git repo following a standardized directory structure
each service has its own pipeline that scans the code and runs a full regression suite before a PR is merged to master/main
each service has a health check that is actually accurate to the state of the service
launch time for the services are minimal
daily pipeline regression suite that updates a global dashboard highlighting all the current repo statuses
SaaS ONLY (onprem is pain)
# Deployment:
terraform code base that is multiplatform, can be deployed to any cloud
deployment process is executed by a pipeline with optional parameters for configure what kind of environment to deploy
all services hosted in GKE/EKS or some other cloud hosting platform
auto scaling horizontally and vertically
all logs routed to logz.io and tagged by environment and service
all deployed infra is tagged by its environment name
# Monitoring:
new relic, all the things go to new relic. (not sponsored, i just used it in the past and it was a dream when setup correctly. was just too expensive for my previous company at the time)
one thing i did learn is that the less you "have" to do. the better. and if your company can afford to outsource something, your life will be much easier.
https://redd.it/1urdygd
@r_devops
we're all striving to optimize our infra, but what would you say is the PERFECT infra configuration in your mind?
I'll go first
# Product:
micro services
each service has its own git repo following a standardized directory structure
each service has its own pipeline that scans the code and runs a full regression suite before a PR is merged to master/main
each service has a health check that is actually accurate to the state of the service
launch time for the services are minimal
daily pipeline regression suite that updates a global dashboard highlighting all the current repo statuses
SaaS ONLY (onprem is pain)
# Deployment:
terraform code base that is multiplatform, can be deployed to any cloud
deployment process is executed by a pipeline with optional parameters for configure what kind of environment to deploy
all services hosted in GKE/EKS or some other cloud hosting platform
auto scaling horizontally and vertically
all logs routed to logz.io and tagged by environment and service
all deployed infra is tagged by its environment name
# Monitoring:
new relic, all the things go to new relic. (not sponsored, i just used it in the past and it was a dream when setup correctly. was just too expensive for my previous company at the time)
one thing i did learn is that the less you "have" to do. the better. and if your company can afford to outsource something, your life will be much easier.
https://redd.it/1urdygd
@r_devops
Logz.io
Logz.io: Modern Observability Powered by AI
Stop chasing alerts. Logz.io's AI-powered observability platform unifies logs, metrics, and traces to cut MTTR, automate root cause analysis, and get ahead of issues before they impact users.
What's up with GitHub runners lately?
https://preview.redd.it/yzae6azo76ch1.png?width=718&format=png&auto=webp&s=6a1dbf5a366b67da18cc08e2b898d8a764cf0580
Too much AI slop to build? If this continues I will probably prefer to just use my self hosted ones for all jobs.
https://redd.it/1urkyux
@r_devops
https://preview.redd.it/yzae6azo76ch1.png?width=718&format=png&auto=webp&s=6a1dbf5a366b67da18cc08e2b898d8a764cf0580
Too much AI slop to build? If this continues I will probably prefer to just use my self hosted ones for all jobs.
https://redd.it/1urkyux
@r_devops