Engineering managers: how do you prevent valuable Slack discussions from disappearing
In my team I notice senior engineers write detailed explanations in Slack, but months later nobody can find them. Curious how others solve this.
https://redd.it/1unvlkj
@r_devops
In my team I notice senior engineers write detailed explanations in Slack, but months later nobody can find them. Curious how others solve this.
https://redd.it/1unvlkj
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Has anyone successfully made the jump from SDET to platform engineer from a Tier 1 company?
Hi everyone,
Iβm currently an SDET (exp 1 year, total exp 2 years) at a Tier 1 tech company and Iβm planning my move into a platform engineer role. I love building tools and want to be closer to product development and feature ownership.
For those who have successfully made this pivot:
Did you find it easier to transfer internally or interview elsewhere?
How did you bridge the gap in System Design if your daily work was focused on automation frameworks?
What was the single most helpful thing you did to prove you were ready?
Appreciate any insights or "traps" to avoid!
https://redd.it/1uo12qx
@r_devops
Hi everyone,
Iβm currently an SDET (exp 1 year, total exp 2 years) at a Tier 1 tech company and Iβm planning my move into a platform engineer role. I love building tools and want to be closer to product development and feature ownership.
For those who have successfully made this pivot:
Did you find it easier to transfer internally or interview elsewhere?
How did you bridge the gap in System Design if your daily work was focused on automation frameworks?
What was the single most helpful thing you did to prove you were ready?
Appreciate any insights or "traps" to avoid!
https://redd.it/1uo12qx
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Reticle - The infrastructure diagram you can operate.
Hi all. I made a (OSS+MIT) tool called Reticle which is an infrastructure diagramming tool, but it uses your host ssh/kconf credentials to access the real running state of your actual system.
It has configurable cron health probes to determine whether your resources are green in realtime, and even gives you access to a terminal on a box for quick fixes without switching to another app.
It feels like a fresh way to get a health overview, and if something is broken, the context/dependencies is immediately obvious.
There's more features like team collaboration and professional PDF export, but I'm really curious whether r/devops (the experts in this space), find it interesting or useful.
So if there's any feedback I'd be incredibly grateful. Thank you.
reticle.live
https://redd.it/1uo6bkc
@r_devops
Hi all. I made a (OSS+MIT) tool called Reticle which is an infrastructure diagramming tool, but it uses your host ssh/kconf credentials to access the real running state of your actual system.
It has configurable cron health probes to determine whether your resources are green in realtime, and even gives you access to a terminal on a box for quick fixes without switching to another app.
It feels like a fresh way to get a health overview, and if something is broken, the context/dependencies is immediately obvious.
There's more features like team collaboration and professional PDF export, but I'm really curious whether r/devops (the experts in this space), find it interesting or useful.
So if there's any feedback I'd be incredibly grateful. Thank you.
reticle.live
https://redd.it/1uo6bkc
@r_devops
Reticle
Reticle: the infrastructure diagram you can operate
Draw your real topology, watch it breathe, fix it from the map. Free MIT desktop app.
Leaving K8s platform engineering for an internal dev-tooling/CI-CD role
I am currently building a central "Cluster-as-a-Service" platform nd deal with automated cluster provisioning, multi-tenant setup, observability work, kubernetes security etc. I also contribute upstream to K8s-sigs. The company I work for is rather unknown and I would love to be at a bigger brand with a real engineering culture. (I have gotten a few faang interviews in the last month, but declined one myself and did not get more offers).
I now have an offer at a well-known fintech where I would work on owning CI/CD pipeline templates (GitHub Actions/Jenkins), and some more terraform templates plus policy work. Dev teams manage their own infra/CI-CD day-to-day, the team has no direct Kubernetes/Linux/scaling ownership and everything's serverless. However it is a well-known tech brand and I would get 35% more salary (which bumps my salary to slightly above market, not much tho).
My worry: My long-term goal is deep infra/platform roles (K8s internals, distributed systems, SRE-adjacent work). I'm concerned 1-2 years here erodes my skills in that area and I navigate myself into a build tooling and devex niche? What is your take on that move? Would I limit myself too much or can I easily move back to infra platform work later, i.e., would owning CI/CD/tooling be seen as equivalent experience to K8s internals, distributed systems, and observability work when I try to move back later?
https://redd.it/1uoatjd
@r_devops
I am currently building a central "Cluster-as-a-Service" platform nd deal with automated cluster provisioning, multi-tenant setup, observability work, kubernetes security etc. I also contribute upstream to K8s-sigs. The company I work for is rather unknown and I would love to be at a bigger brand with a real engineering culture. (I have gotten a few faang interviews in the last month, but declined one myself and did not get more offers).
I now have an offer at a well-known fintech where I would work on owning CI/CD pipeline templates (GitHub Actions/Jenkins), and some more terraform templates plus policy work. Dev teams manage their own infra/CI-CD day-to-day, the team has no direct Kubernetes/Linux/scaling ownership and everything's serverless. However it is a well-known tech brand and I would get 35% more salary (which bumps my salary to slightly above market, not much tho).
My worry: My long-term goal is deep infra/platform roles (K8s internals, distributed systems, SRE-adjacent work). I'm concerned 1-2 years here erodes my skills in that area and I navigate myself into a build tooling and devex niche? What is your take on that move? Would I limit myself too much or can I easily move back to infra platform work later, i.e., would owning CI/CD/tooling be seen as equivalent experience to K8s internals, distributed systems, and observability work when I try to move back later?
https://redd.it/1uoatjd
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
When does an on-premises server become cheaper than AWS, Azure, or a VPS?
Everyone talks about moving to the cloud, but is it always the right choice?
For a small business with stable workloads (website, email, file storage, backups, internal apps), does buying a server once and running it for 5β7 years make more financial sense than paying cloud bills every month?
I'm also thinking from a business perspective. If you were starting a business today with a small budget, what service would you offer that brings recurring monthly income?
I'd love to hear real experiences from people who run infrastructure or do businesses from home network not just theory or marketing.
https://redd.it/1uobn4g
@r_devops
Everyone talks about moving to the cloud, but is it always the right choice?
For a small business with stable workloads (website, email, file storage, backups, internal apps), does buying a server once and running it for 5β7 years make more financial sense than paying cloud bills every month?
I'm also thinking from a business perspective. If you were starting a business today with a small budget, what service would you offer that brings recurring monthly income?
I'd love to hear real experiences from people who run infrastructure or do businesses from home network not just theory or marketing.
https://redd.it/1uobn4g
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
GitLab CI skill for ai agents based on official docs
I use ai agents as helper I talk to, not for blind vibecoding. One thing I kept noticing is asking agent to write or refactor gitlab ci pipeline, and results are often questionable. It creates a god yaml, outdated keywords, no thought about debugging or developer experience.
I looked for existing skills but did not find anything I would actually trust, most looked generated in one shot. So I spent some time and made my own. Used agent help of course, but went through everything myself and checked it against official docs for GitLab 18+
It covers pipeline structure and refactoring, bash in ci jobs, pipelines and other common patterns, debugging failed pipelines, readable logs and naming and many other cases
https://github.com/beeyev/skills/
Works with claude code and anything supporting skills format
I have been using it privately for couple of month and improving constantly, maybe it will useful for someone else too
https://redd.it/1uoab35
@r_devops
I use ai agents as helper I talk to, not for blind vibecoding. One thing I kept noticing is asking agent to write or refactor gitlab ci pipeline, and results are often questionable. It creates a god yaml, outdated keywords, no thought about debugging or developer experience.
I looked for existing skills but did not find anything I would actually trust, most looked generated in one shot. So I spent some time and made my own. Used agent help of course, but went through everything myself and checked it against official docs for GitLab 18+
It covers pipeline structure and refactoring, bash in ci jobs, pipelines and other common patterns, debugging failed pipelines, readable logs and naming and many other cases
https://github.com/beeyev/skills/
Works with claude code and anything supporting skills format
I have been using it privately for couple of month and improving constantly, maybe it will useful for someone else too
https://redd.it/1uoab35
@r_devops
GitHub
GitHub - beeyev/skills: A collection of Agent Skills for Claude Code and other SKILL.md-compatible agents.
A collection of Agent Skills for Claude Code and other SKILL.md-compatible agents. - beeyev/skills
How are you handling query rewrites and schema changes in production databases?
I'm curious how teams are dealing with query performance over time as applications evolve.
A few questions I'd love to hear your experience on:
\\- How do you identify queries that need to be rewritten?
\\- How do you ensure schema changes (indexes, partitions, column changes, etc.) don't negatively impact production?
\\- Is this mostly a manual process, or do you use any tools to analyze query patterns and recommend improvements?
\\- Have you ever had an outage or major performance issue because a query or schema change wasn't optimized?
I'm exploring this space and trying to understand whether this is a pain point worth solving. I'd really appreciate hearing about your workflows, the tools you use, and what's still frustrating today.
https://redd.it/1uoewbr
@r_devops
I'm curious how teams are dealing with query performance over time as applications evolve.
A few questions I'd love to hear your experience on:
\\- How do you identify queries that need to be rewritten?
\\- How do you ensure schema changes (indexes, partitions, column changes, etc.) don't negatively impact production?
\\- Is this mostly a manual process, or do you use any tools to analyze query patterns and recommend improvements?
\\- Have you ever had an outage or major performance issue because a query or schema change wasn't optimized?
I'm exploring this space and trying to understand whether this is a pain point worth solving. I'd really appreciate hearing about your workflows, the tools you use, and what's still frustrating today.
https://redd.it/1uoewbr
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Weekly Self Promotion Thread
Hey r/devops, welcome to our weekly self-promotion thread!
Feel free to use this thread to promote any projects, ideas, or any repos you're wanting to share. Please keep in mind that we ask you to stay friendly, civil, and adhere to the subreddit rules!
https://redd.it/1uopcqy
@r_devops
Hey r/devops, welcome to our weekly self-promotion thread!
Feel free to use this thread to promote any projects, ideas, or any repos you're wanting to share. Please keep in mind that we ask you to stay friendly, civil, and adhere to the subreddit rules!
https://redd.it/1uopcqy
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Interview prep for devops
​
JD says devops, python and linux developer.
I just want focus on scripting and ignore oops.
Is that possible?
Does this heavily depends on project requirements and company?
How did you clear devops interview?
Any suggestions?
My focus area is systems emgineering.
https://redd.it/1uot5eb
@r_devops
​
JD says devops, python and linux developer.
I just want focus on scripting and ignore oops.
Is that possible?
Does this heavily depends on project requirements and company?
How did you clear devops interview?
Any suggestions?
My focus area is systems emgineering.
https://redd.it/1uot5eb
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Help with Devsecops pipeline setup
I am a pentester. My manager had given me a task for cicd integration with checkmarx. I can understand the basic stuff but I am unable to come up with material for the following:
1. How do manage secrets in the pipeline (someone suggested me aws kms)
2. How to run authenticated scans (apps using okta) using pipeline
3. Capturing traffic to run the scans on them.
I would appreciate if someone can help me in this scenario as my job depends on it.
https://redd.it/1uotw74
@r_devops
I am a pentester. My manager had given me a task for cicd integration with checkmarx. I can understand the basic stuff but I am unable to come up with material for the following:
1. How do manage secrets in the pipeline (someone suggested me aws kms)
2. How to run authenticated scans (apps using okta) using pipeline
3. Capturing traffic to run the scans on them.
I would appreciate if someone can help me in this scenario as my job depends on it.
https://redd.it/1uotw74
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
5 YOE in DevOps/SRE at a Tier-1 bank, pivoted after a layoff β now in a support-heavy role. How do I course-correct before it hurts my career?
Hey r/devops β not looking for validation here, looking for people whoβve actually been through this.
The background:
Spent 5 years as a DevOps/SRE engineer at a large global investment bank. Real scope β owned incident response, SLO/SLI definition, Python automation, Kubernetes deployments, Grafana/InfluxDB observability stacks, CI/CD pipelines. High-stakes production trading systems. The kind of work I was proud of.
Earlier this year, my team was restructured. Whole group gone. I was laid off.
The decision I made:
The market was rough. I had a contract offer on the table β onsite at another major financial institution, compensation was actually better than my previous role. I took it. Practically speaking, it was the right call at the time.
What I underestimated: the actual scope of the work. Itβs L1-L2 support. Ticket resolution, runbook execution, escalations. Iβm not owning anything. Iβm not building anything. And every month I stay here, the gap between my resume and my current reps gets wider.
Iβm not bitter about the pay β itβs kept me stable. But I know this lane leads nowhere if I stay in it too long.
Three things I genuinely need help thinking through:
1. How do I keep building real SRE depth when my day job isnβt giving me the reps?
I have a solid foundation β Docker, Kubernetes, Terraform, Ansible, Python, Grafana. But I want to go deeper on observability, chaos engineering, capacity planning, and SRE fundamentals. What projects, home labs, or open-source contributions have actually made you better β not just resume filler?
2. Is the AWS Generative AI Professional cert (AIP-C01) worth pursuing for an SRE career path?
Iβm currently working toward it, specifically because Iβm interested in AIOps β anomaly detection, automated remediation, predictive reliability. Is this a real differentiator for senior SRE roles in 2025-26, or am I chasing the wrong thing? What certs have actually moved the needle for you?
3. How do I frame a support-heavy role in interviews without it undercutting 6 years of solid experience?
My instinct is to move within 6β9 months before the narrative gets harder to control. Is that the right timeline? And how do you position this kind of pivot when youβre targeting L4/L5 SRE or Senior Platform Engineer roles?
Context: Based in Bengaluru. Targeting product companies, global banks, and open to remote/international opportunities.
I know I made a compromise. Iβm not here to relitigate that. Just want to make sure the next move is the right one.
Appreciate any honest perspectives β especially from people whoβve navigated something similar. π
https://redd.it/1uoydmi
@r_devops
Hey r/devops β not looking for validation here, looking for people whoβve actually been through this.
The background:
Spent 5 years as a DevOps/SRE engineer at a large global investment bank. Real scope β owned incident response, SLO/SLI definition, Python automation, Kubernetes deployments, Grafana/InfluxDB observability stacks, CI/CD pipelines. High-stakes production trading systems. The kind of work I was proud of.
Earlier this year, my team was restructured. Whole group gone. I was laid off.
The decision I made:
The market was rough. I had a contract offer on the table β onsite at another major financial institution, compensation was actually better than my previous role. I took it. Practically speaking, it was the right call at the time.
What I underestimated: the actual scope of the work. Itβs L1-L2 support. Ticket resolution, runbook execution, escalations. Iβm not owning anything. Iβm not building anything. And every month I stay here, the gap between my resume and my current reps gets wider.
Iβm not bitter about the pay β itβs kept me stable. But I know this lane leads nowhere if I stay in it too long.
Three things I genuinely need help thinking through:
1. How do I keep building real SRE depth when my day job isnβt giving me the reps?
I have a solid foundation β Docker, Kubernetes, Terraform, Ansible, Python, Grafana. But I want to go deeper on observability, chaos engineering, capacity planning, and SRE fundamentals. What projects, home labs, or open-source contributions have actually made you better β not just resume filler?
2. Is the AWS Generative AI Professional cert (AIP-C01) worth pursuing for an SRE career path?
Iβm currently working toward it, specifically because Iβm interested in AIOps β anomaly detection, automated remediation, predictive reliability. Is this a real differentiator for senior SRE roles in 2025-26, or am I chasing the wrong thing? What certs have actually moved the needle for you?
3. How do I frame a support-heavy role in interviews without it undercutting 6 years of solid experience?
My instinct is to move within 6β9 months before the narrative gets harder to control. Is that the right timeline? And how do you position this kind of pivot when youβre targeting L4/L5 SRE or Senior Platform Engineer roles?
Context: Based in Bengaluru. Targeting product companies, global banks, and open to remote/international opportunities.
I know I made a compromise. Iβm not here to relitigate that. Just want to make sure the next move is the right one.
Appreciate any honest perspectives β especially from people whoβve navigated something similar. π
https://redd.it/1uoydmi
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
I'm starting a new movement
I am officially declaring the start (in my mind) of #MRBA
That stands for "Make Releases Boring Again"
This was prompted by a Release Engineer job posting that was your usual "just be on 24/7 on every communication channel during release windows". So every few months, you over activate my nervous system and it takes until the next release for it to finally calm down only to be activated again? No thanks.
I need to be doing automation, environment config hardening, observability tweaking. Not "monitoring Slack in case someone reports an issue". π
Releases need to be boring. The more boring, the more both dev AND ops sleep. With the added bonus of not over-rewarding heroics. π
Release day hype/fanfare/stress is for shit like clothing, games, etc. Not the newest feature for your internal app with 10 users.
https://redd.it/1up28y5
@r_devops
I am officially declaring the start (in my mind) of #MRBA
That stands for "Make Releases Boring Again"
This was prompted by a Release Engineer job posting that was your usual "just be on 24/7 on every communication channel during release windows". So every few months, you over activate my nervous system and it takes until the next release for it to finally calm down only to be activated again? No thanks.
I need to be doing automation, environment config hardening, observability tweaking. Not "monitoring Slack in case someone reports an issue". π
Releases need to be boring. The more boring, the more both dev AND ops sleep. With the added bonus of not over-rewarding heroics. π
Release day hype/fanfare/stress is for shit like clothing, games, etc. Not the newest feature for your internal app with 10 users.
https://redd.it/1up28y5
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
I built a single-binary TUI that manages Redis, Postgres, SSH, Docker, Git, S3, MySQL, MongoDB, and HTTP β with a built-in MCP server for AI tooling
**Qore** is a single-binary infrastructure orchestrator with a terminal-native UI. You type commands, get results inline β no context switching between redis-cli, psql, ssh, and docker. What makes it different: **8 connection types in one place:**
* Redis (native RESP protocol β no redis-cli needed)
* PostgreSQL / MySQL / MongoDB (full SQL queries, EXPLAIN, slow queries, CSV export)
* S3-compatible (AWS SigV4 β works with MinIO, R2, AWS)
* HTTP/REST (GET, POST, PUT, PATCH, DELETE with auth)
* SSH (exec, SFTP, systemd, Docker Compose, deploy scripts, interactive shell)
* Git (branch graph, merge, rebase, cherry-pick, blame, tags)
**Built-in MCP server:** This is the part I'm most excited about. It exposes 35 tools (SSH, Docker, database queries, system discovery, HTTP) over JSON-RPC 2.0 β so Claude, Cursor, or any MCP-compatible AI can interact with your infrastructure using connection names only. Credentials stay server-side. **Other highlights:**
* Secure vault: AES-256-GCM + scrypt, master password never touches disk
* Docker via Unix socket (no docker CLI dependency)
* Multi-tab: all connections stay mounted, switch with Ctrl+Tab
* Multi-service dashboard with auto-refresh
* Health checks with latency sparklines
* Self-updating (`qore update`)
* Single binary, \~45MB, Linux/macOS/Windows
Install: `curl -fsSL` [`https://github.com/Kodjaoglanian/qore/releases/latest/download/install.sh`](https://github.com/Kodjaoglanian/qore/releases/latest/download/install.sh) `| bash` Code: [https://github.com/Kodjaoglanian/qore](https://github.com/Kodjaoglanian/qore)
Happy to answer questions!
https://redd.it/1up1rh3
@r_devops
**Qore** is a single-binary infrastructure orchestrator with a terminal-native UI. You type commands, get results inline β no context switching between redis-cli, psql, ssh, and docker. What makes it different: **8 connection types in one place:**
* Redis (native RESP protocol β no redis-cli needed)
* PostgreSQL / MySQL / MongoDB (full SQL queries, EXPLAIN, slow queries, CSV export)
* S3-compatible (AWS SigV4 β works with MinIO, R2, AWS)
* HTTP/REST (GET, POST, PUT, PATCH, DELETE with auth)
* SSH (exec, SFTP, systemd, Docker Compose, deploy scripts, interactive shell)
* Git (branch graph, merge, rebase, cherry-pick, blame, tags)
**Built-in MCP server:** This is the part I'm most excited about. It exposes 35 tools (SSH, Docker, database queries, system discovery, HTTP) over JSON-RPC 2.0 β so Claude, Cursor, or any MCP-compatible AI can interact with your infrastructure using connection names only. Credentials stay server-side. **Other highlights:**
* Secure vault: AES-256-GCM + scrypt, master password never touches disk
* Docker via Unix socket (no docker CLI dependency)
* Multi-tab: all connections stay mounted, switch with Ctrl+Tab
* Multi-service dashboard with auto-refresh
* Health checks with latency sparklines
* Self-updating (`qore update`)
* Single binary, \~45MB, Linux/macOS/Windows
Install: `curl -fsSL` [`https://github.com/Kodjaoglanian/qore/releases/latest/download/install.sh`](https://github.com/Kodjaoglanian/qore/releases/latest/download/install.sh) `| bash` Code: [https://github.com/Kodjaoglanian/qore](https://github.com/Kodjaoglanian/qore)
Happy to answer questions!
https://redd.it/1up1rh3
@r_devops
Need guidance !
I have 3 yrs of experience as a cloud engineer/devops worked in same company from start got one promotion and good hike all along . Want to switch to now and thinking on doing CKA certification is it worth having ?
https://redd.it/1up4s58
@r_devops
I have 3 yrs of experience as a cloud engineer/devops worked in same company from start got one promotion and good hike all along . Want to switch to now and thinking on doing CKA certification is it worth having ?
https://redd.it/1up4s58
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
How are you managing dynamic runtime configuration without triggering a full CI/CD deployment?
Hi all!
Iβm currently digging into the "toil" of our (the company i work at's) release process, and Iβm hitting a recurring bottleneck: rate-limiting configuration.
Right now, we have our limits (e.g., token buckets, thresholds) defined as part of our static config, which is baked into our container images. Whenever we need to tune these for a traffic spike or emergency throttle, it forces a full CI/CD deployment aka build, push, wait for rollout, and pray the new pod doesn't have a startup issue.
It feels fundamentally wrong to bounce a production service just to change a numerical threshold.
Iβm looking into moving this "knob-turning" out of the deployment pipeline and into a centralized, runtime-synced store (like Redis), so we can tweak values on the fly without a code push.
Is anyone else using a "Config-as-a-Service" or dynamic sidecar pattern for this, or have we missed a super obvious solution lol thanks guys :)
https://redd.it/1up6pnr
@r_devops
Hi all!
Iβm currently digging into the "toil" of our (the company i work at's) release process, and Iβm hitting a recurring bottleneck: rate-limiting configuration.
Right now, we have our limits (e.g., token buckets, thresholds) defined as part of our static config, which is baked into our container images. Whenever we need to tune these for a traffic spike or emergency throttle, it forces a full CI/CD deployment aka build, push, wait for rollout, and pray the new pod doesn't have a startup issue.
It feels fundamentally wrong to bounce a production service just to change a numerical threshold.
Iβm looking into moving this "knob-turning" out of the deployment pipeline and into a centralized, runtime-synced store (like Redis), so we can tweak values on the fly without a code push.
Is anyone else using a "Config-as-a-Service" or dynamic sidecar pattern for this, or have we missed a super obvious solution lol thanks guys :)
https://redd.it/1up6pnr
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Odd manager behavior - looking for opinions
I've been working at a local startup as a DevOps engineer for the past year. I have a total of around 3 years of experience, which is considered mid-level at best. I've worked with various technologies and I'm not that confident in my technical skills but I do try my best to deliver. That's not my problem though. My problem is that communication with my manager is terrible.
I am rarely assigned any tickets. My work is a mixture of verbally assigned tasks with minimal details and initiatives I take to improve our workflows. My manager rarely joins our weekly one to one meetings so I decided to send him my task updates on teams. He doesn't reply to my messages, sometimes for days. He will only reply fast if he believes the question is important to him. He never reacts or responds to my task updates and won't review my work for months (if at all). Sometimes, when I ask him questions or try to make conversation on meetings he won't respond. When I ask again, he will laugh it off saying he heard me the first time. The kind of questions I ask are usually clarifications on my tasks, or discussion around my technical implementations.
In general, I believe I am fairly independent and I don't burden the team. My tasks are mine to deliver and I am solely responsible for them. I request guidance in a structured way and usually when I don't have enough information to move on. Even when I don't have enough context I will push through instead of waiting for weeks for an answer. The problem is that I am not responsible for the infrastructure architectural decisions, and I need the clarifications in order to do my job effectively.
I will give you a recent example:
I am asked to deploy one of our products to a test environment, but we want this product to be isolated from our other products so it's not the standard procedure.
That's pretty much what I got. I ask my manager to discuss for like 15minutes to show me what he deployed for production. He never does.
I create a list of resources I will need to deploy along with the infra design. I make the deployment. I share that with him. He shows up after a week of being AWOL to tell me some resources should be redeployed. I ask for two very specific clarifications in chat. He says we will discuss in a meeting. We join the meeting. I ask. He stays silent. I ask again. He says he heard me and he just didn't respond.
What am I supposed to do at this point? This whole thing is deeply demoralizing to me. I feel deeply disrespected and looked down upon. I am mad and sad and it's affecting my confidence and will to work and be creative and productive. I've tried a few different things since starting in this company, I've created my own tickets and shared them with him, I've tried texting him I've tried reaching out during meetings to avoid spamming him. Nothing seems to work. I dread going to work every single day. I feel lost about what to do next.
I want to stay professional but I also feel very done. I want them to fuck off but also I want to take technical experience. I want to quit but I know it's a terrible idea. What can I do to continue learning and growing within team and business goals when these are not communicated properly? What could I be doing wrong? Am I needy or is this truly as annoying and disfunctional as I think it is?
Thanks for reaching this far, any opinions, experiences or recommendations are welcome. I can take criticism as long as it is respectful, my mental health is declining fast enough already π
https://redd.it/1upabcd
@r_devops
I've been working at a local startup as a DevOps engineer for the past year. I have a total of around 3 years of experience, which is considered mid-level at best. I've worked with various technologies and I'm not that confident in my technical skills but I do try my best to deliver. That's not my problem though. My problem is that communication with my manager is terrible.
I am rarely assigned any tickets. My work is a mixture of verbally assigned tasks with minimal details and initiatives I take to improve our workflows. My manager rarely joins our weekly one to one meetings so I decided to send him my task updates on teams. He doesn't reply to my messages, sometimes for days. He will only reply fast if he believes the question is important to him. He never reacts or responds to my task updates and won't review my work for months (if at all). Sometimes, when I ask him questions or try to make conversation on meetings he won't respond. When I ask again, he will laugh it off saying he heard me the first time. The kind of questions I ask are usually clarifications on my tasks, or discussion around my technical implementations.
In general, I believe I am fairly independent and I don't burden the team. My tasks are mine to deliver and I am solely responsible for them. I request guidance in a structured way and usually when I don't have enough information to move on. Even when I don't have enough context I will push through instead of waiting for weeks for an answer. The problem is that I am not responsible for the infrastructure architectural decisions, and I need the clarifications in order to do my job effectively.
I will give you a recent example:
I am asked to deploy one of our products to a test environment, but we want this product to be isolated from our other products so it's not the standard procedure.
That's pretty much what I got. I ask my manager to discuss for like 15minutes to show me what he deployed for production. He never does.
I create a list of resources I will need to deploy along with the infra design. I make the deployment. I share that with him. He shows up after a week of being AWOL to tell me some resources should be redeployed. I ask for two very specific clarifications in chat. He says we will discuss in a meeting. We join the meeting. I ask. He stays silent. I ask again. He says he heard me and he just didn't respond.
What am I supposed to do at this point? This whole thing is deeply demoralizing to me. I feel deeply disrespected and looked down upon. I am mad and sad and it's affecting my confidence and will to work and be creative and productive. I've tried a few different things since starting in this company, I've created my own tickets and shared them with him, I've tried texting him I've tried reaching out during meetings to avoid spamming him. Nothing seems to work. I dread going to work every single day. I feel lost about what to do next.
I want to stay professional but I also feel very done. I want them to fuck off but also I want to take technical experience. I want to quit but I know it's a terrible idea. What can I do to continue learning and growing within team and business goals when these are not communicated properly? What could I be doing wrong? Am I needy or is this truly as annoying and disfunctional as I think it is?
Thanks for reaching this far, any opinions, experiences or recommendations are welcome. I can take criticism as long as it is respectful, my mental health is declining fast enough already π
https://redd.it/1upabcd
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
How to structure scoring live traffic
We've had offline evals as part of our CI for a while now, but last month we got hit with something that none of our CI runs flagged. Our production inputs had drifted and users were asking things our test set just didn't cover, and output quality on that section of things had degraded for weeks. So unfortunately, our existing evals gave me a false sense of safety because they can only ever test what I thought to put in them.
So now I'm trying to figure out actually sampling and scoring real production responses, not just CI runs against a fixed dataset.
Main things I'm rubber ducking:
Sampling. Are people scoring all live traffic or some percentage?
Alerting. I want to know when quality drops on live traffic, but Iβm not trying creat another annoying alert channel. And then if/when it goes off who/what owns the response?
https://redd.it/1upa4lk
@r_devops
We've had offline evals as part of our CI for a while now, but last month we got hit with something that none of our CI runs flagged. Our production inputs had drifted and users were asking things our test set just didn't cover, and output quality on that section of things had degraded for weeks. So unfortunately, our existing evals gave me a false sense of safety because they can only ever test what I thought to put in them.
So now I'm trying to figure out actually sampling and scoring real production responses, not just CI runs against a fixed dataset.
Main things I'm rubber ducking:
Sampling. Are people scoring all live traffic or some percentage?
Alerting. I want to know when quality drops on live traffic, but Iβm not trying creat another annoying alert channel. And then if/when it goes off who/what owns the response?
https://redd.it/1upa4lk
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
How do you actually separate CI/CD pipelines for AKS across dev/qa/uat/prod in Azure DevOps?
Hey Folks, need your advice badly ,
I'm building out a CI/CD flow for AKS using Azure DevOps Pipelines (not ArgoCD/GitOps for this one, using native Azure Pipelines +
The MS Learn sample bundles CI and CD into one pipeline (Build stage β Deploy stage, same YAML file), which builds once and deploys straight to the cluster. That seems fine for a single environment, but once you add QA β UAT β Prod with a manual sign-off before prod, it starts to feel like the wrong shape.
Questions:
1. Do you run one CD pipeline with multiple stages (QA β UAT β Prod, each an Azure DevOps Environment with its own approval gates), or separate pipelines per environment (e.g.
2. How do you handle the nonprod β prod ACR promotion? Are you doing
3. If CI only has push access to a nonprod ACR, what triggers the CD pipeline β a pipeline completion trigger (
4. For those who've tried both native Azure Pipelines deploys and ArgoCD/GitOps for AKS was there a specific pain point that pushed you from one to the other?
Not looking for "just use GitOps" as the whole answer (I get the appeal); more interested in how people structure this with plain Azure DevOps pipelines if they're not on ArgoCD, since that's what I'm working with right now.
https://redd.it/1upcl17
@r_devops
Hey Folks, need your advice badly ,
I'm building out a CI/CD flow for AKS using Azure DevOps Pipelines (not ArgoCD/GitOps for this one, using native Azure Pipelines +
KubernetesManifest@1 tasks). Trying to understand what people actually do in production.The MS Learn sample bundles CI and CD into one pipeline (Build stage β Deploy stage, same YAML file), which builds once and deploys straight to the cluster. That seems fine for a single environment, but once you add QA β UAT β Prod with a manual sign-off before prod, it starts to feel like the wrong shape.
Questions:
1. Do you run one CD pipeline with multiple stages (QA β UAT β Prod, each an Azure DevOps Environment with its own approval gates), or separate pipelines per environment (e.g.
cd-nonprod and cd-prod)? What made you choose one over the other?2. How do you handle the nonprod β prod ACR promotion? Are you doing
az acr import to copy the same digest into a separate prod registry, or do you just use one ACR with RBAC-scoped repositories/tags instead of physically separate registries?3. If CI only has push access to a nonprod ACR, what triggers the CD pipeline β a pipeline completion trigger (
resources.pipelines), a manual run with an image tag parameter, or something else?4. For those who've tried both native Azure Pipelines deploys and ArgoCD/GitOps for AKS was there a specific pain point that pushed you from one to the other?
Not looking for "just use GitOps" as the whole answer (I get the appeal); more interested in how people structure this with plain Azure DevOps pipelines if they're not on ArgoCD, since that's what I'm working with right now.
https://redd.it/1upcl17
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community