Best way to restrict AWS/Cloudflare app to specific desktops?
Best way to restrict AWS/Cloudflare app to specific desktops?
We are building a fee payment application for a school organization.
**Stack:** DB/Backend on AWS and frontend on Cloudflare.
**The challenge:** We need to restrict payment work flow used by cashiers to specific systems, while the read fees access should be able to be accessed from anywhere.
The desktops are unmanaged, regular PCs, residing in different branches in different cities. They are all connected via standard consumer ISPs (no static IPs, no company intranet).
As we are already using Cloudflare, is this something that can be achieved with Cloudflare Zero Trust free tier?
I have never worked with this restriction before, SO I am open to any suggestions. And as this is a very low budget project, I'm looking for something that costs as less as possible (Preferably free).
https://redd.it/1unyxz6
@r_devops
Best way to restrict AWS/Cloudflare app to specific desktops?
We are building a fee payment application for a school organization.
**Stack:** DB/Backend on AWS and frontend on Cloudflare.
**The challenge:** We need to restrict payment work flow used by cashiers to specific systems, while the read fees access should be able to be accessed from anywhere.
The desktops are unmanaged, regular PCs, residing in different branches in different cities. They are all connected via standard consumer ISPs (no static IPs, no company intranet).
As we are already using Cloudflare, is this something that can be achieved with Cloudflare Zero Trust free tier?
I have never worked with this restriction before, SO I am open to any suggestions. And as this is a very low budget project, I'm looking for something that costs as less as possible (Preferably free).
https://redd.it/1unyxz6
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Engineering managers: how do you prevent valuable Slack discussions from disappearing
In my team I notice senior engineers write detailed explanations in Slack, but months later nobody can find them. Curious how others solve this.
https://redd.it/1unvlkj
@r_devops
In my team I notice senior engineers write detailed explanations in Slack, but months later nobody can find them. Curious how others solve this.
https://redd.it/1unvlkj
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Has anyone successfully made the jump from SDET to platform engineer from a Tier 1 company?
Hi everyone,
I’m currently an SDET (exp 1 year, total exp 2 years) at a Tier 1 tech company and I’m planning my move into a platform engineer role. I love building tools and want to be closer to product development and feature ownership.
For those who have successfully made this pivot:
Did you find it easier to transfer internally or interview elsewhere?
How did you bridge the gap in System Design if your daily work was focused on automation frameworks?
What was the single most helpful thing you did to prove you were ready?
Appreciate any insights or "traps" to avoid!
https://redd.it/1uo12qx
@r_devops
Hi everyone,
I’m currently an SDET (exp 1 year, total exp 2 years) at a Tier 1 tech company and I’m planning my move into a platform engineer role. I love building tools and want to be closer to product development and feature ownership.
For those who have successfully made this pivot:
Did you find it easier to transfer internally or interview elsewhere?
How did you bridge the gap in System Design if your daily work was focused on automation frameworks?
What was the single most helpful thing you did to prove you were ready?
Appreciate any insights or "traps" to avoid!
https://redd.it/1uo12qx
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Reticle - The infrastructure diagram you can operate.
Hi all. I made a (OSS+MIT) tool called Reticle which is an infrastructure diagramming tool, but it uses your host ssh/kconf credentials to access the real running state of your actual system.
It has configurable cron health probes to determine whether your resources are green in realtime, and even gives you access to a terminal on a box for quick fixes without switching to another app.
It feels like a fresh way to get a health overview, and if something is broken, the context/dependencies is immediately obvious.
There's more features like team collaboration and professional PDF export, but I'm really curious whether r/devops (the experts in this space), find it interesting or useful.
So if there's any feedback I'd be incredibly grateful. Thank you.
reticle.live
https://redd.it/1uo6bkc
@r_devops
Hi all. I made a (OSS+MIT) tool called Reticle which is an infrastructure diagramming tool, but it uses your host ssh/kconf credentials to access the real running state of your actual system.
It has configurable cron health probes to determine whether your resources are green in realtime, and even gives you access to a terminal on a box for quick fixes without switching to another app.
It feels like a fresh way to get a health overview, and if something is broken, the context/dependencies is immediately obvious.
There's more features like team collaboration and professional PDF export, but I'm really curious whether r/devops (the experts in this space), find it interesting or useful.
So if there's any feedback I'd be incredibly grateful. Thank you.
reticle.live
https://redd.it/1uo6bkc
@r_devops
Reticle
Reticle: the infrastructure diagram you can operate
Draw your real topology, watch it breathe, fix it from the map. Free MIT desktop app.
Leaving K8s platform engineering for an internal dev-tooling/CI-CD role
I am currently building a central "Cluster-as-a-Service" platform nd deal with automated cluster provisioning, multi-tenant setup, observability work, kubernetes security etc. I also contribute upstream to K8s-sigs. The company I work for is rather unknown and I would love to be at a bigger brand with a real engineering culture. (I have gotten a few faang interviews in the last month, but declined one myself and did not get more offers).
I now have an offer at a well-known fintech where I would work on owning CI/CD pipeline templates (GitHub Actions/Jenkins), and some more terraform templates plus policy work. Dev teams manage their own infra/CI-CD day-to-day, the team has no direct Kubernetes/Linux/scaling ownership and everything's serverless. However it is a well-known tech brand and I would get 35% more salary (which bumps my salary to slightly above market, not much tho).
My worry: My long-term goal is deep infra/platform roles (K8s internals, distributed systems, SRE-adjacent work). I'm concerned 1-2 years here erodes my skills in that area and I navigate myself into a build tooling and devex niche? What is your take on that move? Would I limit myself too much or can I easily move back to infra platform work later, i.e., would owning CI/CD/tooling be seen as equivalent experience to K8s internals, distributed systems, and observability work when I try to move back later?
https://redd.it/1uoatjd
@r_devops
I am currently building a central "Cluster-as-a-Service" platform nd deal with automated cluster provisioning, multi-tenant setup, observability work, kubernetes security etc. I also contribute upstream to K8s-sigs. The company I work for is rather unknown and I would love to be at a bigger brand with a real engineering culture. (I have gotten a few faang interviews in the last month, but declined one myself and did not get more offers).
I now have an offer at a well-known fintech where I would work on owning CI/CD pipeline templates (GitHub Actions/Jenkins), and some more terraform templates plus policy work. Dev teams manage their own infra/CI-CD day-to-day, the team has no direct Kubernetes/Linux/scaling ownership and everything's serverless. However it is a well-known tech brand and I would get 35% more salary (which bumps my salary to slightly above market, not much tho).
My worry: My long-term goal is deep infra/platform roles (K8s internals, distributed systems, SRE-adjacent work). I'm concerned 1-2 years here erodes my skills in that area and I navigate myself into a build tooling and devex niche? What is your take on that move? Would I limit myself too much or can I easily move back to infra platform work later, i.e., would owning CI/CD/tooling be seen as equivalent experience to K8s internals, distributed systems, and observability work when I try to move back later?
https://redd.it/1uoatjd
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
When does an on-premises server become cheaper than AWS, Azure, or a VPS?
Everyone talks about moving to the cloud, but is it always the right choice?
For a small business with stable workloads (website, email, file storage, backups, internal apps), does buying a server once and running it for 5–7 years make more financial sense than paying cloud bills every month?
I'm also thinking from a business perspective. If you were starting a business today with a small budget, what service would you offer that brings recurring monthly income?
I'd love to hear real experiences from people who run infrastructure or do businesses from home network not just theory or marketing.
https://redd.it/1uobn4g
@r_devops
Everyone talks about moving to the cloud, but is it always the right choice?
For a small business with stable workloads (website, email, file storage, backups, internal apps), does buying a server once and running it for 5–7 years make more financial sense than paying cloud bills every month?
I'm also thinking from a business perspective. If you were starting a business today with a small budget, what service would you offer that brings recurring monthly income?
I'd love to hear real experiences from people who run infrastructure or do businesses from home network not just theory or marketing.
https://redd.it/1uobn4g
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
GitLab CI skill for ai agents based on official docs
I use ai agents as helper I talk to, not for blind vibecoding. One thing I kept noticing is asking agent to write or refactor gitlab ci pipeline, and results are often questionable. It creates a god yaml, outdated keywords, no thought about debugging or developer experience.
I looked for existing skills but did not find anything I would actually trust, most looked generated in one shot. So I spent some time and made my own. Used agent help of course, but went through everything myself and checked it against official docs for GitLab 18+
It covers pipeline structure and refactoring, bash in ci jobs, pipelines and other common patterns, debugging failed pipelines, readable logs and naming and many other cases
https://github.com/beeyev/skills/
Works with claude code and anything supporting skills format
I have been using it privately for couple of month and improving constantly, maybe it will useful for someone else too
https://redd.it/1uoab35
@r_devops
I use ai agents as helper I talk to, not for blind vibecoding. One thing I kept noticing is asking agent to write or refactor gitlab ci pipeline, and results are often questionable. It creates a god yaml, outdated keywords, no thought about debugging or developer experience.
I looked for existing skills but did not find anything I would actually trust, most looked generated in one shot. So I spent some time and made my own. Used agent help of course, but went through everything myself and checked it against official docs for GitLab 18+
It covers pipeline structure and refactoring, bash in ci jobs, pipelines and other common patterns, debugging failed pipelines, readable logs and naming and many other cases
https://github.com/beeyev/skills/
Works with claude code and anything supporting skills format
I have been using it privately for couple of month and improving constantly, maybe it will useful for someone else too
https://redd.it/1uoab35
@r_devops
GitHub
GitHub - beeyev/skills: A collection of Agent Skills for Claude Code and other SKILL.md-compatible agents.
A collection of Agent Skills for Claude Code and other SKILL.md-compatible agents. - beeyev/skills
How are you handling query rewrites and schema changes in production databases?
I'm curious how teams are dealing with query performance over time as applications evolve.
A few questions I'd love to hear your experience on:
\\- How do you identify queries that need to be rewritten?
\\- How do you ensure schema changes (indexes, partitions, column changes, etc.) don't negatively impact production?
\\- Is this mostly a manual process, or do you use any tools to analyze query patterns and recommend improvements?
\\- Have you ever had an outage or major performance issue because a query or schema change wasn't optimized?
I'm exploring this space and trying to understand whether this is a pain point worth solving. I'd really appreciate hearing about your workflows, the tools you use, and what's still frustrating today.
https://redd.it/1uoewbr
@r_devops
I'm curious how teams are dealing with query performance over time as applications evolve.
A few questions I'd love to hear your experience on:
\\- How do you identify queries that need to be rewritten?
\\- How do you ensure schema changes (indexes, partitions, column changes, etc.) don't negatively impact production?
\\- Is this mostly a manual process, or do you use any tools to analyze query patterns and recommend improvements?
\\- Have you ever had an outage or major performance issue because a query or schema change wasn't optimized?
I'm exploring this space and trying to understand whether this is a pain point worth solving. I'd really appreciate hearing about your workflows, the tools you use, and what's still frustrating today.
https://redd.it/1uoewbr
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Weekly Self Promotion Thread
Hey r/devops, welcome to our weekly self-promotion thread!
Feel free to use this thread to promote any projects, ideas, or any repos you're wanting to share. Please keep in mind that we ask you to stay friendly, civil, and adhere to the subreddit rules!
https://redd.it/1uopcqy
@r_devops
Hey r/devops, welcome to our weekly self-promotion thread!
Feel free to use this thread to promote any projects, ideas, or any repos you're wanting to share. Please keep in mind that we ask you to stay friendly, civil, and adhere to the subreddit rules!
https://redd.it/1uopcqy
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Interview prep for devops
​
JD says devops, python and linux developer.
I just want focus on scripting and ignore oops.
Is that possible?
Does this heavily depends on project requirements and company?
How did you clear devops interview?
Any suggestions?
My focus area is systems emgineering.
https://redd.it/1uot5eb
@r_devops
​
JD says devops, python and linux developer.
I just want focus on scripting and ignore oops.
Is that possible?
Does this heavily depends on project requirements and company?
How did you clear devops interview?
Any suggestions?
My focus area is systems emgineering.
https://redd.it/1uot5eb
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Help with Devsecops pipeline setup
I am a pentester. My manager had given me a task for cicd integration with checkmarx. I can understand the basic stuff but I am unable to come up with material for the following:
1. How do manage secrets in the pipeline (someone suggested me aws kms)
2. How to run authenticated scans (apps using okta) using pipeline
3. Capturing traffic to run the scans on them.
I would appreciate if someone can help me in this scenario as my job depends on it.
https://redd.it/1uotw74
@r_devops
I am a pentester. My manager had given me a task for cicd integration with checkmarx. I can understand the basic stuff but I am unable to come up with material for the following:
1. How do manage secrets in the pipeline (someone suggested me aws kms)
2. How to run authenticated scans (apps using okta) using pipeline
3. Capturing traffic to run the scans on them.
I would appreciate if someone can help me in this scenario as my job depends on it.
https://redd.it/1uotw74
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
5 YOE in DevOps/SRE at a Tier-1 bank, pivoted after a layoff — now in a support-heavy role. How do I course-correct before it hurts my career?
Hey r/devops — not looking for validation here, looking for people who’ve actually been through this.
The background:
Spent 5 years as a DevOps/SRE engineer at a large global investment bank. Real scope — owned incident response, SLO/SLI definition, Python automation, Kubernetes deployments, Grafana/InfluxDB observability stacks, CI/CD pipelines. High-stakes production trading systems. The kind of work I was proud of.
Earlier this year, my team was restructured. Whole group gone. I was laid off.
The decision I made:
The market was rough. I had a contract offer on the table — onsite at another major financial institution, compensation was actually better than my previous role. I took it. Practically speaking, it was the right call at the time.
What I underestimated: the actual scope of the work. It’s L1-L2 support. Ticket resolution, runbook execution, escalations. I’m not owning anything. I’m not building anything. And every month I stay here, the gap between my resume and my current reps gets wider.
I’m not bitter about the pay — it’s kept me stable. But I know this lane leads nowhere if I stay in it too long.
Three things I genuinely need help thinking through:
1. How do I keep building real SRE depth when my day job isn’t giving me the reps?
I have a solid foundation — Docker, Kubernetes, Terraform, Ansible, Python, Grafana. But I want to go deeper on observability, chaos engineering, capacity planning, and SRE fundamentals. What projects, home labs, or open-source contributions have actually made you better — not just resume filler?
2. Is the AWS Generative AI Professional cert (AIP-C01) worth pursuing for an SRE career path?
I’m currently working toward it, specifically because I’m interested in AIOps — anomaly detection, automated remediation, predictive reliability. Is this a real differentiator for senior SRE roles in 2025-26, or am I chasing the wrong thing? What certs have actually moved the needle for you?
3. How do I frame a support-heavy role in interviews without it undercutting 6 years of solid experience?
My instinct is to move within 6–9 months before the narrative gets harder to control. Is that the right timeline? And how do you position this kind of pivot when you’re targeting L4/L5 SRE or Senior Platform Engineer roles?
Context: Based in Bengaluru. Targeting product companies, global banks, and open to remote/international opportunities.
I know I made a compromise. I’m not here to relitigate that. Just want to make sure the next move is the right one.
Appreciate any honest perspectives — especially from people who’ve navigated something similar. 🙏
https://redd.it/1uoydmi
@r_devops
Hey r/devops — not looking for validation here, looking for people who’ve actually been through this.
The background:
Spent 5 years as a DevOps/SRE engineer at a large global investment bank. Real scope — owned incident response, SLO/SLI definition, Python automation, Kubernetes deployments, Grafana/InfluxDB observability stacks, CI/CD pipelines. High-stakes production trading systems. The kind of work I was proud of.
Earlier this year, my team was restructured. Whole group gone. I was laid off.
The decision I made:
The market was rough. I had a contract offer on the table — onsite at another major financial institution, compensation was actually better than my previous role. I took it. Practically speaking, it was the right call at the time.
What I underestimated: the actual scope of the work. It’s L1-L2 support. Ticket resolution, runbook execution, escalations. I’m not owning anything. I’m not building anything. And every month I stay here, the gap between my resume and my current reps gets wider.
I’m not bitter about the pay — it’s kept me stable. But I know this lane leads nowhere if I stay in it too long.
Three things I genuinely need help thinking through:
1. How do I keep building real SRE depth when my day job isn’t giving me the reps?
I have a solid foundation — Docker, Kubernetes, Terraform, Ansible, Python, Grafana. But I want to go deeper on observability, chaos engineering, capacity planning, and SRE fundamentals. What projects, home labs, or open-source contributions have actually made you better — not just resume filler?
2. Is the AWS Generative AI Professional cert (AIP-C01) worth pursuing for an SRE career path?
I’m currently working toward it, specifically because I’m interested in AIOps — anomaly detection, automated remediation, predictive reliability. Is this a real differentiator for senior SRE roles in 2025-26, or am I chasing the wrong thing? What certs have actually moved the needle for you?
3. How do I frame a support-heavy role in interviews without it undercutting 6 years of solid experience?
My instinct is to move within 6–9 months before the narrative gets harder to control. Is that the right timeline? And how do you position this kind of pivot when you’re targeting L4/L5 SRE or Senior Platform Engineer roles?
Context: Based in Bengaluru. Targeting product companies, global banks, and open to remote/international opportunities.
I know I made a compromise. I’m not here to relitigate that. Just want to make sure the next move is the right one.
Appreciate any honest perspectives — especially from people who’ve navigated something similar. 🙏
https://redd.it/1uoydmi
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
I'm starting a new movement
I am officially declaring the start (in my mind) of #MRBA
That stands for "Make Releases Boring Again"
This was prompted by a Release Engineer job posting that was your usual "just be on 24/7 on every communication channel during release windows". So every few months, you over activate my nervous system and it takes until the next release for it to finally calm down only to be activated again? No thanks.
I need to be doing automation, environment config hardening, observability tweaking. Not "monitoring Slack in case someone reports an issue". 😒
Releases need to be boring. The more boring, the more both dev AND ops sleep. With the added bonus of not over-rewarding heroics. 😏
Release day hype/fanfare/stress is for shit like clothing, games, etc. Not the newest feature for your internal app with 10 users.
https://redd.it/1up28y5
@r_devops
I am officially declaring the start (in my mind) of #MRBA
That stands for "Make Releases Boring Again"
This was prompted by a Release Engineer job posting that was your usual "just be on 24/7 on every communication channel during release windows". So every few months, you over activate my nervous system and it takes until the next release for it to finally calm down only to be activated again? No thanks.
I need to be doing automation, environment config hardening, observability tweaking. Not "monitoring Slack in case someone reports an issue". 😒
Releases need to be boring. The more boring, the more both dev AND ops sleep. With the added bonus of not over-rewarding heroics. 😏
Release day hype/fanfare/stress is for shit like clothing, games, etc. Not the newest feature for your internal app with 10 users.
https://redd.it/1up28y5
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
I built a single-binary TUI that manages Redis, Postgres, SSH, Docker, Git, S3, MySQL, MongoDB, and HTTP — with a built-in MCP server for AI tooling
**Qore** is a single-binary infrastructure orchestrator with a terminal-native UI. You type commands, get results inline — no context switching between redis-cli, psql, ssh, and docker. What makes it different: **8 connection types in one place:**
* Redis (native RESP protocol — no redis-cli needed)
* PostgreSQL / MySQL / MongoDB (full SQL queries, EXPLAIN, slow queries, CSV export)
* S3-compatible (AWS SigV4 — works with MinIO, R2, AWS)
* HTTP/REST (GET, POST, PUT, PATCH, DELETE with auth)
* SSH (exec, SFTP, systemd, Docker Compose, deploy scripts, interactive shell)
* Git (branch graph, merge, rebase, cherry-pick, blame, tags)
**Built-in MCP server:** This is the part I'm most excited about. It exposes 35 tools (SSH, Docker, database queries, system discovery, HTTP) over JSON-RPC 2.0 — so Claude, Cursor, or any MCP-compatible AI can interact with your infrastructure using connection names only. Credentials stay server-side. **Other highlights:**
* Secure vault: AES-256-GCM + scrypt, master password never touches disk
* Docker via Unix socket (no docker CLI dependency)
* Multi-tab: all connections stay mounted, switch with Ctrl+Tab
* Multi-service dashboard with auto-refresh
* Health checks with latency sparklines
* Self-updating (`qore update`)
* Single binary, \~45MB, Linux/macOS/Windows
Install: `curl -fsSL` [`https://github.com/Kodjaoglanian/qore/releases/latest/download/install.sh`](https://github.com/Kodjaoglanian/qore/releases/latest/download/install.sh) `| bash` Code: [https://github.com/Kodjaoglanian/qore](https://github.com/Kodjaoglanian/qore)
Happy to answer questions!
https://redd.it/1up1rh3
@r_devops
**Qore** is a single-binary infrastructure orchestrator with a terminal-native UI. You type commands, get results inline — no context switching between redis-cli, psql, ssh, and docker. What makes it different: **8 connection types in one place:**
* Redis (native RESP protocol — no redis-cli needed)
* PostgreSQL / MySQL / MongoDB (full SQL queries, EXPLAIN, slow queries, CSV export)
* S3-compatible (AWS SigV4 — works with MinIO, R2, AWS)
* HTTP/REST (GET, POST, PUT, PATCH, DELETE with auth)
* SSH (exec, SFTP, systemd, Docker Compose, deploy scripts, interactive shell)
* Git (branch graph, merge, rebase, cherry-pick, blame, tags)
**Built-in MCP server:** This is the part I'm most excited about. It exposes 35 tools (SSH, Docker, database queries, system discovery, HTTP) over JSON-RPC 2.0 — so Claude, Cursor, or any MCP-compatible AI can interact with your infrastructure using connection names only. Credentials stay server-side. **Other highlights:**
* Secure vault: AES-256-GCM + scrypt, master password never touches disk
* Docker via Unix socket (no docker CLI dependency)
* Multi-tab: all connections stay mounted, switch with Ctrl+Tab
* Multi-service dashboard with auto-refresh
* Health checks with latency sparklines
* Self-updating (`qore update`)
* Single binary, \~45MB, Linux/macOS/Windows
Install: `curl -fsSL` [`https://github.com/Kodjaoglanian/qore/releases/latest/download/install.sh`](https://github.com/Kodjaoglanian/qore/releases/latest/download/install.sh) `| bash` Code: [https://github.com/Kodjaoglanian/qore](https://github.com/Kodjaoglanian/qore)
Happy to answer questions!
https://redd.it/1up1rh3
@r_devops
Need guidance !
I have 3 yrs of experience as a cloud engineer/devops worked in same company from start got one promotion and good hike all along . Want to switch to now and thinking on doing CKA certification is it worth having ?
https://redd.it/1up4s58
@r_devops
I have 3 yrs of experience as a cloud engineer/devops worked in same company from start got one promotion and good hike all along . Want to switch to now and thinking on doing CKA certification is it worth having ?
https://redd.it/1up4s58
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
How are you managing dynamic runtime configuration without triggering a full CI/CD deployment?
Hi all!
I’m currently digging into the "toil" of our (the company i work at's) release process, and I’m hitting a recurring bottleneck: rate-limiting configuration.
Right now, we have our limits (e.g., token buckets, thresholds) defined as part of our static config, which is baked into our container images. Whenever we need to tune these for a traffic spike or emergency throttle, it forces a full CI/CD deployment aka build, push, wait for rollout, and pray the new pod doesn't have a startup issue.
It feels fundamentally wrong to bounce a production service just to change a numerical threshold.
I’m looking into moving this "knob-turning" out of the deployment pipeline and into a centralized, runtime-synced store (like Redis), so we can tweak values on the fly without a code push.
Is anyone else using a "Config-as-a-Service" or dynamic sidecar pattern for this, or have we missed a super obvious solution lol thanks guys :)
https://redd.it/1up6pnr
@r_devops
Hi all!
I’m currently digging into the "toil" of our (the company i work at's) release process, and I’m hitting a recurring bottleneck: rate-limiting configuration.
Right now, we have our limits (e.g., token buckets, thresholds) defined as part of our static config, which is baked into our container images. Whenever we need to tune these for a traffic spike or emergency throttle, it forces a full CI/CD deployment aka build, push, wait for rollout, and pray the new pod doesn't have a startup issue.
It feels fundamentally wrong to bounce a production service just to change a numerical threshold.
I’m looking into moving this "knob-turning" out of the deployment pipeline and into a centralized, runtime-synced store (like Redis), so we can tweak values on the fly without a code push.
Is anyone else using a "Config-as-a-Service" or dynamic sidecar pattern for this, or have we missed a super obvious solution lol thanks guys :)
https://redd.it/1up6pnr
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
Odd manager behavior - looking for opinions
I've been working at a local startup as a DevOps engineer for the past year. I have a total of around 3 years of experience, which is considered mid-level at best. I've worked with various technologies and I'm not that confident in my technical skills but I do try my best to deliver. That's not my problem though. My problem is that communication with my manager is terrible.
I am rarely assigned any tickets. My work is a mixture of verbally assigned tasks with minimal details and initiatives I take to improve our workflows. My manager rarely joins our weekly one to one meetings so I decided to send him my task updates on teams. He doesn't reply to my messages, sometimes for days. He will only reply fast if he believes the question is important to him. He never reacts or responds to my task updates and won't review my work for months (if at all). Sometimes, when I ask him questions or try to make conversation on meetings he won't respond. When I ask again, he will laugh it off saying he heard me the first time. The kind of questions I ask are usually clarifications on my tasks, or discussion around my technical implementations.
In general, I believe I am fairly independent and I don't burden the team. My tasks are mine to deliver and I am solely responsible for them. I request guidance in a structured way and usually when I don't have enough information to move on. Even when I don't have enough context I will push through instead of waiting for weeks for an answer. The problem is that I am not responsible for the infrastructure architectural decisions, and I need the clarifications in order to do my job effectively.
I will give you a recent example:
I am asked to deploy one of our products to a test environment, but we want this product to be isolated from our other products so it's not the standard procedure.
That's pretty much what I got. I ask my manager to discuss for like 15minutes to show me what he deployed for production. He never does.
I create a list of resources I will need to deploy along with the infra design. I make the deployment. I share that with him. He shows up after a week of being AWOL to tell me some resources should be redeployed. I ask for two very specific clarifications in chat. He says we will discuss in a meeting. We join the meeting. I ask. He stays silent. I ask again. He says he heard me and he just didn't respond.
What am I supposed to do at this point? This whole thing is deeply demoralizing to me. I feel deeply disrespected and looked down upon. I am mad and sad and it's affecting my confidence and will to work and be creative and productive. I've tried a few different things since starting in this company, I've created my own tickets and shared them with him, I've tried texting him I've tried reaching out during meetings to avoid spamming him. Nothing seems to work. I dread going to work every single day. I feel lost about what to do next.
I want to stay professional but I also feel very done. I want them to fuck off but also I want to take technical experience. I want to quit but I know it's a terrible idea. What can I do to continue learning and growing within team and business goals when these are not communicated properly? What could I be doing wrong? Am I needy or is this truly as annoying and disfunctional as I think it is?
Thanks for reaching this far, any opinions, experiences or recommendations are welcome. I can take criticism as long as it is respectful, my mental health is declining fast enough already 😋
https://redd.it/1upabcd
@r_devops
I've been working at a local startup as a DevOps engineer for the past year. I have a total of around 3 years of experience, which is considered mid-level at best. I've worked with various technologies and I'm not that confident in my technical skills but I do try my best to deliver. That's not my problem though. My problem is that communication with my manager is terrible.
I am rarely assigned any tickets. My work is a mixture of verbally assigned tasks with minimal details and initiatives I take to improve our workflows. My manager rarely joins our weekly one to one meetings so I decided to send him my task updates on teams. He doesn't reply to my messages, sometimes for days. He will only reply fast if he believes the question is important to him. He never reacts or responds to my task updates and won't review my work for months (if at all). Sometimes, when I ask him questions or try to make conversation on meetings he won't respond. When I ask again, he will laugh it off saying he heard me the first time. The kind of questions I ask are usually clarifications on my tasks, or discussion around my technical implementations.
In general, I believe I am fairly independent and I don't burden the team. My tasks are mine to deliver and I am solely responsible for them. I request guidance in a structured way and usually when I don't have enough information to move on. Even when I don't have enough context I will push through instead of waiting for weeks for an answer. The problem is that I am not responsible for the infrastructure architectural decisions, and I need the clarifications in order to do my job effectively.
I will give you a recent example:
I am asked to deploy one of our products to a test environment, but we want this product to be isolated from our other products so it's not the standard procedure.
That's pretty much what I got. I ask my manager to discuss for like 15minutes to show me what he deployed for production. He never does.
I create a list of resources I will need to deploy along with the infra design. I make the deployment. I share that with him. He shows up after a week of being AWOL to tell me some resources should be redeployed. I ask for two very specific clarifications in chat. He says we will discuss in a meeting. We join the meeting. I ask. He stays silent. I ask again. He says he heard me and he just didn't respond.
What am I supposed to do at this point? This whole thing is deeply demoralizing to me. I feel deeply disrespected and looked down upon. I am mad and sad and it's affecting my confidence and will to work and be creative and productive. I've tried a few different things since starting in this company, I've created my own tickets and shared them with him, I've tried texting him I've tried reaching out during meetings to avoid spamming him. Nothing seems to work. I dread going to work every single day. I feel lost about what to do next.
I want to stay professional but I also feel very done. I want them to fuck off but also I want to take technical experience. I want to quit but I know it's a terrible idea. What can I do to continue learning and growing within team and business goals when these are not communicated properly? What could I be doing wrong? Am I needy or is this truly as annoying and disfunctional as I think it is?
Thanks for reaching this far, any opinions, experiences or recommendations are welcome. I can take criticism as long as it is respectful, my mental health is declining fast enough already 😋
https://redd.it/1upabcd
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community
How to structure scoring live traffic
We've had offline evals as part of our CI for a while now, but last month we got hit with something that none of our CI runs flagged. Our production inputs had drifted and users were asking things our test set just didn't cover, and output quality on that section of things had degraded for weeks. So unfortunately, our existing evals gave me a false sense of safety because they can only ever test what I thought to put in them.
So now I'm trying to figure out actually sampling and scoring real production responses, not just CI runs against a fixed dataset.
Main things I'm rubber ducking:
Sampling. Are people scoring all live traffic or some percentage?
Alerting. I want to know when quality drops on live traffic, but I’m not trying creat another annoying alert channel. And then if/when it goes off who/what owns the response?
https://redd.it/1upa4lk
@r_devops
We've had offline evals as part of our CI for a while now, but last month we got hit with something that none of our CI runs flagged. Our production inputs had drifted and users were asking things our test set just didn't cover, and output quality on that section of things had degraded for weeks. So unfortunately, our existing evals gave me a false sense of safety because they can only ever test what I thought to put in them.
So now I'm trying to figure out actually sampling and scoring real production responses, not just CI runs against a fixed dataset.
Main things I'm rubber ducking:
Sampling. Are people scoring all live traffic or some percentage?
Alerting. I want to know when quality drops on live traffic, but I’m not trying creat another annoying alert channel. And then if/when it goes off who/what owns the response?
https://redd.it/1upa4lk
@r_devops
Reddit
From the devops community on Reddit
Explore this post and more from the devops community