Reddit DevOps
287 subscribers
81 photos
32.5K links
Reddit DevOps. #devops
Thanks @reddit2telegram and @r_channels
Download Telegram
How can I deploy this microservices project for free without a credit card?

Hey everyone,

I'm building a graduation project with this architecture:

Frontend (Next.js)
↓
Traefik Gateway
├── Auth Gateway Service (nestjs)
├── Issues Service (nestjs)
├── Ingestion Service (nestjs)
├── User Service (nestjs)
└── SmartQA Vector Service (fastApi)

I also have the following infrastructure services:

Kafka for event-driven communication
MongoDB for issues and vector data
PostgreSQL for users
MySQL + Keycloak for authentication
Consul for service discovery
Traefik as the reverse proxy/API gateway

The backend services are deployed as separate Docker containers, and I'm currently looking for a way to deploy the whole application for free, if possible.

My main requirements are:

Free tier
No credit card required
Very low traffic — I'm essentially the only user, with at most 1–2 users
Support for Docker containers/multiple services
Ideally, persistent storage for MongoDB/PostgreSQL
I don't mind cold starts since this is mainly a graduation/portfolio project

What platforms or deployment approaches would you recommend for this kind of setup?

Thanks in advance!

https://redd.it/1wqpwys
@r_devops
Indian devs who self-host tools: what actually becomes painful after 6–12 months?

I’m curious about the long-term side of self-hosting from developers in India who actually run things in production.

Setting up a VPS + Docker Compose is usually the easy part.

What I’m more interested in is what starts becoming annoying after a few months:

backups and restore testing
upgrades breaking something
monitoring/alerts
disk usage
SSL/DNS issues
database maintenance
security patches
managing multiple servers
documenting setups so someone else can understand them
getting paged when something dies

Especially for small startups/agencies where there isn’t a dedicated DevOps/SRE person.

At what scale did self-hosting stop feeling “cheap” for you?

And which tools do you still prefer to self-host versus paying for a managed version?

Would be useful to hear actual production experiences rather than just setup tutorials.

https://redd.it/1wr1vsk
@r_devops
How are people verifying unattended/CI agent runs actually did what they claimed?

Running into something I'm curious how others handle: when a coding agent runs unattended (CI, scheduled job, autonomous pipeline), the only account of what happened is the agent's own closing summary. Nobody is watching in real time, and if the summary is wrong or incomplete, there is no independent way to know without going back and manually digging through logs after the fact.

Concrete case that got me thinking about this: had a subagent's tests genuinely fail during a run, and the main agent's final summary still said all tests pass, without mention of the failure. Would have gone completely unnoticed in an unattended context, nobody was reading the transcript live, and the summary gave no reason to suspect anything was wrong.

I am curious what people are actually doing here such as deterministic post-run checks, separate log validation, something else? Are you treating the agent's own report as trustworthy at all in these contexts, or building around the assumption that it might not be?

(For context: this is what pushed me to build a small open-source tool called Rashomon that independently records what actually happened during a Claude Code session, separate from the agent's own account. Happy to share details if it's relevant to what people are describing, but mainly want to hear how others are approaching this.)

https://redd.it/1wr2lng
@r_devops
How to maintain helm charts of multiple micro-services on k8s at scale ?

hey everyone,

how do you guys manage helm charts for multiple micro services.

we have multiple microservices and our CI-CD flow is when the docker image is built and pushed to a registry then the values file will be updated for the image tag and this change is detected by the gitops tools like argo or flux to deploy.

\- My question is we are having these different templates i.e a helm chart like below for each and every service and is there a way to handle this at scale with something like DRY ?

\- Is it possible to maintain a unified helm chart and then have different chart and values files per service to reduce the duplication and maintenance ?

\- and does kustomize helps in any ways to solve this ?

\- how are you maintaining multiple microservices on k8s at scale ?

any insights or suggestions will be highly appreciated.

https://redd.it/1wr7mw9
@r_devops
How We Use Backstage in OpenChoreo Join Our CNCF Community Call

Hey everyone I’m part of the OpenChoreo community.

OpenChoreo is an open source Internal Developer Platform and a CNCF Sandbox project, and for this month’s community call we’re putting the spotlight on Backstage.

A big part of the session is based on our experience building OpenChoreo around Backstage and the questions we’ve seen from teams already using it.

We’ll cover:

* Why OpenChoreo chose Backstage in the first place
* The gaps we encountered when moving from a developer portal toward a complete Internal Developer Platform
* What OpenChoreo adds around Backstage
* How an existing Backstage user can migrate to OpenChoreo
* A live code walkthrough
* How to customize the OpenChoreo Backstage portal
* What we’re working on next

We want this to be useful for people already building with Backstage or trying to figure out how it fits into a broader platform engineering architecture.

If you’re working on Backstage, Kubernetes platforms, Developer Experience, or Internal Developer Platforms, we’d love to have you join and bring your questions.

Community call details + registration:
[https://github.com/openchoreo/openchoreo/discussions/4689](https://github.com/openchoreo/openchoreo/discussions/4689)

Happy to answer questions here as well.

https://redd.it/1wr9j2o
@r_devops
Resume review for devops/cloud engineer roles
https://redd.it/1wrfpsl
@r_devops
Unorganized Knowledge = Noise

Whenever I try learning new technology or troubleshooting without a clear plan, I get overwhelmed. Hopping between random videos, articles, and quick AI summaries leaves me feeling lost and stressed about where to begin or end.

I've realized that having a structured roadmap with clear, step-by-step guidance changes everything. It removes the stress and makes the whole learning process much smoother.

Anyone else feel this way? I’d love to hear how you organize your learning!

https://redd.it/1wrh025
@r_devops
We audited the default Helm charts and Docker configs of popular AI stack tools (Ray, Weaviate, MCP servers, LangChain). The out-of-the-box security defaults are surprisingly bad.

Over the past few weeks, we set up a fresh cluster to test what actually gets deployed when an engineer runs helm install with zero extra configuration on popular AI, vector DB, and MCP charts (Ray, vLLM, LiteLLM, Qdrant, Weaviate, KubeRay, etc.).

Everyone focuses on prompt injection and model jailbreaks, but the platform-level defaults on these workloads look like Kubernetes five to seven years ago:

1. Ray / KubeRay (Distributed Compute)
- The default deployment accepts unauthenticated job submissions over internal cluster networking.
- Worker containers run with passwordless sudo out of the box.
- Any compromised pod on the flat cluster network that can reach the Ray API gets arbitrary remote execution as root with zero barrier.

2. MCP (Model Context Protocol) Servers
- Several popular MCP server implementations implement local protection solely by checking Host: localhost.
- Any basic SSRF or internal proxy passing -H "Host: localhost" bypasses this immediately.
- In live testing, one Kubernetes MCP server mounted a ClusterRole with cluster-wide Secret read and pod exec permissions, exposed over unauthenticated HTTP.

3. LiteLLM Migration Job Credential Leak
- The primary deployment pulls the database password from a Secret using secretKeyRef, but the database migration Job embeds the raw database password in plain-text environment variables.
- Anyone with basic cluster read permissions (devs, CI service accounts, dashboards) can run kubectl get job -o yaml and extract the database password directly from the env block, even without permission to read Secrets.

4. Vector DBs & RAG Stores
- Multiple upstream charts still ship with default values that enable anonymous read and write access.
- Raw embeddings and ingested document payloads are queryable by any internal service without an authentication header.

5. Scanner Blindspots
- 14 of 15 charts provide no default NetworkPolicy objects.
- All 15 charts leave Kubernetes' default automountServiceAccountToken: true enabled.
- Standard static linters (Checkov, Trivy, Kubescape) missed the containers nested inside Custom Resource Definitions like Ray clusters.

The takeaway for platform and security teams:
The AI supply chain conversation is hyper-focused on prompts, but your teams might be deploying unauthenticated root execution endpoints right now. If your cluster is not enforcing Kyverno or OPA gatekeeper admission rules, mandatory NetworkPolicies, and dropping default service account tokens on AI namespaces, the internal perimeter is wide open.

Curious what other teams are doing here. Have you written custom Kyverno/OPA policies or wrapper charts for AI workloads at your org, or are you putting everything behind a service mesh like Istio/Linkerd?

Full write-up, reproduction manifests, and probe logs: https://sorami.com.au/research/ai-kubernetes-helm-chart-security/
Raw probe logs and reproduction repository: https://github.com/Sorami-Consulting-AU/ai-kubernetes-helm-chart-security

https://redd.it/1wrkllb
@r_devops
What's the most annoying "works locally, dies in production" bug you've had?

Not talking big architecture screwups — but rather the small, stupid ones. Things like hardcoded ports, missing env defaults, path separators, whatever. What was it, what stack, and how'd you find it? Would love to hear any feedback. Thanks.

https://redd.it/1ws1iip
@r_devops
Weekly Self Promotion Thread

Hey r/devops, welcome to our weekly self-promotion thread!

Feel free to use this thread to promote any projects, ideas, or any repos you're wanting to share. Please keep in mind that we ask you to stay friendly, civil, and adhere to the subreddit rules!

https://redd.it/1ws736h
@r_devops
Diagram solutions for git / markdown docs and LLM?

Has anyone here got a good solution for maintaining decent diagrams in markdown docs that are being handled with LLM assistance?

Mermaid is the obvious initial answer as they tend to self render, but they're so ugly, right? I don't think I'm missing anything there, but would be keen to know if I might.

So past them I was looking at the python diagrams library. Nice output but quite codey, no fully abstracted definition language, too much python code ingrained in it to feel like it's elegant enough to have LLM's regenerate the output by running a script that way, although I'm currently trying it out.

After that I see Schematex, which, like mermaid, has a proper definition language, and also has an MCP server and other options around it, but I've not yet tried to set up a workflow around that that would see it in action.

Is there a logic in going directly to raw svg?

Are there other options / strategies people are using to keep on top of diagrams? On caveat / excuse I have currently is that I'm mostly doing design work for proposals etc., so 100% accuracy isn't critical yet, so LLM's can have leeway for making things up at times!

https://redd.it/1ws8aqe
@r_devops
What's your default database these days, and why? Yearly survey (mod-approved)

Hey, Robert from MariaDB Foundation here. We are running a survey on how MariaDB is being used, or cases where you went with MySQL, PostgreSQL or another database.

https://mariadb.typeform.com/to/tS69UzVo?utm\_source=reddit\_devops

I split the questions by role this year, so you can skip straight to the operator ones: HA, database versions, upgrade frequency, hosting, etc. Overall I'm curious to know, what's your default database these days, and what made you land on it? Was it your choice, the app or service you run, or whoever set it up before you?

Under 10 minutes, anonymous, results published at mariadb.org/survey.

Thanks for taking the time!
Robert, MariaDB Foundation

https://redd.it/1ws9qlg
@r_devops
Experiences with Signadot

Has anyone here used Signadot in a larger project and had any success with it for ephemeral environments/preview environments? I have been looking into tooling for ephemeral environments after maintaining our own solution for a while and this one keeps popping up as an alternative, so I was hoping to get a few opinions.

https://redd.it/1wse81n
@r_devops
How I cut non-prod cluster costs by scheduling Kubernetes/OpenShift scale-down (with the approach)

Dev and test environments sit idle nights and weekends but still burn compute. Here's how I automated scaling them down and back up on OpenShift/Kubernetes.

The problem: Manual shutdowns were forgotten, and teams complained when environments weren't ready in the morning.

The approach:

Tag namespaces/workloads with a schedule label

A scheduled job scales deployments to 0 at night and restores the previous replica counts in the morning

Exclusions for anything that must stay up, plus alerting if a scale-up fails

Result: [€18572 to €15182\]

https://redd.it/1wsiuzq
@r_devops
Guyssss, so news is 150% hike with mind-blowing dilemma

Interview next week for a role that could potentially give me a 150% hike. Location ? 7 min away by my activa, mode hybrid ( 2 days office, 3 days WFH) and and and office is in a mall you know, like very fancy, tech stack (exactly matching, 3 years of experience they were asking for but selecting me with 1 year of experience) 😎 , not lalu company ,I mean it has 500 employeessss.

so so

Naturally, I checked the company before getting too excited (I was excited already)

And then I found:

Micromanagement.Burnout.Favouritism.Poor work-life balance.

And apparently a whole lot more.

Now I’m sitting here wondering:

Is a 150% hike worth the risk of a toxic work environment?, should I give my 200 percent and grab a offer for the sake of counter

The interview is happening anyway, I mean I will definitely gain confidence while preparing and it's gonna help me with other interviews.

Seniors , help me with your suggestions as it's my first switch and I'm getting unnecessarily excited 🤦

https://redd.it/1wsmzw3
@r_devops
Prioritisation/ Maturity Interview Question

Ok so I flubbed the most basic interview question there is:

"You have a number of pieces of work in flight. How do you choose between them?"

Basically, they are testing your ability to prioritise work.

I tried to answer philosphically by telling them going for a walk was a good idea but my broader point is that it shouldn't have ever gotten to the point where you have loads of pieces of work in flight at any one time.


My question is twofold: what's the accepted 'best' interview answer to this question and how do you deal with this issue in your real life actual job?

https://redd.it/1wsst19
@r_devops
How do you set LLM budgets per team?

We're about 25 people. Engineering, support, and marketing all use llms through the API, and right now it's one shared key with no visibility. Last month someone's test script burned a painful chunk of the budget over a weekend.

ideal world:

\- A separate key per team or project, each with a hard limit

\- A live view of who's spending what, on which models

\- To avoid becoming the person who approves every request

I'm leaning toward a gateway instead of building internal tooling. llmapi.ai has per-key limits and usage analytics, and LiteLLM can do something similar if you self-host it. Has anyone rolled out either one company-wide? What's worked without adding a pile of process?

https://redd.it/1wt1llw
@r_devops
Need guidance on running scripted browser monitoring in production

I've​ been doing monitoring for our web apps for a while and I'm pretty comfortable with HTTP/API checks but scripted browser monitoring is still new territory for me.

A few things I'm trying to understand:

\- Which flows are worth monitoring with a real browser vs just using HTTP checks?

\- ​How do you handle login session and MFA without creating a security hole?

\- How do you keep recorded scripts working accross frontend deploys?

\- How many location and what frequency actually means something vs just noise?

\- How do you tell the difference between an actual app failure and a broken monitoring script?

\- ​Whats goes to on-call vs what just gets logged?

AI is great for explaining how all of this should work, but there's a pretty big gap between that and how people actually run it in production.

Anyone doing scripter browser monitoring regularly? I'm not looking for someone to design it for me, just want to know how people approach it in real setups

https://redd.it/1wt0v99
@r_devops
Your Ray cluster probably opens more ports than its pod spec says (15 of 17 in our test)

We deployed default KubeRay + vLLM on one EKS cluster. 15 of 17 Ray listening sockets were not in the declared container ports, and a pod in another namespace could reach Ray GCS (6379) and raylet RPC (10002-10006) over cleartext gRPC. Four default scanner configurations did not inspect the RayCluster resource.

An ingress default-deny NetworkPolicy blocked every probed Ray port. That was the cheapest fix we measured.

Practical steps: https://sorami.com.au/guides/securing-ray-vllm-on-kubernetes/



Raw data: https://github.com/Sorami-Consulting-AU/distributed-llm-inference-hidden-network




https://redd.it/1wt5kyb
@r_devops