Reddit DevOps
278 subscribers
69 photos
32.2K links
Reddit DevOps. #devops
Thanks @reddit2telegram and @r_channels
Download Telegram
Automated test tool to help in CI/CD

I am new to this CI/CD thing. I know the concept is not new and it is as old as computer science itself but the buzz is new. I learned that for CI to work, you need some kind of automated tool that will test your new code and will tell if it passes or not so you can integrate it with main branch.

My question is can you point me to any such kind of tools that will automate this task? I mean let's say I am working on a project which has three modules; Login, Register, Generate Invoice. If I write a code for login, register and generating invoice than how will that tool be able to test this functionality?

https://redd.it/ed98t5
@r_devops
Migrating from Amazon Web Services (AWS) to Google Cloud Platform (GCP)

Here's a detailed article on how our DevOps team reduced downtime to a mere fraction and could introduce new features faster. [Give it a read.](https://blog.hike.in/migration-to-google-cloud-platform-gcp-b92036429a22)

https://redd.it/ed7349
@r_devops
Are there any good books (or book chapters) or articles on the theory and maths for performance testing?

Most load testing tools produce reports with mean, max, and percentile details. For example [this](https://jmeter.apache.org/usermanual/generating-dashboard.html) is a JMeter report.

But:

* How should we interpret the details in this report?
* What would be the success criteria?
* How can I make sure that my test is actually relevant and capturing the right details?

Are there any good books, papers, or articles that cover these topics?

https://redd.it/ed1p4c
@r_devops
I purchased a domain name from name.com and hosting my site at aws lightsail, where could I configure a forwarding email address?

Say I want an email address like [email protected], which only forwards it to my other account on gmail, do I need to set it up with name.com or with aws? Since AWS is providing nameserver?


Solution : If anyone else, faces the same issue, I used yandex mail service, and added mx record, spf record, and dkim signatures in lisghtsail's domain configuration.

https://redd.it/eczhuh
@r_devops
RollingUpdate Deployment on AKS is killing all POD's

Hi, i'm a DevOps Engineer and i have a issue with the RollingUpdate Deployment on AKS (Azure). When I try to execute a a pipeline in the stage of deploy that execute a RollingUpdate Strategy deployment on AKS the process kill all the POD's.

With this code:

`spec:`
`progressDeadlineSeconds: 60`
`strategy:`
`type: RollingUpdate`
`rollingUpdate:`
`maxUnavailable: 1`
`maxSurge: 1`

​

I see this topic on Stack but I didn't find anything that works
[https://stackoverflow.com/questions/46369100/kubernetes-rolling-update-killing-off-old-pod-without-bringing-up-new-one](https://stackoverflow.com/questions/46369100/kubernetes-rolling-update-killing-off-old-pod-without-bringing-up-new-one)

https://redd.it/ecv3qm
@r_devops
Things to have in a resume for someone who wants to switch from developer to DevOps engineer

I have worked as a developer for more than 3 years. For the last 1 year I have been working with AWS, learning architecture and CICD in cloud. I will go for the solutions architect associate exam very soon. Now I want to apply for DevOps related jobs as I find it more interesting. I want to know what are the dos and don'ts for the resume. Thanks in advance. 🙂

https://redd.it/ecoy41
@r_devops
Serverless approach or Kubernetes Cluster Approach for Banking System

Hi , I am new to serverless architecture. I am about to create a banking application. Is it a good idea to develop my whole system using serverless architecture (using AWS SAM) or should it be a combination of both Serverless and Kubernetes Cluster. Thanks in advance

https://redd.it/ecor5c
@r_devops
Using Node, Git and SFTP for automated deployments

I plan to write a Node script that finds the changed files with Git and then deploys them to the server with SFTP. Is this an okay approach?

https://redd.it/ecokx7
@r_devops
I Want to learn Linux Administration

Hi everyone

I would like to learn Linux administration. i am very new to this.. I have already started with Linux Bible Ninth Edition. In the end I would like to get certified.

Can the experts in the community can help me as in any other books should i refer to.

​

Thank you

Sumesh

https://redd.it/ecohy7
@r_devops
Silly question: What exactly should I do if my build fails in CI/CD process?

So I am a newbie who implemented CI/CD using github actions for a node project. I pushed code which will fail tests and cause a build failure on push to github but the problem is my changes are committed to the branch even if the build failed. Is this expected behavior? How should I handle my build failures?

https://redd.it/efmtep
@r_devops
Issues connecting to consul installed via Helm

I have a Kubernetes cluster that was created using KOPS which I have consul installed on via the helm chart with default values except server count and bootstrapexpect are both set to 1 because of this being a small test cluster - 2 workers and single master

I have a Node.js application that I am trying to connect to Consul. Below are the logs for what I think is relavent information, if I am missing something useful ask and I can provide it likely.

​

Error message given by my applcation seen by doing kubectl logs <pod>

{"message":"Error: Error: connect ECONNREFUSED 10.0.32.42:8500","level":"error"}
/opt/node_app/app/middleware/authorization.js:13
throw err;
^

Error: connect ECONNREFUSED 10.0.32.42:8500
at TCPConnectWrap.afterConnect [as oncomplete] (net.js:1128:14) {
errno: 'ECONNREFUSED',
code: 'ECONNREFUSED',
syscall: 'connect',
address: '10.0.32.42',
port: 8500
}

kubectl describe consul server 0

I'm aware this one is showing rediness probe failed however I can't figure out why

Name: consul-consul-server-0
Namespace: development
Priority: 0
Node: ip-10-0-32-42.ec2.internal/10.0.32.42
Start Time: Wed, 18 Dec 2019 19:32:24 -0500
Labels: app=consul
chart=consul-helm
component=server
controller-revision-hash=consul-consul-server-6c689958b9
hasDNS=true
release=consul
statefulset.kubernetes.io/pod-name=consul-consul-server-0
Annotations: consul.hashicorp.com/connect-inject: false
Status: Running
IP: 100.96.21.3
Controlled By: StatefulSet/consul-consul-server
Containers:
consul:
Container ID: docker://05acd44001b0702aa7219f6c5b2217ab3d9615b5b33b363720c98bc0b8c87ce9
Image: consul:1.6.2
Image ID: docker-pullable://consul@sha256:a167e7222c84687c3e7f392f13b23d9f391cac80b6b839052e58617dab714805
Ports: 8500/TCP, 8301/TCP, 8302/TCP, 8300/TCP, 8600/TCP, 8600/UDP
Host Ports: 0/TCP, 0/TCP, 0/TCP, 0/TCP, 0/TCP, 0/UDP
Command:
/bin/sh
-ec
CONSUL_FULLNAME="consul-consul"

exec /bin/consul agent \
-advertise="${POD_IP}" \
-bind=0.0.0.0 \
-bootstrap-expect=1 \
-client=0.0.0.0 \
-config-dir=/consul/config \
-datacenter=dc1 \
-data-dir=/consul/data \
-domain=consul \
-hcl="connect { enabled = true }" \
-ui \
-retry-join=${CONSUL_FULLNAME}-server-0.${CONSUL_FULLNAME}-server.${NAMESPACE}.svc \
-server

State: Running
Started: Wed, 18 Dec 2019 19:32:33 -0500
Ready: True
Restart Count: 0
Readiness: exec [/bin/sh -ec curl https://127.0.0.1:8500/v1/status/leader 2>/dev/null | \
grep -E '".+"'
] delay=5s timeout=5s period=3s #success=1 #failure=2
Environment:
POD_IP: (v1:status.podIP)
NAMESPACE: development (v1:metadata.namespace)
Mounts:
/consul/config from config (rw)
/consul/data from data-development (rw)
/var/run/secrets/kubernetes.io/serviceaccount from consul-consul-server-token-42ss6 (ro)
Conditions:
Type Status
Initialized True
Ready True
ContainersReady True
PodScheduled True
Volumes:
data-development:
Type: PersistentVolumeClaim (a reference to a PersistentVolumeClaim in the same namespace)
ClaimName: data-development-consul-consul-server-0
ReadOnly: false
config:
Type: ConfigMap (a volume populated by a ConfigMap)
Name: consul-consul-server-config
Optional: false
consul-consul-server-token-
42ss6:
Type: Secret (a volume populated by a Secret)
SecretName: consul-consul-server-token-42ss6
Optional: false
QoS Class: BestEffort
Node-Selectors: <none>
Tolerations: node.kubernetes.io/not-ready:NoExecute for 300s
node.kubernetes.io/unreachable:NoExecute for 300s
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Warning FailedScheduling 5m53s (x7 over 6m9s) default-scheduler pod has unbound immediate PersistentVolumeClaims (repeated 2 times)
Normal Scheduled 5m53s default-scheduler Successfully assigned development/consul-consul-server-0 to ip-10-0-32-42.ec2.internal
Warning FailedAttachVolume 5m51s (x3 over 5m53s) attachdetach-controller AttachVolume.Attach failed for volume "pvc-fd6b0c40-21f6-11ea-9540-020941ae37b7" : "Error attaching EBS volume \"vol-0ff35f1bebdbe2330\"" to instance "i-03ee444d8bb8a0a88" since volume is in "creating" state
Normal SuccessfulAttachVolume 5m47s attachdetach-controller AttachVolume.Attach succeeded for volume "pvc-fd6b0c40-21f6-11ea-9540-020941ae37b7"
Normal Pulled 5m44s kubelet, ip-10-0-32-42.ec2.internal Container image "consul:1.6.2" already present on machine
Normal Created 5m44s kubelet, ip-10-0-32-42.ec2.internal Created container
Normal Started 5m44s kubelet, ip-10-0-32-42.ec2.internal Started container
Warning Unhealthy 5m37s kubelet, ip-10-0-32-42.ec2.internal Readiness probe failed:


&#x200B;

kubectl describe on a agent

Name: consul-consul-hzndl
Namespace: development
Priority: 0
Node: ip-10-0-32-42.ec2.internal/10.0.32.42
Start Time: Wed, 18 Dec 2019 19:32:08 -0500
Labels: app=consul
chart=consul-helm
component=client
controller-revision-hash=5db9fb58cd
hasDNS=true
pod-template-generation=1
release=consul
Annotations: consul.hashicorp.com/connect-inject: false
Status: Running
IP: 100.96.21.2
Controlled By: DaemonSet/consul-consul
Containers:
consul:
Container ID: docker://67cc1d0ae1698732cb944373c79e8685a51ba3899c4ca9d52a9ae934c7f2e479
Image: consul:1.6.2
Image ID: docker-pullable://consul@sha256:a167e7222c84687c3e7f392f13b23d9f391cac80b6b839052e58617dab714805
Ports: 8500/TCP, 8502/TCP, 8301/TCP, 8301/UDP, 8302/TCP, 8300/TCP, 8600/TCP, 8600/UDP
Host Ports: 8500/TCP, 8502/TCP, 0/TCP, 0/UDP, 0/TCP, 0/TCP, 0/TCP, 0/UDP
Command:
/bin/sh
-ec
CONSUL_FULLNAME="consul-consul"

exec /bin/consul agent \
-node="${NODE}" \
-advertise="${ADVERTISE_IP}" \
-bind=0.0.0.0 \
-client=0.0.0.0 \
-hcl="ports { grpc = 8502 }" \
-config-dir=/consul/config \
-datacenter=dc1 \
-data-dir=/consul/data \
-retry-join=${CONSUL_FULLNAME}-server-0.${CONSUL_FULLNAME}-server.${NAMESPACE}.svc \
-domain=consul

State: Running
Started: Wed, 18 Dec 2019 19:32:12 -0500
Ready: True
Restart Count: 0
Readiness: exec [/bin/sh -ec curl https://127.0.0.1:8500/v1/status/leader 2>/dev/null | \
grep -E '".+"'
] delay=0s timeout=1s period=10s #success=1 #failure=3
Environment:
ADVERTISE_IP: (v1:status.podIP)
NAMESPACE: development (v1:metadata.namespace)
NODE: (v1:spec.nodeName)
Mounts:
/consul/config from config (rw)
/consul/data from data (rw)
/var/run/secrets/kubernetes.io/serviceaccount from consul-consul-client-token-xp5q9 (ro)
Conditions:
Type Status
Initialized True
Ready True
ContainersReady True
PodScheduled True
Volumes:
data:
Type: EmptyDir (a temporary directory that shares a pod's lifetime)
Medium:
SizeLimit: <unset>
config:
Type: ConfigMap (a volume populated by a ConfigMap)
Name: consul-consul-client-config
Optional: false
consul-consul-client-token-xp5q9:
Type: Secret (a volume populated by a Secret)
SecretName: consul-consul-client-token-xp5q9
Optional: false
QoS Class: BestEffort
Node-Selectors: <none>
Tolerations: node.kubernetes.io/disk-pressure:NoSchedule
node.kubernetes.io/memory-pressure:NoSchedule
node.kubernetes.io/not-ready:NoExecute
node.kubernetes.io/unreachable:NoExecute
node.kubernetes.io/unschedulable:NoSchedule
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Normal Scheduled 5m44s default-scheduler Successfully assigned development/consul-consul-hzndl to ip-10-0-32-42.ec2.internal
Normal Pulling 5m43s kubelet, ip-10-0-32-42.ec2.internal pulling image "consul:1.6.2"
Normal Pulled 5m40s kubelet, ip-10-0-32-42.ec2.internal Successfully pulled image "consul:1.6.2"
Normal Created 5m40s kubelet, ip-10-0-32-42.ec2.internal Created container
Normal Started 5m40s kubelet, ip-10-0-32-42.ec2.internal Started container
Warning Unhealthy 5m12s (x3 over 5m32s) kubelet, ip-10-0-32-42.ec2.internal Readiness probe failed:


SSH'ing into the worker nodes and doing a curl on localhost (someone said to try it online so I gave it a shot) results in a connection refused.

&#x200B;

In my helm chart (deployment file) for my Node.js application I have this code which is providing the IP for consul to connect to

env:
- name: CONSUL
valueFrom:
fieldRef:
fieldPath: status.hostIP

I got the above from [here](https://www.consul.io/docs/platform/k8s/run.html#accessing-the-consul-http-api) so I expected it to work.

&#x200B;

So from my digging around I am fairly sure it's because the services are not passing the readiness checks even though kubectl get pods is showing 1/1 for all pods

&#x200B;

Suggestions on what I can try to get my application to connect?

https://redd.it/eclflw
@r_devops
How to Package virtualenv as zip/deb/tar and run it on another machine?

Hi,

I have a Django application and I've created a virtualenv for all dependencies for running my Django application and which does work perfectly. So, now when I zip the whole application Django files along with virtualenv. And trying to unpack it to another machine and unzip it and when I tried running the same application, it fails.

On machine X:

cd my_app
virtualenv env_django
source env_django/bin/activate
(env_django)# pip install -r requirements.txt
(env_django)# python manage.py -> this works

Now, Ziping it.

zip my_app.zip my_app

On machine Y:

unzip my_app.zip
source env_django/bin/activate
(env_django)# python manage.py -> this fails

throws out the below site-packages error.


from cryptography.hazmat.backends.openssl.backend import backend
File "/env_django/lib/python2.7/site-packages/cryptography/hazmat/backends/openssl/backend.py", line 46, in <module>
from cryptography.hazmat.bindings._openssl import ffi as _ffi
ImportError: libssl.so.10: cannot open shared object file: No such file or directory

The application really works on Machine X, but not on machine Y. I've tried installing system libraries too. and linked the lib files using [library link](https://askubuntu.com/questions/339364/libssl-so-10-cannot-open-shared-object-file-no-such-file-or-directory), but no luck.

libdev-ssl
libcrypto

I really want to package it and ship to the customers, I want to use only zip/deb/tar.gz package but not as a Docker.

So, is there any way we can ship the virtualenv to different machines via packaging as deb/tar?

https://redd.it/ecigik
@r_devops
Which CI/CD mechanism do you all use?

I was just curious which mechanism did you guys find efficient among the all you used till now.
I know everyone has a different way of doing it, which one did you like.

Thanks.

https://redd.it/efqoz6
@r_devops
Automating Developer Workspaces

Hello,

I'm trying to provision a fair amount of developer workspaces, in all 3 of the major OSes, with a few different types of preset application packs based on the role of the user.

Currently, my strategy is to try out Packer with Terraform or Ansible and try to automatically install these fresh images on empty machines, but I'm not sure exactly if it is possible yet to achieve this with these tools.

Also, would it be possible to somehow connect some laptops to our network and get the OS images to autoinstall themselves?

Anyone from the fields that has any advice? :D Much appreciated

https://redd.it/eclyr7
@r_devops
Alternatives to DockerHub

I'm paying $22/month for DockerHub, which is pretty pricey.

I have my own k8s cluster, and I'm wondering what the best way to either create my own registry or use a cheaper alternative.

- Create a service inside of my cluster.
- Create an EC2 instance outside the cluster.
- Use a 3rd party.

Any insight?

https://redd.it/ech512
@r_devops
Artifactory system level importing and exporting from the command line?

Im trying to replicate an existing Artifactory server into an offline environment. Does anyone have experience with jfrog products? Im looking to automate the export of the systems but I cant find any docs on how to do it other than from the GUI.

https://redd.it/ecirlv
@r_devops
Implementing a CI pipeline using Perforce

Hello there,
I'm working in devops for a game Dev company and our depot size is ~ 200gb - Building the game locally takes ~ 20 mins. We are currently evaluating ways to improve/redevelop our CI pipeline.
Current setup polls the swarm server for new reviews and builds the shelves associated with them but since artists & game designers don't use swarm...it's just semi-helpful.
We have several architectures in our mind (submit trigger, staging depot) but each has some serious drawbacks.
What's the established, usual way to implement a CI pipeline with p4?

https://redd.it/ecipbr
@r_devops
Traffic generation tool that can be used both for baselines and stress tests

Hey everyone, here's one that I think is a relatively straightforward question but I haven't found a solid answer to online. We're trying to run some load testing on our sites, and ideally use one tool for both baseline measuring and stress testing. The trouble I've run into is that some tools out there are great for generating a bunch of traffic and hammering endpoints, thus giving consistent request/response time baselines (like apache bench), and some tools are great at generating a huge load of realistic user behvior (like locust). We've tried running both baseline testing and performance testing with locust, but we've found some marked unreliability in terms of request/response times with locust, indicating to me that there's either some noise in our test infrastructure or the test itself is not keeping all variables the same and is thus unsuitable for baselines.

Have any of you out there used one tool for both use-cases, or are they measured differently inherently and worth using two separate tools for?

https://redd.it/ecfzmh
@r_devops
Where can I start to modernize my DevOps environment?

Please read through my quick overview of our environments at my company, and let me know where you think is a good place to start in order to get this environment into a more modern DevOps culture.

* Production and staging’s technology stack is as follows:
- Legacy .NET Framework APIs running on 8 windows 2008 R2 IIS Servers (EC2 instances).
-Data is managed by a very large monolithic SQL server cluster (2 x i3.16xlarge and 1 i3.8xlarge).
-The new microservices are running on AWS Lambda/DynamoDB/S3 – each with their own Cloudformation templates.
* The development environments:
-20 different dev environments running a single EC2 instance each, and each holding all of the legacy APIs on that 1 EC2.
-Each dev environment also has its own EC2 running SQL server 2008
-The microservices/lambdas for the dev environments are created by me on a per-needed basis.
-Basically if a software developer is working on one of the serverless services, they will come to me to add that service to their development environment.

* Source Control and the CI/CD Pipeline:
-Microsoft’s Team Foundation Server (TFS)
-Developers check in the code to TFS and then manually queue a build from the TFS console
-Deployments are manually initiated on TFS, but then executed by AWS CodeDeploy. CodeDeploy uses the CodeDeploy agent to talk to the dev environment EC2 instances to deploy the code (which was zipped up to S3 via TFS tasks).
-The microservices are deployed via powershell scripts (AWS tools for powershell) from the TFS deploy server

Though our developers try to follow Agile principles, we very much do Waterfall deployments. We are unable to do blue/green or rolling deployments because of the massive legacy SQL server. Every deployment that I’ve run, the SQL server has had very large schema changes which require us to run migration scripts to update the schema.

The goal of the development team is to breakup the endpoints that live in the legacy .NET framework APIs and convert them over to microservices while also removing the dependency of the large SQL server. But in the meantime, every month or so we have a 6 hour outage on a Friday evening to do a large Waterfall deployment.

The things that I want to do in order to change culture here would require such a massive overhaul, I just don’t know if it’s possible without a huge investment from the company.
I guess in the meantime I would like some suggestions on anything you see fit to suggest.
Here are a couple of big-ticket problem items that I can’t figure out:

* How to better manage having a ridiculous 20 dev environments (EC2+Sql server combo) with microservices added to each as-needed.
* Assuming the SQL server dependency doesn’t change for a while- how to better manage that
* Best entry point to start engineering for live production releases rather than a Friday evening downtime once a month.
* Anything else you can think of

https://redd.it/ec3pwl
@r_devops