k8s-infra-cluster/manuals/deploy_k8s_components.md
2026-08-04 06:12:38 +02:00

759 lines
27 KiB
Markdown

### Getting started
kubectl apply -k . # from root project folder
if not detected: kubectl get pods -n db
kubectl apply -f db/yugabytedb.yaml -n db
- Find services and its ports
kubectl get svc -n apps
- Port forwarding local to cosmo-router,argo-cd
kubectl port-forward svc/cosmo-router 3002:3002 -n apps
kubectl port-forward svc/argocd-server -n argocd 8480:443
- Open yugabyte
kubectl exec -it yb-tserver-0 -n db -- ysqlsh -h
yb-tserver-0
- Open nats
kubectl run nats-tools --image=natsio/nats-box -n infra --rm -it -- nats stream info request_stream --server=nats://nats-cluster:4222
if already exist:
kubectl exec -it nats-box -n infra -- nats stream info request_stream -s nats://nats-cluster:4222
- list stream:
kubectl exec -it nats-box -n infra -- nats -s nats://nats-cluster.infra.svc.cluster.local:4222 stream ls
- **add stream**:
kubectl exec -it nats-box nats -n infra -- /bin/sh
nats -s nats://nats-cluster.infra.svc.cluster.local:4222 stream add request_stream --subjects "input_request_logs" --ack --storage file --retention limits --max-msgs=-1 --max-bytes=-1 --max-age=1y --replicas 3
check:
kubectl describe pod nats-cluster-1 -n infra
- Check benthos
kubectl logs -f benthos-556f97988-cgd7f -n infra
- Redeploy
For cosmo-router under namespace apps:
```
to apply changes:
helm upgrade cosmo-router ./cosmo-router -n apps # from folder apps
we can not use this command to apply changes:
kubectl rollout restart deployment cosmo-router -n apps
```
### Install Linkerd on local machine
Linkerd exe is installed in your folder of choice, and set
environment variables path to refer to this.
### Some linkerd commands
Some commands:
linkerd viz dashboard
linkerd version --client
linkerd check --linkerd-namespace apps
linkerd -n apps stat deployments
linkerd check --pre
### Deinstalling Linkerd:
- Deleting Linkerd (depends on namespace) :
```
- kubectl delete deployment linkerd-destination linkerd-identity linkerd-proxy-injector -n apps
- kubectl delete deployment metrics-api prometheus tap tap-injector web -n apps
- Delete the Mutating Webhooks (CRITICAL):
- kubectl delete mutatingwebhookconfiguration linkerd-proxy-injector-webhook-config
- kubectl delete mutatingwebhookconfiguration linkerd-tap-injector-webhook-config
- kubectl delete validatingwebhookconfiguration linkerd-sp-validator-webhook-config
- Remove the annotation from your namespace:
kubectl annotate namespace apps linkerd.io/inject-
- Restart your apps to strip the sidecars:
kubectl rollout restart deployment cosmo-router -n apps
kubectl rollout restart statefulset nats-cluster -n infra
```
- Deleting Linkerd using Helm installation:
```
# 1. Delete the current broken Linkerd installations
helm uninstall linkerd-viz -n apps
helm uninstall linkerd-control-plane -n apps
helm uninstall linkerd-crds -n apps
# 2. IMPORTANT: Delete any leftover Linkerd webhooks (these often block restarts)
kubectl delete mutatingwebhookconfiguration linkerd-proxy-injector-webhook-config
kubectl delete validatingwebhookconfiguration linkerd-sp-validator-webhook-config
```
- Deleting Linkerd using linkerd CLI:
linkerd uninstall --namespace apps | kubectl delete -f -
### Open source Helm chart
Most of the time we do not need to write k8s yaml files.
If we have helm chart available from the provider we generate the files needed using helm template and refer the generated file from kustomization.yaml
#### Using helm chart vs kustomize
if you want to move your cluster to new VMs, you just run:
To run with helm chart: kubectl apply -k .
Kubernetes will create the namespaces, then deploy NATS, then Benthos, then Yugabyte all in one go
|Method|Pro|Con|
|--|---|--|
|helm install|Fast and easy for one-offs.|Hard to remember settings; hard to track in Git.|
|Kustomize + Helm|Single source of truth; everything in Git.|Requires --enable-helm flag; slightly more setup.|
- Manual way with helm chart
We also have 2 ways by using helm template vs helm install
```
helm template yugabytedb yugabytedb/yugabyte `
--namespace db `
--set storage.master.storageClass=longhorn `
--set storage.tserver.storageClass=longhorn `
--set replicas.master=3 `
--set replicas.tserver=3 `
--set enableLoadBalancer=false `
--set gflags.master.max_clock_skew_usec=2000000 `
--set gflags.tserver.max_clock_skew_usec=2000000 `
--set gflags.master.time_source=system `
--set gflags.tserver.time_source=system `
--set gflags.tserver.start_pgsql_proxy=true `
> db/yugabytedb.yaml
OR
helm install yugabytedb yugabytedb/yugabyte `
--namespace db `
--set storage.master.storageClass=longhorn `
--set storage.tserver.storageClass=longhorn `
--set replicas.master=3 `
--set replicas.tserver=4 `
--set gflags.master.max_clock_skew_usec=2000000 `
--set gflags.tserver.max_clock_skew_usec=2000000 `
--set gflags.master.time_source=system `
--set enableLoadBalancer=false `
--set gflags.tserver.start_pgsql_proxy=true `
--set gflags.tserver.time_source=system
```
|Feature|helm template|helm install|
|---|---|----|
|Action|Local Only. Generates raw Kubernetes YAML and prints it |to a file.|Live. Sends the YAML directly to your cluster API.|
|Result|"You get a file (db/yugabytedb.yaml) that you can |inspect| edit| or commit to Git."|The database starts running |immediately in your db namespace.|
|Usage|Best for GitOps/Kustomize. You use the generated file as |a resource in Kustomize.|Best for Quick deployment. Harder to track changes over time in Git.|
- Kustomize way but required installing standalone Kustomization binary file on windows
`./kustomize build . --enable-helm | kubectl apply -f -`
We do not use Kustomize here
#### NATS
```
helm repo add nats https://nats-io.github.io/k8s/helm/charts/
helm repo update
- Basic
nats-values.yaml
config:
jetstream:
enabled: true
fileStore:
pvc:
size: 5Gi # Reserve 5GB for JetStream
or
- HA
# To increase replicas for High Availability
replicaCount: 3
# JetStream Clustering (The "Replica Factor")
config:
cluster:
enabled: true
jetstream:
enabled: true
# This ensures your data survives if x-vm2 goes down
fileStore:
pvc:
enabled: true
storageClassName: "local-path" # k3s default storage
size: 10Gi
Check:
# Create the namespace first
kubectl create namespace infra
# Install NATS using the chart and your values file
helm install nats-cluster nats/nats --namespace infra -f nats-values.yaml
or with helm template
helm template nats-cluster nats/nats `
--namespace infra `
--set replicaCount=3 `
--set config.cluster.enabled=true `
--set config.jetstream.enabled=true `
--set config.jetstream.fileStore.pvc.enabled=true `
--set config.jetstream.fileStore.pvc.storageClassName="longhorn" `
--set config.jetstream.fileStore.pvc.size=10Gi `
> infra/nats-cluster.yaml
Apply the template (without creating cluster)
kubectl apply -f infra/nats-cluster.yaml
Apply the template( with creating cluster pods, service, etc):
helm install nats-cluster nats/nats --namespace infra -f nats-values.yaml
After modification of other files you can apply the changes by:
helm upgrade nats-cluster nats/nats --namespace infra -f nats-values.yaml
Remove nats-cluster
helm uninstall nats-cluster -n infra
Remove pvc of nats-cluster (first uninstall nats-cluster using helm-uninstall)
kubectl delete pvc -n infra -l app.kubernetes.io/instance=nats-cluster
Verify:
kubectl get all -n infra
kubectl get statefulset -n infra
kubectl get pods -n infra -w
kubectl get pvc -n infra
kubectl get sc # verify longhorn is ready and status healthy
To avoid typing -n infra every time, you can switch your "active room" to infra permanently:
kubectl config set-context --current --namespace=infra
Apply changes:
helm upgrade nats-cluster nats/nats -n infra -f nats-values.yaml
Testing:
first terminal:
kubectl run nats-box --image=natsio/nats-box:latest -n infra -it --rm
nats sub -s nats-cluster test.topic
new terminal:
# Instead of 'run', we 'exec' into the pod that is already there
kubectl exec -n infra -it nats-box -- /bin/sh
nats pub -s nats-cluster test.topic "Hello from Country B!"
or
kubectl exec -it nats-box -n infra -- nats -s nats-cluster:4222 pub input_request_logs '{"content": {"hash": "test-123"}, "message": "Hello Benthos!"}'
Delete:
kubectl delete statefulset nats-cluster -n infra
kubectl delete all -l app.kubernetes.io/instance=nats-cluster -n default
```
#### Intrepretation
```
kubectl get all -n infra
NAME READY STATUS RESTARTS AGE
pod/nats-box 1/1 Running 0 13m
pod/nats-cluster-0 2/2 Running 0 87s
pod/nats-cluster-1 2/2 Running 0 87s
pod/nats-cluster-2 2/2 Running 0 87s
pod/nats-cluster-box-868cc6c48b-dlvtn 1/1 Running 0 26m
NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE
service/nats-cluster ClusterIP 10.43.208.167 <none> 4222/TCP 26m
service/nats-cluster-headless ClusterIP None <none> 4222/TCP,6222/TCP,8222/TCP 26m
NAME READY UP-TO-DATE AVAILABLE AGE
deployment.apps/nats-cluster-box 1/1 1 1 26m
NAME DESIRED CURRENT READY AGE
replicaset.apps/nats-cluster-box-868cc6c48b 1 1 1 26m
NAME READY AGE
statefulset.apps/nats-cluster 3/3 88s
Interpretation:
The "Power Trio" (StatefulSet)
pod/nats-cluster-0
pod/nats-cluster-1
pod/nats-cluster-2
Why: The official Helm chart defaults to replicaCount: 3. This is for High Availability. Since you have 4 VMs, Kubernetes likely spread these across your different nodes. If one VM in Country B fails, NATS will keep running because the other two pods have a "quorum" (majority) and won't lose your data.
The "Utility" Pods (Box)
pod/nats-box: This is the temporary pod you created manually with kubectl run to test the connection. It stays there until you delete it or it finishes.
pod/nats-cluster-box-868cc6c48b-dlvtn: This was created automatically by the Helm chart. The NATS team includes a permanent "box" deployment so you always have a toolset inside the cluster to check the status of the NATS stream.
What about the "2/2" READY status?
You'll notice the nats-cluster-x pods say 2/2. This means each pod actually contains two containers:
The NATS Server: The actual engine.
The NATS Config Reloader: A "sidecar" container that watches for changes to your settings and tells the NATS server to refresh without restarting.
```
#### Rules of Engagement
Creating stream with Benthos for how NATS stores your data
```
? Retention Policy Limits, WorkQueue, Interest
? Discard Policy Old vs New
In NATS, -1 means "Infinite" or "No Limit."
? Stream Messages Limit -1
? Per Subject Messages Limit -1
? Total Stream Size -1
? Message TTL -1
? Max Message Size -1
? Duplicate tracking time window 2m0s
? Allow message Roll-ups Yes
? Allow message deletion Yes
? Allow purging subjects or the entire stream (Y/n)
nats -s nats-cluster:4222 stream add request_stream --subjects "input_request_logs" --ack --storage file --retention limits --max-msgs=-1 --max-bytes=-1 --max-age=1y --replicas 3
nats -s nats-cluster:4222 stream add transaction_stream --subjects "input_transaction_logs" --ack --storage file --replicas 3
```
### Benthos
```
# Running namespaces
kubectl apply -f namespaces.yaml # run this command from the root of your repo
# Running Kustomization Deploy configs
kubectl apply -k . -n infra # from infra\benthos folder
kubectl apply -k .\benthos\ -n infra # from infra folder
# Manual creation of config map
kubectl create configmap benthos-config --from-file=./benthos-configs/
# Verify ConfigMaps
Check if your files were successfully "packed" into Kubernetes:
kubectl get configmap -n infra
# To see the actual content of your streams inside K8s:
kubectl describe configmap benthos-streams -n infra
# Verify the 3 Replicas
kubectl get pods -n infra -l app=benthos
# Verify the Dashboard (Port-Forward) To see the Benthos UI on your laptop:
kubectl port-forward svc/benthos-ui -n infra 4195:4195
# Peek inside the running Pod:
kubectl exec -it <pod-name> -n infra -- ls /configs/streams
kubectl exec -it <pod-name> -n infra -- ls -R /configs
# Verify
kubectl get svc -n infra.
# Create Stream:
kubectl exec -n infra -it nats-box -- /bin/sh
```
#### Debugging
```
kubectl describe pod benthos-68cb959d58-47prd -n infra
kubectl logs benthos-86f5c886b7-5fr5t -n infra
If needed delete only deployment and redeploy using kustomization
Depends on the your position
kubectl delete deployment benthos -n infra
kubectl apply -k .\infra\benthos\ -n infra
or only without namespaces
kubectl apply -k .\infra\benthos\
or only (if you are inside \infra\benthos\)
kubectl apply -k .
# You must not delete configmap since it get refreshed everytime you changed. Optional:
kubectl delete configmap --all -n infra (Only if only Benthos is in infra)
# Check output after redeploy
kubectl get pod -n infra -o yaml
kubectl kustomize .
```
### Installing Longhorn on each VMs
- Prerequisite
You need to run the following command on all 4 VMs (Master and all Workers) to ensure Longhorn can communicate with the disks:
```
sudo apt install open-iscsi nfs-common util-linux -y
sudo systemctl enable --now iscsid
Why do we need this? Longhorn creates "Block Devices." To do that, the Linux kernel on your VM needs open-iscsi to "attach" to the virtual disks Longhorn creates. Without this, your pods will be stuck in ContainerCreating forever
1. Add the repo
helm repo add longhorn https://charts.longhorn.io
helm repo update
2. Install into a new namespace
Longhorn is a "System Service" not Application Infrastructure, requires high privileges
helm install longhorn longhorn/longhorn --namespace longhorn-system --create-namespace
- Check : kubectl get pods -n longhorn-system
3. Running the StorageClass Patch from Local Windows Laptop
Once installed, tell K3s to use Longhorn for every database deployed:
kubectl patch storageclass longhorn -p '{"metadata": {"annotations":{"storageclass.kubernetes.io/is-default-class":"true"}}}'
or
kubectl patch storageclass longhorn -p "{\"metadata\": {\"annotations\":{\"storageclass.kubernetes.io/is-default-class\":\"true\"}}}"
or
Run kubectl edit storageclass longhorn
Look for the metadata: section.
Add this line under annotations:
storageclass.kubernetes.io/is-default-class: "true"
Save and exit.
- Notes
Because you have 4 VMs, Longhorn will default to 3 replicas for your data. This is perfect. It means:
Data is written to VM1.
Longhorn clones it to VM2 and VM3.
VM4 stays as a "spare" or handles other data.
If any one VM catches fire, your data is still 100% safe and available.
4. UI Longhorn:
A. Update your Windows Hosts file:
- Open Notepad (Make sure you have right to write)
- Open C:\Windows\System32\drivers\etc\hosts.
- Add the IP of your Master VM:
192.168.x.x longhorn.local (Replace with your actual VM IP)
B. Create longhorn.ingress.yaml and run:
/country-b-cluster-ops
/system <-- NEW FOLDER for cluster-wide tools
longhorn-ingress.yaml
/infra <-- For your messaging/processing (NATS, Benthos)
/db <-- For your databases (Yugabyte)
kubectl apply -f system/longhorn-ingress.yaml
C. Open http://longhorn.local from your laptop
```
#### Additional Configuration?
There is one common issue with K3s and Longhorn. K3s stores its data in /var/lib/rancher/k3s/storage, but Longhorn defaults to /var/lib/longhorn.
Recommendation:
If you have a specific large hard drive or partition on your VMs where you want the data to live, you can configure that in the Longhorn UI. If you just have one big / partition, the default is fine.
### Yugabyte
```
helm repo add yugabytedb https://charts.yugabyte.com
helm repo update
helm install yugabytedb yugabytedb/yugabyte `
--namespace db `
--set storage.master.storageClass=longhorn `
--set storage.tserver.storageClass=longhorn `
--set replicas.master=3 `
--set replicas.tserver=4 `
--set gflags.master.max_clock_skew_usec=2000000 `
--set gflags.tserver.max_clock_skew_usec=2000000 `
--set gflags.master.time_source=system `
--set enableLoadBalancer=false `
--set gflags.tserver.start_pgsql_proxy=true `
--set gflags.tserver.time_source=system
or
helm template yugabytedb yugabytedb/yugabyte `
--namespace db `
--set storage.master.storageClass=longhorn `
--set storage.tserver.storageClass=longhorn `
--set replicas.master=3 `
--set replicas.tserver=4 `
--set gflags.master.max_clock_skew_usec=2000000 `
--set gflags.tserver.max_clock_skew_usec=2000000 `
--set gflags.master.time_source=system `
--set enableLoadBalancer=false `
--set gflags.tserver.start_pgsql_proxy=true `
--set gflags.tserver.time_source=system > yugabytedb.yaml
```
- set gflags.tserver.start_pgsql_proxy=true?
This flag tells the Yugabyte T-Server to turn on the YSQL API layer
- Fully Qualified Domain Name FQDN for communication
- yugabytedb.db.svc.cluster.local
- The pattern is: `[service-name].[namespace].svc.cluster.local`
#### Optimal Node Distribution for 4 VMs
```
For a 4-node cluster, you have enough overhead to ensure High Availability (HA) while maximizing resource use.
Master Count: 3
Why: Yugabyte Masters use the Raft consensus algorithm. You need an odd number to avoid "split-brain" scenarios. 3 masters can tolerate 1 node failure. Since you have 4 VMs, you should stick with 3.
T-Server Count: 3 or 4
Option 3 (Balanced): Keep 3 T-Servers. This leaves 1 VM entirely free for Benthos, NATS, and Redpanda.
Option 4 (High Performance): Run 4 T-Servers. This spreads your data shards across all available hardware.
Recommendation: Go with 3 Masters and 3 T-Servers. This keeps the "Replication Factor" (RF) at 3, which is the standard. Adding a 4th T-Server with RF3 is possible, but 3/3 is more predictable for a small cluster.
```
### Linkerd, Viz, openssl on k8s
```
1. Create the Trust Anchor (Root CA)
# Generate private key for the Root
a.openssl ecparam -name prime256v1 -genkey -noout -out ca.key
b.openssl ecparam -name prime256v1 -genkey -noout -out ca.key
# Generate the self-signed Root Certificate (Valid for 10 years)
a.openssl req -x509 -new -nodes -key ca.key -sha256 -days 3650 \
-out ca.crt -subj "/CN=root.linkerd.cluster.local" \
-addext "basicConstraints=critical,CA:TRUE"
b.openssl req -x509 -new -nodes -key ca.key -sha256 -days 3650 -out ca.crt -subj "/CN=root.linkerd.cluster.local" -addext "basicConstraints=critical,CA:TRUE"
2. Create the Identity Issuer (Intermediate CA)
# Generate private key for the Issuer
a/b. openssl ecparam -name prime256v1 -genkey -noout -out issuer.key
# Create a CSR (Certificate Signing Request)
a/b. openssl req -new -key issuer.key -out issuer.csr -subj "/CN=identity.linkerd.cluster.local"
# Create a config file for the intermediate CA extensions
a. cat > issuer.ext <<EOF
basicConstraints=critical,CA:TRUE,pathlen:0
keyUsage=critical,digitalSignature,keyCertSign,cRLSign
EOF
# Sign the Issuer certificate with your Root CA (Valid for 1 year)
a. openssl x509 -req -in issuer.csr -CA ca.crt -CAkey ca.key \
-CAcreateserial -out issuer.crt -days 365 -sha256 -extfile issuer.ext
b.# Sign Issuer (Critical: must have CA:TRUE)
echo "basicConstraints=critical,CA:TRUE,pathlen:0" > ext.txt
echo "keyUsage=critical,digitalSignature,keyCertSign,cRLSign" >> ext.txt
openssl x509 -req -in issuer.csr -CA ca.crt -CAkey ca.key -CAcreateserial -out issuer.crt -days 365 -sha256 -extfile ext.txt
3. Generate k3s Manifests
helm repo add linkerd https://helm.linkerd.io/stable && helm repo update
4. Generate CRDs
Linkerd Custom Resource Definitions must be installed first.
helm template linkerd-crds linkerd/linkerd-crds --namespace linkerd > linkerd-crds.yaml
5. Generate Control Plane
We will bake your certificates directly into the generated YAML:
```
helm template linkerd-control-plane linkerd/linkerd-control-plane
--namespace linkerd
--set-file identityTrustAnchorsPEM=ca.crt
--set-file identity.issuer.tls.crtPEM=issuer.crt --set-file identity.issuer.tls.keyPEM=issuer.key
--set identity.issuer.scheme=kubernetes.io/tls > linkerd-control-plane.yaml
or
helm install linkerd-control-plane linkerd/linkerd-control-plane
-n linkerd
--set-file identityTrustAnchorsPEM=ca.crt
--set-file identity.issuer.tls.crtPEM=issuer.crt
--set-file identity.issuer.tls.keyPEM=issuer.key
```
6. Generate Viz (Dashboard)
Default namespaces are linkerd and linkerd-viz but we save them under our namespace apps
```
a. helm template linkerd-viz linkerd/linkerd-viz --namespace linkerd-viz > linkerd-viz.yaml (standard namespace)
b. helm template linkerd-viz linkerd/linkerd-viz \
--namespace apps \
--set linkerdNamespace=apps \
--set tap.namespace=apps \
--set dashboard.namespace=apps \
--set prometheus.namespace=apps > linkerd-viz.yaml
or using
c. helm install linkerd-viz linkerd/linkerd-viz -n linkerd-viz --create-namespace
Uninstalling:
helm ls -n apps
helm uninstall linkerd-viz -n apps
helm uninstall linkerd-control-plane -n apps
helm uninstall linkerd-crds -n apps
```
7. Creating k8s Secret in our namespace
the pod is looking for secret in the apps namespace before the pod can start:
kubectl create secret tls linkerd-identity-issuer --cert=issuer.crt --key=issuer.key --namespace=linkerd
or
kubectl create secret tls linkerd-identity-issuer --cert=issuer.crt --key=issuer.key --namespace=apps
8.. Apply CRDs first
kubectl apply -f linkerd-crds.yaml
9. Apply Control Plane
kubectl apply -f linkerd-control-plane.yaml
10. Apply Viz
kubectl apply -f linkerd-viz.yaml
11. Activating the Connection (Injection)
kubectl annotate namespace apps linkerd.io/inject=enabled
12. Check pods status
kubectl get pods -n apps
Redeploy if error
linkerd-destination-6b69957545-9b9tz 0/4 CrashLoopBackOff 720 (2m23s ago) 12h
linkerd-destination-74cc587f5-gvgqd 0/4 CrashLoopBackOff 712 (50s ago) 12h
linkerd-identity-68858f6c75-t2lk9 2/2 Running 0 14h
linkerd-proxy-injector-54bc495b55-sgk6l 0/2 CrashLoopBackOff 309 (2m23s ago) 12h
linkerd-proxy-injector-584f78bf9-86r64 0/2 CrashLoopBackOff 306 (60s ago) 12h
- destination: The main service discovery logic.
-policy: Handles authorization policies.
-linkerd-proxy: The "sidecar" that handles the pod's own traffic.
-linkerd-init: (Finished) Sets up network rules.
13.Rollout Restart:
For existing services (like Yugabyte or your Python subgraph) to join the mesh, they need to be restarted:
kubectl rollout restart deployment cosmo-router -n apps
kubectl rollout restart statefulset yugabytedb -n db
12. How to Verify
To verify the connection and see your Yugabyte/Redpanda traffic:
Check Status: linkerd check (Requires Linkerd CLI: curl -sL https://run.linkerd.io/install | sh).
Open Dashboard: linkerd viz dashboard.
Check mTLS: In the dashboard, look for the shield icon next to your cosmo-router to subgraph traffic. This confirms your OpenSSL certificates are working.
Debugging:
kubectl describe pod -n apps -l linkerd.io/control-plane-component=identity
openssl x509 -in ca.crt -text -noout
kubectl logs -n apps deployment/linkerd-identity
```
#### Linkerd Activation / Deactivation
```
# Run these commands to tell Linkerd that these namespaces are now part of the mesh:
kubectl annotate namespace apps linkerd.io/inject=enabled
kubectl annotate namespace infra linkerd.io/inject=enabled
kubectl annotate namespace db linkerd.io/inject=enabled
kubectl annotate namespace stream linkerd.io/inject=enabled
# Restart everything in those namespaces
kubectl rollout restart deployment -n apps
kubectl rollout restart deployment -n infra
kubectl rollout restart statefulset -n infra # For your NATS cluster
kubectl rollout restart statefulset -n db # For your Yugabyte/DB
kubectl rollout restart deployment -n stream
# Remove the annotation to deactivate
kubectl annotate namespace apps linkerd.io/inject-
kubectl annotate namespace infra linkerd.io/inject-
kubectl annotate namespace db linkerd.io/inject-
kubectl annotate namespace stream linkerd.io/inject-
# Clean the pods
kubectl rollout restart deployment -n apps
kubectl rollout restart deployment -n infra
kubectl rollout restart statefulset -n infra
kubectl rollout restart statefulset -n db
```
### Cosmo router
```
1. Check
cosmo-router/
├── config.yaml <-- Your router config
├── router.json <-- Your composed graph
├── values.yaml <-- Updated with Ingress and Middleware annotations
└── templates/
├── _helpers.tpl <-- Required for names
├── configmap.yaml <-- Uses .Files.Get(our custom config)
├── deployment.yaml <-- Incl Checksum & Linkerd injection (The brain)
├── service.yaml (The external network)
├── ingress.yaml (The internal network)
└── middleware.yaml <-- The Gzip logic
└── hpa.yaml optional
Delete:
a. httproute.yaml (Causes the current error)
b.serviceaccount.yaml (Caused the previous error)
c. tests/ (The default test folder often causes similar issues)
2. Test it
helm template cosmo-router ./cosmo-router
3. Run cosmo-router
# Deploy it
helm install cosmo-router ./cosmo-router -n apps
or Redeploy
helm upgrade cosmo-router ./cosmo-router -n apps
or If needed uninstall
helm uninstall linkerd-crds -n apps
helm uninstall cosmo-router -n apps
Some usefult commands:
kubectl delete pods -n apps -l linkerd.io/extension=viz
kubectl annotate namespace apps linkerd.io/inject=enables
kubectl rollout restart deployment cosmo-router -n apps
kubectl logs -n apps -l app.kubernetes.io/name=cosmo-router
```
#### Helm chart: Which files to Keep, Adapt, or Delete?
Since we ran helm create, you have a lot of "boilerplate."
```
5. Keep and Adapt:
deployment.yaml: Ensure the volumeMounts and volumes match the ConfigMap name above.
service.yaml: Keep it to allow other pods (or the Ingress) to find the router.
values.yaml: Use this for your image tags and environment-specific toggles.
_helpers.tpl: Do not delete. This generates the names (like cosmo-router.fullname) used in every other file.
Keep (but leave disabled):
hpa.yaml: Useful later for auto-scaling, but set autoscaling.enabled: false in values.yaml for now.
ingress.yaml: Keep this if you want to access the router from outside the cluster (e.g., router.example.com).
Delete:
httproute.yaml: This is for the "Gateway API." Unless you have a specific Gateway controller installed, standard ingress.yaml is what you'll use.
NOTES.txt: Usually just contains generic text.
```
#### What ToDo after deleting the primary node that bootstrapped the cluster.
```
ssh to other k3s control plane
sudo systemctl status k3s
sudo systemctl stop k3s
sudo k3s server --cluster-reset
sudo systemctl start k3s
sudo kubectl get nodes
sudo nano /etc/systemd/system/k3s.service
from
ExecStart=/usr/local/bin/k3s server --server https://192.168.3.91:6443 --tls-san 192.168.3.156
to
ExecStart=/usr/local/bin/k3s server --cluster-init --tls-san 192.168.3.156
```
#### Delete stale nodes
```
sudo kubectl delete node invixel-vm1
sudo kubectl delete node invixel-vm6
sudo kubectl delete node invixel-vm7
sudo kubectl delete node invixel-ubuntu1
```