- Contribute to incident resolution and problem management activities for platform services. - Deploying ready-to-use cloud-hosted platforms for running & managing applications (PaaS) - Troubleshooting and solving issues - Deploying, scaling, and maintaining large-scale Kubernetes clusters - Develop and deploy monitoring, logging & alerting Tools
Requirements - 5+ years of experience in DevOps or Cloud computing, or Site Reliability Engineering - Strong documentation, problem-solving, and incident response skills - Strong Linux administration and troubleshooting experience. - Experience with Docker, Kubernetes - Experience with Grafana / Prometheus / VictoriaMetrics - Strong Knowledge about Git and CI/CD and Automation Tools (Gitlab-CI , ArgoCD, GitOps) - Experience with Configuration management Tools (Ansible) - Familiar with Terraform or Pulumi
Bonus points if you have expertise Rancher/RKE2, Software Development (Go or Python)