Job Description

We are looking for a "Software Reliability Engineer" to ensure the reliability, availability, and scalability of our systems. The role works closely with development teams to improve system resilience, automate operations, and respond to production incidents (software updates, bug fixes, and security patches).

Responsibilities

-Monitor system health and assist in identifying and troubleshooting common issues.
-Implement and maintain monitoring and alerting dashboards.
-Participate in on-call rotations, responding to incidents with guidance.
-Contribute to blameless postmortems and recommend systemic improvements.
-Assist in capacity planning and basic automation of manual tasks.
-Collaborate with development teams on deploying updates and patches.
-Good scripting/automation skills (Python, Shell, Go).
-Solid experience with container orchestration (Kubernetes/CCE) and service mesh concepts.
-Deep understanding of network protocols, load balancing, and security best practices.
-Experience with observability stacks and SLI/SLO design.
-Ability to work independently and mentor junior engineers.
-Basic knowledge in at least one field: network, storage, virtualization, containerization, Linux operating systems, cybersecurity, big data, or databases.

Requirements

-Bachelor’s degree in Computer Science, Software Engineering, a related field, or equivalent practical experience.
- +2 years of experience with software development or cloud O&M experience.
-Excellent communication skills.
-Pragmatic, problem-solving attitude.
-A naturally curious and proactive approach to learning and problem-solving.
-Good skills in writing/editing documentation.

Preferred

-Familiarity with OpenStack deployment or operations.
-Familiarity with public cloud deployment or operations.

برای مشاهده‌ی شغل‌هایی که ارتباط بیشتری با حرفه‌ی شما دارد،