Found Description
Kaseya is hiring an experienced Site Reliability Engineer dedicated to keeping our production systems reliable as we scale. You'll oversee services that support thousands of MSPs every day.
This role entails setting SLOs, leading incident responses, and creating automation to ensure system stability in a primarily AWS environment. If you view reliability as a key product feature, you will thrive here.
Key Responsibilities:
• Monitor and enforce SLOs, SLIs, and error budgets
• Lead incident response efforts and write postmortems
• Create automated deployment and configuration systems
• Manage hybrid infrastructure with Terraform or CloudFormation
• Enhance observability through systems monitoring
Requirements:
• 4 to 5 years experience in AWS production
• Solid knowledge of Infrastructure as Code
• Experience with on-call rotations and incident leadership
• Familiarity with SLOs and error budget...
This role entails setting SLOs, leading incident responses, and creating automation to ensure system stability in a primarily AWS environment. If you view reliability as a key product feature, you will thrive here.
Key Responsibilities:
• Monitor and enforce SLOs, SLIs, and error budgets
• Lead incident response efforts and write postmortems
• Create automated deployment and configuration systems
• Manage hybrid infrastructure with Terraform or CloudFormation
• Enhance observability through systems monitoring
Requirements:
• 4 to 5 years experience in AWS production
• Solid knowledge of Infrastructure as Code
• Experience with on-call rotations and incident leadership
• Familiarity with SLOs and error budget...
Ready to Apply?
Submit your application for Experienced Site Reliability Engineer at Kaseya at Kaseya Limited
Apply Now