P

Site Reliability Engineer - AI Infrastructure

Phizenix

remote, remote, Canada Full-time July 09, 2026

Found Description

Lead the reliability of AI infrastructure as a Site Reliability Engineer. Become a key player in automating systems and enhancing observability in a collaborative environment.

This hands-on role requires 5+ years in SRE or infrastructure engineering, focusing on infrastructure reliability across colocation and cloud environments. You'll be responsible for maintaining and automating infrastructure that supports silicon development and customer success. Your expertise in Linux, Kubernetes, and IaC tools like Terraform will greatly impact our performance and reliability.

Key Responsibilities:
• Own reliability of colocation and cloud infrastructure
• Perform hardware troubleshooting and OS configuration
• Automate systems using Terraform and Ansible
• Design monitoring dashboards with Prometheus/Grafana
• Support platform services for customer deployments

Requirements:
• Bachelor’s or Master’s in Computer Science o...

Ready to Apply?

Submit your application for Site Reliability Engineer - AI Infrastructure at Phizenix

Apply Now