Found Description
Lead the reliability of AI infrastructure as a Site Reliability Engineer. Become a key player in automating systems and enhancing observability in a collaborative environment.
This hands-on role requires 5+ years in SRE or infrastructure engineering, focusing on infrastructure reliability across colocation and cloud environments. You'll be responsible for maintaining and automating infrastructure that supports silicon development and customer success. Your expertise in Linux, Kubernetes, and IaC tools like Terraform will greatly impact our performance and reliability.
Key Responsibilities:
• Own reliability of colocation and cloud infrastructure
• Perform hardware troubleshooting and OS configuration
• Automate systems using Terraform and Ansible
• Design monitoring dashboards with Prometheus/Grafana
• Support platform services for customer deployments
Requirements:
• Bachelor’s or Master’s in Computer Science o...
This hands-on role requires 5+ years in SRE or infrastructure engineering, focusing on infrastructure reliability across colocation and cloud environments. You'll be responsible for maintaining and automating infrastructure that supports silicon development and customer success. Your expertise in Linux, Kubernetes, and IaC tools like Terraform will greatly impact our performance and reliability.
Key Responsibilities:
• Own reliability of colocation and cloud infrastructure
• Perform hardware troubleshooting and OS configuration
• Automate systems using Terraform and Ansible
• Design monitoring dashboards with Prometheus/Grafana
• Support platform services for customer deployments
Requirements:
• Bachelor’s or Master’s in Computer Science o...
Ready to Apply?
Submit your application for Site Reliability Engineer - AI Infrastructure at Phizenix
Apply Now