Found Description
Drive innovation in AI compute with Fluidstack as a Compute Fleet Reliability Engineer. Focus on optimizing GPU health metrics and automate repair workflows in a fast-paced environment.
As a key member of Fluidstack's production engineering team, you will take ownership of the compute fleet's health and ensure reliability at scale. This role involves designing automation for deployment and repair, as well as building a GPU qualification platform. Your work directly impacts the performance and functionality of our powerful AI infrastructure.
Key Responsibilities:
• Own compute fleet health metrics and alerting systems
• Automate failure detection, triage, and return to service
• Design GPU qualification processes for production readiness
• Manage firmware telemetry and fleet-scale log collection
• Ensure scalable and reliable operations for GPU infrastructure
Requirements:
• Experience in hardware failure modes a...
As a key member of Fluidstack's production engineering team, you will take ownership of the compute fleet's health and ensure reliability at scale. This role involves designing automation for deployment and repair, as well as building a GPU qualification platform. Your work directly impacts the performance and functionality of our powerful AI infrastructure.
Key Responsibilities:
• Own compute fleet health metrics and alerting systems
• Automate failure detection, triage, and return to service
• Design GPU qualification processes for production readiness
• Manage firmware telemetry and fleet-scale log collection
• Ensure scalable and reliable operations for GPU infrastructure
Requirements:
• Experience in hardware failure modes a...
Ready to Apply?
Submit your application for Compute Fleet Reliability Engineer at Fluidstack at Fluidstack
Apply Now