Found Description
Responsibilities
- Design, build, and maintain highly available, scalable, and fault‑tolerant systems
- Collaborate with software engineering teams to ensure applications are designed with reliability and performance in mind
- Develop and maintain automation procedures to maximize system efficiency, minimize human intervention, and optimize routine tasks
- Monitor and analyze system performance to identify and address bottlenecks before they impact users
- Ensure the infrastructure can handle rapid growth in web traffic and ML data processing
- Participate in 24/7 on‑call rotations (including scheduled shifts and holidays)
- Practice sustainable on‑call response, conduct root‑cause analysis, and lead blameless post‑mortems to prevent recurrence
- Implement monitoring tools (SLIs/SLOs/SLAs) and set up automated alerting and metrics to track system health and performance
- Implement and maintain s...