Found Description
Lead Site Reliability Engineer
Key Responsibilities:
Implement and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to drive reliability efforts.
Develop systems that are resilient to failures and ensure 99.9%+ uptime for critical services.
Lead incident response and post-incident reviews (blameless postmortems), ensuring robust root cause analysis and continuous improvement of systems.
Automate incident detection and response using automated runbooks or predefined workflows.
Write software as needed to support reliability or efficiency needs.
Design and implement full observability across systems using modern tools like Open Telemetry for tracing, metrics, and logging.
Use capacity planning, forecasting, and performance testing to ensure that the...