System Reliability Engineer (San José)

System Reliability Engineer (San José)

04 ago
|
Protection Game
|
San José

04 ago

Protection Game

San José

Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. As the market leader in both data resilience and data security posture management, Veeam is built for the convergence of identity, data, security, and AI risk. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide, who trust Veeam to keep their businesses running. Join us as we go fearlessly forward together, growing, learning, and making a real impact for some of the world’s biggest brands.

About The Role
We are looking for a Site Reliability Engineer (SRE) to join our team and ensure the continuous, reliable operation of company services. This role involves proactive monitoring, incident response, and building resilient observability and escalation practices across our infrastructure.

What You'll Do
- Ensure monitoring and uninterrupted operation of company services
- Write and maintain alerting rules and runbooks




- Perform triage of incoming incidents and initial diagnosis of issues
- Build and maintain escalation chains for incident response
- Perform technical incident resolution activities according to runbooks
- Participate in on-call rotations and post-incident reviews (RCA/postmortems)
- Continuously improve observability coverage and reduce alert noise/false positives
- Collaborate with development and infrastructure teams to identify reliability risks and implement preventive measures

What You'll Bring
- Experience with observability tools (Grafana, ELK, VictoriaMetrics)
- Experience working with Linux
- Experience working with Kubernetes (k8s)
- Experience with AWS and Azure cloud platforms
- Ability to analyze incidents, identify root causes, and propose remediation steps

Bonus Skills
- Experience with Infrastructure as Code (Terraform, Ansible, or similar)
- Scripting skills (Python, Bash) for auto

📌 System Reliability Engineer (San José)
🏢 Protection Game
📍 San José

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: system reliability engineer (san josé) / san josé

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: system reliability engineer (san josé) / san josé