15 ago
|
Infosys
|
Santiago
Domain|Infrastructure-Information Security Management|Information Security Compliance
Domain
Delivery
Interest Group
Infy Chile
Company
IL Chile
Requisition ID
******BR
Job Description – Site Reliability Engineer (SRE)
Role Purpose
The Site Reliability Engineer (SRE) is responsible for ensuring the reliability, availability, and performance of digital services in production, balancing service stability with the ability to deliver change at speed.
The role focuses on strengthening operational resilience through engineering, automation, and proactive reliability practices, working closely with application and platform teams.
Scope of the Role
The locally applied SRE role covers:
Associated infrastructure (cloud, CI/CD pipelines, integrations)
Continuous operations (24/7 reliability mindset, not necessarily shift-based)
Production changes (deployments, configurations)
Incidents, problems, and service degradations
Continuous improvement of stability and operational efficiency
Job Description – Site Reliability Engineer (SRE)
Role Purpose
The Site Reliability Engineer (SRE) is responsible for ensuring the reliability, availability, and performance of digital services in production, balancing service stability with the ability to deliver change at speed.
The role focuses on strengthening operational resilience through engineering, automation, and proactive reliability practices, working closely with application and platform teams.
Scope of the Role
The locally applied SRE role covers:
Production digital services (applications, platforms, data products)
Associated infrastructure (cloud, CI/CD pipelines, integrations)
Continuous operations (24/7 reliability mindset, not necessarily shift-based)
Production changes (deployments, configurations)
Incidents, problems, and service degradations
Continuous improvement of stability and operational efficiency
Key Responsibilities
Service Reliability & Availability
Define, implement, and maintain SLIs and SLOs (availability, latency,
error rates)
Continuously monitor service health and anticipate degradations
Ensure services operate within business‐agreed reliability thresholds
Manage reliability trade‐offs between speed and stability
Incident & Problem Management
Lead or coordinate response to relevant incidents (L2/L3)
Ensure:
Rapid and structured diagnosis
Safe service restoration
Clear and effective communication
Facilitate blameless postmortems
Convert recurring incidents into engineering improvement backlog
Drive long‐term remediation rather than reactive firefighting
Automation & Operational Excellence
Identify repetitive and manual operational tasks
Design and implement automation for:
Deployments
Monitoring and alerting
Health checks
Basic recovery and self‐healing (where applicable)
Reduce toil and increase system resilience through engineering solutions
Change Governance & Production Readiness
Support vendor and internal team change tracking
Ensure changes:
Are traceable
Have defined rollback strategies
Minimize operational risk
Validate operational readiness before production
Participate early in solution and architecture design from a reliability perspective (early involvement)
Metrics, Observability & Continuous Improvement
Define and maintain near real‐time operational KPIs ("service pulse")
Ensure every deviation has:
Clear ownership
Defined corrective actions
Prevent reactive operations by driving data‐driven decision making
Support identification, prioritization, and planning of technical debt remediation
What This Role Is Not
A dedicated incident operator only
An advanced Service Desk
The sole owner of service stability (reliability is shared)
A gatekeeper blocking changes without technical justification
The owner of contractual MOPs
A commercial or account management role
The customer‐side account or delivery lead
Experience & Profile (Indicative)
Proven experience as SRE, Production Engineer, or similar role
Strong background in production systems and reliability engineering
Experience working with:
Cloud platforms
CI/CD pipelines
Monitoring and observability tools
Comfortable operating in product‐oriented or POD‐based team models
Strong problem‐solving, communication, and collaboration skills
Operating Model Alignment
Works embedded or as an enabling function with PODs
Focused on enablement and reliability patterns, not centralized control
Promotes shared ownership of reliability
About Us
Infosys is a global leader in next‐generation digital services and consulting.
We enable clients in more than 50 countries to navigate their digital transformation.
With over four decades of experience in managing the systems and workings of general enterprises, we expertly steer our clients through their digital journey.
We do it by enabling the enterprise with an AI‐powered core that helps prioritize the execution of change.
We also empower the business with agile digital at scale to deliver unprecedented levels of performance and customer delight.
Our always‐on learning agenda drives their continuous improvement through building and transferring digital skills, expertise, and ideas from our innovation ecosystem.
EEO
Infosys provides equal employment opportunities to applicants and employees without regard to race; color; sex; gender identity; sexual orientation; religious practices and observances; national origin; pregnancy, childbirth, or related medical conditions; or disability.
#J-*****-Ljbffr
📌 Site Reliability Engineer (Sre) (Santiago)
🏢 Infosys
📍 Santiago