1. What is a Site Reliability Engineer at HTC Global Services?
As a Site Reliability Engineer (SRE) at HTC Global Services, you sit at the critical intersection of software engineering and systems operations. You are tasked with ensuring that our platforms—particularly those driving our Generative AI initiatives and mission-critical Observability frameworks—remain resilient, performant, and scalable under heavy production loads. Your work directly influences the reliability of the tools our clients depend on every day.
This role is not just about maintaining uptime; it is about engineering solutions that prevent downtime before it occurs. You will be instrumental in defining SLOs (Service Level Objectives), optimizing Kubernetes clusters, and architecting systems that can handle the high-throughput demands of modern AI platforms. It is a position of significant influence, requiring you to bridge the gap between development cycles and production stability in a fast-paced environment.



