
Expert Site Reliability Engineer
Altera Digital Health APAC · Florida, United States
- On site
- Full-time
- $100,000 / year
- Florida, United States
Job highlights
- Ensure reliability of healthcare platforms.
- Lead incident response and root cause analysis.
- Develop proactive monitoring and automation.
- Collaborate with engineering and cloud teams.
- Impact patient care through technology.
About the role
Site Reliability Engineer (SRE) - Remote
Overview
As a Site Reliability Engineer (SRE) at Altera, you will be responsible for ensuring the reliability, scalability, and performance of our hosted healthcare platforms. This role blends software and systems engineering to enhance service availability, automate operations, and improve the customer experience. You will act as a technical leader in monitoring, troubleshooting, incident response, and continuous improvement across our cloud and hybrid environments.
Key Responsibilities
- Maintain and improve the reliability, availability, and performance of our production environments.
- Lead the investigation and resolution of complex application, database, and infrastructure issues.
- Participate in incident management, conduct root cause analysis (RCA), and contribute to post-incident reviews to prevent future occurrences.
- Define and measure Service Level Indicators (SLIs) and Objectives (SLOs) to meet our service commitments.
- Develop proactive monitoring and alerting strategies to identify and resolve issues before they impact customers.
- Automate operational tasks using scripting and Infrastructure-as-Code (IaC) to improve efficiency.
- Partner with engineering and cloud teams to refine deployment, monitoring, and support processes.
- Provide technical leadership during major incidents and act as a key escalation point for critical issues.
Qualifications
Experience:
- 7+ years of experience supporting enterprise applications, infrastructure, or cloud environments.
- Monitoring & Observability: Strong experience with APM tools such as LogicMonitor, AppDynamics, Azure Monitor, SentryOne, Dynatrace, Datadog, or New Relic.
- Microsoft Stack: Deep knowledge of Windows Server administration, IIS, .NET applications, Windows Clustering, MSMQ, Event Logs, and PerfMon.
- Database Skills: Strong SQL Server experience, including performance tuning, query optimization, blocking analysis, and Always On Availability Groups.
- Cloud & Networking: Experience with Azure cloud environments and a solid understanding of networking fundamentals (DNS, TCP/IP, load balancing, firewalls).
- ITSM & ITIL: Familiarity with ServiceNow (or other ITSM platforms) and ITIL principles.
Preferred Skills
- Scripting with PowerShell, Python, or similar languages.
- Infrastructure as Code (Terraform, ARM Templates, Bicep).
- CI/CD pipelines and deployment automation (Azure DevOps, GitHub Actions).
- Experience with Kubernetes and containerized workloads.
- Experience implementing SLOs, SLIs, and Error Budgets.
- Experience in a healthcare technology or patient care environment.
Education
Bachelor's Degree in Computer Science, Information Technology, or Engineering is preferred; equivalent professional experience will be considered.
Working Arrangements
- This is a remote position open to candidates within the United States.
- You will participate in an on-call rotation to support our 24x7 healthcare environment.
- Occasional after-hours work is required for activations, upgrades, and major incidents.
Travel
Travel is not a requirement for this role.
Salary Range
$95,000 - $110,000
Why Altera?
At Altera Digital Health, you will have the opportunity to profoundly impact the lives of patients by empowering healthcare providers to deliver superior care. You will join a passionate and gifted team committed to innovation and excellence. We offer a competitive compensation and benefits package and the opportunity to work in a fast-paced and dynamic environment.
Key skills/competency
- Site Reliability Engineering
- Cloud Computing (Azure)
- Monitoring and Observability
- Infrastructure as Code
- Windows Server Administration
- SQL Server
- Incident Management
- Automation Scripting
- Networking Fundamentals
- Healthcare Technology
Skills & topics
- Site Reliability Engineer
- SRE
- Cloud Engineer
- DevOps Engineer
- System Administrator
- Reliability Engineer
- Healthcare Technology
- Azure
- SQL Server
- Windows Server
How to get hired
- Tailor your resume: Highlight 7+ years of experience in enterprise applications, infrastructure, or cloud environments, emphasizing SRE responsibilities.
- Showcase technical skills: Detail your proficiency with APM tools, Windows Server, SQL Server, Azure, and networking fundamentals.
- Demonstrate problem-solving: Provide examples of leading incident resolution, conducting RCAs, and implementing monitoring strategies.
- Emphasize automation and IaC: Highlight experience with scripting (PowerShell, Python) and Infrastructure as Code tools (Terraform, Bicep).
- Prepare for interviews: Be ready to discuss your experience with ITSM, ITIL, and your approach to ensuring service reliability in a healthcare context.
Technical preparation
Behavioral questions
Frequently asked questions
- What are the key responsibilities for a Site Reliability Engineer at Altera Digital Health APAC?
- As a Site Reliability Engineer at Altera Digital Health APAC, you will focus on ensuring the reliability, scalability, and performance of healthcare platforms. Key responsibilities include leading incident response, conducting root cause analysis, developing proactive monitoring and alerting, automating operational tasks using scripting and Infrastructure-as-Code, and providing technical leadership during critical issues.
- What technical skills are most important for the Site Reliability Engineer role at Altera?
- The most important technical skills for this Site Reliability Engineer role include extensive experience with monitoring and observability tools (like Datadog, New Relic), deep knowledge of Microsoft stack (Windows Server, IIS, .NET), strong SQL Server experience (performance tuning, query optimization), and proficiency with Azure cloud environments and networking fundamentals. Familiarity with ITSM and ITIL principles is also crucial.
- Is the Site Reliability Engineer position remote, and if so, where can candidates be located?
- Yes, the Site Reliability Engineer position at Altera Digital Health APAC is a fully remote role. Candidates can be located anywhere within the United States.
- What is the expected experience level for the Site Reliability Engineer position?
- The role requires a significant level of experience, with a minimum of 7+ years of experience supporting enterprise applications, infrastructure, or cloud environments. This background is essential for handling complex issues and leading initiatives in a critical healthcare technology setting.
- Does Altera Digital Health APAC offer opportunities for professional growth for a Site Reliability Engineer?
- While specific growth paths aren't detailed, Altera Digital Health emphasizes empowering healthcare providers and fostering innovation. Joining a passionate team and working in a fast-paced environment offers opportunities to develop expertise in critical healthcare technology and contribute to significant patient impact.
- What is the salary range for the Site Reliability Engineer role at Altera Digital Health APAC?
- The salary range for the Site Reliability Engineer position at Altera Digital Health APAC is between $95,000 and $110,000 annually. This range is determined by various factors including internal equity, market data, and the applicant's skills and experience.
- What is the role of Infrastructure as Code (IaC) in the Site Reliability Engineer position?
- Infrastructure as Code (IaC) is a preferred skill for the Site Reliability Engineer role. It involves using tools like Terraform, ARM Templates, or Bicep to automate the provisioning and management of infrastructure, which is crucial for improving efficiency and ensuring consistent, reliable environments.
- How does Altera Digital Health utilize monitoring and observability tools in this Site Reliability Engineer role?
- Altera Digital Health relies heavily on monitoring and observability tools for this Site Reliability Engineer role. Candidates are expected to have strong experience with APM tools to develop proactive monitoring and alerting strategies, enabling the identification and resolution of issues before they impact customers.