
Expert Site Reliability Engineer
Altera Digital Health APAC · Colorado, United States
- On site
- Full-time
- $102,500 / year
- Colorado, United States
Job highlights
- Ensure healthcare platform reliability and performance.
- Lead incident response and root cause analysis.
- Automate operations with scripting and IaC.
- Collaborate with engineering and cloud teams.
- Impact patient care through technology innovation.
About the role
Site Reliability Engineer (SRE) - Remote
Overview
As a Site Reliability Engineer (SRE) at Altera, you will be responsible for ensuring the reliability, scalability, and performance of our hosted healthcare platforms. This role blends software and systems engineering to enhance service availability, automate operations, and improve the customer experience. You will act as a technical leader in monitoring, troubleshooting, incident response, and continuous improvement across our cloud and hybrid environments.
Key Responsibilities
- Maintain and improve the reliability, availability, and performance of our production environments.
- Lead the investigation and resolution of complex application, database, and infrastructure issues.
- Participate in incident management, conduct root cause analysis (RCA), and contribute to post-incident reviews to prevent future occurrences.
- Define and measure Service Level Indicators (SLIs) and Objectives (SLOs) to meet our service commitments.
- Develop proactive monitoring and alerting strategies to identify and resolve issues before they impact customers.
- Automate operational tasks using scripting and Infrastructure-as-Code (IaC) to improve efficiency.
- Partner with engineering and cloud teams to refine deployment, monitoring, and support processes.
- Provide technical leadership during major incidents and act as a key escalation point for critical issues.
Qualifications
Experience:
- 7+ years of experience supporting enterprise applications, infrastructure, or cloud environments.
- Monitoring & Observability: Strong experience with APM tools such as LogicMonitor, AppDynamics, Azure Monitor, SentryOne, Dynatrace, Datadog, or New Relic.
- Microsoft Stack: Deep knowledge of Windows Server administration, IIS, .NET applications, Windows Clustering, MSMQ, Event Logs, and PerfMon.
- Database Skills: Strong SQL Server experience, including performance tuning, query optimization, blocking analysis, and Always On Availability Groups.
- Cloud & Networking: Experience with Azure cloud environments and a solid understanding of networking fundamentals (DNS, TCP/IP, load balancing, firewalls).
- ITSM & ITIL: Familiarity with ServiceNow (or other ITSM platforms) and ITIL principles.
Preferred Skills
- Scripting with PowerShell, Python, or similar languages.
- Infrastructure as Code (Terraform, ARM Templates, Bicep).
- CI/CD pipelines and deployment automation (Azure DevOps, GitHub Actions).
- Experience with Kubernetes and containerized workloads.
- Experience implementing SLOs, SLIs, and Error Budgets.
- Experience in a healthcare technology or patient care environment.
Education
Bachelor's Degree in Computer Science, Information Technology, or Engineering is preferred; equivalent professional experience will be considered.
Working Arrangements
This is a remote position open to candidates within the United States. You will participate in an on-call rotation to support our 24x7 healthcare environment. Occasional after-hours work is required for activations, upgrades, and major incidents.
Travel
Travel is not a requirement for this role.
Salary Range
$95,000 - $110,000
Why Altera?
At Altera Digital Health, you will have the opportunity to profoundly impact the lives of patients by empowering healthcare providers to deliver superior care. You will join a passionate and gifted team committed to innovation and excellence. We offer a competitive compensation and benefits package and the opportunity to work in a fast-paced and dynamic environment.
Key skills/competency
- Site Reliability Engineering (SRE)
- Cloud Computing (Azure)
- Monitoring and Observability
- Windows Server Administration
- SQL Server Performance Tuning
- Infrastructure as Code (IaC)
- Scripting (PowerShell, Python)
- Incident Management
- Service Level Objectives (SLOs)
- Healthcare Technology
Skills & topics
- Site Reliability Engineer
- SRE
- Cloud Engineer
- DevOps
- System Administrator
- Azure
- SQL Server
- PowerShell
- Python
- Datadog
- LogicMonitor
- AppDynamics
- Windows Server
- Kubernetes
- Terraform
- Incident Management
- Observability
- Healthcare Technology
- Remote
- US Remote
How to get hired
- Tailor your resume: Highlight 7+ years of experience in enterprise applications, infrastructure, or cloud environments, emphasizing SRE principles.
- Showcase technical skills: Detail your proficiency with APM tools, Microsoft Stack, SQL Server, Azure, and networking fundamentals.
- Demonstrate automation expertise: Provide examples of scripting (PowerShell, Python) and Infrastructure as Code (Terraform, ARM) usage.
- Prepare for interviews: Be ready to discuss incident management, RCA, SLI/SLO definition, and cloud-native technologies like Kubernetes.
- Understand the mission: Articulate how your SRE skills can positively impact patient care in the healthcare technology sector.
Technical preparation
Behavioral questions
Frequently asked questions
- What is the typical career path for a Site Reliability Engineer at Altera Digital Health APAC?
- As a Site Reliability Engineer at Altera Digital Health APAC, your career path can evolve from individual contributor roles focusing on specific infrastructure or application domains to more senior SRE positions. You might eventually move into leadership roles, managing SRE teams, or specializing in areas like cloud architecture, security, or advanced automation. Your growth will be supported by opportunities to tackle complex challenges in the healthcare technology space and continuous learning.
- How does Altera Digital Health APAC prioritize work-life balance for its remote Site Reliability Engineers?
- Altera Digital Health APAC aims to support work-life balance for its remote Site Reliability Engineers by offering a remote-first work environment. While the role requires participation in a 24x7 on-call rotation and occasional after-hours work for critical incidents or upgrades, the company strives to provide flexibility and manage workloads effectively. The focus is on impactful contributions within a structured framework.
- What are the key technologies and tools a Site Reliability Engineer will use at Altera Digital Health APAC?
- A Site Reliability Engineer at Altera Digital Health APAC will extensively use monitoring and observability tools like LogicMonitor, AppDynamics, Azure Monitor, Datadog, or New Relic. You'll work with the Microsoft Stack (Windows Server, IIS, .NET), SQL Server, and Azure cloud environments. Automation will involve scripting languages like PowerShell and Python, and Infrastructure as Code tools such as Terraform. Experience with Kubernetes and CI/CD pipelines (Azure DevOps, GitHub Actions) is also highly valued.
- How important is prior experience in the healthcare industry for this Site Reliability Engineer role?
- While not strictly mandatory, prior experience in a healthcare technology or patient care environment is considered a preferred skill for the Site Reliability Engineer position at Altera Digital Health APAC. Understanding the unique demands and regulatory landscape of healthcare can be a significant advantage in ensuring the reliability and performance of critical patient care platforms.
- What is the interview process like for the Site Reliability Engineer position at Altera Digital Health APAC?
- The interview process for the Site Reliability Engineer role at Altera Digital Health APAC typically involves multiple stages designed to assess your technical expertise and cultural fit. You can expect technical screening, discussions about your experience with monitoring, cloud, databases, and automation, and behavioral questions to gauge your problem-solving and leadership abilities, particularly in incident management scenarios.
- Does Altera Digital Health APAC offer opportunities for professional development for their Site Reliability Engineers?
- Yes, Altera Digital Health APAC is committed to the growth of its employees. For Site Reliability Engineers, this includes opportunities to deepen expertise in cloud technologies, automation, and healthcare-specific platforms. You'll be encouraged to learn and apply new methodologies like SLO implementation and contribute to continuous improvement initiatives.