
Expert Site Reliability Engineer
Altera Digital Health APAC · Michigan, United States
- On site
- Full-time
- $102,500 / year
- Michigan, United States
Job highlights
- Ensure healthcare platform reliability and performance.
- Lead incident response and root cause analysis.
- Automate operations using scripting and IaC.
- Collaborate with engineering and cloud teams.
- Drive continuous improvement in service delivery.
About the role
Site Reliability Engineer (SRE) - Remote
Overview
As a Site Reliability Engineer (SRE) at Altera, you will be responsible for ensuring the reliability, scalability, and performance of our hosted healthcare platforms. This role blends software and systems engineering to enhance service availability, automate operations, and improve the customer experience. You will act as a technical leader in monitoring, troubleshooting, incident response, and continuous improvement across our cloud and hybrid environments.
Key Responsibilities
- Maintain and improve the reliability, availability, and performance of our production environments.
- Lead the investigation and resolution of complex application, database, and infrastructure issues.
- Participate in incident management, conduct root cause analysis (RCA), and contribute to post-incident reviews to prevent future occurrences.
- Define and measure Service Level Indicators (SLIs) and Objectives (SLOs) to meet our service commitments.
- Develop proactive monitoring and alerting strategies to identify and resolve issues before they impact customers.
- Automate operational tasks using scripting and Infrastructure-as-Code (IaC) to improve efficiency.
- Partner with engineering and cloud teams to refine deployment, monitoring, and support processes.
- Provide technical leadership during major incidents and act as a key escalation point for critical issues.
Qualifications
Experience:
- 7+ years of experience supporting enterprise applications, infrastructure, or cloud environments.
- Monitoring & Observability: Strong experience with APM tools such as LogicMonitor, AppDynamics, Azure Monitor, SentryOne, Dynatrace, Datadog, or New Relic.
- Microsoft Stack: Deep knowledge of Windows Server administration, IIS, .NET applications, Windows Clustering, MSMQ, Event Logs, and PerfMon.
- Database Skills: Strong SQL Server experience, including performance tuning, query optimization, blocking analysis, and Always On Availability Groups.
- Cloud & Networking: Experience with Azure cloud environments and a solid understanding of networking fundamentals (DNS, TCP/IP, load balancing, firewalls).
- ITSM & ITIL: Familiarity with ServiceNow (or other ITSM platforms) and ITIL principles.
Preferred Skills
- Scripting with PowerShell, Python, or similar languages.
- Infrastructure as Code (Terraform, ARM Templates, Bicep).
- CI/CD pipelines and deployment automation (Azure DevOps, GitHub Actions).
- Experience with Kubernetes and containerized workloads.
- Experience implementing SLOs, SLIs, and Error Budgets.
- Experience in a healthcare technology or patient care environment.
Education
Bachelor's Degree in Computer Science, Information Technology, or Engineering is preferred; equivalent professional experience will be considered.
Working Arrangements
- This is a remote position open to candidates within the United States.
- You will participate in an on-call rotation to support our 24x7 healthcare environment.
- Occasional after-hours work is required for activations, upgrades, and major incidents.
Travel
Travel is not a requirement for this role.
Salary Range
$95,000 - $110,000
Why Altera?
At Altera Digital Health, you will have the opportunity to profoundly impact the lives of patients by empowering healthcare providers to deliver superior care. You will join a passionate and gifted team committed to innovation and excellence. We offer a competitive compensation and benefits package and the opportunity to work in a fast-paced and dynamic environment.
Key skills/competency
- Site Reliability Engineering
- Cloud Computing (Azure)
- System Administration (Windows Server)
- Database Management (SQL Server)
- Monitoring & Observability
- Automation (Scripting, IaC)
- Incident Management
- Networking Fundamentals
- ITSM/ITIL
- Troubleshooting
Skills & topics
- Site Reliability Engineer
- SRE
- Cloud Engineer
- DevOps Engineer
- Systems Administrator
- Azure
- SQL Server
- Windows Server
- Monitoring
- Automation
- IaC
- Remote Job
- Healthcare Technology
How to get hired
- Tailor your resume: Highlight 7+ years of SRE experience, specific tools (Azure Monitor, Datadog), and Microsoft Stack expertise for this Site Reliability Engineer role.
- Showcase technical skills: Emphasize your experience with Azure, SQL Server, PowerShell/Python scripting, and IaC tools like Terraform in your application.
- Demonstrate problem-solving: Prepare to discuss how you've led incident resolution, conducted RCAs, and defined SLIs/SLOs in past SRE positions.
- Research Altera's mission: Understand their impact on patient care and align your experience with their commitment to innovation and excellence.
- Prepare for remote work: Be ready to discuss your experience with remote collaboration and participation in on-call rotations.
Technical preparation
Behavioral questions
Frequently asked questions
- What are the key responsibilities for a Site Reliability Engineer at Altera Digital Health APAC?
- The key responsibilities for a Site Reliability Engineer at Altera Digital Health APAC include ensuring the reliability, scalability, and performance of hosted healthcare platforms, leading incident response and root cause analysis, automating operational tasks using scripting and Infrastructure-as-Code, and providing technical leadership during critical incidents.
- What specific monitoring and observability tools are preferred for this Site Reliability Engineer role?
- For this Site Reliability Engineer role, strong experience is required with APM tools such as LogicMonitor, AppDynamics, Azure Monitor, SentryOne, Dynatrace, Datadog, or New Relic. Familiarity with these tools is crucial for success.
- Is this Site Reliability Engineer position fully remote, and are there any location restrictions?
- Yes, this Site Reliability Engineer position is fully remote and is open to candidates within the United States. There are no geographical restrictions within the US.
- What kind of technical background is expected for the Site Reliability Engineer role at Altera Digital Health?
- The expected technical background includes 7+ years of experience supporting enterprise applications, infrastructure, or cloud environments, deep knowledge of Windows Server administration and Microsoft Stack, strong SQL Server experience, and familiarity with Azure cloud environments and networking fundamentals.
- Does Altera Digital Health offer opportunities for professional growth in this Site Reliability Engineer position?
- While not explicitly detailed, Altera Digital Health emphasizes innovation and excellence, suggesting a dynamic environment where growth opportunities likely exist for Site Reliability Engineers contributing to their mission of impacting patient lives.
- What is the expected salary range for the Site Reliability Engineer role at Altera Digital Health APAC?
- The expected salary range for the Site Reliability Engineer role at Altera Digital Health APAC is $95,000 to $110,000 annually, with the final offer determined by factors such as experience, skills, and internal equity.
- Will a Site Reliability Engineer be required to participate in an on-call rotation?
- Yes, as a Site Reliability Engineer at Altera Digital Health, you will be required to participate in an on-call rotation to support their 24x7 healthcare environment.
- What are the preferred educational qualifications for a Site Reliability Engineer at Altera Digital Health?
- A Bachelor's Degree in Computer Science, Information Technology, or Engineering is preferred for the Site Reliability Engineer role at Altera Digital Health. However, equivalent professional experience will also be considered.