
Expert Site Reliability Engineer
Altera Digital Health APAC · Minnesota, United States
- On site
- Full-time
- $102,500 / year
- Minnesota, United States
Job highlights
- Ensure platform reliability, scalability, and performance.
- Lead complex issue investigation and resolution.
- Automate operations using scripting and IaC.
- Collaborate with engineering and cloud teams.
- Impact patient lives through healthcare technology.
About the role
Site Reliability Engineer (SRE) - Remote
Overview
As a Site Reliability Engineer (SRE) at Altera, you will be responsible for ensuring the reliability, scalability, and performance of our hosted healthcare platforms. This role blends software and systems engineering to enhance service availability, automate operations, and improve the customer experience. You will act as a technical leader in monitoring, troubleshooting, incident response, and continuous improvement across our cloud and hybrid environments.
Key Responsibilities
- Maintain and improve the reliability, availability, and performance of our production environments.
- Lead the investigation and resolution of complex application, database, and infrastructure issues.
- Participate in incident management, conduct root cause analysis (RCA), and contribute to post-incident reviews to prevent future occurrences.
- Define and measure Service Level Indicators (SLIs) and Objectives (SLOs) to meet our service commitments.
- Develop proactive monitoring and alerting strategies to identify and resolve issues before they impact customers.
- Automate operational tasks using scripting and Infrastructure-as-Code (IaC) to improve efficiency.
- Partner with engineering and cloud teams to refine deployment, monitoring, and support processes.
- Provide technical leadership during major incidents and act as a key escalation point for critical issues.
Qualifications
Experience:
- 7+ years of experience supporting enterprise applications, infrastructure, or cloud environments.
- Monitoring & Observability: Strong experience with APM tools such as LogicMonitor, AppDynamics, Azure Monitor, SentryOne, Dynatrace, Datadog, or New Relic.
- Microsoft Stack: Deep knowledge of Windows Server administration, IIS, .NET applications, Windows Clustering, MSMQ, Event Logs, and PerfMon.
- Database Skills: Strong SQL Server experience, including performance tuning, query optimization, blocking analysis, and Always On Availability Groups.
- Cloud & Networking: Experience with Azure cloud environments and a solid understanding of networking fundamentals (DNS, TCP/IP, load balancing, firewalls).
- ITSM & ITIL: Familiarity with ServiceNow (or other ITSM platforms) and ITIL principles.
Preferred Skills
- Scripting with PowerShell, Python, or similar languages.
- Infrastructure as Code (Terraform, ARM Templates, Bicep).
- CI/CD pipelines and deployment automation (Azure DevOps, GitHub Actions).
- Experience with Kubernetes and containerized workloads.
- Experience implementing SLOs, SLIs, and Error Budgets.
- Experience in a healthcare technology or patient care environment.
Education
Bachelor's Degree in Computer Science, Information Technology, or Engineering is preferred; equivalent professional experience will be considered.
Working Arrangements
- This is a remote position open to candidates within the United States.
- You will participate in an on-call rotation to support our 24x7 healthcare environment.
- Occasional after-hours work is required for activations, upgrades, and major incidents.
Travel
Travel is not a requirement for this role.
Salary Range
$95,000-$110,000
Why Altera?
At Altera Digital Health, you will have the opportunity to profoundly impact the lives of patients by empowering healthcare providers to deliver superior care. You will join a passionate and gifted team committed to innovation and excellence. We offer a competitive compensation and benefits package and the opportunity to work in a fast-paced and dynamic environment.
Key skills/competency
- Site Reliability Engineering
- Cloud Computing (Azure)
- Monitoring and Observability
- Microsoft Stack Administration
- SQL Server Performance Tuning
- Automation (Scripting, IaC)
- Incident Management
- Networking Fundamentals
- ITSM/ITIL
- Problem Solving
Skills & topics
- Site Reliability Engineer
- SRE
- Cloud Engineer
- DevOps
- Azure
- Windows Server
- SQL Server
- Automation
- Monitoring
- Healthcare IT
How to get hired
- Tailor your resume: Highlight 7+ years of experience in enterprise applications, cloud environments, and specific tools like Azure Monitor, Datadog, and SQL Server.
- Showcase technical skills: Emphasize experience with Windows Server, .NET, networking, and ITSM/ITIL principles.
- Demonstrate automation expertise: Detail your proficiency in scripting (PowerShell, Python) and Infrastructure as Code (Terraform).
- Prepare for technical interviews: Be ready to discuss complex troubleshooting scenarios, monitoring strategies, and incident response.
- Research Altera Digital Health: Understand their mission to empower healthcare providers and impact patient lives.
Technical preparation
Behavioral questions
Frequently asked questions
- What are the key responsibilities for a Site Reliability Engineer at Altera Digital Health APAC?
- As a Site Reliability Engineer (SRE) at Altera Digital Health APAC, your primary responsibilities include ensuring the reliability, scalability, and performance of hosted healthcare platforms. This involves leading incident response, conducting root cause analysis, developing proactive monitoring, automating operational tasks with scripting and Infrastructure-as-Code, and collaborating with engineering teams. You'll also be a technical leader in troubleshooting and continuous improvement across cloud and hybrid environments.
- What technical skills are essential for the Remote Site Reliability Engineer role at Altera?
- Essential technical skills include 7+ years of experience supporting enterprise applications, infrastructure, or cloud environments. Strong experience with APM tools (LogicMonitor, Datadog, etc.), deep knowledge of Microsoft Stack (Windows Server, IIS, .NET), robust SQL Server skills (performance tuning, query optimization), and experience with Azure cloud environments and networking fundamentals are crucial. Familiarity with ITSM platforms like ServiceNow and ITIL principles is also important.
- Is the Site Reliability Engineer position at Altera Digital Health remote?
- Yes, the Site Reliability Engineer position at Altera Digital Health is a remote role, open to candidates located within the United States.
- What is the required education for the Site Reliability Engineer role?
- A Bachelor's Degree in Computer Science, Information Technology, or Engineering is preferred for the Site Reliability Engineer role. However, equivalent professional experience will also be considered.
- What kind of on-call duties are expected for this Site Reliability Engineer role?
- As a Site Reliability Engineer, you will participate in an on-call rotation to support Altera's 24x7 healthcare environment. Occasional after-hours work may also be required for system activations, upgrades, and major incidents.
- How does Altera Digital Health leverage technology to impact patient care?
- Altera Digital Health empowers healthcare providers with technology to deliver superior care, directly impacting patient lives. As an SRE, you contribute to this mission by ensuring the stable and high-performing operation of their healthcare platforms, enabling seamless delivery of care.
- What are the preferred skills for the Site Reliability Engineer position?
- Preferred skills include scripting with PowerShell or Python, experience with Infrastructure as Code tools like Terraform, CI/CD pipeline knowledge (Azure DevOps), experience with Kubernetes and containerized workloads, and familiarity with implementing SLOs and SLIs. Experience in healthcare technology is also a plus.