
Expert Site Reliability Engineer
Harris Computer · Minnesota, United States
- On site
- Full-time
- $102,500 / year
- Minnesota, United States
Job highlights
- Ensure platform reliability and scalability in healthcare.
- Lead incident response and root cause analysis.
- Automate operations using scripting and IaC.
- Develop proactive monitoring and alerting strategies.
- Provide technical leadership and support.
About the role
Site Reliability Engineer - Remote
As a Site Reliability Engineer (SRE) at Altera, you will be responsible for ensuring the reliability, scalability, and performance of our hosted healthcare platforms. This role blends software and systems engineering to enhance service availability, automate operations, and improve the customer experience. You will act as a technical leader in monitoring, troubleshooting, incident response, and continuous improvement across our cloud and hybrid environments.
Key Responsibilities
- Maintain and improve the reliability, availability, and performance of our production environments.
- Lead the investigation and resolution of complex application, database, and infrastructure issues.
- Participate in incident management, conduct root cause analysis (RCA), and contribute to post-incident reviews to prevent future occurrences.
- Define and measure Service Level Indicators (SLIs) and Objectives (SLOs) to meet our service commitments.
- Develop proactive monitoring and alerting strategies to identify and resolve issues before they impact customers.
- Automate operational tasks using scripting and Infrastructure-as-Code (IaC) to improve efficiency.
- Partner with engineering and cloud teams to refine deployment, monitoring, and support processes.
- Provide technical leadership during major incidents and act as a key escalation point for critical issues.
Qualifications
- Experience: 7+ years of experience supporting enterprise applications, infrastructure, or cloud environments.
- Monitoring & Observability: Strong experience with APM tools such as LogicMonitor, AppDynamics, Azure Monitor, SentryOne, Dynatrace, Datadog, or New Relic.
- Microsoft Stack: Deep knowledge of Windows Server administration, IIS, .NET applications, Windows Clustering, MSMQ, Event Logs, and PerfMon.
- Database Skills: Strong SQL Server experience, including performance tuning, query optimization, blocking analysis, and Always On Availability Groups.
- Cloud & Networking: Experience with Azure cloud environments and a solid understanding of networking fundamentals (DNS, TCP/IP, load balancing, firewalls).
- ITSM & ITIL: Familiarity with ServiceNow (or other ITSM platforms) and ITIL principles.
Preferred Skills
- Scripting with PowerShell, Python, or similar languages.
- Infrastructure as Code (Terraform, ARM Templates, Bicep).
- CI/CD pipelines and deployment automation (Azure DevOps, GitHub Actions).
- Experience with Kubernetes and containerized workloads.
- Experience implementing SLOs, SLIs, and Error Budgets.
- Experience in a healthcare technology or patient care environment.
Education
Bachelor's Degree in Computer Science, Information Technology, or Engineering is preferred; equivalent professional experience will be considered.
Working Arrangements
This is a remote position open to candidates within the United States. You will participate in an on-call rotation to support our 24x7 healthcare environment. Occasional after-hours work is required for activations, upgrades, and major incidents.
Salary Range
$95,000 - $110,000
Why Altera?
At Altera Digital Health, you will have the opportunity to profoundly impact the lives of patients by empowering healthcare providers to deliver superior care. You will join a passionate and gifted team committed to innovation and excellence. We offer a competitive compensation and benefits package and the opportunity to work in a fast-paced and dynamic environment.
Key skills/competency
- Site Reliability Engineering
- Cloud Computing
- System Administration
- Monitoring and Alerting
- Incident Management
- Root Cause Analysis
- Infrastructure as Code
- Scripting
- Database Management
- Networking
Skills & topics
- Site Reliability Engineer
- SRE
- Cloud Engineer
- DevOps
- System Administrator
- Infrastructure Engineer
- Reliability Engineering
- Azure
- Windows Server
- SQL Server
- Datadog
- New Relic
- PowerShell
- Python
- Terraform
- Kubernetes
- Healthcare IT
How to get hired
- Tailor your resume: Highlight your 7+ years of experience in enterprise applications, infrastructure, or cloud environments, emphasizing SRE principles.
- Showcase technical skills: Detail your expertise in monitoring tools (Datadog, New Relic), Microsoft Stack, SQL Server, and Azure cloud environments.
- Demonstrate problem-solving: Provide examples of leading incident resolution, RCA, and developing proactive monitoring strategies.
- Emphasize automation: Highlight your experience with scripting (PowerShell, Python) and Infrastructure as Code (Terraform, ARM Templates).
- Prepare for technical interviews: Be ready to discuss complex troubleshooting scenarios and your approach to ensuring system reliability.
Technical preparation
Behavioral questions
Frequently asked questions
- What are the key responsibilities for a Site Reliability Engineer at Altera?
- The key responsibilities for a Site Reliability Engineer at Altera include maintaining and improving the reliability, availability, and performance of production environments, leading incident investigations and root cause analysis, developing proactive monitoring and alerting, automating operational tasks with scripting and IaC, and providing technical leadership during critical incidents.
- What technical skills are essential for this Site Reliability Engineer role?
- Essential technical skills include 7+ years of experience, strong proficiency with APM tools (LogicMonitor, Datadog, etc.), deep knowledge of Windows Server administration, IIS, .NET, SQL Server performance tuning, and Azure cloud environments. Familiarity with ITSM and ITIL principles is also important.
- Is this Site Reliability Engineer position remote, and where are candidates located?
- Yes, this is a remote position open to candidates within the United States. The role requires participation in an on-call rotation to support a 24x7 healthcare environment.
- What is the expected salary range for the Site Reliability Engineer position?
- The salary range for the Site Reliability Engineer position is $95,000 to $110,000 annually, with the final offer determined by factors such as experience, skills, and internal equity.
- What are the preferred skills for a Site Reliability Engineer at Altera?
- Preferred skills include scripting with PowerShell or Python, experience with Infrastructure as Code (Terraform, ARM Templates), CI/CD pipelines, Kubernetes, implementing SLOs/SLIs, and experience in the healthcare technology sector.
- What is the educational requirement for this Site Reliability Engineer role?
- A Bachelor's Degree in Computer Science, Information Technology, or Engineering is preferred for the Site Reliability Engineer role. However, equivalent professional experience will also be considered.
- How does Altera Digital Health impact patient care?
- Altera Digital Health empowers healthcare providers to deliver superior care, thereby profoundly impacting the lives of patients. The company fosters a passionate team committed to innovation and excellence in the healthcare technology sector.