
Expert Site Reliability Engineer
Harris Computer · Illinois, United States
- On site
- Full-time
- $102,500 / year
- Illinois, United States
Job highlights
- Ensure healthcare platform reliability and scalability.
- Lead incident response and root cause analysis.
- Develop monitoring and automation strategies.
- Collaborate with engineering and cloud teams.
- Impact patient care through technology.
About the role
Site Reliability Engineer (SRE) - Remote
Overview
As a Site Reliability Engineer (SRE) at Altera, you will be responsible for ensuring the reliability, scalability, and performance of our hosted healthcare platforms. This role blends software and systems engineering to enhance service availability, automate operations, and improve the customer experience. You will act as a technical leader in monitoring, troubleshooting, incident response, and continuous improvement across our cloud and hybrid environments.
Key Responsibilities
- Maintain and improve the reliability, availability, and performance of our production environments.
- Lead the investigation and resolution of complex application, database, and infrastructure issues.
- Participate in incident management, conduct root cause analysis (RCA), and contribute to post-incident reviews to prevent future occurrences.
- Define and measure Service Level Indicators (SLIs) and Objectives (SLOs) to meet our service commitments.
- Develop proactive monitoring and alerting strategies to identify and resolve issues before they impact customers.
- Automate operational tasks using scripting and Infrastructure-as-Code (IaC) to improve efficiency.
- Partner with engineering and cloud teams to refine deployment, monitoring, and support processes.
- Provide technical leadership during major incidents and act as a key escalation point for critical issues.
Qualifications
Experience:
- 7+ years of experience supporting enterprise applications, infrastructure, or cloud environments.
- Monitoring & Observability: Strong experience with APM tools such as LogicMonitor, AppDynamics, Azure Monitor, SentryOne, Dynatrace, Datadog, or New Relic.
- Microsoft Stack: Deep knowledge of Windows Server administration, IIS, .NET applications, Windows Clustering, MSMQ, Event Logs, and PerfMon.
- Database Skills: Strong SQL Server experience, including performance tuning, query optimization, blocking analysis, and Always On Availability Groups.
- Cloud & Networking: Experience with Azure cloud environments and a solid understanding of networking fundamentals (DNS, TCP/IP, load balancing, firewalls).
- ITSM & ITIL: Familiarity with ServiceNow (or other ITSM platforms) and ITIL principles.
Preferred Skills
- Scripting with PowerShell, Python, or similar languages.
- Infrastructure as Code (Terraform, ARM Templates, Bicep).
- CI/CD pipelines and deployment automation (Azure DevOps, GitHub Actions).
- Experience with Kubernetes and containerized workloads.
- Experience implementing SLOs, SLIs, and Error Budgets.
- Experience in a healthcare technology or patient care environment.
Education
Bachelor's Degree in Computer Science, Information Technology, or Engineering is preferred; equivalent professional experience will be considered.
Working Arrangements
- This is a remote position open to candidates within the United States.
- You will participate in an on-call rotation to support our 24x7 healthcare environment.
- Occasional after-hours work is required for activations, upgrades, and major incidents.
Travel
Travel is not a requirement for this role.
Salary Range
$95,000 - $110,000
Why Altera?
At Altera Digital Health, you will have the opportunity to profoundly impact the lives of patients by empowering healthcare providers to deliver superior care. You will join a passionate and gifted team committed to innovation and excellence. We offer a competitive compensation and benefits package and the opportunity to work in a fast-paced and dynamic environment.
Key skills/competency
- Site Reliability Engineering
- Cloud Computing (Azure)
- System Administration (Windows Server)
- Database Management (SQL Server)
- Monitoring and Alerting
- Automation and Scripting
- Infrastructure as Code
- Incident Management
- Networking Fundamentals
- Troubleshooting
Skills & topics
- Site Reliability Engineer
- SRE
- Cloud Engineer
- Systems Administrator
- DevOps
- Azure
- SQL Server
- Windows Server
- Monitoring
- Automation
- Infrastructure as Code
- Incident Management
- Remote
- Healthcare IT
How to get hired
- Tailor your resume: Highlight 7+ years of SRE experience, Azure, SQL Server, and Windows administration, aligning with job description keywords.
- Showcase automation skills: Emphasize experience with scripting (PowerShell, Python) and Infrastructure-as-Code tools (Terraform).
- Demonstrate monitoring expertise: Detail your proficiency with APM tools like Datadog, LogicMonitor, or Azure Monitor.
- Prepare for technical interviews: Be ready to discuss troubleshooting complex issues, SQL Server performance tuning, and cloud/networking concepts.
- Understand the healthcare context: If applicable, mention any experience in healthcare technology or patient care environments.
Technical preparation
Behavioral questions
Frequently asked questions
- What is the salary range for the Site Reliability Engineer role at Harris Computer?
- The salary range for the Site Reliability Engineer position at Harris Computer is $95,000 to $110,000 annually. This range is determined by factors such as experience, skills, and location.
- Is the Site Reliability Engineer position at Harris Computer a remote role?
- Yes, the Site Reliability Engineer position at Harris Computer is a fully remote role, open to candidates within the United States.
- What are the key technical skills required for the Site Reliability Engineer role?
- Key technical skills include 7+ years of experience in enterprise applications/infrastructure/cloud, strong APM tool experience (e.g., Datadog, Azure Monitor), deep knowledge of Windows Server administration, SQL Server performance tuning, and Azure cloud/networking fundamentals.
- Does Harris Computer offer opportunities for professional growth in this Site Reliability Engineer role?
- While not explicitly detailed, SRE roles at companies like Harris Computer typically offer growth through exposure to complex systems, advanced technologies, and leadership opportunities in incident management and process improvement.
- What is the expected commitment for the on-call rotation for the Site Reliability Engineer?
- The Site Reliability Engineer will participate in an on-call rotation to support the 24x7 healthcare environment, which may include occasional after-hours work for critical tasks.
- How does Harris Computer approach Incident Management for their hosted healthcare platforms?
- Harris Computer emphasizes leading investigation and resolution of complex issues, participating in incident management, conducting root cause analysis (RCA), and contributing to post-incident reviews to prevent recurrence.
- What is the educational requirement for the Site Reliability Engineer position?
- A Bachelor's Degree in Computer Science, Information Technology, or Engineering is preferred, but equivalent professional experience will also be considered for the Site Reliability Engineer role.
- What type of monitoring and observability tools does Harris Computer utilize for their SRE role?
- The role requires strong experience with APM tools such as LogicMonitor, AppDynamics, Azure Monitor, SentryOne, Dynatrace, Datadog, or New Relic.