
Expert Site Reliability Engineer
Altera Digital Health APAC · Pennsylvania, United States
- On site
- Full-time
- $102,500 / year
- Pennsylvania, United States
Job highlights
- Ensure platform reliability, scalability, and performance.
- Lead incident response and root cause analysis.
- Automate operations with scripting and IaC.
- Monitor and improve services using APM tools.
- Collaborate with engineering for process refinement.
About the role
Expert Site Reliability Engineer (SRE) - Remote
As a Site Reliability Engineer (SRE) at Altera, you will be responsible for ensuring the reliability, scalability, and performance of our hosted healthcare platforms. This role blends software and systems engineering to enhance service availability, automate operations, and improve the customer experience. You will act as a technical leader in monitoring, troubleshooting, incident response, and continuous improvement across our cloud and hybrid environments.
Key Responsibilities
- Maintain and improve the reliability, availability, and performance of our production environments.
- Lead the investigation and resolution of complex application, database, and infrastructure issues.
- Participate in incident management, conduct root cause analysis (RCA), and contribute to post-incident reviews to prevent future occurrences.
- Define and measure Service Level Indicators (SLIs) and Objectives (SLOs) to meet our service commitments.
- Develop proactive monitoring and alerting strategies to identify and resolve issues before they impact customers.
- Automate operational tasks using scripting and Infrastructure-as-Code (IaC) to improve efficiency.
- Partner with engineering and cloud teams to refine deployment, monitoring, and support processes.
- Provide technical leadership during major incidents and act as a key escalation point for critical issues.
Qualifications
- Experience: 7+ years of experience supporting enterprise applications, infrastructure, or cloud environments.
- Monitoring & Observability: Strong experience with APM tools such as LogicMonitor, AppDynamics, Azure Monitor, SentryOne, Dynatrace, Datadog, or New Relic.
- Microsoft Stack: Deep knowledge of Windows Server administration, IIS, .NET applications, Windows Clustering, MSMQ, Event Logs, and PerfMon.
- Database Skills: Strong SQL Server experience, including performance tuning, query optimization, blocking analysis, and Always On Availability Groups.
- Cloud & Networking: Experience with Azure cloud environments and a solid understanding of networking fundamentals (DNS, TCP/IP, load balancing, firewalls).
- ITSM & ITIL: Familiarity with ServiceNow (or other ITSM platforms) and ITIL principles.
Preferred Skills
- Scripting with PowerShell, Python, or similar languages.
- Infrastructure as Code (Terraform, ARM Templates, Bicep).
- CI/CD pipelines and deployment automation (Azure DevOps, GitHub Actions).
- Experience with Kubernetes and containerized workloads.
- Experience implementing SLOs, SLIs, and Error Budgets.
- Experience in a healthcare technology or patient care environment.
Education
Bachelor's Degree in Computer Science, Information Technology, or Engineering is preferred; equivalent professional experience will be considered.
Working Arrangements
This is a remote position open to candidates within the United States. You will participate in an on-call rotation to support our 24x7 healthcare environment. Occasional after-hours work is required for activations, upgrades, and major incidents.
Key skills/competency
- Site Reliability Engineering
- Cloud Computing (Azure)
- Monitoring and Observability
- Infrastructure as Code
- Scripting (PowerShell, Python)
- SQL Server
- Windows Server Administration
- Incident Management
- Service Level Objectives (SLOs)
- Networking Fundamentals
Skills & topics
- Site Reliability Engineer
- SRE
- Cloud Engineer
- DevOps
- System Administrator
- Reliability Engineering
- Azure
- SQL Server
- PowerShell
- Python
- Datadog
- Windows Server
- Incident Management
- Infrastructure as Code
- Remote
How to get hired
- Tailor your resume: Highlight 7+ years supporting enterprise applications, cloud, and infrastructure, emphasizing your SRE experience.
- Showcase technical skills: Detail your proficiency in Azure, Windows Server, SQL Server, and APM tools like Datadog or New Relic.
- Demonstrate automation expertise: Provide examples of your scripting (PowerShell, Python) and Infrastructure-as-Code (Terraform) experience.
- Emphasize problem-solving: Include instances where you led incident resolution, conducted RCAs, and improved system reliability.
- Research Altera's mission: Understand their impact on patient care and align your application with their values.
Technical preparation
Behavioral questions
Frequently asked questions
- What is the remote work policy for the Expert Site Reliability Engineer role at Altera Digital Health APAC?
- The Expert Site Reliability Engineer position at Altera Digital Health APAC is a remote role open to candidates within the United States. This means you can work from your home location anywhere in the U.S.
- What is the expected experience level for the Expert Site Reliability Engineer position?
- The role requires a minimum of 7 years of experience supporting enterprise applications, infrastructure, or cloud environments. Significant experience with monitoring tools, Microsoft stack, SQL Server, and Azure is also essential.
- Does the Expert Site Reliability Engineer role require on-call availability?
- Yes, as an Expert Site Reliability Engineer at Altera Digital Health APAC, you will be required to participate in an on-call rotation to support their 24x7 healthcare environment. Occasional after-hours work may also be necessary.
- What are the key technical skills needed for the Expert Site Reliability Engineer role?
- Key technical skills include strong experience with APM tools (LogicMonitor, Datadog, etc.), deep knowledge of Windows Server administration, SQL Server performance tuning, Azure cloud environments, and networking fundamentals. Scripting with PowerShell or Python and Infrastructure-as-Code are also highly valued.
- What is the salary range for the Expert Site Reliability Engineer position at Altera Digital Health APAC?
- The salary range for this position is $95,000 to $110,000 annually. The final salary offered will depend on factors such as your experience, skills, and internal equity.
- What kind of impact can I make as a Site Reliability Engineer at Altera Digital Health APAC?
- As a Site Reliability Engineer, you will profoundly impact patients' lives by ensuring healthcare providers can deliver superior care through reliable and high-performing healthcare platforms. You'll contribute to innovation and excellence within the team.