
Site Reliability Engineer
SS&C Technologies · Tennessee, United States
- Hybrid
- Full-time
- $120,000 / year
- Tennessee, United States
Job highlights
- Lead teams to build resilient cloud platforms.
- Enhance application reliability and observability.
- Automate infrastructure and eliminate toil.
- Integrate security and compliance practices.
- Drive continuous improvement with data.
About the role
Site Reliability Engineer
SS&C Technologies is a leading financial services and healthcare technology company with a global presence. We are seeking a motivated Site Reliability Engineer (SRE) to join our Global Technology Infrastructure SRE team.
About the Role
The SRE will lead technology teams in delivering scalable, resilient, and secure infrastructure platforms and services. This role is crucial for enabling business units to innovate, modernize applications, and manage technical debt. You will foster collaboration between product, engineering, and operations teams, embedding reliability, automation, and compliance into our development processes.
What You Will Get To Do
- Collaborate with Technology Infrastructure teams to build and operate cloud-native platforms, abstracting complexity and accelerating delivery.
- Work with business units and technical teams to enhance application availability, observability, and reliability during Private Cloud migration.
- Improve platform reliability through automatic problem detection, self-healing systems, and effective notification protocols.
- Utilize SLOs, SLIs, and KPIs for prioritization, impact measurement, and continuous improvement.
- Eliminate toil through intelligent automation and agentic workflows.
- Conduct blameless retrospectives and share organizational learnings.
- Foster a culture of ownership, learning, and engineering excellence.
- Integrate DevSecOps, zero-trust principles, and policy-as-code into pipelines.
- Produce and promote Architecture Decision Records (ADRs) and Cloud Well-Architected Frameworks.
- Maintain 24x5 active coverage with regional handoffs and weekend escalation protocols.
What You Will Bring
- 5+ years of professional SRE experience, with 3+ years in financial services or regulated industries preferred.
- Minimum Bachelor’s degree in Computer Science, Engineering, or a related field.
- Proven expertise in architecting, designing, and operating private cloud environments (e.g., VMware, OpenStack, OpenShift Virtualization) and Kubernetes clusters.
- Hands-on experience with infrastructure as code, CI/CD pipelines, and observability platforms (e.g., Prometheus, Splunk).
- Strong understanding of modern systems reliability standards, including KPIs, SLAs, and SLOs.
- Familiarity with financial services regulatory frameworks and their infrastructure impact.
- Familiarity with structured naming conventions and asset management for global infrastructure.
- Experience with financial-grade network segmentation, micro-segmentation, and zero-trust architecture.
- Relevant certifications (TOGAF, AWS Certified Solutions Architect, VMware VCP, Red Hat Certified Architect) are a plus.
- Familiarity with ISO 27001, NIST 800-53, and other security frameworks is a plus.
Our Expectations
- Outstanding organization, project management, and attention to detail.
- Tenacious problem solver and continuous learner, adaptable to new technologies.
- Powerful verbal and written communication skills.
- Ability to quickly establish credibility with diverse technical stakeholders.
- Discretion with confidential information and adherence to compliance requirements.
- Commitment to professional and ethical standards.
- Flexibility for non-traditional hours and occasional travel (< 25%).
Why You Will Love It Here!
- Flexibility: Hybrid Work Model & Business Casual Dress Code
- Future Growth: 401k Matching Program, Professional Development Reimbursement
- Work/Life Balance: Flexible Personal/Vacation Time Off, Sick Leave, Paid Holidays
- Wellbeing: Medical, Dental, Vision, Employee Assistance Program, Parental Leave
- Diverse Perspectives: Commitment to celebrating employee backgrounds, talents, and experiences.
- Training: Hands-On, Team-Customized, including SS&C University
- Extra Perks: Discounts on fitness clubs, travel, and more.
Application Instructions
To apply, please visit our careers page at www.ssctech.com/careers. Applications are accepted on an ongoing basis until the position is filled.
Key skills/competency
- Site Reliability Engineering (SRE)
- Cloud-Native Platforms
- Kubernetes
- Infrastructure as Code
- CI/CD Pipelines
- Observability (Prometheus, Splunk)
- SLOs and KPIs
- DevSecOps
- Zero-Trust Architecture
- Financial Services Technology
Skills & topics
- Site Reliability Engineer
- SRE
- Cloud Engineering
- DevOps
- Infrastructure as Code
- Kubernetes
- Automation
- System Reliability
- Financial Services Technology
- Observability
- CI/CD
- Private Cloud
- VMware
- OpenStack
- OpenShift
- Prometheus
- Splunk
- SLO
- KPI
- Zero Trust
How to get hired
- Tailor your resume: Highlight SRE experience, cloud platforms, CI/CD, and financial services exposure.
- Showcase automation skills: Detail your experience with Infrastructure as Code and scripting in your application.
- Emphasize reliability metrics: Quantify your impact using SLOs, SLIs, and KPIs in your experience.
- Prepare for technical deep dives: Be ready to discuss private cloud architecture, Kubernetes, and observability tools.
- Research SS&C values: Understand their focus on innovation, security, and client trust in your interview.
Technical preparation
Behavioral questions
Frequently asked questions
- What is the work arrangement for the Site Reliability Engineer role at SS&C Technologies?
- The Site Reliability Engineer position at SS&C Technologies is a remote role, with specific state affiliations in FL, TX, GA, NC, AZ, TN. While remote, the role operates within a hybrid work model framework, suggesting potential for occasional in-person collaboration or team events, though the primary work location is remote.
- What are the key technical skills required for the Site Reliability Engineer position?
- Key technical skills for the Site Reliability Engineer role include expertise in private cloud environments (VMware, OpenStack, OpenShift Virtualization), Kubernetes, infrastructure as code platforms, CI/CD pipelines, and observability tools like Prometheus and Splunk. A strong understanding of modern systems reliability standards, including SLOs and KPIs, is also essential.
- Does SS&C Technologies prefer candidates with experience in regulated industries for the SRE role?
- Yes, SS&C Technologies prefers candidates with experience in financial services or other regulated industries for the Site Reliability Engineer role. This includes familiarity with financial services regulatory frameworks and their impact on infrastructure design and operations, as well as experience with financial-grade network segmentation and zero-trust architecture.
- What are the expected soft skills for a Site Reliability Engineer at SS&C Technologies?
- SS&C expects Site Reliability Engineers to possess outstanding organization, project management skills, and attention to detail. They should be tenacious problem solvers, continuous learners, powerful communicators, and capable of establishing credibility with diverse technical stakeholders. Adaptability and the ability to work under pressure are also highly valued.
- How does SS&C Technologies support professional development for its employees, particularly for SREs?
- SS&C Technologies offers several avenues for professional development, including a Professional Development Reimbursement program and SS&C University for hands-on, team-customized training. These resources are designed to support continuous learning and skill enhancement for roles like the Site Reliability Engineer.
- What is the typical career path for a Site Reliability Engineer at SS&C Technologies?
- While a specific career path isn't detailed, a Site Reliability Engineer at SS&C Technologies can expect to grow within the Global Technology Infrastructure SRE team. Potential progression could involve leading more complex projects, specializing in specific areas of cloud infrastructure or security, or moving into more senior SRE or architectural roles.
- How important is experience with specific cloud platforms for this SS&C Technologies SRE role?
- Experience with private cloud environments like VMware, OpenStack, or OpenShift Virtualization, and Kubernetes clusters is crucial. While certifications in AWS are a plus, the primary emphasis for this role is on private cloud expertise and Kubernetes orchestration, aligning with SS&C's infrastructure strategy.
- What does 'eliminating toil' mean in the context of the Site Reliability Engineer role at SS&C?
- Eliminating toil refers to automating repetitive, manual, and often low-value tasks that engineers perform. For an SRE at SS&C, this means using intelligent automation and agentic workflows to reduce manual operational work, freeing up time for more strategic initiatives like improving platform reliability and scalability.