PitchMeAI
McGraw Hill

Lead Site Reliability Engineer

McGraw Hill · United States

  • Hybrid
  • Full-time
  • $155,000 / year
  • United States
Tailored resumekeyword-matched to this role.
Hiring managerwe find who's hiring.
Intro emaildrafted to reach them directly.

Job highlights

  • Lead SRE team for K-12 digital learning platforms.
  • Ensure reliability, scalability, and performance of cloud infrastructure.
  • Utilize AWS, Terraform, and observability tools.
  • Drive automation and enhance system resiliency.
  • Mentor engineers and collaborate with cross-functional teams.

About the role

Lead Site Reliability Engineer - McGraw Hill

Impact the Moment

Could your creative thinking build the future? A Lead Site Reliability Engineer at McGraw Hill makes a difference for learners and educators across the world. Our team needs individuals with new ideas who connect with people in innovative ways.

How can you make an Impact?

McGraw Hill, a leading provider of digital educational resources and content, is seeking a Lead Site Reliability Engineer to lead a team of 6 Engineers for our Digital Platform Group in supporting our K–12 learning platforms. These platforms serve millions of students and educators nationwide, and you’ll play a key role in ensuring their reliability, scalability, and performance. Working closely with engineering and product teams, you’ll leverage your expertise in AWS, Terraform, and observability tools to drive automation, enhance resiliency, and maintain the health of our cloud-based infrastructure.

This is a remote position open to applicants authorized to work for any employer within the United States.

What you will be doing:

  • Lead a 6 member SRE team supporting production infrastructure and services
  • Manage backlog, sprint planning, and team velocity
  • Own reliability, uptime, security, cost, and performance of services
  • Define and monitor SLOs for application workloads
  • Plan on-call rotations and work to reduce alert fatigue
  • Forecast seasonal growth and capacity planning
  • Mentor engineers and foster professional growth
  • Report status and issues to leadership monthly
  • Partner with development teams
  • Collaborate with CyberSecurity on risk mitigation
  • Collaborate with FinOps on cost reduction
  • Design and troubleshoot highly-distributed, cloud-based production systems
  • Maintain infrastructure-as-code and monitoring-as-code practices
  • Improve system resiliency through failure injection and chaos testing
  • Participate in on-call rotation and resolve operational issues
  • Optimize existing systems for performance and cost
  • Ensure telemetry provides visibility to application performance
  • Support agile development practices and code reviews

We’re looking for someone with:

  • 5+ years of experience in SRE, DevOps, or Software Engineering roles supporting enterprise applications.
  • Strong problem-solving, triage, and root cause analysis skills with a systems engineering mindset.
  • Deep expertise in the AWS ecosystem, with hands-on experience across core services including primarily EKS, RDS, EKS, IAM, CloudWatch, and networking configurations.
  • Expertise with Terraform for managing and automating scalable cloud infrastructure.
  • Skilled in CI/CD pipelines (e.g., GitHub Actions) and managing end-to-end software delivery lifecycles.
  • Strong familiarity with telemetry and observability tools (e.g., New Relic, Datadog), including querying logs and metrics for performance monitoring.

Why work for us?

The work you do at McGraw Hill will be work that matters. We are collectively designing content that will build the future of education. Play your part and experience a sense of fulfillment that will inspire you to even greater heights.

The pay range for this position is between $124,000 - $155,000 annually. However, base pay offered may vary depending on job-related knowledge, skills, experience, and location. An annual bonus plan may be provided as part of the compensation package, in addition to a full range of medical and/or other benefits, depending on the position offered. Click here to learn more about our benefit offerings.

McGraw Hill recruiters always use a “@mheducation.com” or “@careers.mheducation.com” email addresses and/or from our Applicant Tracking System, iCIMS. Any variation of this email domain should be considered suspicious. Additionally, McGraw Hill recruiters and authorized representatives will never request sensitive information in email.

Key skills/competency

  • Site Reliability Engineering
  • AWS
  • Terraform
  • Observability
  • EKS
  • CI/CD
  • Python
  • Linux
  • Cloud Infrastructure
  • Team Leadership

Skills & topics

  • Site Reliability Engineer
  • Lead
  • SRE
  • DevOps
  • AWS
  • Terraform
  • EKS
  • Cloud
  • Infrastructure
  • Leadership
  • Remote
  • K-12
  • Education

How to get hired

  • Tailor your resume: Highlight SRE experience, AWS, Terraform, and leadership skills for the Lead Site Reliability Engineer role.
  • Showcase technical expertise: Emphasize your experience with EKS, CI/CD, observability tools, and infrastructure-as-code.
  • Demonstrate leadership: Provide examples of mentoring teams, managing backlogs, and improving system reliability.
  • Prepare for technical interviews: Be ready to discuss AWS services, Terraform, troubleshooting distributed systems, and SLO definition.
  • Understand company values: Research McGraw Hill's mission in education and how your role contributes to their impact.

Technical preparation

Master AWS core services (EKS, RDS, IAM).,Practice Terraform for infrastructure automation.,Build CI/CD pipelines with GitHub Actions.,Familiarize with New Relic/Datadog for observability.

Behavioral questions

Describe a time you reduced alert fatigue.,How do you mentor junior engineers?,Share an example of capacity planning.,Explain your approach to system resiliency.

Frequently asked questions

What is the expected experience level for a Lead Site Reliability Engineer at McGraw Hill?
McGraw Hill typically looks for candidates with 5+ years of experience in SRE, DevOps, or Software Engineering roles supporting enterprise applications for their Lead Site Reliability Engineer positions. Demonstrating a strong systems engineering mindset and problem-solving skills is also crucial.
What cloud technologies are most important for the Lead Site Reliability Engineer role at McGraw Hill?
Deep expertise in the AWS ecosystem is essential, with hands-on experience in core services like EKS, RDS, IAM, CloudWatch, and networking configurations. Proficiency with Terraform for infrastructure automation is also a key requirement.
How does McGraw Hill approach on-call rotations and alert fatigue for SRE teams?
The Lead Site Reliability Engineer will be responsible for planning on-call rotations and actively working to reduce alert fatigue. This suggests a focus on improving monitoring and alerting systems to ensure efficient incident response.
What kind of impact can a Lead Site Reliability Engineer have at McGraw Hill?
As a Lead Site Reliability Engineer, you will directly impact millions of students and educators by ensuring the reliability, scalability, and performance of K-12 learning platforms. Your work contributes to the future of education by supporting critical digital resources.
Is this Lead Site Reliability Engineer position remote, and are there location restrictions?
Yes, this Lead Site Reliability Engineer position is fully remote and open to applicants authorized to work for any employer within the United States.
What are the key responsibilities for a Lead Site Reliability Engineer at McGraw Hill?
Key responsibilities include leading an SRE team, managing production infrastructure, defining SLOs, ensuring reliability and performance, mentoring engineers, and collaborating with development and cybersecurity teams.
How does McGraw Hill handle career growth for its SRE team members?
The role includes mentoring engineers and fostering professional growth, indicating a commitment to employee development. McGraw Hill aims to inspire its employees to greater heights through impactful work.