PitchMeAI
Block

Senior Site Reliability Engineer

Block · United States

  • Hybrid
  • Internship
  • $150,000 / year
  • United States
Tailored resumekeyword-matched to this role.
Hiring managerwe find who's hiring.
Intro emaildrafted to reach them directly.

Job highlights

  • Improve platform reliability using AI-driven tools.
  • Enhance observability and incident response.
  • Participate in critical service oncall rotations.
  • Lead incident command and mitigation efforts.
  • Build scalable, reliable distributed platforms.

About the role

About Block

Block is one company built from many blocks, all united by the same purpose of economic empowerment. The blocks that form our foundational teams — People, Finance, Counsel, Hardware, Information Security, Platform Infrastructure Engineering, and more — provide support and guidance at the corporate level. They work across business groups and around the globe, spanning time zones and disciplines to develop inclusive People policies, forecast finances, give legal counsel, safeguard systems, nurture new initiatives, and more. Every challenge creates possibilities, and we need different perspectives to see them all. Bring yours to Block.

The Role

As a member of the SRE team, you will proactively and reactively improve the reliability of Block's platform and critical infrastructure. You are metrics-driven, systems-oriented, and focused on building distributed platforms that enable safe, scalable product development.

You will leverage and continuously improve AI-driven tooling and automation to enhance observability, accelerate incident detection and response, and reduce operational toil. This includes applying AI to incident analysis, alert tuning, and operational workflows.

You will participate in primary platform oncall (12 hours per day, one week every few weeks, depending on team size), supporting Block's most critical (Tier 0) services. In this role, you will lead incident command, coordinate mitigation, and drive effective escalation during high-severity events.

What You Will Do

  • Build and extend platforms to improve system reliability
  • Work on team goals that encompass reliability for the entire company
  • Standardize reliability tools across multiple platforms and organizations
  • Triage, coordinate, and lead stabilization of sev 0–1 incidents
  • Serve as primary oncall, maintaining structured escalation paths and exercising leadership escalation
  • Drive platform-wide reliability improvements, shared operational tooling, and deploy-safety patterns
  • Use AI-driven systems to improve signal detection, reduce noise, and accelerate root cause analysis
  • Design and implement safe deployment patterns (progressive delivery, automated rollback, guardrails)

What You Have

  • Drive to root cause systems with many moving parts and take the necessary steps to fix them
  • Demonstrated technical initiative and leadership on previous projects, especially those with a backend/platform focus
  • Familiarity with AI-driven tooling for observability, incident analysis, or automation
  • A mindset that naturally reaches for AI to accelerate problem-solving and reduce toil
  • Experience running production oncall for high-availability systems
  • Strong incident management skills — structured triage, mitigation under pressure, blameless postmortems
  • Fluency with CI/CD pipelines, progressive rollout strategies, and rollback automation
  • Monitoring & observability expertise — building/tuning alerts for uptime, error rates, latency regression, and resource exhaustion
  • Ability to create and maintain evidence-based maturity assessments using trailing 90-day data windows.
  • Comfort with vendor/dependency management — maintaining validated escalation contacts reachable within ≤ 5 minutes.
  • Boundless curiosity, autonomy, and a strong sense of accountability
  • A strong desire to perform and grow as an engineer
  • 5+ years of software development experience

Technologies We Use and Teach

  • Kotlin, Modern Java (11+)
  • HTTP, JSON, gRPC, and Protocol Buffers
  • MySQL / Vitess / DynamoDB
  • Event driven architectures
  • DataDog
  • LaunchDarkly
  • Terraform, Kubernetes, Istio/Envoy
  • Amazon Web Services

Diversity and Inclusion

We’re working to build an inclusive economy where our customers have equal access to opportunity, and we strive to live by these same values in building our workplace. Block is a proud equal opportunity employer. We work hard to evaluate all employees and job applicants consistently, without regard to identity or other legally protected class. We believe in being fair, and are committed to an inclusive interview experience, including providing reasonable accommodations to disabled applicants throughout the recruitment process. We encourage applicants to share any needed accommodations with their recruiter, who will treat these requests as confidentially as possible. Want to learn more about what we’re doing to build a workplace that is fair and square? Check out our I+D page.

Block is a globally distributed company and this role will require working with other employees in multiple time zones. You may be required to perform work outside of normal business hours as part of this role.

Application Guidelines

Candidates may submit up to 9 active applications within a 60-day period. Reapplications to the same role are accepted 90 days after a previous application has been reviewed.

Use of AI in Our Hiring Process

We may use automated AI tools to evaluate job applications for efficiency and consistency. These tools comply with local regulations, including bias audits, and we handle all personal data in accordance with state and local privacy laws. Contact us here with hiring practice or data privacy questions.

Benefits

Every benefit we offer is designed with one goal: empowering you to do the best work of your career while building the life you want. Remote work, medical insurance, flexible time off, retirement savings plans, and modern family planning are just some of our offering. Check out our other benefits at Block.

About Block, Inc.

Block, Inc. (NYSE: XYZ) builds technology to increase access to the global economy. Each of our brands unlocks different aspects of the economy for more people. Square makes commerce and financial services accessible to sellers. Cash App is the easy way to spend, send, and store money. Afterpay is transforming the way customers manage their spending over time. TIDAL is a music platform that empowers artists to thrive as entrepreneurs. Bitkey is a simple self-custody wallet built for bitcoin. Proto is a suite of bitcoin mining products and services. Together, we’re helping build a financial system that is open to everyone.

Key skills/competency

  • Site Reliability Engineering
  • Platform Infrastructure Engineering
  • AI-driven tooling and automation
  • Observability
  • Incident detection and response
  • Operational toil reduction
  • Incident command and mitigation
  • Scalable product development
  • CI/CD pipelines
  • Kubernetes

Skills & topics

  • Site Reliability Engineer
  • SRE
  • Platform Engineering
  • Infrastructure Engineering
  • Cloud Engineering
  • DevOps
  • Kubernetes
  • AWS
  • Observability
  • Incident Management
  • AI
  • Automation
  • High Availability
  • System Reliability
  • Backend Development

How to get hired

  • Tailor your resume: Highlight your 5+ years of software development and SRE experience, focusing on AI tools, incident management, and CI/CD.
  • Showcase initiative: Emphasize technical leadership and problem-solving skills, especially in backend/platform projects.
  • Demonstrate AI familiarity: Detail your experience using AI for observability, incident analysis, or automation in production environments.
  • Prepare for technical interviews: Be ready to discuss distributed systems, Kubernetes, AWS, and incident response scenarios.
  • Understand Block's mission: Align your application and interview responses with Block's goal of economic empowerment and inclusivity.

Technical preparation

Master Kubernetes and cloud infrastructure.,Deep dive into observability and monitoring tools.,Practice AI/ML for operational efficiency.,Refine incident command and mitigation skills.

Behavioral questions

Describe a complex system you troubleshot.,How do you handle pressure during incidents?,Share an initiative you led for reliability.,How do you approach reducing operational toil?

Frequently asked questions

What are the core responsibilities of a Senior Site Reliability Engineer at Block?
As a Senior Site Reliability Engineer at Block, you'll be responsible for proactively and reactively improving the reliability of Block's platform and critical infrastructure. This includes leveraging AI-driven tooling for observability and incident response, participating in oncall rotations for Tier 0 services, leading incident command, and driving platform-wide reliability improvements.
What technologies are commonly used by the SRE team at Block?
The SRE team at Block utilizes a range of technologies including Kotlin, Modern Java, HTTP, gRPC, Protocol Buffers, MySQL/Vitess/DynamoDB, event-driven architectures, DataDog, LaunchDarkly, Terraform, Kubernetes, Istio/Envoy, and Amazon Web Services.
What is Block's approach to using AI in Site Reliability Engineering?
Block leverages AI-driven tooling and automation to enhance observability, accelerate incident detection and response, and reduce operational toil. This includes applying AI to incident analysis, alert tuning, and optimizing operational workflows for problem-solving.
What is the oncall expectation for this Senior Site Reliability Engineer role at Block?
The oncall expectation involves participating in primary platform oncall, which typically means 12 hours per day for one week every few weeks, depending on team size. This supports Block's most critical Tier 0 services and involves leading incident command during high-severity events.
How does Block approach diversity and inclusion in its hiring process for SRE roles?
Block is committed to building a workplace that reflects its values of economic empowerment and inclusivity. As an equal opportunity employer, they ensure fair evaluation of all candidates and offer inclusive interview experiences, including reasonable accommodations for disabled applicants.
What kind of growth opportunities are available for a Senior Site Reliability Engineer at Block?
Block emphasizes empowering employees to do their best work and build the lives they want. For SREs, this means opportunities to work with cutting-edge technologies, drive platform-wide reliability improvements, and grow their careers within a globally distributed company that values continuous learning and development.