
Senior Site Reliability Engineer
Visa · Austin, TX
- On site
- Full-time
- $171,800 / year
- Austin, TX
Job highlights
- Own core platform components lifecycle and design.
- Apply SRE principles for resilience and fault tolerance.
- Lead infrastructure orchestration and GitOps practices.
- Promote SRE operational excellence and collaboration.
- Work with cloud, Kubernetes, and service mesh technologies.
About the role
About Us
Visa is a world leader in payments technology, facilitating transactions between consumers, merchants, financial institutions and government entities across more than 200 countries and territories, dedicated to uplifting everyone, everywhere by being the best way to pay and be paid.
At Visa, you’ll have the opportunity to create impact at scale — tackling meaningful challenges, growing your skills and seeing your contributions impact lives around the world. Join Visa and do work that matters – to you, to your community, and to the world. Progress starts with you.
Job Description
What You’ll Do:
- Own the end‑to‑end lifecycle (design, provisioning, upgrades, maintenance, and decommissioning) of core platform components, including: Cloud infrastructure primitives, Kubernetes clusters and cluster services, Networking, ingress, and service discovery, Service Mesh and supporting data‑plane components.
- Design platform components to be resilient by default, applying SRE principles such as: Fault isolation and graceful degradation, Capacity planning and saturation control, Reduced operational toil and clear failure modes.
- Lead the design and implementation of infrastructure bootstrap orchestration, including: Automated cluster and environment provisioning, Deterministic, repeatable platform bring‑up and teardown, Dependency‑aware orchestration across cloud, network, and Kubernetes layers.
- Drive Infrastructure‑as‑Code and GitOps‑first practices to ensure: Platform components are reproducible and auditable, Changes are automated, testable, and reversible, Manual intervention is minimized or eliminated.
- Identify automation gaps and lead initiatives that reduce human effort, onboarding time, and operational risk.
- Apply And Promote SRE Operational Excellence Practices, Including: Clear ownership and runbooks for platform components, Participation in on‑call rotation as a platform reliability escalation point, Incident response, post‑incident reviews, and problem management.
- Improve day‑2 operations by standardizing upgrade/rollback strategies and reducing MTTD/MTTR.
- Ensure platform operations align with security, compliance, and internal control requirements.
- Collaborate with engineering teams across the organization to influence platform adoption, reliability standards, and cloud‑native best practices.
This is a hybrid position. Expectation of days in office will be confirmed by your hiring manager.
Qualifications
Basic Qualifications:
- 2+ years of relevant work experience and a Bachelors degree, OR 5+ years of relevant work experience
Preferred Qualifications:
- 3 or more years of work experience with a Bachelor’s Degree or more than 2 years of work experience with an Advanced Degree (e.g. Masters, MBA, JD, MD)
- Experience in creating and updating documentation for infrastructure and operational procedures.
- Experience in providing first-level support for infrastructure and deployment issues.
- Experience in automating repetitive tasks and suggesting workflow improvements.
- Experience in learning and applying DevOps and SRE best practices.
- Experience in supporting implementation and management of containerization technologies.
- Language Skills: Proficiency in English at B2 level or above (Upper-Intermediate)
Technical Skills:
- Strong hands‑on experience with public cloud platforms (Azure mandatory, AWS preferred).
- Proven experience operating and administering Kubernetes at scale in production environments.
- Strong experience with container orchestration platforms and cloud architecture fundamentals (networking, IAM/security concepts, and reliability patterns).
- Experience with Infrastructure as Code (Terraform preferred) and automation‑first workflows.
- Familiarity with GitOps practices and CI/CD pipelines.
- Strong troubleshooting skills for distributed systems, including root‑cause analysis and reliability improvements.
- Experience with observability concepts and practices (monitoring, logging, alerting, tracing).
- Experience with Service Mesh technologies (Istio preferred, App Mesh or Linkerd).
- Experience working with critical or mission‑critical systems.
- Strong background applying SRE principles (operational readiness, incident management, runbooks, toil reduction).
U.S. Applicants Only
The estimated salary range for this position is $110,700.00 to $ 171,800.00 USD per year, which may include potential sales incentive payments (if applicable). Salary may vary depending on job-related factors which may include knowledge, skills, experience, and location. In addition, this position may be eligible for bonus and equity. Visa has a comprehensive benefits package for which this position may be eligible that includes Medical, Dental, Vision, 401(k), FSA/HSA, Life Insurance, Paid Time Off, and Wellness Program.
Work Hours: Varies upon the needs of the department.
Travel Requirements: This position requires travel 5-10% of the time.
Mental/Physical Requirements: This position will be performed in an office setting. The position will require the incumbent to sit and stand at a desk, communicate in person and by telephone, frequently operate standard office equipment, such as telephones and computers.
Visa is an EEO Employer
Qualified applicants will receive consideration for employment without regard to race, color religion, sex, national origin, sexual orientation, gender identity, disability or protect veteran status. Visa will also consider for employment qualified applicants with criminal histories in a manner consistent with the EEOC guidelines and applicable local law.
Key skills/competency
- Site Reliability Engineering
- Kubernetes
- Cloud Infrastructure
- Infrastructure as Code
- DevOps
- Automation
- Distributed Systems
- Service Mesh
- Observability
- Incident Management
Skills & topics
- Site Reliability Engineer
- SRE
- Kubernetes
- Cloud
- Azure
- AWS
- Terraform
- DevOps
- Automation
- Infrastructure as Code
- Service Mesh
- Istio
- Observability
- Incident Management
- Distributed Systems
- System Administration
- Platform Engineering
- Production Support
- Reliability Engineering
- Senior Engineer
How to get hired
- Tailor your resume: Highlight experience with Kubernetes, Azure, Terraform, and SRE principles.
- Showcase your skills: Quantify achievements in automation, incident response, and system reliability.
- Prepare for technical interviews: Review distributed systems, cloud architecture, and containerization concepts.
- Understand Visa's values: Align your responses with Visa's commitment to innovation and customer service.
- Ask insightful questions: Demonstrate your engagement with the team and technical challenges.
Technical preparation
Behavioral questions
Frequently asked questions
- What are the key technical skills required for the Senior Site Reliability Engineer role at Visa?
- The Senior Site Reliability Engineer role at Visa requires strong hands-on experience with public cloud platforms (Azure mandatory, AWS preferred), Kubernetes administration at scale, container orchestration, Infrastructure as Code (Terraform preferred), GitOps practices, and observability tools. Experience with Service Mesh technologies like Istio is also highly valued.
- What is the expected work arrangement for this Senior Site Reliability Engineer position at Visa?
- This position is a hybrid role at Visa. The specific expectation for days in the office will be confirmed by your hiring manager. Travel requirements are estimated at 5-10% of the time.
- Can you provide details on the salary range for the Senior Site Reliability Engineer at Visa?
- The estimated annual salary range for this Senior Site Reliability Engineer position at Visa is $110,700.00 to $171,800.00 USD. This range may vary based on factors like knowledge, skills, experience, and location. Potential sales incentive payments, bonus, and equity may also apply.
- What is Visa's approach to Site Reliability Engineering for their core platform components?
- Visa applies SRE principles to ensure platform components are resilient by default, focusing on fault isolation, graceful degradation, capacity planning, and reducing operational toil. They drive Infrastructure-as-Code and GitOps-first practices for reproducibility and automation, aiming to minimize manual intervention and operational risk.
- What are the essential qualifications for a Senior Site Reliability Engineer at Visa?
- Basic qualifications include 2+ years of relevant experience with a Bachelor's degree, or 5+ years of experience. Preferred qualifications include 3+ years with a Bachelor's or 2+ years with an advanced degree, along with experience in documentation, first-level support, automation, DevOps/SRE best practices, and containerization.
- How does Visa handle incident response and operational excellence for their platform?
- Visa promotes SRE operational excellence through clear ownership and runbooks, participation in on-call rotations, and robust incident response processes including post-incident reviews and problem management. They focus on improving day-2 operations by standardizing upgrade/rollback strategies to reduce Mean Time To Detect (MTTD) and Mean Time To Resolve (MTTR).