PitchMeAI
Oracle

Senior Site Reliability Engineer

Oracle · United States

  • Hybrid
  • Full-time
  • $150,000 / year
  • United States
Tailored resumekeyword-matched to this role.
Hiring managerwe find who's hiring.
Intro emaildrafted to reach them directly.

Job highlights

  • Manage major incidents in Oracle Cloud.
  • Ensure high availability of OCI services.
  • Automate tasks for continuous availability.
  • Collaborate with leaders across Oracle.
  • Improve OCI service availability iteratively.

About the role

About the Role

OCI Incident Response is the first line of defense in maintaining the high availability of Oracle’s cloud. We minimize customer-impacting events by making them shorter, less frequent, and less impactful through large-scale incident management. We are at the forefront of reducing event duration by leveraging our operational experience, knowledge of best practices, and ability to develop tools that automate incident management.

We are looking for a Senior Site Reliability Engineer to join our OCI team. This role is part of a globally distributed team responsible for detecting, triaging, and mitigating OCI service-impacting events as quickly as possible. You will be part of one of these regional teams and will be responsible for minimizing the downtime of OCI services. You will achieve this by delivering excellent major incident management and operating systems with high scalability, performance, and security that help prevent incidents from occurring.

Oracle’s Cloud is state-of-the-art and constantly evolving. When issues arise, your team will respond within minutes to ensure customer impact is minimized. This role will expose you to the inner workings of OCI’s systems and organization. You will interact with and influence leaders across Oracle and drive broad, cross-organization programs aimed at iteratively improving OCI-wide service availability. We are an agile team with significant impact. If you want to be part of a fast-moving team breaking new ground, we would love to speak with you!

We are looking for candidates who are flexible to work AMER shift hours (9:30 AM to 5:30 PM PST) on a rotating roster, including occasional weekends and public holidays.

Responsibilities

  • Solve complex problems related to infrastructure cloud services and automate common tasks to ensure continuous availability with minimal human intervention.
  • Command and coordinate SMEs and service leaders to restore services as quickly as possible during major incidents, while keeping accurate and timely data on the progress of such incidents.
  • Utilize a deep understanding of cloud computing design patterns and their dependencies to mitigate complex major incidents.
  • Embed a methodical approach to troubleshoot large, complex, interconnected systems used in incident detection and orchestration.
  • Document pertinent information related to incidents that aids process improvement, identifies deviations, and enables the creation of an incident knowledge base.
  • Monitor and evaluate high-level service and infrastructure dashboards, taking action to address identified anomalies.
  • Identify opportunities and take ownership of automation and/or continuous improvement of incident management process steps and best practices.
  • Define and document the technical architecture of large-scale distributed systems.
  • Understand the end-to-end configuration, technical dependencies, and overall behavioral characteristics of production services.
  • Be responsible for the design and delivery of the mission-critical stack, with a focus on security, resiliency, scalability, and performance.
  • Partner with development teams to define operational requirements for product roadmaps.
  • Articulate the technical characteristics of services and technology areas, and guide development teams to engineer and add premier capabilities to the Oracle Cloud service portfolio.
  • Act as the ultimate escalation point for complex or critical issues that have not yet been documented as Standard Operating Procedures (SOPs).

Minimum Qualifications

  • Bachelor’s degree or higher in Computer Science or relevant work experience.
  • 3+ years’ experience in Site Reliability Engineering, DevOps, or System Engineering.
  • Must have public cloud operations experience (e.g., AWS, Azure, GCP, OCI).
  • Extensive experience with Major Incident Management in a cloud-based environment.
  • Demonstrate clear understanding of automation and orchestration principles.
  • Experience having worked in at least one modern object-oriented programming language.
  • Experience with professional software engineering standard methodologies such as Agile project management, coding standards, code reviews, source control management, build processes, testing, and operations.
  • Familiarity with infrastructure automation tools such as Chef, Ansible, Jenkins, Terraform
  • Excellent expertise with several of following technologies: Infrastructure-as-a-Service, CI/CD systems, Docker, RESTful APIs, log analysis tools, debugging tools

Key skills/competency

  • Site Reliability Engineering
  • DevOps
  • Cloud Operations
  • Incident Management
  • Automation
  • Orchestration
  • Object-Oriented Programming
  • Agile Methodologies
  • Infrastructure as a Service (IaaS)
  • CI/CD

Skills & topics

  • Site Reliability Engineer
  • SRE
  • DevOps
  • Cloud Engineer
  • Incident Management
  • Automation
  • Oracle Cloud
  • OCI
  • Systems Engineering
  • Infrastructure

How to get hired

  • Tailor your resume: Highlight your SRE, DevOps, and cloud experience, emphasizing major incident management and automation skills.
  • Showcase cloud expertise: Detail your experience with public cloud platforms (AWS, Azure, GCP, OCI) and infrastructure automation tools.
  • Demonstrate problem-solving: Provide examples of how you've solved complex infrastructure problems and automated common tasks.
  • Prepare for technical interviews: Be ready to discuss your understanding of cloud computing, distributed systems, and incident response.
  • Understand Oracle's values: Research Oracle's commitment to innovation, cloud solutions, and employee empowerment.

Technical preparation

Master incident response frameworks.,Practice automation with scripting languages.,Deepen cloud infrastructure knowledge.,Study distributed systems architecture.

Behavioral questions

Describe a major incident you managed.,How do you handle high-pressure situations?,Give an example of successful automation.,How do you collaborate with development teams?

Frequently asked questions

What is the typical career progression for a Senior Site Reliability Engineer at Oracle?
As a Senior Site Reliability Engineer (IC3) at Oracle, career progression often involves moving into more senior technical roles, specializing in specific cloud technologies, or transitioning into management positions. You might become a Principal SRE, a technical lead for a specific cloud service, or a manager overseeing an incident response team, leveraging your deep understanding of Oracle Cloud Infrastructure (OCI) and incident management.
What kind of technical challenges can I expect as a Senior Site Reliability Engineer at Oracle?
You can expect to tackle complex challenges related to the scalability, performance, and security of Oracle's Cloud Infrastructure. This includes mitigating major incidents in large-scale distributed systems, automating operational tasks, defining technical architectures, and partnering with development teams to enhance OCI services. Your role will involve deep dives into system dependencies and driving improvements across the OCI platform.
Does Oracle offer opportunities for professional development for Site Reliability Engineers?
Yes, Oracle emphasizes continuous learning and professional development. As a Senior Site Reliability Engineer, you'll have opportunities to work with cutting-edge cloud technologies, influence product roadmaps, and gain exposure to various facets of OCI. Oracle also offers resources for skill enhancement and encourages employees to grow within the company.
How does Oracle handle on-call rotations and work-life balance for SRE roles?
This Senior Site Reliability Engineer role requires flexibility for AMER shift hours (9:30 AM to 5:30 PM PST) on a rotating roster, including occasional weekends and public holidays. While incident response is critical, Oracle aims to balance this with benefits and time-off policies, including flexible vacation and paid sick leave, to support employee well-being.
What specific technologies are most important for a Senior Site Reliability Engineer at Oracle?
Key technologies include public cloud operations (OCI, AWS, Azure, GCP), major incident management, automation and orchestration principles, object-oriented programming, and professional software engineering methodologies. Experience with infrastructure automation tools like Chef, Ansible, Jenkins, Terraform, and familiarity with IaaS, CI/CD systems, Docker, RESTful APIs, and log analysis tools are highly valued.