
Site Reliability Engineer
OneStream Software · United States
- Hybrid
- Full-time
- $148,000 / year
- United States
Job highlights
- Ensure platform reliability, performance, and availability.
- Automate infrastructure deployments and processes.
- Implement observability and APM solutions.
- Collaborate with product and engineering teams.
- Mentor others and contribute to team success.
About the role
Site Reliability Engineer
OneStream is seeking a passionate and experienced Site Reliability Engineer to ensure the reliability, performance, and high availability of our platform and services. This remote role, based in the United States, is crucial for our Cloud Services team. You will leverage your expertise in designing, implementing, and monitoring scalable and secure cloud services, working closely with internal teams and customers. If you thrive on staying at the forefront of technology, automating infrastructure, and fostering a culture of learning, this is an excellent opportunity for you.
Primary Duties And Responsibilities
- Implement application/infrastructure observability solutions to ensure desired application availability, reliability, and performance.
- Participate in regular On-Call rotations and share details related to incidents and their resolution through post-mortem reports and regular review meetings.
- Proactively partner with Product and Engineering teams to identify, develop, deploy, and maintain reliable systems and services.
- Influence and create new designs, architectures, standards, and methods for large-scale systems.
- Sustain a high level of reliability for key services and automated systems.
- Automate processes to improve reliability, performance, and availability.
- Update technical documentation, workflows, and knowledge base articles.
- Provide feedback in pull requests and peer coding reviews.
- Implement codified automated solutions that build integrations between Dynatrace, Azure DevOps and Jira.
- Solid knowledge in focused areas of OneStream Software.
- Ability to mentor others in several technical areas.
- Understanding practical use of SOC/FedRAMP controls to assist Compliance and Security teams.
Required Education And Experience
- BS/BA in computer science, engineering, or technology-related field (or equivalent work experience).
- Proven work experience as a Site Reliability Engineer or in a similar role.
- 6+ years of cloud infrastructure and software development experience.
- 2+ years hands-on experience with Azure Kubernetes Services (AKS) with container-based deployment skills or other platforms such as OpenShift, GKS, EKS.
- Advanced understanding of APM and observability tools such as Dynatrace, AppInsights, DataDog, Log Analytics, New Relic, Prometheus and Grafana.
- Advanced understanding of Infrastructure-as-Code (IaC) concepts and tooling (Terraform, CloudFormation templates, Bicep or ARM templates) on Microsoft Azure, Amazon Web Services (AWS), or Google Cloud Platform (GCP).
- Deep knowledge of Configuration Management/Orchestration utilities such as Ansible, PowerShell DSC, Chef, and Puppet.
- Advanced understanding of cloud concepts including elasticity, security, and identity management.
- Well-versed familiarity with Agile Development methodologies utilizing Jira or Azure DevOps Boards.
- 6+ years of hands-on experience with: Automating processes using PowerShell, Bash, CLI, REST APIs, python, ARM Templates or other scripting languages.
- Comfortable leveraging source control tools such as Git, Azure DevOps, or GitHub.
- Knowledge of container orchestration platforms such as Kubernetes, OpenShift, AKS, GKS or helm.
- Microsoft Azure, Amazon Web Services (AWS) or Google Cloud (GCP).
Preferred Education And Experience
- Experience working for a cloud service provider (CSP), managed service provider (MSP), or SaaS provider.
- 6+ years of relevant Azure experience deploying and managing leveraging Infrastructure-as-Code (IAC) concepts.
- Experience with Microsoft and .NET (.NET, C#, SQL).
- Experience writing efficient and reliable code in a development environment.
- Debian, Ubuntu, Alpine or other distributions of the Linux operating systems.
- Deep knowledge and understanding of containerized applications, with special attention to reliability and monitoring of those containerized applications.
Knowledge, Skills, And Abilities
- Deal well with ambiguous/undefined problems.
- Ability to self-motivate and work independently.
- Strong organizational and prioritization skills.
- Ability to find and apply effective solutions to emerging problems and challenges.
- Strong attention to detail.
- Comfortable communicating with all levels of management and engineering.
- Ability to get up to speed quickly with modern technologies and services.
- Ability to multitask on a variety of projects.
Who We Are
OneStream is how today’s Finance teams can go beyond just reporting on the past and Take Finance Further™ by steering the business to the future. It’s the only enterprise finance platform that unifies financial and operational data, embeds AI for better decisions and productivity, and empowers the CFO to become a critical driver of business strategy and execution. Our vision is to be the operating system for modern finance, digitizing core financial functions and empowering the CFO to become a critical driver of business strategy. To learn more visit www.onestream.com.
Why Join The OneStream Team
- Transparency around corporate structure, salary, and benefits.
- Core value of customer success.
- Variety of project work (not industry-specific).
- Strong culture and camaraderie.
- Multiple training opportunities.
Benefits At OneStream
OneStream employees are passionate, hardworking individuals who go above and beyond to keep our customers happy and follow through on our mission statement. They consistently deliver the best and in turn, we make every effort to keep them cared for and happy. A sample of the benefits we provide are:
- Excellent Medical Plan.
- Dental & Vision Insurance.
- Life Insurance.
- Short & Long Term Disability.
- Vacation Time.
- Paid Holidays.
- Professional Development.
- Retirement Plan.
Key skills/competency
- Site Reliability Engineering
- Cloud Infrastructure
- Automation
- Observability
- Kubernetes (AKS)
- Infrastructure as Code (IaC)
- APM Tools (Dynatrace)
- Scripting (Python, PowerShell, Bash)
- Containerization
- Azure
Skills & topics
- Site Reliability Engineer
- SRE
- Cloud Engineer
- DevOps
- Automation
- Observability
- Kubernetes
- Azure
- Infrastructure as Code
- Reliability Engineering
- Software Development
- System Administration
- Performance Monitoring
- Incident Management
- Scripting
- Python
- PowerShell
- Bash
- Dynatrace
- Terraform
- AWS
- GCP
How to get hired
- Tailor your resume: Highlight your 6+ years of cloud infrastructure and software development experience, specifically mentioning Azure Kubernetes Services, observability tools, and Infrastructure-as-Code.
- Showcase automation skills: Emphasize your experience with scripting languages (PowerShell, Python, Bash) and configuration management tools (Ansible, Chef) in your application.
- Address experience requirements: Clearly state your BS/BA in a relevant field or equivalent work experience, and your familiarity with Agile methodologies.
- Prepare for technical interviews: Be ready to discuss your advanced understanding of cloud concepts, container orchestration, and APM tools like Dynatrace.
- Demonstrate soft skills: Highlight your ability to work independently, handle ambiguity, and communicate effectively with diverse teams.
Technical preparation
Behavioral questions
Frequently asked questions
- What are the key technologies for the Site Reliability Engineer role at OneStream Software?
- The Site Reliability Engineer role at OneStream Software heavily utilizes Azure Kubernetes Services (AKS), observability tools like Dynatrace, Infrastructure-as-Code (IaC) with tools like Terraform, and scripting languages such as Python, PowerShell, and Bash. Familiarity with container orchestration and cloud platforms like Azure, AWS, or GCP is also essential.
- What is the expected salary range for the Site Reliability Engineer position at OneStream Software?
- The Gross Annual Base Salary for the Site Reliability Engineer position at OneStream Software ranges from USD 114,000 to 148,000. Additional variable compensation and benefits may apply, with total compensation dependent on experience, skills, and location.
- Does OneStream Software offer remote work for the Site Reliability Engineer role?
- Yes, the Site Reliability Engineer position at OneStream Software is a remote role based in the United States. This allows for flexibility in where you work while contributing to the company's cloud services team.
- What level of experience is required for the Site Reliability Engineer role?
- OneStream Software requires a proven work experience as a Site Reliability Engineer or in a similar role, with a minimum of 6 years of cloud infrastructure and software development experience. Additionally, 2+ years of hands-on experience with Azure Kubernetes Services (AKS) is expected.
- What are the primary responsibilities of a Site Reliability Engineer at OneStream Software?
- The primary responsibilities include implementing observability solutions, participating in on-call rotations, partnering with engineering teams to build reliable systems, automating processes for improved performance and availability, and maintaining technical documentation. You will also provide feedback on code reviews and mentor other team members.
- What benefits does OneStream Software offer its employees?
- OneStream Software offers a comprehensive benefits package including Medical, Dental, Vision, and Life Insurance, Short & Long Term Disability, Vacation Time, Paid Holidays, Professional Development opportunities, and a Retirement Plan (401K).
- How does OneStream Software approach transparency regarding compensation and structure?
- OneStream Software emphasizes transparency around its corporate structure, salary, and benefits as a core value. This approach is part of their commitment to their employees and fostering a positive work environment.