PitchMeAI
Jobgether

Site Reliability Engineer

Jobgether · Spain

Location
Spain
Posted
9 days ago
Type
Full-time / Remote
Salary
Not listed — worth asking early.
Tailored resumekeyword-matched to this role.
Hiring managerwe find who's hiring.
Intro emaildrafted to reach them directly.

What the job really is

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineer based in Spain.

This is a senior reliability engineering role focused on defining and advancing the infrastructure standards behind a globally scaled, AI-native platform.
You will take ownership of reliability strategy across production infrastructure, with a particular focus on AWS, Kubernetes, event-driven systems, and AI agent workloads.
The role combines deep hands-on engineering with architectural leadership, incident management, observability, and technical mentorship.
You will design systems that remain resilient under increasing transaction volumes while establishing measurable standards for reliability across engineering teams.
A key part of the role will be evolving synchronous architectures toward durable asynchronous communication and strengthening the platform through resilience testing and chaos engineering.
You will also help shape how AI-assisted tools are used for automation, incident analysis, runbooks, and root-cause investigations.
Success means becoming the trusted technical authority for complex reliability decisions while creating practices that make reliability scalable across the organization.

Similar roles

Open positions we recommend based on this role.