Site Reliability Engineer
Jobgether · Germany
- Location
- Germany
- Posted
- 9 days ago
- Type
- Full-time / Remote
- Salary
- Not listed — worth asking early.
What the job really is
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineer based in Germany.
This is a senior reliability engineering role focused on defining and advancing the infrastructure standards behind a globally scaled, AI-native platform.
You will take ownership of reliability strategy across production infrastructure, with a particular focus on AWS, Kubernetes, event-driven systems, and AI agent workloads.
The role combines deep hands-on engineering with architectural leadership, incident management, observability, and technical mentorship.
You will design systems that remain resilient under increasing transaction volumes while establishing measurable standards for reliability across engineering teams.
A key part of the role will be evolving synchronous architectures toward durable asynchronous communication and strengthening the platform through resilience testing and chaos engineering.
You will also help shape how AI-assisted tools are used for automation, incident analysis, runbooks, and root-cause investigations.
Success means becoming the trusted technical authority for complex reliability decisions while creating practices that make reliability scalable across the organization.
Similar roles
Open positions we recommend based on this role.