Login Enter

gridscale GmbH

1 day ago

Site Reliability Engineer - Openstack (m/f/d)

Sign up free Log in

Sign up to save this job, get alerts, and apply with an optimized CV.

Company information

Company
gridscale GmbH
Location
Köln, Nordrhein-Westfalen, Deutschland Germany
Posted
1 day ago
View all jobs at gridscale GmbH

Job description

Your Role

You will support us in further developing, operating, and industrializing OVHcloud's on-premise cloud platform. As part of a small, experienced team, you will work on our OpenStack-based infrastructure as well as the Kubernetes and GitOps stack on which our customer-oriented cloud platform runs. AI-Assisted Engineering is an integral part of our daily engineering practice – from Spec-Driven Development and agentic coding workflows to incident response and automation. The platform is actively under construction, giving you direct influence on architecture, automation strategy, and the use of AI in Platform Engineering. As a Senior, you will take ownership of your areas and shape your professional focus according to your strengths – with a focus on Automation, Compute Lifecycle, Platform Engineering, and AI Substrate.

Our Tech Stack 🚀

· OpenStack · Kubernetes · KVM · Linux · Bare Metal · Ansible · Terraform · Go
· FluxCD / ArgoCD · Git· Python · Claude Code · Cursor · Agentic Coding Tooling

Your Tasks

  • You will conceptualize and develop our OpenStack-based on-premise cloud infrastructure with the goal of being able to roll out and commission complete cloud environments on bare metal in a highly automated manner.

  • You will develop and operate Infrastructure as Code with Ansible and Terraform, as well as our Kubernetes and GitOps workflows with FluxCD / ArgoCD – supported by LLMs, agentic workflows, and automated test and review processes.

  • You will be responsible for the lifecycle of our compute infrastructure – from bare metal, firmware, and provisioning to hypervisors and virtual compute nodes. This includes, among other things, patching, migrations, host evacuations, capacity rebalancing, and the automation of stable platform operations.

  • You will further develop our AI Substrate and our Self-Healing Strategy – from structured knowledge bases and agentic workflows for incident triage and capacity planning to the gradual automation of current runbooks.

  • You will design tests for non-regression, performance, and security, document and package solutions, and continuously improve the platform based on telemetry, operational experience, and user feedback.

  • As a Senior Engineer, you will be the technical point of contact and sparring partner for colleagues regarding automation, platform engineering, and AI tooling.

What We Offer You

  • A platform in true Build Mode with plenty of scope for development, visible influence on architectural decisions, and an experienced senior team where autonomy and ownership count.

  • AI-Augmented Engineering as an integral part of our way of working – with Claude Code and comparable agentic tools, Markdown-based knowledge bases, and room to actively develop our engineering practice.

  • Exceptional team spirit across all departments and national borders – we live #OneTeam.

  • Exciting tasks in a highly innovative and international environment with state-of-the-art technologies.

  • 32 vacation days, which increase with increasing length of service.

  • Flexible working hours, home office options, and a secure, permanent employment contract with market- and performance-based compensation.

  • Employer-funded pension plan and an attractive insurance package.

  • OVHcloud covers 50% of the costs for public transport.

  • Up to €400 annual subsidy for sports activities, such as gym memberships or sports courses.

  • Attractive discounts at numerous shops and companies through Corporate Benefits.

  • We contribute to the leasing of your Cargo Bike.

  • Regular company events and free cold and hot drinks.


  • You bring several years of practical experience as an SRE, Platform Engineer, or DevOps Engineer in operating productive infrastructure and have in-depth experience with OpenStack, Kubernetes, and Linux, including on bare metal.

  • You have operated compute infrastructure end-to-end – from firmware/BIOS rollouts, bare-metal provisioning, and hardware diagnostics to hypervisors, migrations, host evacuations, graceful drains, and capacity rebalancing.

  • AI-Assisted Engineering is already part of your daily work. You use LLMs and agentic tools specifically where they meaningfully support development, testing, reviews, or operations, and can accurately assess where AI provides real added value and where sound engineering expertise remains crucial.

  • You work confidently with Ansible, Terraform, and GitOps workflows like FluxCD or ArgoCD and have experience operating automated processes in a stable and reproducible manner in production environments.

  • You have experience with Go and/or Python, as well as with Claude Code, Cursor, Aider, or comparable agentic coding environments.

  • Observability, networking, compute tuning, auto-remediation, security-critical infrastructures, and multi-site cloud environments also belong to your technical experience spectrum.

  • You work independently, possess a strong ownership mindset, and enjoy not only operating things but also continuously improving them. At the same time, you are happy to share your knowledge and can clearly communicate complex technical contexts.

  • We work in an international environment in English. Therefore, you feel comfortable discussing, documenting, and driving technical topics in English together with the team.

Required skills

Interested in this position?

Create your free account and tailor your CV to match this job.