Sign up to save this job, get alerts, and apply with an optimized CV.
Senior Site Reliability Engineer - AI Platform
Job description
About the opportunity We are seeking a Senior Site Reliability Engineer to join the Platform Engineering Domain in the AI Platform Team. The mission of Platform Engineering is to provide trusted, performant, self-service platforms that empower product teams to build 'the bank the world loves to use.' The AI Platform team contributes to this mission by creating scalable, secure, and compliant infrastructure solutions that support MLOps and GenAI capabilities. The ideal candidate is not only a seasoned SRE but also has experience with machine learning infrastructure, AI/ML platforms, and distributed systems. What you will do: * Design, build, and maintain the infrastructure that supports our AI/ML platforms, including model training, serving, and monitoring. * Automate infrastructure provisioning, deployment, and management using tools like Terraform, Kubernetes, and CI/CD pipelines. * Implement and maintain monitoring, alerting, and logging systems to ensure the reliability and performance of our AI/ML services. * Collaborate with ML engineers, data scientists, and other stakeholders to understand their needs and provide solutions that meet their requirements. * Participate in on-call rotations and incident response to ensure the availability of our AI/ML platforms. * Identify and resolve performance bottlenecks, scalability issues, and other operational challenges. * Write clear, concise documentation and contribute to knowledge sharing within the team. * Stay up-to-date with the latest trends and technologies in SRE, AI/ML, and cloud computing. What you bring: * Bachelor's or Master's degree in Computer Science, a related field, or equivalent experience. * 5+ years of experience as a Site Reliability Engineer or DevOps Engineer. * Experience with infrastructure-as-code tools such as Terraform or CloudFormation. * Strong experience with containerization and orchestration tools like Docker and Kubernetes. * Experience with CI/CD pipelines and tools (e.g., Jenkins, GitLab CI, CircleCI). * Experience with monitoring and logging tools (e.g., Prometheus, Grafana, ELK stack). * Experience with cloud platforms (e.g., AWS, GCP, Azure). * Experience with Linux operating systems. * Experience with scripting languages such as Python or Bash. * Strong understanding of networking concepts (e.g., TCP/IP, DNS, load balancing). * Experience with machine learning infrastructure and AI/ML platforms is a plus. * Excellent communication, collaboration, and problem-solving skills. * Ability to work independently and as part of a team. What we offer: * A collaborative and innovative work environment. * Opportunities for professional development and growth. * Competitive salary and benefits package. * The chance to work on cutting-edge AI/ML technologies. * The opportunity to make a real impact on the future of banking.
Required skills
Sign up to apply
Create a free account to apply for this job and get access to:
- AI-powered CV optimization for this specific job
- Save jobs and create custom alerts
- See your CV match score for each job
Company information
- Company
- N26 GmbH
- Location
-
Deutschland, Berlin
Germany - Posted
- 10 months ago
Interested in this position?
Create your free account and tailor your CV to match this job.