Deeptech Recruitment
2 hours ago
Director AI Platform Engineering
Sign up to save this job, get alerts, and apply with an optimized CV.
Company information
- Company
- Deeptech Recruitment
- Location
- Germany Germany
- Posted
- 2 hours ago
Job description
Our client is building a new AI Engineering & Platform team in Germany to support the next generation of AI-powered media, localization, automation, and workflow capabilities. The Director, AI Engineering & Platform will lead this team and will be responsible for building, operating, and evolving the technical foundation for production AI services across our client.
This is a senior hands-on leadership role for someone who can build and run a high-performing engineering team, deliver complex AI platform and integration projects, and provide technical direction across cloud infrastructure, model deployment, AI services, agentic workflows, and production operations. The role will lead a hybrid engineering team based primarily in Germany while also managing and collaborating with engineers in other our client locations, including the United States.
In addition to leading delivery and operations, this person will continuously evaluate emerging AI technologies, model providers, open-source frameworks, agentic workflow patterns, and industry capabilities. They will be expected to make clear recommendations on where our client should build, buy, partner, or integrate in order to improve speed, quality, cost, automation, and product differentiation.
RESPONSIBILITIES:
Build, lead, mentor, and develop a distributed AI engineering team across Germany and other our client locations. Experience recruiting, hiring, and standing up the full AI/ML engineering organization from scratch: applied ML engineers, MLOps/platform engineers, data engineers, integration engineers, etc.
Lead development, hosting and ongoing improvement of our client’s cloud-native training and inference platform (compute, storage, orchestration, model serving, monitoring, cost management) for our client’s NLP technology for MT, ASR, and TTS in production.
Own technical delivery for our client’s AI platform, model-serving, workflow automation, and applied AI engineering initiatives.
Lead architecture, execution, deployment and SLA performance for scalable AI services, including inference platforms, training environments, deployment pipelines, observability, and operational support.
Own infrastructure cost efficiency (GPU utilization, training vs. inference cost trade-offs, autoscaling) as a first-class, ongoing responsibility rather than a one-time setup task.
Evaluate internal and external AI technologies on an ongoing basis, including commercial models, open-source models, agentic frameworks, orchestration tools, and vendor platforms.
Recommend when our client should build, buy, partner, or integrate with third-party AI capabilities.
Guide teams responsible for cloud infrastructure, DevOps, AI platform operations, applied AI engineering, and model integration.
Establish engineering standards, operating rhythms, delivery practices, documentation expectations, and technical review processes.
Establish and drive MLOps practices: continuous improvement across AI engineering practices, reproducible training pipelines, model versioning and registry, automation, CI/CD, monitoring, automated quality evaluation, testing, and release management.
Support adoption of AI capabilities across our client applications and workflows through reusable platforms, services, APIs, and technical patterns.
Create clarity around priorities, risks, technical tradeoffs, delivery timelines, and resourcing needs.
Ensure the team can operate production AI systems independently through strong documentation, runbooks, monitoring, incident response, and knowledge sharing.
Partner with U.S. and European stakeholders across engineering, product, operations, localization, media workflows, and business leadership. Partner closely with the leaders of our client’s localization business lines to understand where current MT/ASR/TTS quality falls short and translate that into a prioritized technical roadmap.
Ensure AI systems are reliable, secure, maintainable, cost-conscious, and aligned with production business needs.
Represent AI/ML technology capability and constraints to executive leadership; contribute to the company's overall technology and product strategy.
QUALIFICATIONS:
[10+] years in software/ML engineering, including [4+] years in a technical leadership role owning both people management and architecture/system-design decisions.
Direct, hands-on experience with model training/fine-tuning, evaluation methodology, and production deployment (not solely API-consumer experience).
Proven experience designing and operating production ML infrastructure at scale: training pipelines, model serving/inference, GPU resource management, and MLOps tooling, on at least one major cloud provider (AWS, GCP, or Azure). Experience in distributed training, Slurm/HPC, Kubernetes, Docker, Terraform, or similar infrastructure technologies.
Track record of building a technical team from scratch to a functioning, multi-workstream organization, including recruiting, leveling, and distributed team design across locations and time zones.
Demonstrated ability to translate business/product requirements from non-technical stakeholders into a technical roadmap, and to communicate technical trade-offs to executives in business terms.
Significant experience leading engineering teams responsible for cloud platforms, AI/ML systems, data platforms, workflow automation, or production software systems.
Strong technical background in cloud architecture, software engineering, platform engineering, DevOps, or AI/ML engineering.
Experience with agentic workflows, LLM orchestration, model evaluation, RAG patterns, or AI-enabled automation platforms.
Experience evaluating commercial and open-source AI models or vendor platforms.
Experience delivering production systems that require reliability, monitoring, automation, security, and operational ownership.
Strong understanding of modern AI architectures, model-serving patterns, APIs, data pipelines, and applied AI integration.
Ability to evaluate technical options, make pragmatic recommendations, and communicate tradeoffs to both technical and non-technical audiences.
Experience leading complex technical programs from planning through production deployment and ongoing operations.
Ability to work effectively with U.S.-based leadership while operating independently from Germany.
PREFERRED EXPERIENCE:
Experience with speech, language, localization, dubbing, media workflows, ASR, MT, TTS, or content-processing systems.
Prior experience in media, localization, entertainment technology, or a related content-production industry — direct exposure to subtitling, dubbing, or content compliance workflows is a strong plus.
Familiarity with content-security frameworks used by major studios and broadcasters (e.g., MPA content security best practices, Trusted Partner Network) and what they require of vendor/internal systems.
Prior experience taking over or migrating away from a third-party AI/ML vendor relationship, including contract wind-down and knowledge-transfer.
Experience in media, entertainment, localization, post-production, or content supply chain environments.
Required skills
- docker
- communication
- documentation
- ci/cd
- operations
- security
- monitoring
- incident response
- reliability
- cost management
- planning
- aws
- kubernetes
- nlp
- automation
- software engineering
- cloud-native
- testing
- gcp
- azure
- product strategy
- product management
- apis
- devops
- terraform
- machine learning
- architecture
- technical leadership
- people management
- team building
- observability
- cloud infrastructure
- stakeholder management
- recruiting
- vendor management
- partnerships
- mt
- knowledge sharing
- data pipelines
- localization
- operational support
- workflow automation
- mlops
- language
- release management
- system design
- content production
- orchestration
- autoscaling
- platform engineering
- ai services
- leveling
- maintainability
- executive communication
- production operations
- asr
- cloud architecture
- priorities
- business requirements
- production deployment
- training platform
- product requirements
- hpc
- orchestration tools
- fine-tuning
- speech
- deployment pipelines
- executive leadership
- model deployment
- slurm
- rag patterns
- knowledge transfer
- operating rhythms
- tts
- subtitling
- distributed teams
- delivery timelines
- open-source models
- agentic workflows
- engineering standards
- data platforms
- post-production
- technical roadmap
- risks
- distributed training
- model serving
- technology strategy
- llm orchestration
- mpa
- content supply chain
- ml infrastructure
- model training
- dubbing
- time zones
- content compliance
- vendor platforms
- business leadership
- media industry
- non-technical audiences
- ai-enabled automation
- runbooks
- third-party integration
- commercial models
- technical review
- technical audiences
- cost-conscious
- model versioning
- model integration
- delivery practices
- agentic frameworks
- media workflows
- localization industry
- model registry
- build vs buy
- production ai systems
- inference platform
- pragmatic recommendations
- evaluation methodology
- gpu utilization
- technical tradeoffs
- reusable platforms
- ai platform operations
- technical programs
- applied ai engineering
- ai platform engineering
- scalable ai services
- reproducible training pipelines
- automated quality evaluation
- technical patterns
- resourcing needs
- production business needs
- applied ai integration
- content processing systems
- entertainment technology
- content security
- trusted partner network
- contract wind-down
Interested in this position?
Create your free account and tailor your CV to match this job.