Sign up to save this job, get alerts, and apply with an optimized CV.

Head of AI Engineering (f/m/x)

Full Time Lead

Job description


Your mission

Own and evolve our AI engineering function — transforming a 15–20 person ML team from research-heavy to a high-throughput, production-grade organization. You’ll partner with the CTO on strategy, build the platform that unifies LLM access, RAG, and backend services, and ship reliable, scalable AI features that change how banks work.
 
Key responsibilities 

  • Team leadership and org build
    • Hire, mentor, and develop a high-performing team; set the technical bar, operating rhythms, and code/research review practices
    • Organize sub-teams (e.g., Core Modeling, AI Platform/Infra, Integrations) with clear ownership, SLOs, andon-call
    • Manage roadmap, capacity planning, and delivery across parallel initiatives
  • Architecture and platform
    • Own the LLM gateway: unified APIs and proxy layers for multi-provider routing (OpenAI, Gemini, Bedrock), with rate limits, fallbacks, and cost tracking
    • Build high-performance RAG pipelines (ingestion, embeddings, vector stores, caching) with robust observability and safety guardrails
    • Partner with Java/NestJSteams to define clean async contracts, schemas, andeventingpatterns; drive low-latency, scalable inference
  • Model lifecycle and operations
    • Lead end-to-end model and prompt lifecycle: data curation, training/fine-tuning, evaluation, deployment, rollback
    • EstablishLLMOps/MLOps: model/prompt registries, CI/CD, canary/A/B tests, offline/online evals, drift and cost monitoring
    • Optimizeinference throughput and cost (autoscaling, batching, quantization/distillation, caching)
  • Strategy and collaboration
    • Translate company goals into an AI/ML roadmap with measurable outcomes; balance exploration with reliability and cost
    • Own build-vs-buy/vendor strategy for models, infrastructure, and data services; manage budgets and SLAs
  • Governance and security
    • Implement data privacy, security, and compliance practices (RBAC, secrets, auditability); track prompt/model lineage and reproducibility
    • Define incident response, runbooks, and postmortems for AI features


Your profile

  • 5+ years as a backend engineer and 4+ years leading AI/ML engineering in production (10+ years total experience ideal)
  • Deep architectureexpertisein Java (JVM) and/or Node.js(NestJS), distributed systems, APIs, microservices, and messaging/streaming
  • Hands-on with LLM stacks: orchestration (e.g.,LangChain/LlamaIndexor custom), vector DBs (Pinecone,Qdrant, FAISS), cloud AI (e.g., AWS Bedrock)
  • Proven operation of systems at scale (millions of daily API calls) with strong SLOs, observability, and incident management
  • MLOpsfoundations: model registries, experiment tracking, CI/CD, Kubernetes,IaC(e.g., Terraform), security best practices
  • Excellent communication and stakeholder management; strong product sense focused on shipping user-facing feature 
  • Fluent German and English for daily team collaboration, stakeholder management, and technical documentation
Nice to have 
  • Experience with GPU/accelerator serving and optimization (vLLM, TGI, Triton, ONNX Runtime)
  • Cost optimization for LLM workloads (token budgets, dynamic routing, caching)
  • Evaluation and safety/red-teaming for generative systems; startup/high-growth experience
Impact metrics 
  • Platform: adoption of a unified LLM gateway; standardized observability and cost reporting
  • Delivery: 2–3 user-facing AI features shipped with clear SLOs and measurable impact
  • Reliability/cost: reduced average latency and cost per request; autoscaling and caching in place
  • Org: sub-team structureestablished; improved code quality and on-time delivery; targeted hiring completed
Our stack  
  • Backend: Java (JVM), Node.js(NestJS); event-driven microservices; API gateways/proxies
  • AI platform: Python,PyTorch, LLM orchestration, prompt pipelines/registry; vector DBs (Pinecone,Qdrant); RAG services
  • Infra/DevOps: AWS (incl. Bedrock), Kubernetes, Terraform, CI/CD, Observability (OpenTelemetry, Prometheus/Grafana)


Why us

  • Because we value talent more than hierarchy.
  • Because at neoshare, responsibility isn't delegated - it's owned.
  • Because we use modern AI and technology as a lever for exceptional results.
  • Because we develop people who want to learn, grow, and deliver.
  • Because performance, quality, and impact belong together for us.
  • Because we are working together towards building a European tech champion.


What You Can Expect

  • Performance-driven, above-average compensation that rewards outstanding commitment.
  • High-end offices designed to support collaboration, wellbeing, and peak performance - including great health and fitness benefits.
  • Legendary team events where we celebrate our wins together and strengthen team spirit.
  • State-of-the-art AI tools, first-class equipment, and an environment that fosters ownership and personal growth.
  • Concentration of top talent, fast decision-making, and the chance to make a real impact early on.
Candidates must have the right to work in the EU; visa sponsorship is not provided for this role. 

Find more English Speaking Jobs in Germany on Arbeitnow

Required skills

english communication technology german responsibility training safety quality java ci/cd security incident response governance team leadership strategy python aws kubernetes llm pytorch delivery collaboration caching evaluation learning node.js embeddings apis terraform prometheus grafana code quality infrastructure equipment architecture wellbeing microservices observability gemini personal growth backend services capacity planning messaging budgets stakeholder management impact backend engineer team spirit performance incident management platform growth code review team events distributed systems nestjs iac cost optimization organizational development rag openai startup deployment compensation slas llamaindex hiring ai tools commitment distillation mlops proxies orchestration langchain autoscaling ownership ai engineering health benefits security best practices streaming cost tracking feature delivery talent rbac opentelemetry visa sponsorship fine-tuning data privacy a/b tests faiss pinecone cost reporting top talent cost monitoring slos reproducibility aws bedrock llmops operating rhythms cto jvm vllm roadmap management delivery management secrets management bedrock on-call high-growth scalable inference data services batching platform adoption model registries vector dbs api gateways hierarchy ingestion model lineage schemas fitness benefits measurable outcomes triton on-time delivery postmortems auditability peak performance qdrant runbooks exceptional results quantization tgi safety guardrails fast decision-making rate limits product sense rollback ml engineering user-facing features data curation vector stores event-driven microservices onnx runtime multi-provider routing model lifecycle dynamic routing build vs buy systems at scale low-latency inference technical bar fallbacks generative systems cloud ai prompt pipelines vendor strategy canary tests drift monitoring sub-teams research review llm gateway proxy layers high-performance rag async contracts eventing patterns prompt lifecycle prompt registries offline evals online evals inference throughput ai/ml roadmap compliance practices prompt lineage production ai/ml llm stacks millions of daily api calls gpu serving accelerator serving token budgets red-teaming unified llm gateway standardized observability average latency cost per request sub-team structure prompt registry rag services modern ai european tech champion high-end offices right to work in eu

Sign up to apply

Create a free account to apply for this job and get access to:

  • AI-powered CV optimization for this specific job
  • Save jobs and create custom alerts
  • See your CV match score for each job

Company information

Company
neoshare
Location
Berlin
Germany
Posted
10 hours ago

Find similar jobs

Explore more opportunities like this one.

Interested in this position?

Create your free account and tailor your CV to match this job.