Login Enter

WeSort.AI GmbH

3 months ago

Machine Learning Engineer (m/w/d) - Foundation Models

Sign up free Log in

Sign up to save this job, get alerts, and apply with an optimized CV.

Company information

Company
WeSort.AI GmbH
Location
Würzburg Germany
Posted
3 months ago
View all jobs at WeSort.AI GmbH

Job description

This is your new passion

  • You develop and train our domain-specific vision models based on current state-of-the-art architectures and use our waste image data basis for pretraining and continued pretraining.
  • You design our complete ML training pipeline: from data preparation to distributed training (PyTorch FSDP/DDP, Mixed Precision) to model versioning.
  • You build and maintain our Eval Suite – the central infrastructure that measures whether our models are really getting better: Linear Probing, k-NN-Probing, Few-Shot-Detection, Cross-Domain-Generalization, Anomaly Detection.
  • You fine-tune and distill our models for specific downstream tasks and edge hardware (sorting plants, GPU inference).
  • You systematically analyze training runs, identify problems like feature collapse or domain shift, and develop sustainable solutions instead of quick fixes.
  • You work closely with the cloud backend team to bring models efficiently into deployment (ONNX, TensorRT, OpenVINO).
  • You actively follow research development in the field of computer vision and translate relevant papers into productive solutions.
  • You think beyond the model and have a view on how your work will affect real-world operations – for sorting plants, customers, and the overall system.

This is what excites us

  • You bring several years of experience in developing and training computer vision models with modern vision transformer architectures and self-supervised learning methods.
  • You master PyTorch securely – including distributed training (DDP, FSDP), mixed precision (bf16/fp16), and performance optimization (torch.compile, profiling).
  • You understand not only how to train a model but also how to evaluate it. You know that a weak eval suite makes any pretraining worthless.
  • You have experience with modern ML tooling stacks for configuration management, experiment tracking, data versioning, and backbone libraries.
  • You use modern AI tools (e.g., Claude, Copilot) to accelerate routine coding and focus on the really hard research and architecture questions.
  • You have a good understanding of data pipelines with large datasets (millions of images): efficient data formats, GPU augmentations, I/O bottlenecks.
  • Experience with common detection/segmentation and anomaly detection frameworks is an advantage.
  • You are familiar with inference optimization and model distillation and have ideally already deployed models on edge hardware.
  • You possess exceptional problem-solving skills, analytical thinking, and scientific rigor – you work hypothesis-driven and not by the try-and-error principle.
  • You handle cloud GPU infrastructure (AWS, Azure, GCP or On-Premise H100/A100 cluster) securely.
  • You are expected to have fluent German and good English skills.
  • Ideal candidates have their own research experience (papers, open-source contributions, conference talks) or a Ph.D. – not a must but a plus.

This is what you can look forward to

  • Work on the

Interested in this position?

Create your free account and tailor your CV to match this job.