WeSort.AI GmbH
3 months ago
Machine Learning Engineer (m/w/d) - Foundation Models
Sign up to save this job, get alerts, and apply with an optimized CV.
Company information
- Company
- WeSort.AI GmbH
- Location
- Würzburg Germany
- Posted
- 3 months ago
Job description
This is your new passion
- You develop and train our domain-specific vision models based on current state-of-the-art architectures and use our waste image data basis for pretraining and continued pretraining.
- You design our complete ML training pipeline: from data preparation to distributed training (PyTorch FSDP/DDP, Mixed Precision) to model versioning.
- You build and maintain our Eval Suite – the central infrastructure that measures whether our models are really getting better: Linear Probing, k-NN-Probing, Few-Shot-Detection, Cross-Domain-Generalization, Anomaly Detection.
- You fine-tune and distill our models for specific downstream tasks and edge hardware (sorting plants, GPU inference).
- You systematically analyze training runs, identify problems like feature collapse or domain shift, and develop sustainable solutions instead of quick fixes.
- You work closely with the cloud backend team to bring models efficiently into deployment (ONNX, TensorRT, OpenVINO).
- You actively follow research development in the field of computer vision and translate relevant papers into productive solutions.
- You think beyond the model and have a view on how your work will affect real-world operations – for sorting plants, customers, and the overall system.
This is what excites us
- You bring several years of experience in developing and training computer vision models with modern vision transformer architectures and self-supervised learning methods.
- You master PyTorch securely – including distributed training (DDP, FSDP), mixed precision (bf16/fp16), and performance optimization (torch.compile, profiling).
- You understand not only how to train a model but also how to evaluate it. You know that a weak eval suite makes any pretraining worthless.
- You have experience with modern ML tooling stacks for configuration management, experiment tracking, data versioning, and backbone libraries.
- You use modern AI tools (e.g., Claude, Copilot) to accelerate routine coding and focus on the really hard research and architecture questions.
- You have a good understanding of data pipelines with large datasets (millions of images): efficient data formats, GPU augmentations, I/O bottlenecks.
- Experience with common detection/segmentation and anomaly detection frameworks is an advantage.
- You are familiar with inference optimization and model distillation and have ideally already deployed models on edge hardware.
- You possess exceptional problem-solving skills, analytical thinking, and scientific rigor – you work hypothesis-driven and not by the try-and-error principle.
- You handle cloud GPU infrastructure (AWS, Azure, GCP or On-Premise H100/A100 cluster) securely.
- You are expected to have fluent German and good English skills.
- Ideal candidates have their own research experience (papers, open-source contributions, conference talks) or a Ph.D. – not a must but a plus.
This is what you can look forward to
- Work on the
Interested in this position?
Create your free account and tailor your CV to match this job.