Are you passionate about Kubernetes, platform engineering, and modern AI technologies? Do you want to do more than just deploy individual models, but create the infrastructure and platforms on which innovative AI applications can be operated reliably, securely, and scalably?
Then you've come to the right place.
As a (Senior) AI Platform Engineer (m/f/d), you will develop, operate, and optimize modern AI and ML platforms for our clients. You will work with Kubernetes, GPU infrastructures, Infrastructure as Code, and modern CI/CD, MLOps, and LLMOps approaches, supporting companies in successfully deploying Generative AI and Machine Learning in production.
You can expect exciting projects in a wide variety of industries, technological creative freedom, and a team of cloud, platform, data, and AI specialists who not only discuss innovation but implement it daily.
Tasks
- Planning, development and operation of highly available AI/ML platforms and services, especially based on Kubernetes in on-premises, hybrid and cloud environments.
- Design and implementation of scalable GPU infrastructures within Kubernetes clusters – including GPU scheduling and sharing (e.g. Kueue, KAI-Scheduler, Run:ai, NVIDIA MIG/Time-Slicing) with a focus on utilization and cost efficiency.
- Development and operation of model serving and inference services (e.g., vLLM, NVIDIA Triton, KServe, NVIDIA NIM, Ray) for classical ML models and Large Language Models.
- Development and operation of MLOps/LLMOps workflows for productive GenAI services – including guardrails & policies, evaluation, regression tests, and cost and quality monitoring.
- Provision of fine-tuning and re-training workflows (e.g., LoRA, continuous re-training) on the platform.
- Monitoring, troubleshooting and performance tuning of AI platforms (including GPU utilization, inference latency, throughput) to ensure maximum performance and availability.
- Advising our clients on best practices, AI governance, data security and privacy in the context of AI platforms on Kubernetes.
- Close collaboration with data science, data platform and development teams to ensure seamless integration of AI solutions.
qualification
- At least 3 years of experience in operating and optimizing platforms and services at scale in Kubernetes environments, whether on-premises or in the cloud.
- Solid knowledge of the architecture and operation of scalable solutions using Kubernetes, ideally including GPU workloads.
- Programming skills in Python, ideally complemented by Golang or Java.
- Experience with at least one model serving framework (e.g., vLLM, NVIDIA Triton, KServe) and initial or in-depth experience with MLOps tools (e.g., MLflow, Kubeflow, ClearML).
- Experience with Infrastructure-as-Code tools such as Ansible and Terraform, as well as CI/CD and GitOps tools (e.g., Argo CD, Flux).
- Excellent communication skills in German and English. Ideally, you have already advised clients and can explain technical topics in a way that is appropriate for the target group – from the engineering team to management.
For the senior role, additionally:
- At least 5 years of relevant experience in platform engineering.
- Architecture Ownership: You are responsible for the end-to-end design of AI platforms and make fundamental technological decisions.
- Technical leadership in customer projects (lead role) and mentoring of colleagues.
Ideally, you should also have:
- Experience with cloud AI services (e.g., AWS SageMaker/Bedrock, Azure ML/Azure OpenAI/Azure AI Foundry) and their integration with Kubernetes workloads.
- Knowledge of operating LLM-based architectures, e.g., RAG pipelines, vector databases and embedding services, as well as agentic AI approaches and Model Context Protocol (MCP).
- Understanding of system-level fundamentals of LLM serving (rate limiting, token streaming, load balancing) and LLM concepts such as reasoning, tool calling, and prompt templates.
- Experience with new serving components such as llm-d or Kubernetes Inference Gateway.
- Knowledge of securing AI platforms, e.g. TLS, RBAC and network policies within Kubernetes environments, as well as a basic understanding of AI governance (e.g. EU AI Act).
- Certifications in relevant technologies (e.g., Certified Kubernetes Administrator/Developer, NVIDIA certifications).
- Willingness to travel occasionally within the DACH region (Germany, Austria, Switzerland).
Benefits
We offer you, among other things:
- Occupational pension scheme
- Team-oriented corporate culture with regular team events
- Targeted support for your professional development
- Corporate Benefits Program
- Flexible working hours
- Modern workplace with contemporary equipment
- Remote work options
We look forward to receiving your online application and learning more about you.
We get it done.