Job Drop BerlinYOUR WAY INTO BERLIN TECH
NewsletterLinkedIn
AboutTermsImpressumPrivacy
Browse jobs
Engineering jobs in BerlinProduct jobs in BerlinDesign jobs in BerlinMarketing jobs in BerlinSales jobs in BerlinData jobs in BerlinOperations jobs in BerlinFinance jobs in BerlinCustomer success jobs in BerlinPeople & HR jobs in BerlinEnglish-speaking jobs in BerlinRemote jobs at Berlin startupsStartup internships in Berlin

ML Engineer, Infrastructure

PPrior Labs
Seniority
Midweight
Model
In-Office
Sector
AI-native
Salary
Undisclosed
Contract
Full-Time

You own the full stack of multi-cluster GPU infrastructure, from cluster operations and cost optimization to distributed training performance and the tooling layer that keeps researchers moving fast. You work directly with the research team and understand what they're doing well enough to make infrastructure decisions that actually help them.

What you'll do

  • Own and evolve multi-cluster GPU infrastructure. Slurm on GCP today, multi-provider and new hardware tomorrow. Architecture, scheduling, reliability, cost optimization
  • Drive GPU utilization and training throughput: profiling, memory optimization, communication bottlenecks, systems-level debugging of distributed training across large runs
  • Architect the next generation of our infrastructure: multi-cluster orchestration, new GPU generations, provider diversification, capacity planning against growing compute demands
  • Build the developer productivity layer: CI pipelines, experiment tracking, model registry, data processing, and internal tooling that keeps research iteration speed high
  • Own the compute budget. You understand cost per FLOP across providers and hardware, and you hate wasted compute

What you'll need

  • 5+ years building and operating production GPU infrastructure or distributed training systems at scale. At a major AI lab, a well-funded ML startup, or an HPC environment
  • Deep hands-on experience with Slurm and cluster management. You've debugged scheduling failures, optimized utilization across multi-tenant GPU workloads, and operated infrastructure where downtime has real cost
  • Expert-level systems thinking: memory bandwidth, GPU profiling. You reason about hardware, not configs
  • Strong Python and genuine fluency with PyTorch internals. Enough to profile a training run and tell whether the bottleneck is data loading, communication, or compute
  • Track record of making infrastructure decisions that measurably improved training throughput or cost efficiency
  • Strong AI tooling skills. You use Claude Code, Cursor, or similar fluently to move fast without sacrificing quality

Nice to have

  • Experience operating at tens-of-millions-scale GPU spend
  • Multi-cloud or hybrid HPC/cloud infrastructure experience
  • Triton, CUDA, or custom kernel experience
  • Experience scaling from single cluster to multi-cluster orchestration
  • Background building experiment tracking, model registry, or ML pipeline tooling
APPLY →

ABOUT PRIOR LABS

AI-native · Pre-Seed stage

20+ employees

9 more open roles at Prior Labs

This role is English-speaking — no German required.

SIMILAR ROLES THIS WEEK

More roles like this, every Thursday →

No spam. Unsubscribe anytime.