VLM Research Engineer
Almetra
Seniority
Midweight
Model
In-Office
Sector
Salary
Undisclosed
Contract
Full-Time
Design and adapt vision-language and video models for scene understanding, temporal reasoning and activity / action recognition. You'll work on applied, state-of-the-art multimodal models and turn them into production pipelines used by customers.
What you'll do
- Design and adapt vision-language and video models for scene understanding, temporal reasoning and activity / action recognition
- Build and maintain large-scale training and evaluation pipelines on GPU clusters
- Curate and augment video-text and action datasets, including synthetic labels and retrieval-based augmentation
- Develop robust benchmarks for video QA, instruction following and temporal understanding, and use them to drive iterative model improvements
- Cut and refactor model architectures for efficiency and deployability (compression, pruning, distillation)
- Deliver production-ready inference pipelines to product and customer teams, working closely with CV, platform and robotics engineers
What you'll need
- Completed PhD (or equivalent research track record) in computer vision, machine learning, robotics or a related field
- Strong background in video-centric deep learning: scene understanding, temporal / activity / action recognition, or video generation
- Experience training and adapting large vision or VLM models (e.g. InternVL, Qwen-VL, DeepSeek-VL, similar stacks)
- Proven work with multi-GPU training (PyTorch, distributed, mixed precision) and large-scale datasets
- Solid engineering habits: clean Python, reproducible experiments, reliable data and training pipelines
- Track record of moving research into usable systems (demos, internal tools, or productised features) in fast-moving teams
Nice to have
- Publications at top-tier venues (CVPR, ICCV, ECCV, NeurIPS, ICLR, etc.) on video, multimodal learning or scene understanding
- Experience with 3D/4D scene representations, action generation or embodied / sense-plan-act style projects
- Inference optimisation: quantisation, TensorRT, model distillation, or deployment on constrained hardware
- Prior experience in a startup or applied research lab environment
What they offer
- Employee Share Options Program for all permanent employees
- Urban Sports club and quarterly team retreats
- Flat hierarchy, open 360° feedback, and flexible working hours

