Data Platform Engineer – Data Operations
Stark
Seniority
Midweight
Model
In-Office
Sector
Salary
Undisclosed
Contract
Full-Time
You own the software backbone of STARK's data platform: the metadata systems, ETL pipelines, data contracts, catalogs, databases, and internal tools that let engineers find, understand, validate, and reuse terabytes of multi-sensor field data in minutes, not days. You treat data context as a product: structured, searchable, version-aware, documented, and traceable from raw recording to processed asset.
What you'll do
- Design, write, and maintain robust Python software and ETL pipelines that ingest, process, and merge complex field recordings (flight data, rosbags, video) and external deliveries into GCP cloud storage.
- Build the programmatic tools that bridge raw data and downstream usage, including services to automatically sub-sample video feeds, extract valuable frames, and package datasets.
- Design, implement, and maintain metadata databases and data catalogs (tracking datasets, recordings, sensors, labels, and lineage) using solid SQL and schema design principles.
- Own the end-to-end technical workflows for data curation and labeling, building operational tooling and coordinating with external labeling subcontractors to ensure high-quality data deliveries and run automated QA.
- Develop backend APIs, self-service dashboards, and data-access tools that allow the entire AI and engineering organization to quickly search, filter, and understand terabytes of multi-sensor data.
- Establish and enforce good software engineering hygiene in the codebase: rigorous testing, typing, documentation, logging, and CI/CD pipelines.
What you'll need
- Strong software engineering in Python: you write clean, typed, tested, and maintainable code with a software developer's mindset.
- Proven experience designing, building, and operating robust ETL and data ingestion pipelines.
- Solid SQL (PostgreSQL preferred) with a strong grasp of schema design, data modeling, and metadata systems.
- Hands-on experience with object storage (GCP/GCS, S3), containerization (Docker), and CI/CD.
- Strong familiarity with Unix-based systems, shell scripting, and standard Unix tools; comfortable working natively in a Linux environment.
- Ability to balance quick fixes with long-term architectural solutions, and comfort with changing priorities in a startup environment.
- Comfortable working cross-functionally and coordinating with external vendors and non-technical stakeholders.
Nice to have
- Experience with synthetic data generation or GenAI-assisted workflows (auto-labeling, data augmentation, foundation-model-based curation).
- Familiarity with data versioning (DVC, LakeFS, FiftyOne), computer vision annotation formats (e.g., COCO), or large-scale data curation workflows.
- Experience with GCP services beyond basic storage (BigQuery, Cloud Run, IAM).
- Exposure to robotics data formats (ROS bags, MCAP, PX4 logs) or handling heavy, multi-modal data streams (video, lidar).

