Job Description

Your mission & challenges

As an AI Data Annotation Specialist, you work at the intersection of data ingestion, annotation workflow design, and machine learning—with a strong focus on building training datasets for multimodal foundation models that combine perception, language, and robotic action. Your main responsibility is the design and maintenance of scalable workflows for automated and human-in-the-loop annotation, ensuring that datasets are correctly labeled, curated, validated, and prepared for efficient model training and evaluation.

You play a crucial role in enabling high-quality embodied AI systems by transforming raw multimodal robot data into structured, reliable, and semantically rich training datasets.

  • You design, build, and maintain pipelines for automated, semi-automated, and human-in-the-loop data annotation, with a focus on subtask labeling for long-horizon robot demonstrations.

  • You will develop annotation workflows for language-based robot learning and visual question answering (VQA) datasets, including instruction generation, subtask decomposition, and the anchoring of natural language in visual and action data.

  • You capture multimodal data (video, depth, proprioception, gripper states, speech instructions) and integrate it into structured annotation workflows.

  • You apply pre-labeling techniques with foundation models (e.g., Vision Language Models, LLMs) to speed up annotation and reduce manual effort.

  • You define and implement data quality checks: Inter-Annotator Agreement, Label Consistency, Coverage Analysis, and Annotation Drift Detection.

  • You drive data curation initiatives forward: dataset balancing, deduplication, failure case mining, task diversity analysis, and targeted data collection to close capability gaps.

  • You identify and resolve data quality issues, labeling inconsistencies, distribution bias, and gaps in task/skill coverage.

What you bring

  • Degree in Computer Science, Data Science, Engineering or a related field

  • 4+ years of experience in Machine Learning Operations, AI or Software Engineering

  • Practical experience with data annotation tools and labeling workflows (e.g., Encord, CVAT, Label Studio or similar platforms)

  • Experience with subtask annotation, ontology/taxonomy design, and VQA-style labeling

  • Familiarity with robotics dataset formats and multimodal data structuring (e.g., LeRobot, RLDS, Open X-Embodiment)

  • Experience with cloud platforms (AWS, GCP, Azure) is an advantage.

  • Experience in writing annotation guidelines / taxonomies (complements the above-mentioned responsibilities)

  • Familiarity with VLMs/LLMs for auto-labeling (e.g., open-vocabulary detectors, Gemini/GPT class models)

  • Basic knowledge of data engineering — confident handling of large-scale/streaming data and formats such as ROS Bags/MCAP and S3-like object storage

  • Experience in building backend services and tooling (e.g., Node.js/TypeScript, REST APIs) for integrating annotation platforms into internal data infrastructure

  • Advantageous: direct experience with embodied AI/teleoperations data