R&D-038 Data Engineer
- Employment
- Full Time
- Posted
- 5d ago
(日本語が下部に続きます) About AIRoA The AI Robot Association (AIRoA) is launching a groundbreaking initiative: collecting one million hours of humanoid robot operation data with hundreds of robots, and leveraging it to train the world’s most powerful Vision-Language-Action (VLA) models. What makes AIRoA unique is not only the unprecedented scale of real-world data and humanoid platforms, but also our commitment to making everything open and accessible. We are building a shared “robot data ecosystem” where datasets, trained models, and benchmarks are available to everyone. Researchers around the world will be able to evaluate their models on standardized humanoid robots through our open evaluation platform. For researchers, this means an opportunity to: Work on fundamental challenges in robotics and AI: multimodal learning, tactile-rich manipulation, sim-to-real transfer, and large-scale benchmarking. Access state-of-the-art infrastructure: hundreds of humanoid robots, GPU clusters, high-fidelity simulators, and a global-scale evaluation pipeline. Collaborate with leading experts across academia and industry, and publish results that will shape the next decade of robotics. Contribute to an initiative that will redefine the future of embodied AI—with all results made open to the world. Key Responsibilities You will play a critical role in building the data backbone powering next-generation robotics foundation models: Design and implement large-scale data pipelines that cover the full lifecycle of high-quality datasets for robotics foundation models—collection, processing, curation, and publishing. Design, build, and maintain data schemas, storage solutions, and query interfaces to enable VLA researchers to efficiently discover, query, and consume curated datasets. Collaborate closely with VLA researchers to capture evolving data requirements and continuously improve data pipelines through analysis and experimentation. Design and scale distributed data-processing pipelines capable of handling petabyte-scale multimodal datasets (e.g., RGB/Depth, point clouds) with full lineage and reproducibility. Define data-quality metrics and build feedback loops to continuously monitor and improve data quality. AIRoAについて AI Robot Association(AIRoA)は、数百台のヒューマノイドロボットを用いて100万時間分のロボット操作データを収集し、それを活用して世界最高水準のVision-Language-Action(VLA)モデルを学習させるという、画期的な取り組みを進めています。 AIRoAの特徴は、実世界データとヒューマノイドプラットフォームの規模がこれまでにないものであるだけでなく、すべてをオープンかつ誰もが利用できる形にすることを目指している点にあります。私たちは、データセット、学習済みモデル、ベンチマークを誰もが利用できる共通の「ロボットデータ・エコシステム」を構築しています。また、世界中の研究者が、オープンな評価プラットフォームを通じて、標準化されたヒューマノイドロボット上で自身のモデルを評価できる環境を提供します。 研究者にとって、これは以下のような機会を意味します。 マルチモーダル学習、触覚情報を活用したマニピュレーション、Sim-to-Real転移、大規模ベンチマーキングなど、ロボティクスおよびAIにおける本質的な課題に取り組む。 数百台のヒューマノイドロボット、GPUクラスター、高精度シミュレーター、グローバル規模の評価パイプラインといった最先端のインフラを利用する。 学術界・産業界を代表する専門家と協働し、今後10年のロボティクスを形作る研究成果を発表する。 Embodied AIの未来を再定義する取り組みに貢献し、その成果をすべて世界にオープンにする。 主な業務内容 次世代のロボティクス基盤モデルを支えるデータ基盤の構築において、重要な役割を担っていただきます。 ロボティクス基盤モデル向けの高品質なデータセットについて、収集、処理、キュレーション、公開までのライフサイクル全体をカバーする大規模データパイプラインを設計・実装する。 VLA研究者がキュレーション済みデータセットを効率的に発見・検索・利用できるよう、データスキーマ、ストレージソリューション、クエリインターフェースを設計・構築・運用する。 VLA研究者と密接に連携し、変化するデータ要件を把握するとともに、分析や実験を通じてデータパイプラインを継続的に改善する。 RGB/Depth、ポイントクラウドなどのペタバイト規模のマルチモーダルデータセットを扱い、完全なデータリネージと再現性を確保できる分散データ処理パイプラインを設計し、スケールさせる。 データ品質指標を定義し、データ品質を継続的に監視・改善するためのフィードバックループを構築する。 Requirements (日本語が下部に続きます) Required Qualifications 【1. Academic & Professional】 Bachelor's degree in Computer Science, Engineering, or related field (or equivalent practical experience). 5+ years professional experience in data engineering / data platform development. Proven record of delivering production-grade, distributed data systems. 【2. Large-Scale Multimodal / Unstructured Data Processing】 Experience personally designing and operating pipelines that process unstructured data such as video, images, point clouds, or sensor time-series at scale (100TB+ in total, or TB+/day throughput). Working understanding of storage formats (Parquet, WebDataset, MCAP/rosbag, etc.) or codecs (H.264, H.265, AV1, etc.) for such data. Relevant domains include autonomous driving, robotics, video/streaming, industrial IoT, and video analytics. Equivalent experience from other domains is also welcome. 【3. ETL / Distributed Data Processing】 3+ years designing and operating large-scale ETL / ELT pipelines using a distributed engine such as Spark, Flink, or Ray. Experience personally building and operating pipelines with orchestration tools such as Airflow or Dagster. 【4. Collaboration & Data Democratization】 Worked with research, product, or analytics teams and designed pipelines and platforms around their goals. Built tools, interfaces, and documentation that let users who are not data-engineering specialists (such as researchers and analysts) find and use data on their own. Preferred Qualifications Experience analyzing and evaluating robot data: statistics and visualization of trajectories and manipulation logs, task success/failure assessment, and metrics for dataset diversity and quality. Understanding of sensor characteristics (camera, depth, IMU, force/tactile), time synchronization, and calibration; knowledge of robot learning dataset formats such as LeRobot and Open X-Embodiment (RLDS). Proven optimization of workloads at 10TB+/day or petabyte scale. Experience operating data-processing workloads on Kubernetes (e.g., EKS, GKE). Experience building pipelines that feed ML training jobs, including dataloader/sharding optimization (WebDataset, Mosaic StreamingDataset, Lance, etc.) and dataset versioning for reproducibility. Experience building dataset search and discovery, using full-text search and vector search (e.g., OpenSearch, FAISS, pgvector). Experience building and operating annotation platforms (e.g., CVAT, Label Studio, Labelbox), designing human-in-the-loop workflows, and managing annotation quality. Experience with large-scale telemetry or streaming ingestion from devices and fleets (e.g., Kafka, Kinesis, MQTT). Experience building lakehouses