HFlow is an open source SDK for scalable multimodal data pipelines in robotics and physical AI. Define your own transforms, checks, labels, and enrichments. HFlow handles orchestration, storage, versioning, provenance, and curation, so you can inspect data quality and reproduce how each dataset was built. HFlow is built by Hebbian Robotics and is part of Y Combinator’s Summer 2026 batch.
We built HFlow because robotics teams are collecting more hours of video data than they can handle. Each episode can contain camera streams, robot state, actions, timestamps, and metadata. As a corpus grows, one-off scripts make it hard to answer basic questions: Did a camera freeze? Did streams drift out of sync? Which version of a check ran? Can we reproduce the dataset we used?
HFlow turns that work into a pipeline. Transforms, checks, labels, and enrichments stay yours. HFlow handles orchestration, storage, versioning, and curation around them. It writes canonical MCAP episodes with provenance, records quality evidence in a Parquet catalog, and builds version-pinned manifests with DuckDB.
The core lifecycle works end to end today. You can run the included quickstart locally without Docker, Airflow, robot hardware, or an external service.
HFlow is Apache-2.0 licensed and built in public by Hebbian Robotics (YC S26). If you work with robotics or physical AI data, try the quickstart and tell us where your current data workflow loses the most context. Is it ingestion, quality control, provenance, or curation?
HFlow
Hi Product Hunt,
We built HFlow because robotics teams are collecting more hours of video data than they can handle. Each episode can contain camera streams, robot state, actions, timestamps, and metadata. As a corpus grows, one-off scripts make it hard to answer basic questions: Did a camera freeze? Did streams drift out of sync? Which version of a check ran? Can we reproduce the dataset we used?
HFlow turns that work into a pipeline. Transforms, checks, labels, and enrichments stay yours. HFlow handles orchestration, storage, versioning, and curation around them. It writes canonical MCAP episodes with provenance, records quality evidence in a Parquet catalog, and builds version-pinned manifests with DuckDB.
The core lifecycle works end to end today. You can run the included quickstart locally
without Docker, Airflow, robot hardware, or an external service.
HFlow is Apache-2.0 licensed and built in public by Hebbian Robotics (YC S26). If you work with robotics or physical AI data, try the quickstart and tell us where your current data workflow loses the most context. Is it ingestion, quality control, provenance, or curation?
Kingston and Brandon
HFlow
@kstonekuan better data smarter robots!