We've been dealing with exactly the mess they describe - a robotics data corpus where nobody could say for sure which version of a labeling script produced which episode. Ran the quickstart locally with no Docker or Airflow setup required, which was a relief since half the tools in this space assume you already have an orchestration stack running. Writing canonical MCAP episodes with a DuckDB-queryable manifest is the right call - being able to just query your dataset lineage with SQL instead of grepping through log files is a genuinely useful change in workflow.