RayOrch keeps track of source records while batching preprocessing across GPUs
single source· 1 articles · confidence: medium · first seen 2026-09-15 20:00 UTC
What this means for you
Worth a look only if you run large document or video preprocessing jobs across many GPUs. The code is public on GitHub, but every figure comes from one paper, measured on Nvidia H20s against Ray Data and Daft on three named pipelines. Nothing else was tested.
RayOrch, a programming model and execution engine for preparing foundation-model training data, keeps track of which parent document or video each output record came from while spreading that work across GPUs. The paper, posted to arXiv on 15 September, reports 15.14× speedup when scaling the MinerU document pipeline from 4 to 64 Nvidia H20 GPUs, and 7.82× when scaling a video pipeline from 8 to 64. End-to-end time falls 13.1% against Ray Data and 29.0% against Daft on MinerU, and 16.0% against Ray Data on Docling. Code is on GitHub.
Key facts
- ·RayOrch is a programming model and distributed execution engine that preserves parent-child relationships throughout a data pipeline's execution source
- ·RayOrch reports 15.14× speedup scaling the MinerU pipeline from 4 to 64 Nvidia H20 GPUs, and 7.82× scaling a video pipeline from 8 to 64 source
- ·End-to-end time is reduced 13.1% versus Ray Data and 29.0% versus Daft on MinerU, and 16.0% versus Ray Data on Docling source
- ·The runtime records child membership, immediate parents, immutable ordinals and terminal states, and gathers results by declared membership and ordinals rather than batch boundaries source
- ·Code is available at https://github.com/OpenDCAI/RayOrch source
- ·The paper was posted to arXiv on 15 September 2026 source
What the sources say
- Hugging Face Daily Papers (research) — Describes a compiler and runtime that track which source document each output record came from.
Sources
The original reporting. Follow these — they did the work.
- Hugging Face Daily PapersRayOrch: Programming and Executing Lineage-Controlled Multi-Grain Dataflows for Foundation-Model Data Preparation2026-09-15