RayOrch keeps track of source records while batching preprocessing across GPUs

single source· 1 articles · confidence: medium · first seen 2026-09-15 20:00 UTC

What this means for you

Worth a look only if you run large document or video preprocessing jobs across many GPUs. The code is public on GitHub, but every figure comes from one paper, measured on Nvidia H20s against Ray Data and Daft on three named pipelines. Nothing else was tested.

RayOrch, a programming model and execution engine for preparing foundation-model training data, keeps track of which parent document or video each output record came from while spreading that work across GPUs. The paper, posted to arXiv on 15 September, reports 15.14× speedup when scaling the MinerU document pipeline from 4 to 64 Nvidia H20 GPUs, and 7.82× when scaling a video pipeline from 8 to 64. End-to-end time falls 13.1% against Ray Data and 29.0% against Daft on MinerU, and 16.0% against Ray Data on Docling. Code is on GitHub.

Key facts

  • ·RayOrch is a programming model and distributed execution engine that preserves parent-child relationships throughout a data pipeline's execution source
  • ·RayOrch reports 15.14× speedup scaling the MinerU pipeline from 4 to 64 Nvidia H20 GPUs, and 7.82× scaling a video pipeline from 8 to 64 source
  • ·End-to-end time is reduced 13.1% versus Ray Data and 29.0% versus Daft on MinerU, and 16.0% versus Ray Data on Docling source
  • ·The runtime records child membership, immediate parents, immutable ordinals and terminal states, and gathers results by declared membership and ordinals rather than batch boundaries source
  • ·Code is available at https://github.com/OpenDCAI/RayOrch source
  • ·The paper was posted to arXiv on 15 September 2026 source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire