Consult this skill when designing or debugging distributed data processing pipelines on Ray Data. Trigger when choosing between Ray Data vs Ray Core actors vs Spark, tuning CPU/GPU resource ratios, handling fault tolerance for native code (C++/CUDA), sizing blocks and partitions, integrating with storage layers (Parquet/Lance/Iceberg/S3), or diagnosing back-pressure and spilling issues. NOT for Ray Serve inference endpoints, Ray Tune hyperparameter search, or general distributed systems design.