Ray
Library-reference skill for production, open-source Ray — 26 rules across 6 categories covering the path from training to serving. Ray's API surface churned hard through the 2.x line (Train V2 became the default, Serve removed parameters and handle classes outright, Ray Data reversed a deprecation), so a model fluent in the older corpus produces code that warns, errors, or silently means something else. Each rule names the wrong default it corrects; there is no rule for things a capable model already gets right.
Scope is classic-ML Ray on self-hosted/KubeRay clusters. LLM serving and batch inference (ray.serve.llm, ray.data.llm) are the sibling ray-llm skill.
Pinned to ray 2.57.0 (Python ≥ 3.10). API claims were verified against the unpacked 2.57.0 wheel; classic-ML examples were exercised on a live local Ray 2.57.0 runtime.
When to Apply
- Writing or reviewing distributed training code —
TorchTrainer,ScalingConfig, checkpointing, fault tolerance - Building data pipelines with Ray Data — reads,
map_batches, GPU inference pools, training ingest - Running hyperparameter sweeps with Ray Tune, especially combined with Ray Train
- Writing or reviewing Ray Serve deployments — scaling, handles, composition, production config
- Using Ray Core primitives directly — tasks, actors, object store, retries
- Standing up or reviewing production Ray infrastructure — KubeRay CRDs, job submission, fault tolerance, observability
Rule Categories
| # | Category | Prefix | Covers |
|---|---|---|---|
| 1 | Ray Train | train- | Train V2 as the default (deprecated config fields), ray.train.report over ray.air session, config imports and elastic scaling, the prepare_model/prepare_data_loader wrappers |
| 2 | Ray Serve | serve- | max_ongoing_requests (old name removed), current autoscaling fields, DeploymentResponse handles, serve.run/build/deploy lifecycle, replica placement options |
| 3 | Ray Data | data- | compute= strategies (the concurrency reversal), override_num_blocks, streaming execution replacing pipelines, torch ingest, zero-copy read-only batches |
| 4 | Ray Core | core- | The classic anti-patterns' non-obvious residue, retry/restart defaults, object store & /dev/shm sizing, ray.util.state |
| 5 | Ray Tune | tune- | Tuner as canonical (with tune.run's true status), ray.tune.RunConfig imports, the driver-function Train integration |
| 6 | Production & Clusters | prod- | KubeRay CRD choice, Jobs API submission, GCS fault tolerance with Redis, baked images vs runtime_env, Prometheus/Grafana wiring |
Quick Reference
1. Ray Train
train-v2-default-changed-semantics— V2 is the default since 2.51;sync_config/verbose/fail_fastare gonetrain-report-not-air-session—ray.train.report(metrics, checkpoint=);ray.airis a legacy shimtrain-configs-from-ray-train-elastic— V2 config imports, elasticnum_workers=(min, max), case-sensitive resource keystrain-prepare-model-and-loader— withoutprepare_model/prepare_data_loader, N workers train N unsynchronized copies
2. Ray Serve
serve-max-ongoing-requests—max_concurrent_queriesis removed;max_ongoing_requests(default 5) +max_queued_requestsserve-autoscaling-current-fields—target_ongoing_requests,*_factorfields,num_replicas="auto", scale-to-zeroserve-deployment-response-handles—DeploymentResponse.result()/await, neverray.get; composition via.bind()serve-run-build-deploy—serve.run+serve build/deploy;Deployment.deploy()era is removedserve-placement-controls—max_replicas_per_node, per-replica placement groups, gang scheduling, request routers
3. Ray Data
data-compute-not-concurrency—concurrency=is deprecated (again);compute=ActorPoolStrategy/TaskPoolStrategydata-override-num-blocks—parallelismis deprecated in read APIsdata-streaming-replaced-pipelines—DatasetPipeline/window/repeatremoved; execution streams on consumptiondata-iter-torch-batches-not-to-torch—to_torchis gone;iter_torch_batchesand Train dataset shardsdata-zero-copy-read-only-batches— batches are read-only views by default; set an explicitbatch_size
4. Ray Core
core-classic-traps-residue— the non-obvious residue of the classic anti-patterns:ray.waitdraining, closure/global copies, few-ms task floorcore-retries-system-failures-only—retry_exceptionsis opt-in; actors don't restart by defaultcore-object-store-sizing— 30% default,/dev/shmin containers, spilling to local diskcore-state-api-not-ray-state—ray.stateis gone;ray.util.stateis the introspection API
5. Ray Tune
tune-tuner-canonical—Tuneris the API;tune.runis legacy but not removedtune-runconfig-from-ray-tune— Tuner takesray.tune.RunConfig, notray.train's orray.air'stune-driver-function-not-trainer— Trainer-inside-Tuner is deprecated; use a driver function +with_resources
6. Production & Clusters
prod-kuberay-crd-choice— RayCluster vs RayJob (shutdownAfterJobFinishesdefaults false) vs RayServiceprod-jobs-api-not-head-driver—ray job submitfor long-lived clusters; Ray Client is a dev toolprod-gcs-ft-redis— head-crash survival for RayService needs external Redis GCSprod-baked-images-not-runtime-env— images carry prod dependencies;runtime_env(now incl.uv,image_uri) is for iterationprod-metrics-wiring— the dashboard needs external Prometheus/Grafana to show time series
How to Use
Read a reference file when its decision comes up. Each rule names the wrong default it corrects, then shows the canonical way (with an incorrect/correct contrast only where the wrong way is a real trap).
- Section definitions — category structure
- Rule template — for adding new rules
- AGENTS.md — auto-built table of contents across all rules
Related Skills
ray-llm— the sibling rule pack for LLM serving (ray.serve.llm) and batch inference (ray.data.llm) on Raymlflow-3— experiment tracking and model registry; pairs with Ray Train for the tracking side of the MLOps cycle
Reference Files
| File | Description |
|---|---|
| references/_sections.md | Category definitions and ordering |
| assets/templates/_template.md | Template for new rules |
| metadata.json | Version and source references |