ray-llm

v2026.09.24

LLM workloads on open-source Ray (pinned to 2.57) — OpenAI-compatible serving with ray.serve.llm (vLLM-backed LLMConfig + build_openai_app) and batch inference with ray.data.llm (build_processor). Corrects the stale defaults a model produces — the archived ray-llm repo and its YAML configs, hand-rolled vLLM engines inside plain Serve deployments, the removed build_llm_processor name, deprecated boolean stage flags, top-level LLMServer/LLMRouter imports, free-form accelerator strings, one-deployment-per-LoRA-adapter designs — with the 2.57 idioms (stage configs, placement_group_config over hand-rolled PGs, deployment_config autoscaling, dynamic LoRA multiplexing, prefix-cache-affinity routing, the full OpenAI endpoint surface). Use when writing, reviewing, or productionizing LLM serving or batch inference on Ray. Classic-ML Ray (Train/Tune/Data/Serve/clusters) lives in the sibling ray skill.

GitHub
Install command
npx skhub add pproenca/ray-llm
Markdown
SKILL.md

Ray LLM

Library-reference skill for LLM workloads on open-source Ray — 13 rules across 5 categories covering ray.serve.llm (OpenAI-compatible, vLLM-backed serving) and ray.data.llm (batch inference). This surface churned faster than any other part of Ray — a standalone repo was absorbed and archived, entry points were renamed, and config shapes restructured — so the examples a model learned from mostly no longer run. Each rule names the wrong default it corrects; there is no rule for things a capable model already gets right.

Scope is the LLM-specific layer. Generic Serve/Data/cluster decisions (deployment lifecycle, autoscaling semantics, KubeRay) are the sibling ray skill — the two compose.

Pinned to ray 2.57.0 (ray[llm] extra, which pins its matching vLLM). API claims were verified against the unpacked 2.57.0 wheel and the installed package source, and every config example in the rules was constructed under CPU-only pydantic validation (including the traps, which fail exactly as described); engine/GPU runtime behavior is source-verified only — no model was actually served.

When to Apply

  • Standing up or reviewing an OpenAI-compatible LLM serving deployment on Ray
  • Writing batch LLM inference over datasets — summarization, embedding, scoring at scale
  • Sizing or placing multi-GPU models — tensor/pipeline parallelism, accelerator selection
  • Scaling LLM deployments — replica autoscaling, ingress sizing, request routing
  • Serving families of LoRA fine-tunes of a shared base model
  • Migrating code that uses the archived ray-llm repo, hand-rolled vLLM engines, or pre-2.5x ray.data.llm names

Rule Categories

#CategoryPrefixCovers
1Serving Setupserve-LLMConfig + build_openai_app over hand-rolled engines and the archived repo; model_id vs model_source; the ray[llm]↔vLLM version pin; relocated LLMServer/OpenAiIngress imports
2Batch Inferencebatch-build_processor (old name removed), stage configs over boolean flags, CPU-default accelerator_type and autoscaling concurrency, HTTP/Serve processor alternatives
3Placement & Acceleratorsplace-The engine's own TP×PP placement group (and when to override its strategy), validated accelerator_type names
4Autoscaling & Routingscale-deployment_config.autoscaling_config with engine-sized replicas, ingress replica sizing, prefix-cache-affinity routing
5LoRA & API Surfaceapi-Dynamic LoRA multiplexing over per-adapter deployments; the full OpenAI endpoint surface; GPU-free config validation

Quick Reference

1. Serving Setup

2. Batch Inference

3. Placement & Accelerators

4. Autoscaling & Routing

5. LoRA & API Surface

How to Use

Read a reference file when its decision comes up. Each rule names the wrong default it corrects, then shows the canonical way (with an incorrect/correct contrast only where the wrong way is a real trap).

Related Skills

  • ray — the sibling rule pack for classic-ML Ray (Train, Tune, Data, Serve, Core, KubeRay production topology); generic Serve and cluster decisions live there

Reference Files

FileDescription
references/_sections.mdCategory definitions and ordering
assets/templates/_template.mdTemplate for new rules
metadata.jsonVersion and source references
Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.24

Published

Sep 24, 2026

Category

Uncategorized

License

MIT

Source path

skills/.experimental/ray-llm

Default branch

master

Latest commit

cf93c57

Tree SHA

afbb575