Ray LLM
Library-reference skill for LLM workloads on open-source Ray — 13 rules across 5 categories covering ray.serve.llm (OpenAI-compatible, vLLM-backed serving) and ray.data.llm (batch inference). This surface churned faster than any other part of Ray — a standalone repo was absorbed and archived, entry points were renamed, and config shapes restructured — so the examples a model learned from mostly no longer run. Each rule names the wrong default it corrects; there is no rule for things a capable model already gets right.
Scope is the LLM-specific layer. Generic Serve/Data/cluster decisions (deployment lifecycle, autoscaling semantics, KubeRay) are the sibling ray skill — the two compose.
Pinned to ray 2.57.0 (ray[llm] extra, which pins its matching vLLM). API claims were verified against the unpacked 2.57.0 wheel and the installed package source, and every config example in the rules was constructed under CPU-only pydantic validation (including the traps, which fail exactly as described); engine/GPU runtime behavior is source-verified only — no model was actually served.
When to Apply
- Standing up or reviewing an OpenAI-compatible LLM serving deployment on Ray
- Writing batch LLM inference over datasets — summarization, embedding, scoring at scale
- Sizing or placing multi-GPU models — tensor/pipeline parallelism, accelerator selection
- Scaling LLM deployments — replica autoscaling, ingress sizing, request routing
- Serving families of LoRA fine-tunes of a shared base model
- Migrating code that uses the archived ray-llm repo, hand-rolled vLLM engines, or pre-2.5x
ray.data.llmnames
Rule Categories
| # | Category | Prefix | Covers |
|---|---|---|---|
| 1 | Serving Setup | serve- | LLMConfig + build_openai_app over hand-rolled engines and the archived repo; model_id vs model_source; the ray[llm]↔vLLM version pin; relocated LLMServer/OpenAiIngress imports |
| 2 | Batch Inference | batch- | build_processor (old name removed), stage configs over boolean flags, CPU-default accelerator_type and autoscaling concurrency, HTTP/Serve processor alternatives |
| 3 | Placement & Accelerators | place- | The engine's own TP×PP placement group (and when to override its strategy), validated accelerator_type names |
| 4 | Autoscaling & Routing | scale- | deployment_config.autoscaling_config with engine-sized replicas, ingress replica sizing, prefix-cache-affinity routing |
| 5 | LoRA & API Surface | api- | Dynamic LoRA multiplexing over per-adapter deployments; the full OpenAI endpoint surface; GPU-free config validation |
Quick Reference
1. Serving Setup
serve-builtin-not-handrolled—LLMConfig+build_openai_app; the ray-llm repo is archived, hand-rolled engines re-implement lessserve-model-id-vs-source—model_idis the client-facing name;model_sourceis where weights liveserve-ray-llm-extra-pins-vllm—ray[llm]pins its exact vLLM; don't mix independently chosen versionsserve-deprecated-server-router—LLMServer/LLMRoutermoved; the ingress class isOpenAiIngress(exact casing)
2. Batch Inference
batch-build-processor-renamed—build_llm_processoris removed;build_processor+vLLMEngineProcessorConfigbatch-stage-configs-not-flags— booleanapply_chat_template/tokenize/detokenizegave way to stage configsbatch-concurrency-and-alternatives—accelerator_type=Nonemeans CPU;(min, max)concurrency; HTTP/Serve processors
3. Placement & Accelerators
place-engine-builds-pg— the engine builds its TP×PP placement group; override strategy, don't hand-rollplace-accelerator-type-validated— canonical accelerator constants;"A10"normalizes, most typos raise
4. Autoscaling & Routing
scale-autoscaling-deployment-config— standard Serve autoscaling nested indeployment_config; replicas are whole enginesscale-prefix-cache-routing—PrefixCacheAffinityRouterkeeps same-prefix requests on warm KV caches
5. LoRA & API Surface
api-lora-dynamic-multiplexing— dynamic adapter loading from cloud storage, not one deployment per adapterapi-endpoint-surface-cpu-validation— embeddings/transcription/score/tokenize are served too; configs validate without GPUs
How to Use
Read a reference file when its decision comes up. Each rule names the wrong default it corrects, then shows the canonical way (with an incorrect/correct contrast only where the wrong way is a real trap).
- Section definitions — category structure
- Rule template — for adding new rules
- AGENTS.md — auto-built table of contents across all rules
Related Skills
ray— the sibling rule pack for classic-ML Ray (Train, Tune, Data, Serve, Core, KubeRay production topology); generic Serve and cluster decisions live there
Reference Files
| File | Description |
|---|---|
| references/_sections.md | Category definitions and ordering |
| assets/templates/_template.md | Template for new rules |
| metadata.json | Version and source references |