vastai-core-workflow-b

v2026.09.24

Build and roll out a Vast.ai Serverless endpoint with measured autoscaling, a canary template, and a zero-downtime worker update. Use when production inference capacity or model bytes change. Trigger with: "deploy Vast.ai Serverless", "tune a Vast.ai worker group", "roll a model without downtime".

GitHub
安装命令
npx skhub add jeremylongshore/vastai-core-workflow-b
Markdown
SKILL.md

Vast.ai Serverless Endpoint Rollout

Overview

Replace ad hoc multi-instance orchestration with the provider Serverless control plane. Prove a template independently, establish worker and queue bounds from load evidence, then let the workergroup perform a graceful rolling update.

Prerequisites

  • Latency, error-rate, queue-time, concurrency, and cost objectives
  • Immutable model/template candidate and a separate canary endpoint
  • Initial, minimum, maximum, cold-worker, and inactivity policy

Instructions

Step 1: Prove the candidate template

Launch the new model or environment on a non-production endpoint and verify load, readiness, response schema, and representative outputs.

Step 2: Define scaling bounds

Set min_load, min_workers, max_workers, cold_workers, inactivity_timeout, target_queue_time, and max_queue_time from explicit SLO and budget assumptions.

Step 3: Exercise convergence

During initial rollout, drive representative load up to roughly twice expected capacity and back down three times so the engine can learn GPU cost/performance.

Step 4: Establish the pre-update baseline

Record endpoint latency, queue time, error rate, active/inactive workers, model identity, and spend before changing production.

Step 5: Trigger the rolling update

Save the new template, update the workergroup reference, and monitor inactive workers updating first while active workers drain in-flight requests.

Step 6: Accept or roll back

Verify every worker is on the candidate and compare SLOs. If it fails, point the workergroup back to the last verified template and observe the reverse rollout.

Authentication

Use a scoped key with the documented misc Serverless permissions and no billing-write authority. Keep model registry credentials in approved environment variables, separate from the Vast.ai key.

Tool Discipline

Use Read and Grep to inspect manifests, configuration, provider output, and existing tests before proposing a mutation. Use Write or Edit only for the approved plan, implementation, test, or redacted receipt; do not create, update, destroy, or fund Vast.ai resources without explicit operator approval.

Output

  • Canary and production template identities
  • Scaling policy plus load-test and rollout timeline
  • SLO comparison, worker convergence, and rollback decision

Return endpoint/workergroup IDs, old and new template identities, scaling bounds, load profile, SLO delta, and final rollout state.

Examples

A vLLM endpoint validates a new model on a canary, applies bounded queue targets, then updates its workergroup; active requests drain while new requests move to updated workers, with the old template retained for rollback.

Error Handling

FailureResponse
Canary cannot load the modelDo not update production; fix image, model, or environment configuration.
Queue time breaches during rolloutPause acceptance, increase safe capacity within budget, or roll back the template.
Workers do not convergeInspect workergroup logs and configuration; do not claim zero-downtime completion.
New output contract regressesRoll back to the last verified template and preserve comparison evidence.

Resources

发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

Sep 24, 2026

分类

未分类

许可证

MIT

源路径

skills/.curated/vastai-core-workflow-b

默认分支

main

最新提交

e5a6c3b

Tree SHA

c2dc8e8