openrouter-performance-tuning

v2026.09.24

Optimize OpenRouter request latency and throughput. Use when building real-time applications, reducing TTFT, or scaling request volume. Triggers: 'openrouter performance', 'openrouter latency', 'openrouter speed', 'optimize openrouter throughput'.

GitHub
安装命令
npx skhub add jeremylongshore/openrouter-performance-tuning
Markdown
SKILL.md

OpenRouter Performance Tuning

Overview

OpenRouter adds minimal overhead (~50-100ms) to direct provider calls. Most latency comes from the upstream model. Key levers: model selection (smaller = faster), streaming (lower TTFT), parallel requests, prompt size reduction, and provider routing to faster infrastructure. This skill covers benchmarking, streaming optimization, concurrent processing, and connection tuning.

Prerequisites

  • An OpenRouter API key (sk-or-v1-...) exported as OPENROUTER_API_KEY — see the openrouter-install-auth skill for setup
  • Python 3.8+ with the OpenAI SDK (openai package) — the examples use both the sync OpenAI client and AsyncOpenAI for parallel processing
  • Credits on the key if you benchmark paid models like anthropic/claude-3.5-sonnet; a :free model is enough to validate the benchmark harness itself
  • HTTP-Referer / X-Title header values for your app (set in every client constructor here)

Instructions

  1. Establish a baseline: run benchmark_model() from Benchmark Latency against your candidate models (e.g. openai/gpt-4o-mini vs anthropic/claude-3.5-sonnet) and record p50/p95.
  2. Check the results against the Model Speed Tiers table to confirm each candidate sits in the right tier for your latency budget (200-500ms TTFT fastest tier; 5-30s for reasoning models).
  3. Switch user-facing paths to stream_completion() per Streaming for Lower TTFT and verify ttft_ms drops (typically 2-10x).
  4. Move batch workloads to parallel_completions() per Parallel Request Processing, capping concurrency with asyncio.Semaphore (max_concurrent=5-10).
  5. Apply Connection Optimization — one shared client with timeout=30.0 and max_retries=2 instead of a new client per request.
  6. Work through the Performance Optimization Checklist (set max_tokens, shrink prompts, consider :nitro variants and provider routing), then re-run the benchmark to quantify each change.

Benchmark Latency

import os, time, statistics
from openai import OpenAI

client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key=os.environ["OPENROUTER_API_KEY"],
    default_headers={"HTTP-Referer": "https://my-app.com", "X-Title": "my-app"},
)

def benchmark_model(model: str, prompt: str = "Say hello", n: int = 5) -> dict:
    """Benchmark a model's latency over N requests."""
    latencies = []
    for _ in range(n):
        start = time.monotonic()
        response = client.chat.completions.create(
            model=model,
            messages=[{"role": "user", "content": prompt}],
            max_tokens=50,
        )
        latencies.append((time.monotonic() - start) * 1000)

    return {
        "model": model,
        "p50_ms": round(statistics.median(latencies)),
        "p95_ms": round(sorted(latencies)[int(len(latencies) * 0.95)]),
        "avg_ms": round(statistics.mean(latencies)),
        "min_ms": round(min(latencies)),
        "max_ms": round(max(latencies)),
    }

# Compare fast vs slow models
for model in ["openai/gpt-4o-mini", "anthropic/claude-3-haiku", "anthropic/claude-3.5-sonnet"]:
    result = benchmark_model(model)
    print(f"{result['model']}: p50={result['p50_ms']}ms p95={result['p95_ms']}ms")

Streaming for Lower TTFT

def stream_completion(messages, model="openai/gpt-4o-mini", **kwargs):
    """Stream response for lower time-to-first-token."""
    start = time.monotonic()
    first_token_time = None
    full_content = []

    stream = client.chat.completions.create(
        model=model, messages=messages, stream=True,
        stream_options={"include_usage": True},  # Get token counts at end
        **kwargs,
    )

    for chunk in stream:
        if chunk.choices and chunk.choices[0].delta.content:
            if first_token_time is None:
                first_token_time = (time.monotonic() - start) * 1000
            full_content.append(chunk.choices[0].delta.content)

    total_time = (time.monotonic() - start) * 1000
    return {
        "content": "".join(full_content),
        "ttft_ms": round(first_token_time or 0),
        "total_ms": round(total_time),
    }

Parallel Request Processing

import asyncio
from openai import AsyncOpenAI

async def parallel_completions(prompts: list[str], model="openai/gpt-4o-mini",
                                max_concurrent=10, **kwargs):
    """Process multiple prompts concurrently."""
    semaphore = asyncio.Semaphore(max_concurrent)
    client = AsyncOpenAI(
        base_url="https://openrouter.ai/api/v1",
        api_key=os.environ["OPENROUTER_API_KEY"],
        default_headers={"HTTP-Referer": "https://my-app.com", "X-Title": "my-app"},
    )

    async def process(prompt):
        async with semaphore:
            response = await client.chat.completions.create(
                model=model,
                messages=[{"role": "user", "content": prompt}],
                **kwargs,
            )
            return response.choices[0].message.content

    return await asyncio.gather(*[process(p) for p in prompts])

# 10 requests in parallel instead of sequential
results = asyncio.run(parallel_completions(
    ["Summarize: " + text for text in documents],
    max_concurrent=5,
    max_tokens=200,
))

Performance Optimization Checklist

OptimizationImpactEffort
Use streamingTTFT drops 2-10xLow
Use smaller models for simple tasks2-5x fasterLow
Reduce prompt sizeProportional to reductionMedium
Set max_tokensCaps response timeLow
Parallel requestsN requests in ~1 request timeMedium
Use :nitro variantFaster inference (where available)Low
Provider routing to fastest10-30% latency reductionLow
Connection keep-aliveSaves TCP/TLS handshakeLow

Model Speed Tiers

SpeedModelsTypical TTFT
Fastestopenai/gpt-4o-mini, anthropic/claude-3-haiku200-500ms
Fastopenai/gpt-4o, google/gemini-2.0-flash-001500ms-1s
Standardanthropic/claude-3.5-sonnet1-3s
Slowopenai/o1, reasoning models5-30s

Connection Optimization

# Reuse client instance (connection pooling)
# BAD: creating new client per request
for prompt in prompts:
    c = OpenAI(base_url="https://openrouter.ai/api/v1", ...)  # New TCP connection each time
    c.chat.completions.create(...)

# GOOD: reuse single client
client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key=os.environ["OPENROUTER_API_KEY"],
    timeout=30.0,           # Set appropriate timeout
    max_retries=2,          # Built-in retry with backoff
    default_headers={"HTTP-Referer": "https://my-app.com", "X-Title": "my-app"},
)
for prompt in prompts:
    client.chat.completions.create(...)  # Reuses HTTP connection

Output

  • A latency benchmark table per model from benchmark_model(): p50_ms, p95_ms, avg_ms, min_ms, max_ms over N sample requests
  • Streaming metrics from stream_completion(): the full content plus ttft_ms and total_ms for each request
  • A list of completions from parallel_completions() produced in roughly one request's wall-clock time instead of N sequential round-trips
  • A prioritized tuning plan drawn from the Performance Optimization Checklist (lever, expected impact, effort)

Examples

Benchmark two fastest-tier candidates before committing to one:

for model in ["openai/gpt-4o-mini", "anthropic/claude-3-haiku"]:
    r = benchmark_model(model, n=5)
    print(f"{r['model']}: p50={r['p50_ms']}ms p95={r['p95_ms']}ms avg={r['avg_ms']}ms")
# openai/gpt-4o-mini: p50=430ms p95=610ms avg=455ms
# anthropic/claude-3-haiku: p50=395ms p95=580ms avg=418ms

Both land in the fastest tier (200-500ms typical TTFT), so choose on cost or quality — then stream_completion() cuts perceived latency further for user-facing paths. More worked examples: references/examples.md.

Error Handling

ErrorCauseFix
High TTFT (>5s)Model cold-starting or overloadedSwitch to :nitro variant or different provider
Timeout errorsmax_tokens too high or model too slowReduce max_tokens; use streaming; increase timeout
Throughput bottleneckSequential processingUse async + semaphore for concurrent requests
Inconsistent latencyProvider load variesUse provider.order to pin to fastest provider

Enterprise Considerations

  • Benchmark models in your infrastructure, not just locally -- network path matters
  • Use streaming for all user-facing requests to minimize perceived latency
  • Set max_tokens on every request to bound response time and cost
  • Reuse client instances to benefit from HTTP connection pooling
  • Use asyncio.Semaphore to control concurrency and avoid overwhelming the API
  • Monitor P95 latency, not just average -- tail latencies indicate provider issues
  • Consider :nitro model variants for latency-critical paths

References

发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

2026年9月24日

分类

未分类

许可证

MIT

源路径

skills/.curated/openrouter-performance-tuning

默认分支

main

最新提交

e5a6c3b

Tree SHA

c2dc8e8