gin

Gin / Docs

Gin docs

Gin is a fetch wrapper for the AI clients you already have. Your calls go where they went before. After each response, Gin records what it cost.

Wrap one call

Start from a plain OpenRouter call with the OpenAI SDK. The only change is one line: pass Gin as the client's fetch.

app.ts
import OpenAI from 'openai';
import { gin } from '@winding-labs/gin';

const openrouter = new OpenAI({
  baseURL: 'https://openrouter.ai/api/v1',
  apiKey: process.env.OPENROUTER_API_KEY,
  fetch: gin({ useCase: 'review' }),   // + the only new line
});

const res = await openrouter.chat.completions.create({
  model: 'deepseek/deepseek-v4.1-flash',
  messages,
  usage: { include: true },
});
  • Your call still goes straight to your router, with your key. Gin is not a proxy and never sees the key.
  • After the response, Gin reads a copy and posts one event to https://trygin.ai/api/v1/events in the background. It is never in the request path.
  • No GIN_KEY, or any Gin error, means plain fetch. If Gin is down, your app doesn't notice.

When the response comes back

  1. Your app gets the router's response as it arrives. Gin passes every chunk straight through and keeps a copy on the side. It never buffers a stream.
  2. When the body ends, Gin reads the copy. JSON is parsed as is. A stream's data: chunks are folded into one response: text, tool calls, the last usage, any error.
  3. It pulls out the fields that matter: model, provider, tokens in/out/cached, cost, finish reason, error. Each router puts them somewhere else; the tabs below show exactly where.
  4. It adds what it measured itself: status, time to headers, time to first token, total duration, and why a call failed (timeout, cancelled, network, rate limit, provider error).
  5. It posts one event to POST /api/v1/events after your response is done, and that's what the daily email is built from.

Every router

The wrap is the same everywhere. What changes is where Gin finds the provider, the cost and the cached tokens in each response.

OpenRouter

Base URL https://openrouter.ai/api/v1, OpenAI chat format.

1Your app calls the router through Gin's fetch. Same request, your key, no proxy.

OpenAI SDK
const openrouter = new OpenAI({
  baseURL: 'https://openrouter.ai/api/v1',
  apiKey: process.env.OPENROUTER_API_KEY,
  fetch: gin({ useCase: 'review' }),
});

await openrouter.chat.completions.create({
  model: 'deepseek/deepseek-v4.1-flash',
  messages,
  usage: { include: true },   // exact cost in usage.cost
});

2The router returns. Your app gets this response untouched, chunk by chunk. Gin keeps a copy.

HTTP 200 · application/json
{
  "id": "gen-1760051234-a8Qx",
  "provider": "Together",
  "model": "deepseek/deepseek-v4.1-flash",
  "choices": [{ "message": { "role": "assistant", "content": "…" }, "finish_reason": "stop" }],
  "usage": {
    "prompt_tokens": 52310,
    "completion_tokens": 1840,
    "prompt_tokens_details": { "cached_tokens": 49152 },
    "cost": 0.00312
  }
}

3Gin posts one event to POST /api/v1/events in the background, after the body ends. Same numbers, same fields.

event
{
  "ts": "2026-10-09T17:04:11.020Z",
  "url": "https://openrouter.ai/api/v1/chat/completions",
  "use_case": "review",
  "model": "deepseek/deepseek-v4.1-flash",
  "served_model": "deepseek/deepseek-v4.1-flash",
  "provider": "Together",
  "status": 200,
  "tokens_in": 52310, "tokens_out": 1840,
  "tokens_cached": 49152,
  "cost_usd": 0.00312,
  "finish_reason": "stop",
  "latency_ms": 455, "duration_ms": 7660   // measured by Gin
}
  1. 1provider: who served it
  2. 2model: what answered
  3. 3usage: tokens in and out
  4. 4cached_tokens: the caching angle
  5. 5usage.cost: exact, because the request asked for usage: { include: true }
  6. 6finish_reason

!If it fails mid-stream. OpenRouter can send an error chunk inside an HTTP 200. Gin still records it as an error.

last SSE chunk → event
data: { "id": "gen-…", "error": { "code": 502, "message": "Network connection lost." }, "choices": [{ "delta": {}, "finish_reason": "error" }] }

{ "status": 200, "error": "502 Network connection lost.", "error_type": "stream", "ttft_ms": 500 }

Router.com

Base URL https://api.router.com/v1. OpenAI Responses API only: POST /v1/responses.

1Your app calls the router through Gin's fetch. Same request, your key, no proxy.

OpenAI SDK
const router = new OpenAI({
  baseURL: 'https://api.router.com/v1',
  apiKey: process.env.ROUTER_API_KEY,
  fetch: gin({ useCase: 'summaries' }),
});

await router.responses.create({ model, input });

2The router returns. Your app gets this response untouched, chunk by chunk. Gin keeps a copy.

HTTP 200 · application/json (Responses API)
{
  "id": "resp_68e7c2",
  "object": "response",
  "model": "gpt-6-astra",
  "status": "completed",
  "output": [{ "type": "message", "content": [{ "type": "output_text", "text": "…" }] }],
  "usage": {
    "input_tokens": 1200,
    "input_tokens_details": { "cached_tokens": 1024 },
    "output_tokens": 85
  }
}

3Gin posts one event to POST /api/v1/events in the background, after the body ends. Same numbers, same fields.

event
{
  "url": "https://api.router.com/v1/responses",
  "use_case": "summarize",
  "served_model": "gpt-6-astra",
  "status": 200,
  "tokens_in": 1200, "tokens_out": 85,
  "tokens_cached": 1024,
  "finish_reason": "completed",
  "latency_ms": 610, "duration_ms": 1480
}
// no cost in the body: Gin prices the tokens from list prices
// no provider in the body: spend is grouped under Router.com
  1. 2model
  2. 3input_tokens / output_tokens (cached tokens are part of input)
  3. 4input_tokens_details.cached_tokens
  4. 6status becomes the finish reason

Vercel AI Gateway

Base URL https://ai-gateway.vercel.sh/v1 for OpenAI chat and Responses; Anthropic Messages at https://ai-gateway.vercel.sh.

1Your app calls the router through Gin's fetch. Same request, your key, no proxy.

OpenAI SDK
const gateway = new OpenAI({
  baseURL: 'https://ai-gateway.vercel.sh/v1',
  apiKey: process.env.AI_GATEWAY_API_KEY,
  fetch: gin({ useCase: 'chat' }),
});

await gateway.chat.completions.create({
  model: 'anthropic/claude-sonnet-5',
  messages,
});

2The router returns. Your app gets this response untouched, chunk by chunk. Gin keeps a copy.

HTTP 200 · text/event-stream (last two chunks)
data: { "id": "chatcmpl-9x", "model": "anthropic/claude-sonnet-5", "choices": [{ "delta": { "content": "…" } }] }

data: { "id": "chatcmpl-9x", "choices": [{ "delta": { "provider_metadata": { "gateway": {
          "routing": { "resolvedProvider": "anthropic", "finalProvider": "bedrock" },
          "cost": "0.0031" } } }, "finish_reason": "stop" }],
        "usage": { "prompt_tokens": 940, "completion_tokens": 210 } }

data: [DONE]

3Gin posts one event to POST /api/v1/events in the background, after the body ends. Same numbers, same fields.

event
{
  "url": "https://ai-gateway.vercel.sh/v1/chat/completions",
  "use_case": "chat",
  "served_model": "anthropic/claude-sonnet-5",
  "provider": "bedrock",
  "status": 200, "streamed": true,
  "tokens_in": 940, "tokens_out": 210,
  "cost_usd": 0.0031,
  "finish_reason": "stop",
  "latency_ms": 380, "ttft_ms": 512, "duration_ms": 2950
}
  1. 1finalProvider: who actually served it, after any fallback
  2. 2model
  3. 3usage on the last chunk
  4. 5gateway.cost: a decimal string, read as a number
  5. 6finish_reason

Streams: Gin folds every data: chunk into one response (text, tool calls, the last usage, any error) and reads it the same way as JSON. ttft_ms is when the first chunk arrived.

Cloudflare AI Gateway

URL https://gateway.ai.cloudflare.com/v1/<account>/<gateway>/<provider>/…. The body is the provider's own format.

1Your app calls the router through Gin's fetch. Same request, your key, no proxy.

OpenAI SDK, /compat/ endpoint
const cf = new OpenAI({
  baseURL: `https://gateway.ai.cloudflare.com/v1/${ACCOUNT_ID}/${GATEWAY_ID}/compat`,
  apiKey: process.env.GROQ_API_KEY,
  fetch: gin({ useCase: 'triage' }),
});

await cf.chat.completions.create({
  model: 'groq/llama-3.3-70b-versatile',
  messages,
});

2The router returns. Your app gets this response untouched, chunk by chunk. Gin keeps a copy.

HTTP 200 · application/json (Anthropic format, passed through)
{
  "id": "msg_01XF",
  "type": "message",
  "model": "claude-sonnet-5",
  "content": [{ "type": "text", "text": "…" }],
  "stop_reason": "end_turn",
  "usage": {
    "input_tokens": 24,
    "cache_read_input_tokens": 3100,
    "cache_creation_input_tokens": 0,
    "output_tokens": 340
  }
}

3Gin posts one event to POST /api/v1/events in the background, after the body ends. Same numbers, same fields.

event
{
  "url": "https://gateway.ai.cloudflare.com/v1/acct/gw/anthropic/v1/messages",
  "use_case": "triage",
  "served_model": "claude-sonnet-5",
  "provider": "anthropic",   // from the URL path
  "status": 200,
  "tokens_in": 3124,      // 24 + 3100 cached + 0 written
  "tokens_cached": 3100,
  "tokens_out": 340,
  "finish_reason": "end_turn",
  "latency_ms": 820, "duration_ms": 4100
}
  1. 1the provider is the path segment after your gateway name
  2. 2model
  3. 3input = new + cache read + cache write tokens
  4. 4cache_read_input_tokens
  5. 6stop_reason

Direct APIs

OpenAI, Anthropic, Gemini, Groq, Together, Fireworks, DeepSeek, xAI and Mistral. The provider is the host.

1Your app calls the router through Gin's fetch. Same request, your key, no proxy.

Anthropic SDK
import Anthropic from '@anthropic-ai/sdk';
import { gin } from '@winding-labs/gin';

const anthropic = new Anthropic({
  apiKey: process.env.ANTHROPIC_API_KEY,
  fetch: gin({ useCase: 'support-chat' }),
});

await anthropic.messages.create({
  model: 'claude-sonnet-5',
  max_tokens: 1024,
  messages,
});

2The router returns. Your app gets this response untouched, chunk by chunk. Gin keeps a copy.

HTTP 200 · application/json (Anthropic)
{
  "id": "msg_01AB",
  "model": "claude-sonnet-5",
  "stop_reason": "end_turn",
  "usage": {
    "input_tokens": 12,
    "cache_read_input_tokens": 1800,
    "output_tokens": 210
  }
}

3Gin posts one event to POST /api/v1/events in the background, after the body ends. Same numbers, same fields.

event
{
  "url": "https://api.anthropic.com/v1/messages",
  "use_case": "review",
  "served_model": "claude-sonnet-5",
  "provider": "anthropic",   // the host is the provider
  "status": 200,
  "tokens_in": 1812, "tokens_cached": 1800, "tokens_out": 210,
  "finish_reason": "end_turn",
  "latency_ms": 640, "duration_ms": 3900
}
// cost: Gin prices the tokens from list prices (cache reads at the cache rate)
  1. 1direct APIs: the host is the provider
  2. 2model
  3. 3tokens in and out
  4. 4cache reads
  5. 6stop_reason

Other clients and runtimes

  • Vercel AI SDK. createOpenAICompatible, createOpenRouter and createAnthropic take the same fetch: gin(…).
  • Cloudflare Workers. Pass key: env.GIN_KEY and waitUntil: ctx.waitUntil.bind(ctx) so the event post outlives the response.
  • One shared client. Tag each call with the header x-gin-use-case. Gin strips every x-gin-* header before the provider sees it.
Vercel AI SDK, Workers, one shared client
// Vercel AI SDK
const openrouter = createOpenRouter({
  apiKey: process.env.OPENROUTER_API_KEY,
  fetch: gin({ useCase: 'chat' }),
});

// Cloudflare Workers, inside fetch(request, env, ctx)
const client = new OpenAI({
  baseURL: 'https://openrouter.ai/api/v1',
  apiKey: env.OPENROUTER_API_KEY,
  fetch: gin({
    key: env.GIN_KEY,
    useCase: 'review',
    waitUntil: ctx.waitUntil.bind(ctx),
  }),
});

// One client, many features: tag per call
await client.chat.completions.create(
  { model, messages },
  { headers: { 'x-gin-use-case': 'summaries' } },
);

What Gin records per call

One event per call, posted to POST /api/v1/events.

Event fields
FieldWhat it holds
tsWhen the call started
urlThe router or API that was called
use_caseThe product feature, from useCase or x-gin-use-case
repo, envWhere the call came from
modelThe model you asked for
served_modelThe model that answered
providerWho served it
status, streamedHTTP status; whether the response streamed
tokens_in, tokens_out, tokens_cached, tokens_reasoningToken counts
cost_usdCost of the call
latency_msTime to response headers
ttft_msTime to the first streamed token
duration_msTime to the end of the response
finish_reasonWhy the model stopped
error, error_typeThe error, and one of timeout, aborted, network, rate_limit, server, client, stream
request, responseThe bodies, encrypted per workspace and deleted after 7 days. capture: 'metadata' never sends them.

Never captured: headers, API keys, cookies.

At high volume the SDK sends a weighted sample of calls in full and exact counters for the rest (POST /api/v1/metrics), so totals stay exact.

How the coding agent installs it

You don't wire Gin in by hand. One command sets up your agent, and the agent does the edits in a PR you review.

How Gin gets installed and what happens after Run npx @winding-labs/gin: it signs you in with GitHub, mints a key for the repo, and installs the gin skill into your coding agents such as Claude Code, Codex and Cursor. You ask the agent to use the gin skill to wrap the app's AI calls. The agent finds every AI call site, writes .gin/spec.json, adds fetch: gin with a use case to each client, stores GIN_KEY as a secret, and runs your tests plus one real call. Afterwards your app calls the router as before, and after each response the SDK posts one event to Gin. Gin extracts provider, cost, cache and errors, grades cheaper models, and sends the daily email. $ npx @winding-labs/gin signs in with GitHub · mints a per-repo key installs the gin skill into your agents your coding agents Claude Code Codex Cursor … you ask “Use the gin skill to wrap this app’s AI calls” The agent afinds every AI call site bwrites .gin/spec.json cadds fetch: gin({ useCase }) dstores GIN_KEY as a secret eruns your tests + one real call Your app Your router as before one event, after each response Gin extracts provider, cost, cache, errors grades cheaper models on your calls sends the daily email
  1. Run npx @winding-labs/gin. It signs you in with GitHub, mints a key for this repo, and installs the gin skill into your coding agents (Claude Code, Codex, Cursor, …).
  2. Ask your agent: “Use the gin skill to wrap this app’s AI calls.”
  3. The agent finds every AI call site, writes the Gin spec, adds fetch: gin({ useCase }) to each client, stores GIN_KEY as a secret, then runs your tests and makes one real call.
  4. After that, your app's calls go to the router as before. After each response the SDK posts the event to Gin, which extracts provider, cost, cache and errors, grades cheaper models, and sends the daily email.

The Gin spec

.gin/spec.json maps each call site to a use case, router, client and model. One name per product feature, reused wherever that feature calls AI. It is reviewed in the PR like any other file, so labels stay stable as the code changes. It holds no secrets.

.gin/spec.json
{
  "version": 1,
  "repo": "Winding-Labs/dash",
  "useCases": [
    {
      "name": "code-review",
      "router": "openrouter",
      "client": "openai-sdk",
      "model": "deepseek/deepseek-v4.1-flash",
      "files": ["agents/ionwarp/review/lanes.ts"]
    },
    {
      "name": "chat",
      "router": "vercel",
      "client": "ai-sdk",
      "model": "anthropic/claude-sonnet-5",
      "files": ["apps/web/lib/ai/chat.ts"]
    }
  ]
}

Reference

SDK options

gin(options) returns a fetch. Every option is optional.

OptionDefaultWhat it does
keyGIN_KEYYour Gin key. No key means plain fetch.
useCase—The product feature, short kebab-case: review, chat.
env—production, staging, development (free text).
repothe key's repoOwner/name.
waitUntil—Keeps the event post alive after the response. On Workers: ctx.waitUntil.bind(ctx).
capture'full''metadata' never sends request or response bodies.
fetchglobal fetchThe fetch to wrap.
baseUrlhttps://trygin.aiThe Gin server.
gin.flush()
Waits for queued event posts and counter batches. Use it in scripts and in serverless handlers without waitUntil.

Endpoints

POST /api/v1/events
One event per call.
POST /api/v1/metrics
Counter batches for calls not sent in full.
GET /api/v1/config
The workspace's settings for the SDK.
https://trygin.ai/api/v1/openrouter
No SDK: point an OpenAI-compatible client here and send your Gin key in the x-gin-key header.

CLI

npx @winding-labs/gin
Installs the gin skill into your coding agents and signs you in.
npx @winding-labs/gin login
Signs in again and gets a key for the current repo.
npx @winding-labs/gin key
Prints the key for the current repo, to pipe into a secret store.
npx @winding-labs/gin status
Shows agents, sign-in and key for the current repo.
npx @winding-labs/gin uninstall
Removes exactly what install wrote.