Gin / Docs
Gin docs
Gin is a fetch wrapper for the AI clients you already have. Your calls go where they went before. After each response, Gin records what it cost.
Wrap one call
Start from a plain OpenRouter call with the OpenAI SDK. The only change is one line: pass Gin as the client's fetch.
import OpenAI from 'openai';
import { gin } from '@winding-labs/gin';
const openrouter = new OpenAI({
baseURL: 'https://openrouter.ai/api/v1',
apiKey: process.env.OPENROUTER_API_KEY,
fetch: gin({ useCase: 'review' }), // + the only new line
});
const res = await openrouter.chat.completions.create({
model: 'deepseek/deepseek-v4.1-flash',
messages,
usage: { include: true },
});
- Your call still goes straight to your router, with your key. Gin is not a proxy and never sees the key.
- After the response, Gin reads a copy and posts one event to
https://trygin.ai/api/v1/eventsin the background. It is never in the request path. - No
GIN_KEY, or any Gin error, means plainfetch. If Gin is down, your app doesn't notice.
When the response comes back
- Your app gets the router's response as it arrives. Gin passes every chunk straight through and keeps a copy on the side. It never buffers a stream.
- When the body ends, Gin reads the copy. JSON is parsed as is. A stream's
data:chunks are folded into one response: text, tool calls, the lastusage, anyerror. - It pulls out the fields that matter: model, provider, tokens in/out/cached, cost, finish reason, error. Each router puts them somewhere else; the tabs below show exactly where.
- It adds what it measured itself: status, time to headers, time to first token, total duration, and why a call failed (timeout, cancelled, network, rate limit, provider error).
- It posts one event to
POST /api/v1/eventsafter your response is done, and that's what the daily email is built from.
Every router
The wrap is the same everywhere. What changes is where Gin finds the provider, the cost and the cached tokens in each response.
OpenRouter
Base URL https://openrouter.ai/api/v1, OpenAI chat format.
1Your app calls the router through Gin's fetch. Same request, your key, no proxy.
const openrouter = new OpenAI({
baseURL: 'https://openrouter.ai/api/v1',
apiKey: process.env.OPENROUTER_API_KEY,
fetch: gin({ useCase: 'review' }),
});
await openrouter.chat.completions.create({
model: 'deepseek/deepseek-v4.1-flash',
messages,
usage: { include: true }, // exact cost in usage.cost
});
2The router returns. Your app gets this response untouched, chunk by chunk. Gin keeps a copy.
{
"id": "gen-1760051234-a8Qx",
"provider": "Together",
"model": "deepseek/deepseek-v4.1-flash",
"choices": [{ "message": { "role": "assistant", "content": "…" }, "finish_reason": "stop" }],
"usage": {
"prompt_tokens": 52310,
"completion_tokens": 1840,
"prompt_tokens_details": { "cached_tokens": 49152 },
"cost": 0.00312
}
}
3Gin posts one event to POST /api/v1/events in the background, after the body ends. Same numbers, same fields.
{
"ts": "2026-10-09T17:04:11.020Z",
"url": "https://openrouter.ai/api/v1/chat/completions",
"use_case": "review",
"model": "deepseek/deepseek-v4.1-flash",
"served_model": "deepseek/deepseek-v4.1-flash",
"provider": "Together",
"status": 200,
"tokens_in": 52310, "tokens_out": 1840,
"tokens_cached": 49152,
"cost_usd": 0.00312,
"finish_reason": "stop",
"latency_ms": 455, "duration_ms": 7660 // measured by Gin
}
- 1
provider: who served it - 2
model: what answered - 3
usage: tokens in and out - 4
cached_tokens: the caching angle - 5
usage.cost: exact, because the request asked forusage: { include: true } - 6
finish_reason
!If it fails mid-stream. OpenRouter can send an error chunk inside an HTTP 200. Gin still records it as an error.
data: { "id": "gen-…", "error": { "code": 502, "message": "Network connection lost." }, "choices": [{ "delta": {}, "finish_reason": "error" }] }
{ "status": 200, "error": "502 Network connection lost.", "error_type": "stream", "ttft_ms": 500 }
Router.com
Base URL https://api.router.com/v1. OpenAI Responses API only: POST /v1/responses.
1Your app calls the router through Gin's fetch. Same request, your key, no proxy.
const router = new OpenAI({
baseURL: 'https://api.router.com/v1',
apiKey: process.env.ROUTER_API_KEY,
fetch: gin({ useCase: 'summaries' }),
});
await router.responses.create({ model, input });
2The router returns. Your app gets this response untouched, chunk by chunk. Gin keeps a copy.
{
"id": "resp_68e7c2",
"object": "response",
"model": "gpt-6-astra",
"status": "completed",
"output": [{ "type": "message", "content": [{ "type": "output_text", "text": "…" }] }],
"usage": {
"input_tokens": 1200,
"input_tokens_details": { "cached_tokens": 1024 },
"output_tokens": 85
}
}
3Gin posts one event to POST /api/v1/events in the background, after the body ends. Same numbers, same fields.
{
"url": "https://api.router.com/v1/responses",
"use_case": "summarize",
"served_model": "gpt-6-astra",
"status": 200,
"tokens_in": 1200, "tokens_out": 85,
"tokens_cached": 1024,
"finish_reason": "completed",
"latency_ms": 610, "duration_ms": 1480
}
// no cost in the body: Gin prices the tokens from list prices
// no provider in the body: spend is grouped under Router.com
- 2
model - 3
input_tokens/output_tokens(cached tokens are part of input) - 4
input_tokens_details.cached_tokens - 6
statusbecomes the finish reason
Vercel AI Gateway
Base URL https://ai-gateway.vercel.sh/v1 for OpenAI chat and Responses; Anthropic Messages at https://ai-gateway.vercel.sh.
1Your app calls the router through Gin's fetch. Same request, your key, no proxy.
const gateway = new OpenAI({
baseURL: 'https://ai-gateway.vercel.sh/v1',
apiKey: process.env.AI_GATEWAY_API_KEY,
fetch: gin({ useCase: 'chat' }),
});
await gateway.chat.completions.create({
model: 'anthropic/claude-sonnet-5',
messages,
});
2The router returns. Your app gets this response untouched, chunk by chunk. Gin keeps a copy.
data: { "id": "chatcmpl-9x", "model": "anthropic/claude-sonnet-5", "choices": [{ "delta": { "content": "…" } }] }
data: { "id": "chatcmpl-9x", "choices": [{ "delta": { "provider_metadata": { "gateway": {
"routing": { "resolvedProvider": "anthropic", "finalProvider": "bedrock" },
"cost": "0.0031" } } }, "finish_reason": "stop" }],
"usage": { "prompt_tokens": 940, "completion_tokens": 210 } }
data: [DONE]
3Gin posts one event to POST /api/v1/events in the background, after the body ends. Same numbers, same fields.
{
"url": "https://ai-gateway.vercel.sh/v1/chat/completions",
"use_case": "chat",
"served_model": "anthropic/claude-sonnet-5",
"provider": "bedrock",
"status": 200, "streamed": true,
"tokens_in": 940, "tokens_out": 210,
"cost_usd": 0.0031,
"finish_reason": "stop",
"latency_ms": 380, "ttft_ms": 512, "duration_ms": 2950
}
- 1
finalProvider: who actually served it, after any fallback - 2
model - 3
usageon the last chunk - 5
gateway.cost: a decimal string, read as a number - 6
finish_reason
Streams: Gin folds every data: chunk into one response (text, tool calls, the last usage, any error) and reads it the same way as JSON. ttft_ms is when the first chunk arrived.
Cloudflare AI Gateway
URL https://gateway.ai.cloudflare.com/v1/<account>/<gateway>/<provider>/…. The body is the provider's own format.
1Your app calls the router through Gin's fetch. Same request, your key, no proxy.
const cf = new OpenAI({
baseURL: `https://gateway.ai.cloudflare.com/v1/${ACCOUNT_ID}/${GATEWAY_ID}/compat`,
apiKey: process.env.GROQ_API_KEY,
fetch: gin({ useCase: 'triage' }),
});
await cf.chat.completions.create({
model: 'groq/llama-3.3-70b-versatile',
messages,
});
2The router returns. Your app gets this response untouched, chunk by chunk. Gin keeps a copy.
{
"id": "msg_01XF",
"type": "message",
"model": "claude-sonnet-5",
"content": [{ "type": "text", "text": "…" }],
"stop_reason": "end_turn",
"usage": {
"input_tokens": 24,
"cache_read_input_tokens": 3100,
"cache_creation_input_tokens": 0,
"output_tokens": 340
}
}
3Gin posts one event to POST /api/v1/events in the background, after the body ends. Same numbers, same fields.
{
"url": "https://gateway.ai.cloudflare.com/v1/acct/gw/anthropic/v1/messages",
"use_case": "triage",
"served_model": "claude-sonnet-5",
"provider": "anthropic", // from the URL path
"status": 200,
"tokens_in": 3124, // 24 + 3100 cached + 0 written
"tokens_cached": 3100,
"tokens_out": 340,
"finish_reason": "end_turn",
"latency_ms": 820, "duration_ms": 4100
}
- 1the provider is the path segment after your gateway name
- 2
model - 3input = new + cache read + cache write tokens
- 4
cache_read_input_tokens - 6
stop_reason
Direct APIs
OpenAI, Anthropic, Gemini, Groq, Together, Fireworks, DeepSeek, xAI and Mistral. The provider is the host.
1Your app calls the router through Gin's fetch. Same request, your key, no proxy.
import Anthropic from '@anthropic-ai/sdk';
import { gin } from '@winding-labs/gin';
const anthropic = new Anthropic({
apiKey: process.env.ANTHROPIC_API_KEY,
fetch: gin({ useCase: 'support-chat' }),
});
await anthropic.messages.create({
model: 'claude-sonnet-5',
max_tokens: 1024,
messages,
});
2The router returns. Your app gets this response untouched, chunk by chunk. Gin keeps a copy.
{
"id": "msg_01AB",
"model": "claude-sonnet-5",
"stop_reason": "end_turn",
"usage": {
"input_tokens": 12,
"cache_read_input_tokens": 1800,
"output_tokens": 210
}
}
3Gin posts one event to POST /api/v1/events in the background, after the body ends. Same numbers, same fields.
{
"url": "https://api.anthropic.com/v1/messages",
"use_case": "review",
"served_model": "claude-sonnet-5",
"provider": "anthropic", // the host is the provider
"status": 200,
"tokens_in": 1812, "tokens_cached": 1800, "tokens_out": 210,
"finish_reason": "end_turn",
"latency_ms": 640, "duration_ms": 3900
}
// cost: Gin prices the tokens from list prices (cache reads at the cache rate)
- 1direct APIs: the host is the provider
- 2
model - 3tokens in and out
- 4cache reads
- 6
stop_reason
Other clients and runtimes
- Vercel AI SDK.
createOpenAICompatible,createOpenRouterandcreateAnthropictake the samefetch: gin(…). - Cloudflare Workers. Pass
key: env.GIN_KEYandwaitUntil: ctx.waitUntil.bind(ctx)so the event post outlives the response. - One shared client. Tag each call with the header
x-gin-use-case. Gin strips everyx-gin-*header before the provider sees it.
// Vercel AI SDK
const openrouter = createOpenRouter({
apiKey: process.env.OPENROUTER_API_KEY,
fetch: gin({ useCase: 'chat' }),
});
// Cloudflare Workers, inside fetch(request, env, ctx)
const client = new OpenAI({
baseURL: 'https://openrouter.ai/api/v1',
apiKey: env.OPENROUTER_API_KEY,
fetch: gin({
key: env.GIN_KEY,
useCase: 'review',
waitUntil: ctx.waitUntil.bind(ctx),
}),
});
// One client, many features: tag per call
await client.chat.completions.create(
{ model, messages },
{ headers: { 'x-gin-use-case': 'summaries' } },
);
What Gin records per call
One event per call, posted to POST /api/v1/events.
| Field | What it holds |
|---|---|
ts | When the call started |
url | The router or API that was called |
use_case | The product feature, from useCase or x-gin-use-case |
repo, env | Where the call came from |
model | The model you asked for |
served_model | The model that answered |
provider | Who served it |
status, streamed | HTTP status; whether the response streamed |
tokens_in, tokens_out, tokens_cached, tokens_reasoning | Token counts |
cost_usd | Cost of the call |
latency_ms | Time to response headers |
ttft_ms | Time to the first streamed token |
duration_ms | Time to the end of the response |
finish_reason | Why the model stopped |
error, error_type | The error, and one of timeout, aborted, network, rate_limit, server, client, stream |
request, response | The bodies, encrypted per workspace and deleted after 7 days. capture: 'metadata' never sends them. |
Never captured: headers, API keys, cookies.
At high volume the SDK sends a weighted sample of calls in full and exact counters for the rest (POST /api/v1/metrics), so totals stay exact.
How the coding agent installs it
You don't wire Gin in by hand. One command sets up your agent, and the agent does the edits in a PR you review.
- Run
npx @winding-labs/gin. It signs you in with GitHub, mints a key for this repo, and installs theginskill into your coding agents (Claude Code, Codex, Cursor, …). - Ask your agent: “Use the gin skill to wrap this app’s AI calls.”
- The agent finds every AI call site, writes the Gin spec, adds
fetch: gin({ useCase })to each client, storesGIN_KEYas a secret, then runs your tests and makes one real call. - After that, your app's calls go to the router as before. After each response the SDK posts the event to Gin, which extracts provider, cost, cache and errors, grades cheaper models, and sends the daily email.
The Gin spec
.gin/spec.json maps each call site to a use case, router, client and model. One name per product feature, reused wherever that feature calls AI. It is reviewed in the PR like any other file, so labels stay stable as the code changes. It holds no secrets.
{
"version": 1,
"repo": "Winding-Labs/dash",
"useCases": [
{
"name": "code-review",
"router": "openrouter",
"client": "openai-sdk",
"model": "deepseek/deepseek-v4.1-flash",
"files": ["agents/ionwarp/review/lanes.ts"]
},
{
"name": "chat",
"router": "vercel",
"client": "ai-sdk",
"model": "anthropic/claude-sonnet-5",
"files": ["apps/web/lib/ai/chat.ts"]
}
]
}
Reference
SDK options
gin(options) returns a fetch. Every option is optional.
| Option | Default | What it does |
|---|---|---|
key | GIN_KEY | Your Gin key. No key means plain fetch. |
useCase | — | The product feature, short kebab-case: review, chat. |
env | — | production, staging, development (free text). |
repo | the key's repo | Owner/name. |
waitUntil | — | Keeps the event post alive after the response. On Workers: ctx.waitUntil.bind(ctx). |
capture | 'full' | 'metadata' never sends request or response bodies. |
fetch | global fetch | The fetch to wrap. |
baseUrl | https://trygin.ai | The Gin server. |
gin.flush()- Waits for queued event posts and counter batches. Use it in scripts and in serverless handlers without
waitUntil.
Endpoints
POST /api/v1/events- One event per call.
POST /api/v1/metrics- Counter batches for calls not sent in full.
GET /api/v1/config- The workspace's settings for the SDK.
https://trygin.ai/api/v1/openrouter- No SDK: point an OpenAI-compatible client here and send your Gin key in the
x-gin-keyheader.
CLI
npx @winding-labs/gin- Installs the gin skill into your coding agents and signs you in.
npx @winding-labs/gin login- Signs in again and gets a key for the current repo.
npx @winding-labs/gin key- Prints the key for the current repo, to pipe into a secret store.
npx @winding-labs/gin status- Shows agents, sign-in and key for the current repo.
npx @winding-labs/gin uninstall- Removes exactly what install wrote.