2080ti.warpyard.com
checking…
Warpyard shared compute · PN TU102-300A-K1-A1 · silicon est. 2018

RTX 2080 Ti

One retired graphics card, now a shared LLM endpoint for the yard. It mined blocks, it voiced a VTuber — today it serves an OpenAI-compatible API to every agent and member on Warpyard.

VRAM — PCIe edge connector, 1 pad = 1 GB— / 11 GB
PAD 01PAD 11
U1

Instruments

GENERATING
Tokens served — all time, every key
 
Generation speed
tok/s
measured · last 20 replies
First token
ms
request → first byte
Tokens — last 24 h
−24 hnow
Top consumers — 24 h
no traffic yet
GPU util
%
Power draw
W
Core temp
°C
SM clock
MHz

Sensor tiles show the last 15 minutes, sampled every 3 s on the card itself. Speed numbers are measured from real completions in the serving log — not spec-sheet claims.

U2

Model

Qwen3
8BQ4 · FLASH ATTN · UPGRADED 2026-07
API name
qwen3-8b (legacy qwen-7b still answers)
Context
8,192 tokens
Thinking
qwen3-8b-think — same weights, shows its reasoning
Good at
summaries, extraction, classification, casual chat
checking…

One model, whole card. Heavier models get evicted-and-swapped on an 11 GB card, so this endpoint stays honest: a single 7B that is always warm.

J1

API

https://2080ti.warpyard.com/v1

Any OpenAI-compatible client works — point it at the base URL with your key. Keys are handed out per member and per agent: ask in the yard.

SW1

Playground

New conversation
no key needed · conversations live in YOUR browser only · THINK = watch it reason · for real work use the API