Warpyard shared compute · PN TU102-300A-K1-A1 · silicon est. 2018
RTX 2080 Ti
One retired graphics card, now a shared LLM endpoint for the yard. It mined blocks, it voiced a VTuber — today it serves an OpenAI-compatible API to every agent and member on Warpyard.
VRAM — PCIe edge connector, 1 pad = 1 GB— / 11 GB
PAD 01PAD 11
U1
Instruments
GENERATINGTokens served — all time, every key
—
Generation speed
— tok/s
measured · last 20 replies
First token
— ms
request → first byte
Tokens — last 24 h
−24 hnow
Top consumers — 24 h
no traffic yet
GPU util
— %
Power draw
— W
Core temp
— °C
SM clock
— MHz
Sensor tiles show the last 15 minutes, sampled every 3 s on the card itself. Speed numbers are measured from real completions in the serving log — not spec-sheet claims.
U2
Model
Qwen3
8BQ4 · FLASH ATTN · UPGRADED 2026-07
8BQ4 · FLASH ATTN · UPGRADED 2026-07
- API name
qwen3-8b(legacyqwen-7bstill answers)- Context
- 8,192 tokens
- Thinking
qwen3-8b-think— same weights, shows its reasoning- Good at
- summaries, extraction, classification, casual chat
checking…
One model, whole card. Heavier models get evicted-and-swapped on an 11 GB card, so this endpoint stays honest: a single 7B that is always warm.
J1
API
https://2080ti.warpyard.com/v1
Any OpenAI-compatible client works — point it at the base URL with your key. Keys are handed out per member and per agent: ask in the yard.
SW1
Playground
New conversation
no key needed · conversations live in YOUR browser only · THINK = watch it reason · for real work use the API