Your GPU, everyone's model.

Deference turns idle graphics cards into a shared inference network. Join with one command, serve a model your hardware can run, and earn credit you spend on anyone else's machine — from the chat here, from an OpenAI SDK, or from Claude Code.

Models in the catalog

Browse models →

Machines serving now

See the network →

Your balance

credits available
Ledger →

How it works

1. Join

Create a join code, run the one-line installer. The agent profiles your GPU, picks the best model that fits, and dials out to the network — no port forwarding.

2. Earn

Every token your machine generates for someone else is credited to you, minus a small network fee. Prompts are metered server-side; completions are bounded against what was actually relayed.

3. Spend

Use your credit on any machine in the network. Point an OpenAI client at /v1 or Claude Code at /anthropic with an API key.