Open models.
Rented GPUs.
One API.
dazz.ai stands between AI marketplaces and the graphics cards that run their models. One request comes in, it finds the right GPU, and the tokens stream back — no matter whose card is doing the work.
No spam — just a note when the first model goes live.
How a request travels
Every call takes the same path, whether it lands on our first card or the fiftieth one a provider adds next month.
A request arrives
Someone calls a model through the marketplace where it's listed. The call carries a key and a model name, nothing more.
It's matched to a card
A small router checks which GPUs are awake, which are already busy, and picks the one that can answer fastest.
The model runs
The chosen card, wherever it's rented from, loads the weights already in memory and starts generating tokens.
Tokens stream back
Output flows back through the same router, the same marketplace, the same request — usually before the caller notices any of this happened.
Compute doesn't need to come from us
dazz.ai is a routing layer, not a data center. The cards behind it are rented by the hour, and over time some of them will belong to other people entirely.
-
Anyone can serve a model
If you're already renting a GPU, you can list it, take a share of the requests it handles, and get paid for the hours it's actually working.
-
Only open-source weights
Every model we route to is one anyone could download and run themselves. We're just making that easier and cheaper to reach.
-
Priced by demand, not by us
GPU rental prices move with the market. Our job is to route around the cheapest healthy card, not to set a margin on top of it.
-
Built to fail loudly, not silently
A card that goes quiet gets pulled from rotation immediately. Slow is one thing; a hung request is another, and we watch for both.
Where things stand today
We'd rather say this plainly than dress it up: dazz.ai hasn't launched yet. Here's what's actually being decided right now.
Marketplace — comparing a handful of open inference marketplaces that let anyone register as a provider, rather than the larger platforms with closed onboarding.
First model — looking for something with real demand and thin supply, so the first card we rent pays for itself quickly.
Hardware — sizing GPU memory against model weights plus room for concurrent requests, starting small with one to four cards.
Cloud provider — weighing price against reliability across a few GPU rental markets before committing hours to any one of them.
Be first to know when a model goes live
One email, sent once, when the first GPU starts answering real requests.