Open models.
Rented GPUs.
One API.

dazz.ai stands between AI marketplaces and the graphics cards that run their models. One request comes in, it finds the right GPU, and the tokens stream back — no matter whose card is doing the work.

No spam — just a note when the first model goes live.

How a request travels

Every call takes the same path, whether it lands on our first card or the fiftieth one a provider adds next month.

Marketplace

A request arrives

Someone calls a model through the marketplace where it's listed. The call carries a key and a model name, nothing more.

Router

It's matched to a card

A small router checks which GPUs are awake, which are already busy, and picks the one that can answer fastest.

GPU

The model runs

The chosen card, wherever it's rented from, loads the weights already in memory and starts generating tokens.

Response

Tokens stream back

Output flows back through the same router, the same marketplace, the same request — usually before the caller notices any of this happened.

Compute doesn't need to come from us

dazz.ai is a routing layer, not a data center. The cards behind it are rented by the hour, and over time some of them will belong to other people entirely.

  • Anyone can serve a model

    If you're already renting a GPU, you can list it, take a share of the requests it handles, and get paid for the hours it's actually working.

  • Only open-source weights

    Every model we route to is one anyone could download and run themselves. We're just making that easier and cheaper to reach.

  • Priced by demand, not by us

    GPU rental prices move with the market. Our job is to route around the cheapest healthy card, not to set a margin on top of it.

  • Built to fail loudly, not silently

    A card that goes quiet gets pulled from rotation immediately. Slow is one thing; a hung request is another, and we watch for both.

Where things stand today

We'd rather say this plainly than dress it up: dazz.ai hasn't launched yet. Here's what's actually being decided right now.

Marketplace — comparing a handful of open inference marketplaces that let anyone register as a provider, rather than the larger platforms with closed onboarding.

First model — looking for something with real demand and thin supply, so the first card we rent pays for itself quickly.

Hardware — sizing GPU memory against model weights plus room for concurrent requests, starting small with one to four cards.

Cloud provider — weighing price against reliability across a few GPU rental markets before committing hours to any one of them.

Be first to know when a model goes live

One email, sent once, when the first GPU starts answering real requests.