Open-weight vs frontier AI models, and when each one wins

This platform runs two kinds of model side by side, and almost everything else about it follows from that one fact — the two wallets, the plan gates, what fails when something goes wrong, and how long you wait. The split is not about quality and it is not about age. It is about where the weights live: some of these models are files on a disk in this building, and the rest exist for us only as an HTTPS call somebody sends us a bill for.

What the split actually is

"Open-weight" here means one thing and nothing more: the checkpoint is a file we hold, and we run it ourselves on a graphics card in this building. "Frontier" means the opposite in the only sense that matters to a platform — the model exists for us solely as a request to somebody else's server, which answers with a picture and an invoice. We are not making a claim about licences or about which lab is ahead. We are describing which of two very different things happens when you press Generate.

In our catalogue the distinction is not editorial, it is a field. An own-GPU row names a workflow_file and a workstation key; a hosted row names a provider and nothing else. Everything a user experiences downstream — the wallet, the tier gate, the failure mode, the wait — is derived from which of those two shapes the row has.

The same product, two supply chains
Open weights, on our GPUsFrontier, behind an API
Wallet⚡ green✦ gold
What one output costs usGPU seconds and electricityA per-call invoice
Reachable on the free planYes, from ⚡1No, at any size or length
What is scarceOne card, one job at a timeYour gold balance
CeilingThis card's memoryThe provider's fleet
What failure looks likeA queue you can watch"Not enough credits", or a provider outage
Who can change the priceUs, at runtimeThem, and we pass it through

Why this is the whole reason a free tier exists

A free tier is only possible where the marginal cost of the next unit is near zero. On a machine we already bought, one more picture is a few seconds of electricity, so we can hand out a wallet that refills on a clock and not go out of business. On a hosted endpoint, one more picture is money leaving the account the moment the request is made, and no refilling wallet survives contact with that. Green credits therefore buy own-GPU work and nothing else — not as a restriction, but because that is the only place the economics allow.

The neatest demonstration of the whole argument is a single model. Qwen Image Edit runs here twice: once as a checkpoint on our own GPU from ⚡1, and once as the same family served by a hosted endpoint at ✦1.4 in gold. Same lineage, same job, two supply chains, two prices, one picker. Z-Image Turbo carries the same arrangement inside a single row: when the own-GPU lane cannot serve it, the run is routed to its hosted twin and billed ✦2.5 in gold instead of ⚡1 in green — the same trade, made automatically.

What open weights are genuinely better at

Three things, and none of them is "the pictures are better". We have not run a controlled comparison of output quality and neither has anyone else who is selling you a subscription, so we are not going to publish a scoreboard. What we can point at is capability that follows structurally from holding the weights.

The quirks are visible, which is itself the feature

Running a model yourself means its awkward edges are yours to explain rather than a provider's to hide. Four of ours, all of which are on the page beside the model:

What frontier models are genuinely better at

Also three things, and they are not small. The first is ceilings. Our own video lanes are bounded by one card's memory, which is why the local envelopes top out where they do; a hosted model is bounded by a provider's fleet. Seedance 2.0 Mini renders up to 4K, Kling 3.0 Pro renders its native 1080p tier with no resolution control to get wrong, and Veo 3.1 produces a native audio track with the picture. Nothing on our hardware does any of those.

The second is inputs. The hosted lanes carry reference systems built by the labs that trained them — up to six reference images on Nano Banana Pro and Seedream 5.0 Pro, three Ingredients on Veo 3.1, nine reference items including video and audio on Seedance. Those are real capability differences and they are the reason a hosted call is worth ✦4.7 when a local one is worth ⚡1.

The third is simply throughput. There is one GPU behind every own-GPU model on this platform and it runs one job at a time. That is the honest cost of a free tier on hardware we own: the electricity is cheap, the machine is not duplicable. A hosted model has no queue of ours to sit in.

Watch out: The own-GPU queue is the real price of the green lane. Credits refill; the card does not multiply. Every queued job shows its position and can be boosted into a faster lane — Fast doubles the run's cost, Express quadruples it, and the multiplier is snapshotted at submit so a busier minute later cannot change what you agreed to.

How to choose, as a rule you can actually follow

  1. Draft on the open-weight lane. Z-Image Turbo or Qwen Image Edit for stills, LTX 2.3 for motion. These are the attempts you are allowed to waste, and wasting attempts is how the shot gets found.
  2. Fix the prompt before you change the model. Switching to a frontier model to escape a bad prompt spends gold proving the prompt was the problem. The own-GPU models are fast enough that a second attempt costs a minute.
  3. Move up only for a ceiling you have actually hit. 4K, a native audio track, six references, or a length our card will not render. If none of those is the reason, the hosted call is buying a difference you have not defined.
  4. Spend gold once, on the frame you intend to keep. One deliberate hosted render at the end of a green session is the cheapest good workflow this platform supports, and it is the one we use ourselves.

Open the model picker — Every model shows its wallet and its plan gate in the picker, so the menu is also the price list.

Are open-weight models worse than frontier ones?

We do not publish a quality scoreboard, because we have not run a controlled comparison and would be guessing. What is measurable is that the frontier models here reach ceilings ours cannot — 4K, native audio, more references — and that they cost more per output. The tables above give the real numbers rather than a single multiple, because green and gold credits do not convert into each other and any ratio between them would be inventing an exchange rate.

Why can't my free credits pay for Nano Banana Pro?

Because that call invoices us in dollars the moment it is made, and a wallet that refills on a clock cannot pay an invoice. Green credits buy GPU time on hardware we already own; gold buys everything a third party bills for.

Can I run my own trained model here?

Yes, on the open-weight lane. A LoRA trained on our GPUs can be loaded beside the checkpoint it was trained against and generated with. Training is metered per GPU-hour and the entitlement opens at Creator.

What happens if your GPU is down?

Where the model has a hosted twin, the run is routed to it and billed in gold instead of green — ✦2.5 for Z-Image Turbo. Where it does not, the job waits for the card. Reachability is checked before anything is quoted, so a run is refused or queued rather than silently failing.

Which should I use for a client deliverable?

Draft on the open-weight lane, then render the final frame or clip on the hosted model whose ceiling you need. Also check the watermark: Free and Starter output always carries it, and switching it off opens at Creator.