Chat here is built around one question: which model actually answers this better? You tick up to four models, write the prompt once, and each of them answers in its own column at the same time. Regenerating a reply makes a sibling rather than a replacement, so the answer you are trying to beat is still there to beat. This page covers which models your plan can pick, what each one is metered against, and where it stops.
The model control in the chat composer is a tray, not a dropdown. Tick a model and it joins the turn; tick a second and both answer the same prompt. Four models a turn is the ceiling — one row of reply columns plus one. On send, each ticked card flies out of the tray to the head of its own column and that column starts writing.
The columns are independent. A model that fails, refuses or runs out of its daily allowance carries that error in its own column and the others keep going. Partial replies are persisted on the server while they are still being written, so reloading mid-answer recovers what has arrived rather than losing the turn.
Which plan: Chat requires an account: an anonymous request to the chat API is refused outright, not served a trimmed-down demo. A free signed-in account gets 5 replies a day, shared across GPT-4o mini, GPT-3.5 Turbo and GPT-5.6 Luna.
Three kinds of model share the tray, and they are metered three different ways — which is the part that surprises people. One group is billed to us in dollars per token, one rides a subscription we pay for, and one runs on a GPU in this building.
| Model | Runs on | Who can pick it |
|---|---|---|
| GPT-5.6 Luna | OpenAI | Every signed-in account; one of the three in the free plan's daily replies |
| GPT-4o mini | OpenAI | Every signed-in account; in the free three |
| GPT-3.5 Turbo | OpenAI | Every signed-in account; in the free three |
| GPT-4o, GPT-5.6 Terra, GPT-5.6 Sol | OpenAI | Paid plans, from the daily budget; outside the free three |
| Claude Haiku | Our own Claude subscription | Every plan — the free plan's 5 Claude replies a day are Haiku |
| Claude Sonnet, Opus, Fable | Our own Claude subscription | Paid plans only; a free account is blocked from these outright |
| Aeon 32B MoE | Our own GPU | Pro and Ultra only |
Claude runs on our own Claude subscription rather than a metered API key, so what is shared out is one subscription's rate limit across everyone on the site — which is why the allowance is counted in replies rather than in dollars. A free plan gets five of those replies a day, and they must be Claude Haiku: the free tier is deliberately barred from Sonnet, Opus and Fable, because the same five replies spent on the most expensive model would be the owner's priciest capacity handed to accounts that pay nothing.
Paid plans get a larger pool and may spend it on any Claude model: 30 replies a day on Starter, 80 on Creator, 200 on Pro, 600 on Ultra, counted in Sonnet-equivalents. Running the pool out is not a dead end on any plan. When the day's allowance is spent, the smallest useful amount is bought automatically from your wallet — green credits first, gold for any remainder — and the reply goes through. A Sonnet reply is worth about one credit and an Opus reply about two, and the purchase is capped at ten credits at a time, or as little as one reply's worth when that is all the wallet holds. Bought allowance does not expire.
When neither wallet can fund even one reply, the original refusal stands and reads as what it has become: add credits. Nothing is charged silently, and nothing is charged twice — the top-up is spent only after the day's included allowance is gone.
Aeon 32B MoE is the one chat model running on hardware we own, and it carries the hardest gate on the page: Pro and Ultra. Below that the server refuses the request rather than quietly routing you to something else. Aeon also sleeps when the GPU is handed over to image and video work, so the first request after a swap waits for a cold load. That wait is priced and shown rather than hidden behind a spinner, and a small green load fee buys the load itself.
Open the chat and pick your models — Opens the chat with the model tray. Nothing is sent until you press send.
Regenerating a reply does not throw the old one away. The new reply is created as a sibling of the old one under the same message, so a conversation is a tree rather than a list, and every version stays reachable. That is the whole point when you are comparing: an answer you are trying to improve on has to still exist.
The same structure is what makes a multi-model turn work at all — four replies to one prompt are four siblings of each other. It survives a reload too: the server holds the tree, the client rebuilds it, and the branch you were reading stays the branch you are on rather than jumping to whichever reply happened to finish last.
Chat does not charge credits per reply. The bound is a daily allowance, and it comes in two shapes. OpenAI models share a per-day budget measured in dollars of model time, because dollars per token is what they cost us. Claude models share a pool of replies measured in Sonnet-equivalents, so an Opus turn — worth about 1.7 Sonnets — draws more of the pool than a Sonnet turn does, and a small model draws less.
| Plan | OpenAI models | Claude replies a day | Aeon 32B |
|---|---|---|---|
| Free | 5 replies a day across three models | 5 Claude Haiku replies a day. Haiku is the only Claude a free plan may pick | Not selectable |
| Starter | $0.20 a day of model time | 30 | Not selectable |
| Creator | $0.95 a day | 80 | Not selectable |
| Pro | $2.70 a day | 200 | Included |
| Ultra | $24.50 a day | 600 | Included |
What it costs: A daily allowance running out is answered rather than final: the smallest useful allowance is bought from your wallet, green credits first and gold for the remainder, and the request is retried. Models metered in dollars rather than replies keep a ten-credit minimum on that purchase. If neither wallet can fund it, the refusal stands.
Two other things in chat can cost credits, both optional and both quoted before you accept them. Aeon's model-load fee is green and buys a cold load instead of a wait. A queue boost on the Aeon lane is green as well, priced fresh against the line as it stands at that second, and charged only when the GPU is genuinely contended — a boost on an idle GPU buys nothing and costs nothing. If someone outbids you while you are deciding, the request is refused rather than billed at the higher number.
Watch out: Free chat is 5 replies a day across those three models. Some of our own older documentation says 20; 5 is the number the server enforces, and the server is what answers.
You can select it on a free account, but it is not free. A free plan includes no Claude allowance, so each reply buys a small amount from your wallet — green credits first, roughly a credit a Sonnet reply. If both wallets are empty, the reply is refused rather than run.
In allowance, yes: each column is a separate reply metered against that model's own allowance. In credits, usually nothing — chat spends credits only once a daily allowance has run out.
It stays. Regenerating creates a sibling branch under the same message and both versions remain reachable. Nothing in a conversation is overwritten.
It runs on our own GPU and needs Pro or Ultra. The gate is enforced on the server, so there is nothing to work around in the browser.
No. The chat endpoint refuses anonymous callers. Signing up is free and grants the daily reply allowance immediately.
Compare across families rather than within one. Two sizes of the same model usually differ in speed and length; a hosted model against a model on our own GPU differs in the things you are actually choosing between.