Latency is the number most people feel first. The first render of a session, while you are still deciding whether this tool is any good, sets how the tool feels for a long time afterwards.
It is also a number we only partly own. ockeeper does not train or host any model. When you press Generate, we send the request to a third-party model through OpenRouter, and most of the time between your click and the image is spent on somebody else's hardware.
Where the seconds go
A render has three parts: getting your request out of the studio, the model doing its work, and getting the result back into your Vault. The first and last are ours. The middle one — almost always the biggest — belongs to the model provider, including whatever queueing and warm-up happens on their side.
So we stopped talking about cold start as if it were one number we could tune. The part we can make faster is the part around the model.
We cannot make someone else's model faster. We can stop making you wait on anything else.
— engineering note
What we do not run
There is no ockeeper GPU fleet, no warm pool and no weight cache to tune, because the weights are not ours. The image models are made by Google and OpenAI, the video models by ByteDance and MiniMax, and OpenRouter routes each request to them.
Anything we claimed about warming those machines would be a claim about somebody else's infrastructure, so we do not make one.
Picking the model is picking the wait
The biggest lever you have is the model. Each one in the picker shows a typical wait next to its price and maximum output size, and they really are that different:
Nano Banana 2 Lite — About 3 seconds. 1 credit, up to 1024 px. The Draft tier, and the one on the free plan.
Nano Banana 2 — About 5 seconds. 2 credits, up to 2048 px.
GPT-5 Image Mini — About 8 seconds. 10 credits, up to 4096 px.
Nano Banana Pro — About 20 seconds. 30 credits, up to 4096 px.
Video — One to two minutes per clip, depending on the model. The estimate is shown before you render.
Those are typical figures for display, not guarantees — the provider's load on the day decides the rest.
What we did change
Within our own part of the path, three things make the wait shorter or at least easier to sit through. Generate returns straight away and the render carries on in the background, so the studio never freezes on a request. A batch of 2, 4 or 8 goes out as parallel requests rather than one after another. And each image appears in the thread as soon as it lands, instead of waiting for the slowest one in the batch.
A failed render does not cost you anything either. If a model errors, or its safety filter refuses the prompt, that image is not charged — you pay for each image that actually lands in your Vault.
~3s — typical wait on the Draft tier
~20s — typical wait on Nano Banana Pro
0 — credits for a failed or refused image
What we cannot fix
Provider-side slowdowns are real. Some days a model runs slower than its typical figure, and occasionally a request fails outright. When that happens you are not charged, and switching to a faster model is one click in the picker.
We also cannot change how long a model takes to work. If a provider makes a model faster or slower, the figure in the picker is what we update — not a story about our own servers.
If a render ever takes far longer than its estimate, tell us. We respond within 1 business day, and a model that is running slow is worth knowing about.