xAI's newest image model, built for images you can put into real work — designed typography, layouts that hold together, and the same subject carried across generations. Run it here at 1K or 2K with up to five reference images, from a browser.
Headlines, credits and small print planned like a designer would set them.
Feed it the character, the product and the palette in one generation.
Run your first prompt in the box above without an account.
Generated with Grok Imagine Image 2.0 on this site on 2026-08-16, unretouched. A gig poster where every credit line reads, a product shot with an embossed label, and one character drawn three times without drifting. Nothing was typeset or retouched afterwards.



xAI's image model, released on 2026-08-07 as the Quality Mode inside Grok. It was built around following instructions closely — planning typography and layout, and holding what you feed it across generations and edits.
xAI states the design goal plainly in its announcement: make images you can use in real work. That shows up in three places. It follows instructions down to the details rather than treating them as mood. It plans typography and layout the way a designer would, so a dense multi-part visual holds together and small text comes out sharp instead of dissolving into letter-shaped noise.
And it preserves what you put in — the same character, the same product, the same palette — across a run of generations, which is what turns a single lucky image into a set you can ship.
The company reports it ranking second in the world on both the text-to-image and image-editing Arena leaderboards as of its launch day, where xAI's entries are listed under SpaceXAI. On Veida AI it runs from the browser at 1K or 2K, in five aspect ratios, with up to five reference images in a single generation and up to four images per run.
Credits are shared with every other model here, so trying it costs nothing extra beyond the run itself.
Creative engine
Prompt in, 1K or 2K out, with optional reference images. The credit estimate updates before you generate.
The first generation is rarely the final asset, so the model is trained to change what you name and keep the rest.
Up to five input images per generation, which removes a round of manual compositing.
Pick it when the picture has a job to do and a second version is coming.
Posters, packaging, ads, thumbnails, UI mockups — anything where a misspelt headline makes the whole render useless.
A character across locations, a product across formats, an icon family. Consistency across generations is what it was trained for.
A face, a garment and a setting can go in together rather than being composited by hand afterwards.
What the model accepts here, and what each control does.
Start from a prompt, or upload something you already have and describe the change you want made to it.
16:9, 1:1, 2:3, 3:2 and 9:16 — the frame is decided before generation rather than cropped afterwards.
1K for drafts and feeds, 2K when the result is going to print or needs to survive a crop.
Four variations of one prompt in a single run, which is the cheapest way to find the composition you want.
Measured against the live API this model runs on here, on 2026-08-16.
Four habits that decide whether you get a picture or something you can actually ship.

Put the exact headline, the exact date line and the exact small print in the prompt, in quotation marks. Guessed wording is the one thing it cannot get right for you.

Say where the headline sits, what goes underneath it and what fills the footer. Layout instructions are what this model was built to follow.

Upload the character, the product or the palette rather than describing them. Five slots is enough to pin down a whole look in one pass.

Once a composition works, re-run it with a single clause changed. Holding the rest steady is exactly what it is good at, and the comparison stays readable.
Mostly work with a deadline attached.

Where the credit block, the date and the venue all have to be readable at the size it prints.

One bottle, one lighting setup, and a label that stays legible across every crop the listing needs.

Game sprites, mascots and icon families that have to look like they came from one hand.

Five ratios cover the feed formats, so the same idea is generated to fit rather than cropped until the text falls off.
Every model on this page draws from the same balance, and the credit cost appears before you press generate. Failed jobs are never charged.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
One-time top-ups — buy extra credits any time you run low.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
Charges appear as “VEIDA AI” on your card statement. You can cancel any time from Settings → Billing; cancellation takes effect at the end of the period you already paid for. Operator details, the full model list, and the refund window are on the about page.
Powered by
Use one balance across every model listed below, and check the credit cost before each request. Nothing is charged for a job that fails.
Model credit guide
You're only charged for successful generations. The exact estimate in the generator varies by model, length, resolution, audio, and number of images.
| Type | Model | Credit cost |
|---|---|---|
| GPT Image 2 | Text to image and description-based edits. 1K ≈ 20 credits, 2K ≈ 90, 4K ≈ 160. | |
| Nano Banana 2 | Generate or edit from up to 14 reference photos. Same 20 / 40 / 80 ladder by resolution. | |
| Nano Banana 2 Lite | The free tier's editor. Lightest cost per edit on the site. | |
| Seedream 5 Lite | Fast drafts. Priced at the 1K tier for iteration passes. | |
| Qwen Image 3 | Legible text and page layout rather than photoreal. Priced by resolution. |
Common questions about generating with this model online.
You can start in the generator on this page, and every new account gets a credit balance to spend on it. Credits are shared across every model on Veida AI, so nothing is locked to one of them.
Yes — Image 2.0 is the model xAI shipped on 2026-08-07 as the Quality Mode in its own apps. What differs here is the surface around it: a browser generator with the aspect ratio, resolution and reference slots exposed as controls, and one credit balance across every model on the site.
Five in a single generation. That is enough to pin a character, a garment, a prop and a palette at once, which is the case where compositing by hand used to be the only route.
The three examples on this page are the honest answer: a gig poster with four credit lines, a date line and a street address, all generated here on 2026-08-16, none of it typeset afterwards. Write the exact wording into your prompt and it renders that wording.
30 credits at 1K and 40 at 2K. The aspect ratio does not change either number, and a batch of four costs four runs.
Yes. Upload a source image in the generator and describe the change. Editing was treated as a first-class capability in this generation rather than as an afterthought bolted onto a text-to-image model.
xAI reports it second in the world on both the text-to-image and image-editing Arena leaderboards as of its launch day, with OpenAI's GPT Image 2 first in both. Both models run here, so you can put the same prompt through each and judge for your own use case rather than take either company's word for it.
Grok renders at medium quality by default and has a distinct, less airbrushed look that suits meme-adjacent and editorial images. The Banana line is more controllable for edits. Use Grok when you want its texture, not as a substitute.
Write the exact words, name the layout, bring your references, and get back an asset rather than a mood board.