Alibaba's image model for pictures that contain writing. Type a prompt with the exact words you want on the page, pick a ratio, and it comes back with those words spelled correctly — posters, infographics, storyboards, signage and UI mockups.
Headlines, body copy and captions, spelled the way you typed them.
Grids, numbered steps, panels and rules — not just a picture with a word on it.
Run your first prompt in the box above without an account.
Generated with the Pro tier on this site on 2026-08-05, unretouched, at 2K. Every character you can read in them came out of the model — a poster headline, a full page of infographic body copy, and six handwritten storyboard captions. Nothing was typeset afterwards.



This is the third generation of Alibaba's Qwen image line, and the thing it is built around is writing. Most image models treat text as decoration: ask for a shop sign and you get shapes that look like letters from a distance and fall apart up close. This model treats the words as content. Give it a headline, a set of numbered steps, a caption under every panel, and it lays the page out and spells the words.
That difference changes what you can ask for. A poster stops being a background you have to add type to in another tool, and becomes one generation. An infographic with four steps and a paragraph under each one becomes a prompt. The same capability covers newspapers, exam papers, menus, product packaging, game UI and interface mockups — anything where the information in the picture is the point of the picture.
It also renders more than one writing system, so the words do not have to be English. The tier that runs here is Pro, the top rung of the family, at 1K or 2K with a choice of seven aspect ratios.
Creative engine
Two resolutions, seven ratios. The ratio never changes the price.
Headlines, paragraphs and captions come back spelled, not suggested.
Grids, numbered sequences, panels and headers hold their structure.
Start from a prompt, or hand it a picture and describe the change.
Pick it the moment your prompt contains a word in quotation marks. If you are describing a scene — a fox in autumn woodland, a portrait in mixed light — a photoreal model like [Seedream 5.0 Pro](/image/seedream-5-pro) will serve you better, and [Nano Banana Pro](/image/nano-banana-pro) is the one to reach for when a design has to look art-directed. But the second the picture has to say something, the ranking inverts. A model that renders a beautiful café and puts MOFFEE CAFE over the door has not made you a usable image, and no amount of re-rolling fixes it reliably. This is the model that gets the sign right, and then gets the opening hours under it right too. It is also the model to use when the layout carries meaning: step one above step two, a header separated by a rule, a caption tied to the panel it belongs to. Those are structural decisions, and it makes them.
Posters, covers, flyers, packaging — where the type is the design.
Infographics, menus, exam papers, documents held in frame.
In-scene lettering that has to survive being looked at closely.
What the model accepts here, and what each control does.
A prompt in, a finished page out, with the wording carried through.
Hand it a picture and describe the change you want made to it.
1:1, 16:9, 9:16, 4:3, 3:4, 3:2 and 2:3 — portrait, landscape and square.
2K is the one to use when the smallest text on the page still has to read.
Measured against the live API this model runs on here, on 2026-08-05.
Four habits that decide whether the words come out right.

Write the exact string you want rendered — reading 'MARCH 14-22', not a description of it.

Headline at the top, a second line beneath, credits along the bottom edge.

Condensed grotesk capitals, handwritten script, a thin rule under the header.

Body copy and captions hold together at 2K in a way they cannot at 1K.
The pictures that have something to say.

A headline, a date line and a credit block, finished in one generation.

Numbered steps with real body copy under each one, laid out on a grid.

Panelled sheets where every frame carries its own caption.

Shopfronts, menus, packaging and interface screens with legible labels.
Every model on this page draws from the same balance, and the credit cost appears before you press generate. Failed jobs are never charged.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
One-time top-ups — buy extra credits any time you run low.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
What's included
Secure checkout by Stripe. Card details never touch our servers.
Charges appear as “VEIDA AI” on your card statement. You can cancel any time from Settings → Billing; cancellation takes effect at the end of the period you already paid for. Operator details, the full model list, and the refund window are on the about page.
Powered by
Use one balance across every model listed below, and check the credit cost before each request. Nothing is charged for a job that fails.
Model credit guide
You're only charged for successful generations. The exact estimate in the generator varies by model, length, resolution, audio, and number of images.
| Type | Model | Credit cost |
|---|---|---|
| GPT Image 2 | Text to image and description-based edits. 1K ≈ 20 credits, 2K ≈ 90, 4K ≈ 160. | |
| Nano Banana 2 | Generate or edit from up to 14 reference photos. Same 20 / 40 / 80 ladder by resolution. | |
| Nano Banana 2 Lite | The free tier's editor. Lightest cost per edit on the site. | |
| Seedream 5 Lite | Fast drafts. Priced at the 1K tier for iteration passes. | |
| Qwen Image 3 | Legible text and page layout rather than photoreal. Priced by resolution. |
Common questions about generating with this model online.
You can start in the generator on this page without an account, and every new account gets a credit balance to spend on it. Credits are shared across every model on Veida AI, so nothing is locked to one of them.
Yes. The tier wired here is the Pro one, the top rung of the family, which is the one that holds small text together. It runs at 1K or 2K.
The three examples on this page are the honest answer: a poster headline, four paragraphs of infographic body copy, and six handwritten captions, all generated here on 2026-08-05 and none of them typeset afterwards. Write the exact wording into your prompt and it renders that wording.
Yes — multilingual rendering is one of the things the 3.0 generation was built for. Put the exact characters you want in the prompt.
20 credits at 1K and at 2K — the flat rate is the same at both sizes, and the aspect ratio does not change it.
Yes. Upload a source image in the generator and describe the change — the same text-rendering strength applies when you are adding or replacing writing in a picture you already have.
Qwen's flat rate is its trick — 20 credits whether 1K or 2K, so 2K edits cost less than anywhere else on the page. Banana 2 recovers better when an edit goes wrong. Bulk 2K work → Qwen; tricky single edits → Banana.
Write the words you want on the page, pick a ratio, and let it typeset the thing for you.