You generated a hundred AI images for a campaign. Every one of them looked great until you examined the headline closely. Extra letters. A word that did not exist. A CTA button that was supposed to say “Shop Now” but instead read “Shpo Nwo.”
That problem became far less common on April 21, 2026, when OpenAI introduced ChatGPT Images 2.0, powered by its new image-generation model, GPT Image 2. The biggest upgrade was not just photorealism it was text rendering. Infographics, posters, and banner ads with specific copy became a much more realistic use case instead of a five-attempt gamble.
Here is what changed, how to prompt it, what it costs, and the limits worth knowing before you promise anything to a client.
⚠️ Note: information current as of August 2026. OpenAI updates model behavior, pricing, and availability regularly. Check the official documentation before relying on any specific number.

If you are still assembling a stack rather than choosing between two image models, start with our broader roundup of AI tools for marketers and come back here when text on the image is the specific problem.
What Changed With GPT Image 2
OpenAI describes GPT Image 2 as its state-of-the-art model for image generation and editing, supporting flexible image sizes and high-fidelity image inputs through both the Images and Responses APIs. It replaced the GPT Image 1.5 pipeline as the default across ChatGPT’s free and paid tiers. Three changes matter most for marketers, plus one caveat about sizes.
Two modes instead of one. Instant mode is the default for everyone, including free accounts. It renders quickly and includes the improved text and multilingual rendering, so it is not a stripped-down version. Thinking mode adds a reasoning pass before the model draws: it can plan a layout and check its own output first. Launch coverage consistently reports that Thinking mode inside ChatGPT can return a batch of up to eight images from one prompt while holding style and subject steady across the set. Thinking mode is noticeably slower, and OpenAI’s release notes confirm it requires a paid plan while Images 2.0 itself is on all plans.

Text rendering that survives contact with real work. Short headlines, labels, callouts, and CTA copy are now far more likely to render legibly within the first few attempts. OpenAI’s own launch materials lean heavily on text-dense examples: magazine-style infographic spreads, academic posters, campaign brochures, even a print-ready bookmark with bleed and trim guides. But OpenAI still lists text placement and clarity as a known limitation, so this is a large improvement rather than a solved problem.
Multilingual text. The model renders non-Latin scripts including Japanese, Korean, Hindi, and Bengali far more reliably than earlier image models. Relevant if your brand runs localized creative.
One interface change worth knowing. Every image you generate is saved automatically under a dedicated Images section, available on web and mobile, so past assets stay browsable without scrolling back through chat history. Useful when you need to find the version a client approved three weeks ago.
One capability worth flagging separately, because it changes what a prompt can be: OpenAI frames the model as a “visual thought partner” and demonstrates Thinking mode researching a topic before drawing it. One of its own examples generates a product mockup grid from a query about current merchandise. In practice that means a prompt can ask for a visual about something the model looks up, rather than only something you fully specify.
The size caveat. Through the API, OpenAI documents flexible custom sizes and lists 4K dimensions (3840×2160 and 2160×3840) among its popular options, so 4K is real. But output above 2560×1440 is officially labeled experimental, which matters for a client-facing hero asset. Community testing reports that the ChatGPT interface exposes far fewer options than the API, effectively 1024×1024, 1536×1024, and 1024×1536. So check before you build a process around generating a 3:1 banner in ChatGPT.
For API users: the Images API lets you control the model, prompt, size, quality, output format, compression, and number of generated assets. Custom sizes must keep the longest edge at 3840px or under, both edges as multiples of 16, a long-to-short ratio no greater than 3:1, and total pixels between 655,360 and 8,294,400. There is a pinned snapshot, gpt-image-2-2026-04-21 if you need behavior to stay stable, and the free API tier does not include this model. Some guides quote “4096×4096” as the ceiling; that size breaks two of those rules and is a useful tell for how carefully a guide was checked. If you want the full developer-side picture rather than the marketing-side one, including the whole model family, both API surfaces, and code samples, we cover that separately in What is GPT Image?
What It Actually Costs
Third-party guides quote wildly different figures because many are reseller prices. These are OpenAI’s own published prices per generated image, which is the number that matters if you use the API:
| Quality | 1024×1024 | 1024×1536 | 1536×1024 |
|---|---|---|---|
| Low | $0.006 | $0.005 | $0.005 |
| Medium | $0.053 | $0.041 | $0.041 |
| High | $0.211 | $0.165 | $0.165 |
Two things worth noticing. A low-quality draft costs well under a cent, so there is no reason to be precious about iterating.
And among these three reference sizes, the square costs more despite being the smaller image. That is not a typo. A 1024×1024 square is about 1.05 million pixels; a 1024×1536 portrait is about 1.57 million, half again as many, and it costs less at every tier. At high quality the square runs roughly 28 percent more expensive than the rectangle with two thirds of the pixels.
The reason is that you are billed for output tokens rather than for pixels, and OpenAI notes directly that a larger non-square resolution can sometimes produce fewer output tokens than a smaller or square one at the same quality. Worth separating two things that sound like they should agree: OpenAI notes square images are typically the fastest to generate, and they are also the most expensive of these three. Speed and cost do not move together here. The practical takeaway: if the asset works in landscape or portrait, defaulting everything to square is quietly costing you money. This holds across the three sizes OpenAI publishes for comparison; the model supports thousands of resolutions, so check the cost calculator for any size you plan to use in volume.
One more lever if you generate in volume: OpenAI’s Batch tier prices image tokens at roughly half the standard rate. For a set of localized variants that does not need to exist in the next five minutes, that halves the bill.
If you only use ChatGPT rather than the API, none of this applies; generation is covered by your plan.
Prices verified against OpenAI’s official API pricing on August 4, 2026. Pricing changes; check the current pricing page and cost calculator before budgeting.

None of these replaces the others. A realistic workflow uses Midjourney or Nano Banana Pro for a hero product photo, and GPT Image 2 whenever the copy on the image has to be exactly right.
Two honest footnotes. If text accuracy is your only concern, typography-first generators like Ideogram are worth a look for high-volume batch work. And if you need a logo or scalable brand asset, vector-native generators such as Recraft solve a problem GPT Image 2 cannot: editable SVG rather than a flat raster.
How to Prompt GPT Image 2 for Text-Heavy Visuals
Structure controls the visual. Exact wording controls the text.
Most AI image prompts focus only on what the image should look like. For marketing assets, that is not enough. A strong GPT Image 2 prompt works like a creative brief: it defines the purpose, composition, exact copy, style, and limits.
Before you write: choose the right mode
Use Instant mode for exploration:
- testing ideas
- trying headline variations
- checking layouts and colors
Use Thinking mode for final assets:
- multiple text blocks
- complex layouts
- important marketing visuals
For API workflows, use lower quality settings for drafts and higher quality for final assets, especially when the image contains dense text or detailed layouts.
The 6-Part GPT Image 2 Prompt Structure
A production-ready prompt has six parts:
- Purpose and format, why the asset exists and where it will be used
- Scene and subject, the environment and main visual focus
- Layout and hierarchy, where elements appear and what matters most
- Exact text and typography, the words that must appear and how they should look
- Style and visual direction, the overall visual language
- Constraints, what must not change or appear
One example runs through all six parts
Purpose: Website hero banner for a premium flower delivery service.
Audience: customers looking for fresh flowers for gifts, events, and celebrations.
Format: 1536×1024 landscape website banner.
Scene: bright studio with a elegant bouquet arrangement on a marble table.
Subject: fresh flower bouquet with roses, peonies, and eucalyptus in soft morning light. Layout: bouquet on the right side, headline and CTA on the left.
Hierarchy: headline is the largest element, supporting text below, CTA button underneath. Spacing: clean negative space around text and flowers.
Headline: “Fresh Flowers Delivered Same Day” Subhead: “Hand-picked bouquets for every occasion, delivered to your door”
CTA text: “Shop Now”
Typography: Bold sans-serif headline. Left aligned. Modern boutique style. CTA in a soft pink accent button. Do not rewrite, shorten, translate, or add any text.
Style: Premium floral brand, elegant, fresh, romantic.
Palette: soft pink, cream, sage green, white. Lighting: natural morning light.
Constraints: No additional text. No fake logos. No watermark. No distorted flowers. Keep the bouquet realistic and fresh-looking.

Purpose: Visual storyboard for a short emotional animated film about a son trying to bake a birthday cake for his mother.
Audience: families, young adults, and viewers who enjoy warm, humorous, emotional stories.
Format: 1536×1024 horizontal storyboard sheet with 8 clearly separated panels arranged in two rows of four.
Story Title: “Storyboard: Ming’s Birthday Cake”
Subtitle: “A clumsy but heartfelt gift.”
Main Character:
Ming, a young Asian office worker in his late twenties, with short black hair, expressive brown eyes, and a light blue office shirt. He is kind, slightly clumsy, sincere, and emotionally expressive. His appearance and clothing must remain consistent across all eight panels.



Two rules that apply to every part
Replace vague adjectives with visual facts. “Stunning,” “professional,” and “high quality” carry no visual information. Concrete parameters do. This matters more here than with some competitors, because an under-specified prompt is a reported trigger for stray textures.
| Instead of | Write |
|---|---|
| beautiful lighting | golden hour side lighting, soft shadows |
| professional | studio three-point lighting, seamless white backdrop |
| cinematic | wide framing, teal and orange grade, low sun angle |
| photorealistic | Contax T2 film look, Kodak Portra 800 grain, natural light falloff |
| high quality | visible fabric texture, shallow depth of field |
| modern | flat design, generous white space, single accent color |
Know that your prompt may be rewritten, and check what was actually sent. OpenAI documents that when image generation runs as a tool inside a conversation, the main model revises your prompt “for improved performance” and exposes what it sent in a field called revised_prompt. For most images that is harmless. For an asset where you specified exact headline text at a specific weight and alignment, it is precisely what you do not want.
The route matters here. The rewrite happens on the tool-inside-a-conversation path; a direct Images API call sends your prompt as written, which makes it the predictable choice for text-critical work. Inside ChatGPT, add a short line instead: (format 1536×1024.) (don’t change the prompt, send it as it is.)
You can also ask ChatGPT to show you the prompt it sent. Fastest way to check whether your typography instructions survived the trip.
Editing an approved image: Change / Preserve / Constraints
The six-part structure is for creating an asset. Fixing one is a different and much shorter formula. Do not re-describe the whole scene. Write direct commands, be terse, and state what changes, what stays identical, and what must not happen.
Change: replace the headline text “Simple Pricing” with “New Pricing, Same Simplicity.” Preserve: background, logo placement, color palette, font style, overall layout. Constraints: no added elements, no watermark.
Two details improve reliability. Quote both the old and the new text when swapping copy. And keep it to one or two changes per request, because five changes in one prompt is how you lose an image you had already approved.
If you need to confine a change to one region, the API supports a mask. OpenAI notes it must match the source in format and size, stay under 50MB, and include an alpha channel, and that masking is prompt-guided rather than pixel-exact.
Inside ChatGPT you can skip the mask entirely. Open the image, use the selection tool to highlight the exact area that needs to change, then describe the change in the conversation panel. OpenAI documents both paths: highlight first, or skip the selection and describe the edit directly. It also notes that highlights are not always precise and an edit can extend past the area you selected, so keep the same one-or-two-changes-per-request discipline even with a selection in place.
After you generate: verify before you ship
Text rendering is much better, not solved. Run the same five checks every time:
- Read every word in the image out loud, including the CTA
- Check alignment and hierarchy against Part 3
- Confirm no text was added that you did not ask for
- Confirm the protected negative space is still clear
- Confirm the delivered pixel dimensions match what you requested
Three Problems the Reviews Rarely Mention
Every “best AI image generator” list will tell you GPT Image 2 handles text well. The following comes from OpenAI’s own community forum, where practitioners document reproducible issues, and matches what independent hands-on reviewers describe. Treat these as widely reported observations rather than published specifications.
1. Output tends to degrade the longer you stay in one chat
The ChatGPT image generator appears to carry data over from images already made in the session. Testers report output picking up a visible noise or grid-like texture that compounds with each generation, often becoming obvious after roughly three to five images in a row. At least one independent reviewer describes the same symptom as stray random textures requiring repeated reruns to clear.
The reported fix is simple: start a fresh chat. Reload the page and begin a new session. Prompt wording does not appear to help, because the cause is not in the prompt.
This cuts against the common advice to stay in one thread and refine iteratively. That advice is good for holding creative context and bad for image quality. Our suggested compromise: iterate in one thread until the composition is approved, then open a clean session and run the approved prompt once for the final asset.
2. Re-running a prompt in the same chat may return a near-identical image
Because the session appears to reuse prior image data, testers report that running an identical prompt again in the same chat returns an effectively identical image rather than a fresh variation. If you want a genuine second option, start a new session. Worth knowing before you spend twenty minutes clicking regenerate and wondering why nothing changes.
3. Reference images and style references can introduce artifacts
Uploading a reference image, or asking for something “in the style of” a named artist, has been observed to trigger the same texture artifacts even on the first generation of a clean session. In the second case, forum contributors attribute it to the model retrieving reference material before drawing, though this is an inference from observed behavior rather than documented mechanics.
Two mitigations that testers report working. A follow-up instruction along the lines of “remove the noise from the image, keep all the lines” repairs a good share of it. And asking for less detail reduces the complexity that triggers the pattern, at the cost of some richness. When you do not need a reference, leaving it out produces the cleanest output.
Where GPT Image 2 Still Falls Short
Four of these are OpenAI’s own documented limitations, which is worth reading before you promise anything to a client. Two land directly on marketing work.
Composition control. In OpenAI’s words, the model may have difficulty placing elements precisely in structured or layout-sensitive compositions. That is a description of an infographic. Expect to iterate on anything with a strict grid.
Text placement and clarity. Documented as significantly improved and still imperfect.
Consistency across generations. The model can struggle to hold recurring characters or brand elements steady across separate generations. Within one Thinking-mode batch it is much better, which is why batching matters for a campaign set.
Latency. Complex prompts can take up to two minutes.
Three more come from the output format rather than the model:
No transparent backgrounds. Requests with a transparent background are not supported on this model. Generate on a flat solid background and remove it afterwards with a dedicated tool. GPT Image 1.5 is still listed and priced on OpenAI’s API pricing page as of August 2026, so an API workflow that depends on transparency has somewhere to go, but transparency support varies by model and older models get retired, so verify current behavior before building on it. Inside ChatGPT it is not a practical route at all.
Raster only. PNG by default, with JPEG and WebP available plus a compression setting for those two; JPEG generates faster if latency matters. No vector output, so logos and print work need rebuilding or a vector-native tool.
Brand marks are unreliable. The model will render something that is almost your logo. Composite real brand assets in afterwards.
Frequently Asked Questions
Is GPT Image 2 the same as DALL-E 3?
No. GPT Image 2 replaced DALL-E 3 as the default image model in ChatGPT. It is a separate model with stronger text rendering, flexible sizing, and an optional reasoning pass called Thinking mode.
Do I need a paid ChatGPT plan to use GPT Image 2?
Instant mode is on the free tier and includes the improved text rendering, so quality is not what you give up. Volume and Thinking mode are. Free allowances change, so check your account rather than planning a campaign around one.
How much does GPT Image 2 cost per image?
Through the API, roughly half a cent for a low-quality draft, four to five cents at medium, and 17 to 21 cents at high quality. Square formats cost more than landscape or portrait at the same tier. Inside ChatGPT, generation is covered by your plan.
Why do my GPT Image 2 images get worse the longer I work?
Testers report the session reusing data from earlier images, producing a noise or grid texture that compounds, often visible after three to five generations. Reload the page and start a fresh chat.
Can GPT Image 2 keep a design consistent across multiple images?
Within a single Thinking-mode batch in ChatGPT, yes: up to eight images holding colors, subject, and layout style steady. Across separate generations it is weaker, and OpenAI lists consistency for recurring brand elements as a known limitation.
Can GPT Image 2 make a transparent PNG for a logo?
No. Transparent backgrounds are not supported. Generate on a flat solid background, then remove the background separately. For production, rebuild the mark as a vector file.
Is GPT Image 2 better than Midjourney for marketing images?
It depends on the asset. GPT Image 2 is stronger when the image needs exact readable text. Midjourney keeps the edge for cinematic and editorial imagery where no text renders inside the image.
Further Reading
- Nano Banana Pro: The Complete Guide for Marketers, for the other strong option when your visual needs believable materials rather than exact text
- Our Midjourney prompt guide, for photorealistic product and campaign imagery. Written against V7; the prompt patterns still apply to the current version
- NotebookLM Infographic: Turning Your Data Into Visual Stories, for deciding what an infographic should actually say
- AI Tools for Marketers in 2026, for the wider stack
- What is GPT Image? on StackNova, the technical companion to this article: the full model family, the Image API versus the Responses API, masks, streaming, and code
About the author
Serafima Osovitny is a marketing manager at Nova Express. Passionate about turning complex marketing tactics into simple, actionable guides, she shares insights about AI search visibility and generative engine optimization.
Explore her work at serafima.digital and follow her on X: @OSerafimaA




Leave a Comment