Nano Banana Pro - What It Is and How Google's AI Image Generator Works

Version 7, as of September 1, 2026. Prices and limits in Google services change several times a year, so check the current pricing before implementation.
Nano Banana Pro (gemini-3-pro-image) is Google DeepMind's most powerful image model, combining Gemini 3 Pro reasoning with image generation. It was created for tasks where earlier generators failed: legible text within the frame, consistent characters across a series of shots, and informational graphics that must match reality.
What Nano Banana Pro Is and How the Gemini 3 Pro Image Architecture Works
Nano Banana is the marketing name for Google's entire visual model family, not a single product. As of September 2026, it consists of four variants:
| Variant | Model Identifier | What It Is Used For |
|---|---|---|
| Nano Banana Pro | gemini-3-pro-image |
Most complex compositions, brand consistency, reasoning, and grounding |
| Nano Banana 2 | gemini-3.1-flash-image |
Workhorse model, generation up to 4K, good quality-to-cost ratio |
| Nano Banana 2 Lite | gemini-3.1-flash-lite-image |
Fastest and cheapest, 1K only |
| Nano Banana (older) | gemini-2.5-flash-image |
Deprecated version, Google recommends upgrading to Lite |
The Pro variant undergoes a reasoning phase before generating an image: it analyzes the prompt, organizes spatial relationships within the scene, and only then creates the graphic. The result is evident in the control over the physics of the frame. Google describes this as production-level control over lighting, camera, focus, and color grading.
It is worth setting expectations right away. Google does not publish a detailed description of this model's internal architecture, so the descriptions of a "two-stage LLM plus diffusion pipeline" circulating online are reconstructions, not documentation. We know what the model does. How exactly it is built has not been disclosed by the manufacturer.
Google Search Grounding in Image Generation: Fact-Based Graphics
Nano Banana Pro can be connected to Google Search. Enabling grounding gives the model access to real-time web content during generation, which matters for infographics, diagrams, and data-driven visualizations.
This is an optional feature, enabled intentionally on the query side. It does not run in the background for every graphic.
In our view, this is the most overhyped feature in the entire suite. In the model description, Google explicitly points out that its knowledge of the world is vast, but not infallible. Grounding reduces the amount of nonsense in a technical diagram, but it does not replace a review by someone who knows the subject. If you are building a workflow where an infographic goes straight from prompt to a client's website, grounding will not save that process.
The second thing that genuinely changes the workflow is typography. The model renders multilingual text along with Polish diacritics and handles text applied to curved surfaces, such as a bottle label. Google reports that in single-line text benchmarks, the error rate stays below 10 percent in most languages. This is a manufacturer measurement without independent confirmation, so treat it as a ballpark estimate. Accurate text plus fact-checking form the foundation of consistent branding materials and content prepared for pozycjonowanie AI.
Working with Reference Images: 14 Inputs, 6 Shots, 5 Characters
Reference limits are often distorted, so we provide them based on Google's developer announcement. In a single task, Nano Banana Pro can:
- preserve the recognizable appearance of up to five people,
- maintain the fidelity of six high-resolution shots,
- combine up to fourteen standard inputs into a single composition.
The Gemini API documentation describes this same limit slightly differently: as 6 objects, 5 characters, and 3 styles for the Pro variant. By comparison, Nano Banana 2 supports 10 objects, 4 characters, and 3 styles. The numbers therefore depend on what a given reference represents to the model, and there is no single "maximum of X images" figure.
Method for Combining Multiple References for Character and Style
- Separate anatomy from the environment. Facial features and body proportions remain, while the camera angle and lighting change. This is the use case where Pro pulls ahead of the rest of the market most significantly.
- Selective style transfer transfers grain, grading, and exposure from the reference to new objects without transferring the actual content of the photo.
- Low-quality references degrade the output faster than having none at all. It is worth filling the six high-fidelity slots with material you would publish yourself.
Prompting for Visual Consistency
The most effective habit is explicit role mapping. State in the prompt which attachment accounts for product geometry and which for the background style. Without this, the model mixes features, and after three shots, the series stops looking like a single campaign.
The second habit: separate fixed parameters from variable ones. Brand color codes, lighting type, and typeface remain unchanged across every prompt. Pose, facial expression, and perspective are the variables.
Let's check your website's potential
Share your website and email - we'll get back to you with a real analysis, no strings attached.
Technical Specifications, Conversational Editing, and SynthID
| Parameter | Specification |
|---|---|
| Resolution | Generation in 1K, 2K, and 4K. Lite variant 1K only. |
| Aspect ratios | Pro supports 1:1, 4:3, 5:3, 16:9, 9:16, as well as cinematic formats like 1.85:1, 2.39:1, 2.75:1, and 4:1, among others. The list differs across variants, so check the documentation for the specific model you use. |
| Editing and retouching | Changes applied across subsequent conversation turns without rendering the entire composition from scratch. Background replacement, object removal and addition, changing the time of day. |
| SynthID | An invisible watermark embedded into the pixels of every image generated or edited by the model. It survives compression, cropping, and format changes. |
| Provenance verification | A suspicious file can be uploaded to the Gemini app to ask whether it was created using Google tools. The feature launched with support for English prompts. |
On top of SynthID, there is a second, visible watermark. The Gemini sparkle appears on graphics for free users and Google AI Pro plan subscribers. Google removes it for Google AI Ultra subscribers and for images created in Google AI Studio.
This has an immediate operational consequence, and for us, it is a rule written into our workflow: we do not generate final client assets in the Gemini app on the Pro plan. Any creative intended for a campaign is created in AI Studio or via the API. Otherwise, someone will spot the sparkle in the corner during the approval stage, and the whole batch gets sent back for revisions.
Nano Banana Pro Compared to Nano Banana 2
| Model | Consistency and references | In-frame text | Grounding | Availability |
|---|---|---|---|---|
| Nano Banana Pro | 5 characters, 6 high-fidelity shots, up to 14 inputs | Strong point of the model, including Polish diacritics | Optional, via Google Search | Gemini, AI Studio, Gemini API, Vertex AI, Slides, Vids, Ads, NotebookLM |
| Nano Banana 2 | 10 objects, 4 characters, 3 styles | Good, weaker with complex typographic layouts | None in this variant | Default model in the Gemini app, including for free accounts |
We do not recommend picking a variant based on aesthetics, as that leads to poor purchasing decisions. The difference lies in discipline: version 2 usually makes a pretty image; Pro makes an image that matches the brief. For sketches and draft graphics, Nano Banana 2 is sufficient, while Lite covers the lower end of the pricing tier for simple, flat layouts. Paying extra for Pro pays off in one specific scenario: for creative work where legible text and a consistent character must make it into client assets without manual touch-ups.
What Nano Banana Pro Does Not Do Well
Google's materials describe what the model excels at. The rest is covered below, ordered by what most frequently disrupts production work.
The brief can be interpreted arbitrarily. In the test described below, none of the three variants returned the 4:5 aspect ratio specified in the prompt; all of them output a square. Set the format via parameters, not through a sentence in the prompt. The same applies to genre: Nano Banana 2 treated the word "poster" as a mood rather than a format, creating a photograph of a coffee shop interior instead. With vector graphics, all three left the bottom half of the frame blank, requiring the file to be cropped anyway.
The visible watermark extends to the Pro plan. The Gemini sparkle disappears only on Ultra and in Google AI Studio. Anyone paying $19.99 per month thinking they are purchasing clean files is still getting them with a signature in the corner.
Calculate the cost separately. A 4K image via the API costs $0.24, meaning five hundred creative variations for A/B testing cost $120 for a single iteration. With weekly rotation in a performance campaign, this turns into a budget line item, not pocket change.
The model can quietly fall back to a weaker variant. In the Gemini app, once the daily Pro quota is exhausted, generation shifts to version 2 without any clear indication. You will only notice by comparing files from the beginning and end of the session. This is the most frustrating aspect of the entire tool, and the sole reason we stick to AI Studio for client work, where the model is explicitly selected from a list.
Two things are left for the sections below because they cannot be summarized in a single sentence: grounding reduces the number of factual errors, but does not eliminate them, and the rights to generated graphics work differently than most people assume.
How to Use Nano Banana Pro Step by Step
Prompt Construction Geared Toward Photorealism
- Specify optical parameters. Focal length, aperture, lighting type, material texture. The model uses these when planning the scene rather than treating them as mere embellishments.
- Switch to Pro. In the Gemini app, use the option to regenerate the image with the Pro model. It is available to users on paid plans, starting with Google AI Plus.
- Refine across subsequent turns. Conversational editing modifies the specified element without recalculating the entire composition, so do not start from scratch when a single detail is off.
Two Prompts, Three Model Variants, One Generation Each
Instead of describing AI typography based on press releases, we ran two identical prompts through three variants of the family. The first tests a single piece of text with Polish diacritics integrated into a photographic scene. The second is more demanding: multiple text blocks, numbering, and layout - the elements that marketing materials are actually made of.
Method. One generation per prompt and model, six images in total. We did not pick the best attempt from a series, and we did not retouch anything after generation. Everything was run in Google AI Studio, 4:5 aspect ratio, grounding disabled.
This approach eliminates the accusation of cherry-picking results to fit a thesis, but it does not control for variance. Image models can produce distinctly different results with the exact same prompt, so treat this as a snapshot of performance on a specific day, not a definitive benchmark. If you are making a purchasing decision, repeat the test using your own materials.
One caveat regarding resolution. Nano Banana 2 Lite generates exclusively in 1K, so we conducted the entire comparison in 1K. The Pro variant and version 2 perform better in 2K and 4K than what is shown below, but using those resolutions would have prevented a side-by-side comparison.
| Variant | Model in AI Studio | Test Resolution |
|---|---|---|
| Nano Banana Pro | gemini-3-pro-image |
1K |
| Nano Banana 2 | gemini-3.1-flash-image |
1K |
| Nano Banana 2 Lite | gemini-3.1-flash-lite-image |
1K |
What to look for in both prompts: the shape of the letters Ś, ż, ń, and ó, the acute accent above ń in the word "Gdańsk", font consistency within a single text element, and the absence of blur along letter edges. In the infographic, additional factors include the ogonek on ą in the heading "Jak wygląda proces SEO", step numbering, and whether captions stay aligned with their corresponding icons.
Prompt 1: Text Embedded in a Photographic Scene
Minimalist advertising poster for a Polish artisan coffee shop. In the foreground, a ceramic coffee mug; in the background, warm wood and a neon sign on the wall with clear, flawless Polish lettering: "Świeżo palona kawa - Gdańsk Wrzeszcz". Style: product photography 50mm f/1.8, natural morning light.
Nano Banana Pro

The only one of the three that took the word "poster" literally, building the framing around the text instead of simply taking a photo of a coffee shop interior. Ś, ż, and ń are flawless, the typeface is consistent across both lines, and the neon glow reflects accurately off the ceramics and wood. This file is ready to publish without edits.
Nano Banana 2

Ignored the word "poster" and delivered an interior photograph instead. Nevertheless, the text is error-free, and the model added coffee bags on its own featuring a legible "ŚWIEŻO PALONA" label - successfully handling a second, much smaller tier of text within the same scene.
Nano Banana 2 Lite

Here lies the difference we ran this test to find. The first line of the neon sign is correct, but the second part falls apart completely: mismatched font, smaller type size, and the ending "WRZESZCZ" blurred beyond legibility. On top of that, an unprompted hand appears in the frame.
Prompt 2: Infographic with Four Captions
Minimalist infographic on a light background. Four process steps arranged horizontally, each with a simple line icon and a caption in Polish: "1. Badanie słów kluczowych", "2. Optymalizacja treści", "3. Budowa linków", "4. Pomiar wyników". Heading at the top: "Jak wygląda proces SEO". Color palette: navy blue, white, one orange accent. Flat vector style, high legibility.
Nano Banana Pro

The strongest visual hierarchy of the three: an all-caps heading, four error-free captions, arrows guiding the eye between steps, and a palette matching the prompt exactly. The icons correspond to their descriptions - something that is not a given with image generators. The only element left to fix is an empty band across the bottom of the frame.
Nano Banana 2

All diacritics are in place. The icons were placed inside circles, unrequested decorative elements appeared, and the heading lost some visual hierarchy compared to the Pro version. The composition remains legible, just less disciplined.
Nano Banana 2 Lite

A pleasant surprise. Four error-free captions, and the model even added correct microtext inside the icon: "Słowo 1", "Słowo 2", "Słowo 3", and an "SEO" badge on a shield. That is a harder task than lettering on a neon sign, and the cheapest variant pulled it off. However, it had the weakest composition of the three, leaving the bottom half of the frame blank.
What Does This Mean?
For flat vector graphics, the gap between variants turned out to be narrower than Google's marketing suggests. All three models correctly spelled "Optymalizacja treści", "Budowa linków", and "Pomiar wyników" - including the one that costs $0.0336 per image. If you are creating simple infographics in Polish, the Pro variant is not necessary.
The difference becomes obvious with text embedded in a photorealistic scene. That is where Lite broke down, version 2 built a different composition than requested, and Pro was the only one that nailed both the brief and the typography. This is the boundary where paying extra makes sense: product and brand visuals with in-frame text.
One issue affected all three: none respected the requested 4:5 aspect ratio. All returned a square, and Lite added white letterbox bars to the sides of its vertical frame. For production workflows targeting specific ad formats, this must be enforced via parameters rather than prompt descriptions.
Availability, Pricing, and Integration via Gemini API and Vertex AI
Nano Banana Pro has rolled out across most Google surfaces. Beyond the Gemini app, it works in Google Slides and Vids for Workspace customers, in Google Ads globally, in NotebookLM, and for Ultra subscribers in the Flow tool.
Consumer plans (US list prices; billing in Poland is in PLN with different rates):
- Free: generation using Nano Banana 2, limited access pool for Pro.
- Google AI Plus: $7.99 per month: image regeneration with Nano Banana Pro.
- Google AI Pro: $19.99 per month: expanded access to generation and editing.
- Google AI Ultra: starting at $99.99 per month: highest limits, no visible watermark.
API pricing (standard rates; batch mode is 50% cheaper):
| Model | 1K | 2K | 4K |
|---|---|---|---|
| Nano Banana Pro | $0.134 | $0.134 | $0.24 |
| Nano Banana 2 | $0.067 | $0.101 | $0.151 |
| Nano Banana 2 Lite | $0.0336 | n/a | n/a |
For enterprise customers, the appropriate channel is Vertex AI on Google Cloud, and not just because of SLAs. The main reason is legal, detailed below.
Image Rights and AI Content Labeling Requirements
Google claims no ownership over generated images, and commercial use is permitted across all plans, including the free tier. However, this is not the same as the "full license" often promised in marketing materials.
Two things you need to know before implementation:
Images generated entirely by AI are not protected by copyright, because they lack a human creator. The US Court of Appeals for the District of Columbia Circuit reaffirmed this in 2025 in Thaler v. Perlmutter. The practical consequence: competitors can copy your creative work, and you have no legal grounds to stop them. Protection applies only when there is substantial human contribution to the asset.
Third-party indemnification depends on the channel. Consumer plans and the Gemini API provide none. Workspace offers limited protection. Google's full copyright indemnification commitment is exclusive to Vertex AI. If you generate assets for clients in regulated industries, this factor matters far more than the price per image.
On top of that comes EU regulation. As of August 2, 2026, transparency rules under Article 50 of the AI Act take effect. Model providers must ensure machine-readable marking of synthetic content, which SynthID handles. A separate obligation applies to deployers: anyone publishing a deepfake or AI-generated informational content must disclose it. For systems placed on the market before August 2, 2026, a transitional period runs until December 2, 2026. What matters is the date the content was generated, not the date of publication.
An agency releasing creatives under a client's brand cannot assume that the model provider has taken care of these compliance obligations for them.
FAQ
Can images from Nano Banana Pro be used commercially?
Yes, across all plans, including the free one. Google does not claim ownership of the generated files. A separate issue is that artwork generated entirely by AI is not protected by copyright, so you will not be able to prevent competitors from copying it.
How does Nano Banana Pro differ from Nano Banana 2?
Pro is based on Gemini 3 Pro Image; it features reasoning, optional Search grounding, and higher character consistency limits. Nano Banana 2 is Gemini 3.1 Flash Image, a faster and cheaper model that serves as the default in the Gemini app. Version 2 is sufficient for exploring ideas, but for final creative assets with in-frame text, it is worth reaching for Pro.
Is Nano Banana Pro available for free?
A free account generates with the Nano Banana 2 model and gets a limited pool of Pro generations. Once it is exhausted, the app falls back to the less powerful model. Full access requires a paid Google AI plan or pay-as-you-go billing in the Gemini API.
How many reference images does Nano Banana Pro support?
Up to fourteen standard inputs in a single task, six of which can be high-fidelity shots. The model maintains a consistent appearance for up to five people.
How does the SynthID watermark work?
SynthID modifies pixels in a way that is invisible to the human eye and survives compression, cropping, and format changes. In addition to this, Google applies a visible Gemini sparkle to images generated by free users and Google AI Pro plan subscribers. It disappears on the Ultra plan and in Google AI Studio.
Do I need to label AI-generated images in publications?
Starting August 2, 2026, Article 50 of the EU Artificial Intelligence Act imposes a disclosure obligation on any entity publishing deepfakes and AI-generated informational content. Machine-readable labeling is provided by the model provider, and SynthID meets this requirement, but this does not exempt the publisher from their own obligations.
Sources
- Google DeepMind, Gemini 3 Pro Image - model specifications, consistency limits, typography benchmark.
- Google, Build with Nano Banana Pro, our Gemini 3 Pro Image model - developer announcement, grounding, SynthID.
- Google, Nano Banana Pro - product availability, visible watermark policy.
- Google AI for Developers, Image generation - model variants and identifiers.
- Google AI for Developers, Gemini Developer API pricing - rates per image.
- Regulation (EU) 2024/1689, Article 50 - transparency obligations applicable from August 2, 2026.