Google released Nano Banana 2.1 on 6 October, making its latest image-generation and editing model generally available through the Gemini API. Its listed image-output charges are lower than Nano Banana 2’s. Early tests also show a useful distinction: following a visual brief more closely does not guarantee that every face, cap or pose survives an edit.

Based on 9 primary sources · See sources

A live API model, with product rollouts under way

Google’s release notes confirm the API launch under the model name gemini-nano-banana-2.1. The company’s original rollout post separately names Gemini, AI Mode in Search, AI Studio, Flow, Stitch, Google Ads and Gemini Enterprise Platform. Those product rollouts should not be read as simultaneous access to every feature on every account.

The model documentation lists 1K, 2K and 4K output, up to 14 reference images, web and image-search grounding, and configurable reasoning levels. It describes improvements to text, subject consistency and panoramic images. Those are Google’s capability claims; the practical question is how they behave on a particular brief.

From the source · Google ·

Google introduces Nano Banana 2.1 and describes its image-generation and editing improvements.

View the original post on X ↗

For European developers, Google’s supported-region list includes Belgium, France and Germany, along with the UK and Switzerland. The service has account and age requirements, and image generation has no free API tier. This establishes a developer route, rather than a blanket promise about each consumer or advertising product.

The saving depends on resolution—and what the bill includes

Google’s standard API pricing roughly halves the listed 1K and 2K image-output charges. The 4K equivalent falls from $0.151 to $0.113, a smaller reduction of about 25%. The chart translates the listed per-image amounts into 1,000 outputs at each resolution.

These are image-output charges. Input tokens, text or reasoning output, and chargeable search queries can add to a request. The new model’s input rate is $1.50 per million tokens, versus $0.50 for Nano Banana 2. Reference-heavy requests therefore need their own cost comparison.

The comparison

Image-output charges for 1,000 images

Google’s listed standard Gemini API equivalents, USD. Same resolution compared within each pair.

1K · Nano Banana 2.1
$33.60
1K · Nano Banana 2
$67.00
2K · Nano Banana 2.1
$50.40
2K · Nano Banana 2
$101.00
4K · Nano Banana 2.1
$113.00
4K · Nano Banana 2
$151.00
Calculated as 1,000 × Google’s listed per-image equivalents, retaining their rounding. Excludes inputs, text/thinking output, search fees and retries; not an all-in campaign budget.Source: Google Gemini API pricing · 7 Oct 2026

The winner changes with the task

A same-prompt comparison by Fuser, a commercial creative-workflow platform, makes the trade-offs tangible. It compared 2.1 with Nano Banana 2 and the premium Nano Banana Pro, using identical inputs at 2K. Each model got one attempt at each of five tasks, with no retries.

Three selected cases show why the distinction matters. Getting the lighting right, preserving a packshot and keeping a person recognisable are separate jobs. An improvement in one does not establish an improvement in the others.

The detail

Three tasks from Fuser’s five-test comparison

Same prompts and inputs; one run per model per task, 2K output, 6 October 2026. Reported judgments, not numerical scores.

Three tasks from Fuser’s five-test comparison. Same prompts and inputs; one run per model per task, 2K output, 6 October 2026. Reported judgments, not numerical scores.
TaskReported resultWhat to inspect
Product lightingNano Banana 2.1 preferredOnly 2.1 followed the left-light/right-shadow brief.
Move a serum bottle into a new settingTieAll kept the label, but gave a smooth cap ring a ribbed texture.
Reference person in a new sceneNano Banana Pro preferredCloser likeness and readable signs; 2.1 softened distinctive features.
Selected examples from a model-access vendor’s own test. One output per condition cannot establish a success rate.Source: Fuser’s original prompts, outputs and judgments · 7 Oct 2026

Good local edits can coexist with awkward geometry

Himanshu Goel’s Segmind walkthrough reports 18 paid generations through that provider’s endpoint, with prompts, output dimensions and billed costs. He found a label edit could survive a subsequent change of background. But a two-object scene placed a camera below the requested position and lost a robot’s arm.

His conclusion pinpoints the gap between recognising an object and arranging it correctly. Like Fuser, Segmind sells model access; the useful evidence is in the published prompts and outputs.

“Object identity is strong, physical pose instruction is weaker.”

Himanshu Goel, writing for Segmind · Read the original ↗

Arena’s launch snapshot gives a broader comparison: fourth in multi-image editing, fifth in text-to-image and sixth in single-image editing. These are three separate task rankings, rather than positions in one overall table. They also do not answer whether a particular product photograph remains exact.

A cheaper candidate, with details still worth checking

Google has deprecated Nano Banana 2, but the reviewed API release notes give no shutdown date. That is a reason to plan migration without inventing a deadline.

For product and campaign imagery, the useful comparison is the approved brief: exact label text, the untouched parts of a product, a recognisable reference face and the actual bill after retries. Ideogram’s emphasis on preserving unedited areas offers another approach to this same editing problem.

Nano Banana 2.1 is a credible candidate for that comparison. The documented savings are real at the listed image-output rates, and the early examples identify both gains and specific failure modes. They give a team better questions to test, rather than permission to skip the final visual check.

Sources

Written with AI assistance from the sources above. Analysis and practical implications are our interpretation. Our editorial approach.