Google releases Nano Banana 2.1 with 4K editing and cheaper image generation

Google’s new hosted image model adds stronger text rendering, multi-image consistency and search grounding, while its published context limits conflict across official documentation.

2 min read

A computer monitor showing a photo-editing interface, illustrating AI-assisted image generation and editing.

Google released Gemini Nano Banana 2.1 on October 6 as a generally available image-generation and conversational-editing model. The API model ID is gemini-nano-banana-2.1, and Google positions it as the faster, more cost-efficient counterpart to Gemini 3 Pro Image.

The new version accepts text, images, video and PDFs and can return images plus text. It generates at 1K, 2K or 4K resolution, supports wide panoramic aspect ratios, and can combine as many as 14 reference images—up to four people and ten objects, according to Google. It also adds configurable minimal, medium and high thinking levels and can ground requests with Google Search.

Google’s model card reports stronger vendor-run results than Gemini 3.1 Flash Image and Gemini 3 Pro Image. With thinking enabled, Nano Banana 2.1 scored 1050 ±14 on Google’s text-to-image overall-preference Elo, versus 990 ±7 and 935 ±8 respectively. Its infographic factuality score was 0.521, compared with 0.179 and 0.265. These are not independent benchmarks: Google used a mix of internal and public tasks, side-by-side human preference judgments and an automated factuality rater. The tested prompts and full evaluation data are not published, so teams should validate quality and consistency on their own workloads.

The model is available through the Gemini app, AI Studio, Gemini API, Google Search AI Mode, Flow, Stitch, Ads and Google’s enterprise platform. Google Cloud lists standard global pricing at $1.50 per million input tokens, $7.50 per million text-output tokens and $30 per million image-output tokens; batch or off-peak pricing is lower. This is a hosted proprietary service, not a downloadable weight release.

One documentation issue remains: the Gemini API model page lists a 131,072-token input limit and 32,768-token output limit, while Google’s model card describes a one-million-token input window and up to 64K text output. Developers should rely on the limit enforced by their chosen endpoint until Google reconciles those pages.

A computer monitor showing a photo-editing interface, illustrating AI-assisted image generation and editing. This is an editorial illustration, not a Google product screenshot.

Sources: Gemini API model page · Google DeepMind model card · Google Cloud availability and limits · Google Cloud pricing · Unsplash image license

Googleimage editingGeminiimage generationNano Banana 2.1