Introduction
Google Gemini is now one of the easiest ways to generate and edit images for free. With the latest models — including Nano Banana 2 (officially Gemini 3.1 Flash Image), launched in February 2026 — you can create near-Pro quality images at Flash-level speed without paying a cent. The catch is the same as always: the better your prompt, the better your picture.
A vague prompt like "a cat in a park" returns a forgettable, generic result. A structured prompt that names the subject, composition, action, location, and style returns something you'd actually want to share. This guide breaks down how to write great Gemini image generation prompts, gives you copy-and-paste templates for common scenarios, and shows you how to fix the results that miss.
Key Takeaways
- Google's official recipe has six parts: subject, composition, action, location, style, and editing instructions.
- Using photography language — camera angle, lens, lighting, depth of field — dramatically improves realism and mood.
- Gemini's conversational editing means you can refine in follow-up turns instead of starting over.
- Iterate: a great image usually takes a few rounds of small prompt edits, not one perfect sentence.
What Gemini Can Do for Image Generation
Gemini's image tools go beyond simple text-to-image. As of 2026, three main models cover most needs: Gemini 3.1 Flash Image (fast, free, for volume and everyday use), Gemini 3 Pro Image (professional quality for demanding work), and Gemini 2.5 Flash Image (efficient, lightweight).
The standout features are character consistency and conversational editing. You can keep the same character across multiple images, make precise edits to one part of a picture in a follow-up message, and combine photos into a completely new creation. That turns image generation from a one-shot dice roll into an actual design conversation.
It's available across the Gemini app, AI Studio, and Vertex AI — so whether you're a hobbyist or building an app, the same prompting skills carry over.
The 6-Part Prompt Formula
Google's own guidance boils a strong image prompt down to six elements. You don't need all six every time, but the more you include, the more control you get.
Subject. Who or what is the picture about? Be specific: "a tabby cat wearing round glasses" beats "a cat."
Composition. How is it framed? Close-up, wide shot, head-and-shoulders, rule of thirds.
Action. What is happening? A cat "reaching for a mug" tells the model far more than "a cat."
Location. Where is the scene? A cluttered kitchen, a misty forest, a rooftop at dusk.
Style. What does it look like? Photorealistic, watercolor, 3D render, 1980s anime, cinematic.
Editing instructions. What changes if you're editing? "Change the background to a night sky, keep the cat unchanged."
Weave these into a single descriptive sentence rather than a checklist. Example: "A close-up portrait of a tabby cat wearing round glasses, reaching for a steaming mug, in a cozy sunlit kitchen, watercolor style."
Best Gemini Prompts for Image Generation
Here are copy-and-paste prompts for the most useful scenarios. Adjust the bracketed details to fit your idea.
1. Photorealistic Portrait
For lifelike people and faces.
A photorealistic head-and-shoulders portrait of a [description of person], soft natural window light, 85mm lens, shallow depth of field, blurred [color] background, realistic skin texture, gentle smile, looking at camera.
The camera-language here — "85mm lens, shallow depth of field" — is what pushes Gemini from an illustration to something that looks shot on a real camera.
2. Cinematic Scene
For dramatic, story-like images.
A cinematic wide shot of [scene], dramatic golden-hour lighting, volumetric light rays, moody atmosphere, film grain, anamorphic look, deep shadows, cinematic color grading.
Terms like "film grain" and "volumetric light" tell Gemini to imitate a film still rather than a flat render, which makes the result feel intentional and cinematic.
3. Illustrative / Watercolor Style
For soft, artistic visuals.
A watercolor illustration of [subject], loose brushstrokes, soft pastel palette, paper texture, minimal and elegant, subtle background wash, no text.
Specifying the medium (watercolor), the texture (paper), and the palette (pastel) gives the model a clear art direction it can reliably follow.
4. Product Mockup
For blog images and product shots.
A clean product mockup of [product] on a neutral [color] studio background, soft diffused lighting, subtle shadow, commercial photography style, sharp focus on the product, minimal props.
"Commercial photography style" plus "soft diffused lighting" produces a professional-looking product shot you can use in articles or social posts.
5. Consistent Character Across Images
Keep the same character in a series.
Create a character: [detailed description of appearance, outfit, age]. In the next images, keep this exact character unchanged while I change only the background and action. Start with: [scene 1].
For consistency, describe the character once in detail, then ask Gemini to keep them unchanged across scenes. Character consistency is one of Gemini's headline strengths, and a good character description is what unlocks it.
6. Text in the Image
When you need the image to include readable text.
Generate an image of [scene] with the exact text "[your text]" displayed prominently. Render the text clearly and correctly spelled, in a [style] font, and do not add any other text.
Spelling in AI images used to be unreliable, but Gemini handles in-image text far better now. State the exact text and "do not add any other text" to keep it clean.
How to Iterate and Fix Results
Edit conversationally. Instead of starting over, send a follow-up: "change the background to a night sky, keep the character unchanged." Gemini is built for multi-turn edits.
Add lens language. If the image looks flat, add "85mm portrait lens, shallow depth of field, soft bokeh." If it looks too stiff, add "candid, natural motion."
Fix perspective errors. AI still slips on physics and perspective. Call it out directly: "fix the hand proportions" or "correct the reflection."
Avoid watermark noise. If you get unwanted artifacts, specify "no watermark, no signature, no text."
Choose the right aspect ratio. A wide banner needs a wide ratio, a portrait needs a tall one. Set it before you generate to avoid awkward cropping.
Frequently Asked Questions
Is Gemini image generation free?
Yes — Gemini 3.1 Flash Image (Nano Banana 2) is free to all users and offers near-Pro quality at fast speed. Higher-end Pro models and higher-volume use may have limits or come with paid tiers.
Can I edit an existing photo with Gemini?
Yes. Upload a photo and prompt Gemini to change specific parts — the background, lighting, clothing, or a single object — while keeping the rest intact. Conversational editing lets you refine in follow-up messages.
What does Nano Banana mean?
"Nano Banana" is the nickname for Gemini's image generation models. Nano Banana 2 is the official Gemini 3.1 Flash Image model — free, fast, and high quality. It's a fun codename for a genuinely capable model.
How do I keep the same character in multiple images?
Describe the character once in rich detail, then tell Gemini to keep that character unchanged while you alter only the background and action. This leverages Gemini's character-consistency feature.
Can I use the images commercially?
Check the model's terms of service. Gemini-generated images are generally usable, but usage rights and attribution rules vary by plan and platform, so review the policy for your specific tier.
Conclusion: Master Gemini Image Prompts
Gemini has made high-quality, free image generation accessible to everyone — and with the right prompts, you can turn simple ideas into polished visuals. Remember the six-part formula: subject, composition, action, location, style, and editing instructions. Layer in photography language for realism, and lean on conversational editing to refine instead of restarting.
Start with the six templates above — photorealistic portraits, cinematic scenes, watercolor illustration, product mockups, consistent characters, and in-image text — then iterate until each one lands. With a little practice, Gemini becomes one of the most capable and generous image tools on the market, and it costs you nothing to master it.

