How to Use Flow AI
Eight steps from a blank project to a sequence that cuts together. Flow costs credits per generation, so this walkthrough is written around spending them well rather than iterating blindly until something works.
Flowain · Last reviewed 2026-09-03
What you need before your first generation
Three things worth settling first, because each one costs credits to discover later.
A Google account
Flow is free to start. Any signed-in Google account gets 50 credits a day, plus 100 starter credits. That is two Veo 3.1 Fast generations daily, or five on Lite. Paid plans add monthly credits on top rather than unlocking features.
One shot, described
Not a scene, not a story. One continuous camera move with a clear subject. If your description contains the word “then”, you are describing two shots.
An editor for afterwards
Flow produces clips, not finished pieces. Assembly, grading, and titles happen in whatever editing software you already use.
Eight steps to your first Flow sequence
Follow these in order. Steps seven and eight are where most first sessions go wrong.
Decide the shot before you open the tool
The single biggest source of wasted credits is opening Flow with a vague idea and prompting your way toward one. Decide four things first: who or what is in frame, what they are doing, where the camera is, and how it moves. If you cannot answer those, you are not ready to spend a generation.
Write it as one sentence. If your sentence describes more than one camera move, it is more than one shot, and it should be more than one generation.
Open Flow and check your plan
Flow lives on Google Labs. Sign in with any Google account and you get 50 credits a day free, plus 100 starter credits. You only need a paid plan if the daily allowance is too small for your pace.
Check which model is selected before generating, because it decides the price. Veo 3.1 Lite is 10 credits, Fast is 20, Quality is 100. Gemini Omni Flash runs from 4 to 15 depending on length and resolution.
Choose your generation mode
Text to Video takes a written description. Frames to Video takes a first frame, or a first and last, and generates the motion between them. Ingredients to Video reuses a saved character or object across shots. Extend appends footage to a Veo clip, but only on Lite.
Mode choice is a cost decision as much as a creative one. Frames to Video usually needs fewer attempts because you have already fixed the composition, so it burns fewer credits on retries.
Write the prompt in camera language
Veo 3.1 responds to the vocabulary of physical filmmaking far more reliably than to adjectives. Shot size, lens behaviour, camera movement, and light direction all steer the output. Words like beautiful, stunning, and cinematic mostly do not.
A workable structure is: subject, action, setting, light, then camera. "A woman in a red coat walks toward camera through a rain soaked street at night, neon reflections underfoot, slow dolly back, shallow depth of field." Every clause in that sentence is doing work.
Set camera, aspect ratio, and quality
Camera controls cover angle, perspective, and movement as direct settings rather than prompt text. Use them instead of describing the same thing twice, since the explicit control is more reliable than the written instruction.
Pick your model with the credit cost in mind. On 50 free credits a day, Lite at 10 lets you try five ideas where Fast at 20 gives you two. Save Quality at 100 credits for a shot you have already committed to.
Generate and judge honestly
Watch the clip twice before deciding. First for whether the composition and motion are what you asked for, second for the failure modes Veo is prone to: hands, background faces, and any text that should be readable in frame.
If the shot is wrong in a way that a prompt tweak will not fix, change the mode rather than rewording. Three failed text to video attempts usually means the composition needs pinning down with Frames to Video.
Reuse the character, do not re-describe it
For continuity across separate shots, save the character as an ingredient and call it with @name in the next prompt. Google describes Ingredients to Video as the most reliable way to reuse the same character and objects from shot to shot. To lengthen one clip instead, use Extend, remembering it only works on Veo 3.1 Lite at 8 seconds.
This is the step people skip, and skipping it is why so much AI video looks like disconnected fragments. Describing the same person again in a fresh prompt will give you a different person; referencing the saved ingredient will not.
Assemble, then upscale last
Upscale only once a sequence is locked. 1080p is free on paid plans and unavailable on the free tier; 4K costs 50 credits and is Ultra only. Upscaling a shot you later cut is spending on footage that does not ship.
Assemble in Scenebuilder: arrange the clips, reorder them, trim heads and tails, preview, and download the scene. For multi track work, grading, or titles, take the export into a real editor.
The camera terms Veo actually responds to
Trading adjectives for filmmaking vocabulary is the fastest single improvement you can make.
Shot size
Wide, medium, close up, extreme close up. This sets how much of the subject fills the frame and is the term the model follows most reliably.
Camera movement
Dolly in or out, pan, tilt, tracking, handheld, locked off. Naming the move gives you far more control than asking for something dynamic.
Lens behaviour
Shallow depth of field, deep focus, wide angle, telephoto compression. These change the feel of a shot more than any style word will.
Light direction
Backlight, side light, top light, golden hour, practical sources in frame. Direction matters more than mood adjectives like moody or dramatic.
Compare two prompts for the same idea. “A beautiful cinematic shot of a woman walking in the rain, stunning atmosphere, 4K” gives the model almost nothing to act on. Every word in it is a value judgement rather than an instruction.
“Medium shot, woman in a red coat walks toward camera through a rain soaked street at night, neon signs reflecting in puddles, slow dolly back, shallow depth of field” specifies framing, subject, action, setting, light source, camera move, and lens behaviour. The second prompt is not longer for the sake of it, it is longer because each clause removes a decision the model would otherwise make for you.
Resolution belongs in the settings, not the prompt. Asking for 4K in text does nothing, because output quality is a generation setting.
Worked shot prompts are on the Flow AI examples page.
Where beginners waste credits
Five patterns that cost people generations in their first week.
Re-describing a character instead of saving it
The most expensive mistake. A fresh prompt cannot inherit your character. Save it as an ingredient and reference it with @name.
Describing a whole scene in one prompt
Eight seconds holds one camera move. A prompt containing a sequence of events produces a muddle, and you pay full price for it.
Rewording when the mode is wrong
If three text attempts have missed the composition, the fourth will too. Switch to Frames to Video and fix the endpoints yourself.
Generating at high quality while exploring
Explore on Lite or Omni Flash. Quality costs 100 credits, which is two full days of the free allowance for one clip.
Upscaling before the edit is locked
Upscaling a clip you later cut is pure waste. Lock the sequence first, then upscale what survived.
Asking for readable text in frame
Signage, labels, and captions remain unreliable. Add text in your editor afterwards rather than burning generations hoping it resolves.
How to get more out of a month of credits
Practical habits that stretch a fixed credit pool further.
Keep a note of prompts that worked and what the output looked like. Veo responds consistently enough to phrasing that a personal library of working structures saves real money over a few months. Most people rebuild the same prompt from scratch every time and pay for the rediscovery.
Batch your exploration. Deciding on four shots and generating them in one session is cheaper than returning to the tool eight times, because you carry context between attempts instead of starting cold.
Generate your establishing shot last. It is the one most likely to change once you see how the rest of the sequence cuts together, and generating it first usually means generating it twice.
Where a shot exists mainly to bridge two compositions you already have, Frames to Video is almost always cheaper than describing the bridge in text and hoping.
A full feature by feature breakdown is on the Flow AI features page, and what is Flow AI covers the model and pricing background. If you are bringing Flow into an existing design or previsualization pipeline rather than starting from scratch, Flow AI for designers covers where it sits alongside the tools you already run.
Questions about using Flow AI
The practical questions that come up in a first session.
How long does a first Flow session take?+
Budget about fifteen minutes for a first useful result. A single generation returns in a couple of minutes, but you will spend two or three attempts learning how the model reads your prompt. The free tier gives you 50 credits a day, which is two Fast generations, so a first session may well use your whole daily allowance.
How many credits will I spend learning?+
Veo 3.1 Lite costs 10 credits, Fast costs 20, Quality costs 100. Most people spend 60 to 100 credits before landing a shot worth keeping. On the free 50 a day that is two or three days of learning, which is why starting on Lite, or on Gemini Omni Flash at 4 to 15 credits, stretches the allowance much further than starting on Fast.
Why does my character change between clips?+
Because you described the character again instead of referencing a saved one. A fresh prompt has no memory of the previous clip. Save the character as an ingredient and call it with @name, or use your own avatar with @me. Google describes Ingredients to Video as the most reliable way to reuse the same character and objects from shot to shot, and it fixes almost every continuity complaint.
Should I use text to video or Frames to Video?+
Use Frames to Video whenever you already know the composition you want, because fixing both endpoints removes most of the uncertainty and therefore most of the wasted retries. Use text to video when you want the model to propose a composition you have not settled on yet.
Can I get a video longer than eight seconds?+
Not from one generation. Veo 3.1 offers 4, 6, and 8 seconds; Gemini Omni Flash reaches 10. Extend appends more footage to an existing Veo clip, but only on the Lite model and only at 8 seconds, which catches out anyone working on Fast. Longer pieces come from generating several shots, keeping the character consistent with ingredients, then assembling them in Scenebuilder.
Flowain
Writer and editor, FlowAIFX
Flowain writes and maintains FlowAIFX
Every page is checked against Google's published specifications for Flow and Veo 3.1 on the date shown, and corrected when Google changes them. Where a limit is undocumented or a credit cost is easy to misjudge, it is written down rather than left out.
Last reviewed 2026-09-03 · Report a correction
Open Flow and generate your first shot
Write one shot in camera language, generate it on a Fast pass, then extend it once. That single loop teaches you more than reading another walkthrough will.
