Video model

Veo by Google — video with sound in one render

Google DeepMind Veo models in GPTunneL: video with sound in the same render, spoken lines in frame, clips from a photo. Pay per second, no VPN.

What the Veo family is good at

Traits Google DeepMind keeps from one generation to the next — you can count on them whichever model you pick.

Sound together with the picture

This is what the family is chosen for: dialogue, footsteps, street noise and music arrive in the same render as the footage. There is no separate audio pass.

Lines written into the prompt

What the character says on screen goes into the same request: the model speaks the line and the lips match the audio.

A cinematic frame

Light, depth of field and camera movement in Veo are closer to filming than to animation — which is why the family is used for ad inserts and presentation clips.

Video from text and from a frame

The model works both ways: it builds a scene from a description from scratch, or brings a finished image to life — a photo, a concept, a product shot.

Which model to pick

The family has three tiers — for different tasks and budgets.

Veo 3.1 [Lite]Drafts and idea sweepsVeo 3.1 [Fast]The all-round choiceVeo 3.1Maximum quality
Best forSketching a scene, testing an idea, collecting takesClips for social media and ad insertsComplex scenes, spoken lines, a cinematic frame
Sound in the renderYesYesYes
Frame detailMediumHighMaximum
SpeedHighestHighMedium
Price per secondLowestLowHighest

Veo is billed per second of finished video. Exact prices for every model are on the Pricing

Every Google video model

Release dates follow Google DeepMind's official announcements.

May 2026

Gemini omni

Google's new family alongside Veo: one model instead of a set of separate ones. It takes text, images, audio and video in a shared context, and the finished clip can be edited further in plain words without rebuilding the scene from scratch.

October 2025

Veo 3.1Veo 3.1 [Fast]Veo 3.1 [Lite]

The current generation: it follows the script more precisely, sounds richer and can lean on reference frames. Alongside it come the faster Fast and the lighter Lite — the same generation for noticeably less per second.

May 2025

Veo 3Veo 3 [Fast]

The generation where the family started generating sound together with the picture — speech, ambience and effects in the same render as the footage. Nothing else could do that at the time.

December 2024

Veo 2

A clear step up in motion physics and resolution: the model stopped melting objects on long takes and learned to hold the camera move it was given.

May 2024

Veo

Google DeepMind's first video model, shown at the I/O conference. It is where the company entered the video generation race.

How billing works

There is no subscription: you top up one balance and spend it on any model on the platform.

Pay per second of video

You are charged for the length of the finished clip and the model you chose. Not using it costs nothing: no limits, no monthly fee.

One balance for every model

Veo, ChatGPT, Claude, image and music generation — all out of the same wallet. No separate subscription per service.

Top up the way you prefer

International cards, Apple Pay, Google Pay or crypto. The minimum top-up is $5.

How to get started with Veo

Sign in to GPTunneL
One account for every model on the platform.
A woman calmly reads a book on a sofa while water pours through the bright room from every side — down the walls, the windows and the floor
8 seconds1080pVeo 3.1

Sign in whichever way suits you.

The model switches right inside the input — you can change it between clips.

The clip arrives with its soundtrack — on this page it plays muted.

Try Veo in GPTunneL

Signing up takes a minute, and $5 on the balance is enough to see whether the model fits your task.

Frequently asked questions

No. Requests go through GPTunneL's infrastructure, so Google models open like any ordinary website — no VPN and no proxy. You do not need a Google account either.

Veo is Google DeepMind's family of video models. The model builds a short clip from a written description, or brings a finished frame to life: a photograph, a concept, a product shot.

The family's main distinction is sound. Veo generates speech, footsteps, street noise and music in the same render as the picture, so the clip arrives already scored and the character's lips match the line. The line itself can be written straight into the scene description. In terms of the frame, the family is closer to filming than to animation: light, depth of field and camera movement look like a cinematographer's work, and scenes like these are used for ad inserts, presentation clips and storyboards.

Veo models are available in GPTunneL — worldwide, without a VPN and without a Google account, with one balance and billing for the seconds of video you generate.