Grom Zvuk
The first music model in the Grom lineup.
How a track is made
First the model floods the whole length of the future song with noise and marks out where the verses go, where the chorus lands and where the bridge sits. Then, over several passes, it turns that noise into sound worth listening to.
There is no partial result: the track arrives whole, once it is finished.
How the pipeline works
Two experts under shared attention: the first writes music as symbols, the second unfolds those symbols into sound.
- yPrompt
- sScore
- cSemantics
- zAcoustic latents
- 48 kHzFinished record
Walks through the music step by step: decides where the melody leads and where the harmony turns.
Works across the full length at once: refines the sound as a whole, not piece by piece.
The diagram is simplified: it explains the principle rather than repeating the model's internals.
Where Grom Zvuk sits among music models
Two axes at once: how the track sounds and how closely the model sings the given lyrics. On lyric accuracy Grom Zvuk is in the top three; on sound it holds its place among the closed services.
sound / lyrics
- Suno v588 / 90
- Mureka 990 / 78
- Grom Zvuk83 / 85
- Suno v680 / 85
- Suno v5.582 / 82
- Suno v4.584 / 78
- ACE-Step 1.567 / 74
- MiniMax Music 370 / 64
- Muse59 / 62
- LeVo 265 / 42
- DiffRhythm 237 / 53
- SongBloom9 / 9
Blind comparison across 150 prompts: identical briefs, finished tracks rated on sound and on how well they follow the given lyrics, with model names hidden from the raters. Values are normalised to a 0–100 scale. Hover a dot to see its name.
What sits behind a generation
48 kHz, stereo
Not a compressed draft but a studio sample rate and a real stereo image.
Tracks up to four minutes
A full song with verses, choruses and an ending, not a fifteen-second loop.
A single pass
Every part of the song is computed together rather than in turn — hence the unified sound.
Our own hardware
The model runs on our own fleet of accelerators, not in someone else's cloud.
Sings in any language
From English to Chinese and Japanese. The model reads the lyrics, the stress and the phrasing — the vocal never turns into accented mush.
The list is open-ended: these are not all the languages the model sings in.
Create, cover, edit
Create
Describe the style and the mood, hand over your lines — the model writes the melody, builds the arrangement and sings it.
Cover
Upload a song and name the new sound: the melody stays recognisable while the arrangement changes completely.
Edit
The intent of the track lives in the score, so it can be changed surgically — swap a section or an instrument without rewriting the song.
Need your own stream of music?
Get in touch — we'll go over volumes, limits and pricing ahead of launch. The model runs on our own hardware, so load and timelines are settled by contract rather than by a queue in someone else's cloud.
Questions about Grom Zvuk
Our other models
Our flagship text model: copy, code, documents and answers to questions.
Images from scratch by description: photorealism, illustration, art.
Changes, removes and adds objects in a photo from a text instruction.
Upscaling to 100 megapixels with sharpness and texture restored.
Video from text and photos: clips for social, ads and presentations.









