On September 18 we released Grom Zvuk - our own music model that makes a whole song: vocals, arrangement, mix. The first messages in the comments and in my inbox were not questions but an accusation: "you scoundrels, taking bread from musicians who are poor enough already". I'm answering at length because we have something most people in this argument don't - data. Tens of thousands of generations a day go through GPTunneL's music models, and they show clearly what people actually do with a music model. Short answer: AI won't replace musicians. The long one is below: four tracks, some comparisons, and the story of how I played in a band myself.
Who generates music tens of thousands of times a day, and why
The requests fall into four groups, and none of them looks like "fire the musician".
The first and biggest - music where there never was a musician. A song for a colleague's birthday, a greeting for parents, an anthem for a group chat, a track for a short video, background for a presentation. None of these people were going to hire an arranger and pay for a studio. Before, they took someone else's track from a library or went without music; now they have their own.
The second - a sketch of an idea. These are the musicians and producers. There's a lyric and a feeling of how it should sound, but it's unclear where to go: ballad or punk, guitars or synths. You generate five versions in five styles and listen for where the idea comes alive. It's a sketch, not a release: from there you go to the studio or your DAW your own way.
The third - covers and rearrangements. Hearing your own song in reggae or with a big band, or taking a well-known melody and putting your own words to it. One of the most popular scenarios, and judging by the wording, people here are mostly playing, not working.
The fourth - applied music: jingles for podcasts and small businesses, an instrumental for a video, ambient for an app. This used to be the stock-library niche, and yes, generation is squeezing it.
Notice: there's no "write a hit instead of the artist" on this list. Not because the model is forbidden to, but because that's not what the requests look like. And that's no accident: a hit takes talent. Typing "make a hit" into a model is not enough - it will return a song, but what makes it a hit is what isn't in the prompt: an idea, taste, and a person with something to say.
What the model can do and what a musician can do
Grom Zvuk works in two stages: a planner writes the score in symbols (notes and structure - where the verse is, where the chorus is, what the harmony is), and a synthesizer unfolds it into 48 kHz stereo sound, up to five minutes in a single pass. That's craft: arrangement, performance, mixing. The model does it fast and predictably. What the song is about, for whom and why - that it doesn't decide.
| Task | AI model | Musician |
|---|---|---|
| Decide what the song is about and what matters in it | No: works only with what it's given | Yes, that's what authorship is |
| Write an arrangement for a lyric | 1–2 minutes | Hours to days |
| Move a finished song into another genre | Minutes | Days of rehearsal and re-recording |
| Hear that a take is bad and explain why | No, it has no taste | Yes - that takes years to learn |
| Play on stage in front of a crowd | No | Yes |
| Put your own name behind the result | No | Yes |
The stages that got cheap are the ones musicians delegated anyway: to session players, arrangers, sound engineers. What stayed expensive is the idea, taste and selection. The model hands you options, and a person decides which one is the song. If there's nothing to decide, the tool won't help - you get slop, more on that below.
How it was in the 2000s and how it is now

In the 2000s I played electric guitar in clubs, so when I say "before", it's not from articles. A rehearsal room by the hour, or more often a garage - half the price of a proper studio. We recorded, at best, at someone's flat through the line-in of an old Pentium 4. A MacBook on Mac OS X Leopard was an unthinkable luxury, but it let you record a track without mains hum. More often we just recorded onto a cassette deck in one take - so we'd have something to show the club manager. A studio was out of our budget, and it would have made everything many times slower anyway: recording, then mixing. In my whole history we recorded exactly one track in a studio, and laying down all the parts for three minutes took almost a full day without breaks. The recordings didn't survive, sadly, but it was quite the mess. My own guitar part alone I re-recorded hundreds of times: every take meant missed notes, mistakes, and starting over.
Now in Grom Zvuk the song's idea lives in the score, and you can edit it in parts: change the genre, drop the vocals, swap an instrument, put in different lines. The result comes back in a minute or two as a finished track. We would have dreamed of this, if we could have imagined it would ever be possible: press a button and hear the same song without the wall of guitars but with brass. Back then the hard drive inside an iPod already felt like technology from another planet.
The songs themselves, the riffs and lyrics we wrote in the kitchen, wouldn't have gotten better from a button. The button shortens the road from idea to check; it doesn't replace the idea.
Four tracks: one model, four deliveries
Four examples from the Grom Zvuk page, all made by the model end to end, no manual mixing. They're one-shot - the very first take as it came out of the model, so there are rough edges: a wrong stress, an off pronunciation. All of that is fixable, but the tracks are here to show what you get by just throwing in an idea and pressing a button.
"Thunder with a Pepsi Flavour" is an original lyric, and the melody and arrangement with a touch of Lana Del Rey the model built from a reference I fed it. "Do You Know?", "Shade" and "Smart Assistant" are covers: the melody and the recognizable core of the original are there, but everything else is fully rethought - different instruments, a different approach to delivery. Same score, a completely different song. That is the practical point of generative music today: test an idea in several styles and deliveries in a matter of minutes, not weeks of rehearsals. Which version is "right", the model doesn't know. That's decided by whoever had the idea.
Technology has scared musicians every twenty years
In the 1930s, when sound came to cinemas, the American Federation of Musicians ran a campaign against "canned music": the orchestras that played under silent films really did lose their jobs. Film music didn't go anywhere, though - it started being recorded, and more film composers were needed, not fewer.
In the early 1980s the British Musicians' Union tried to restrict synthesizers and drum machines in recordings: drummers were told their end was near. The Roland TR-808 came out in 1980, and its sound became the foundation of hip-hop and electronic music - genres that gave work to a whole generation of new musicians. Live drummers stayed; they just stopped being the only way to get a beat.
In 1998 Cher released "Believe" with deliberately audible Auto-Tune, and the argument "is this still singing" ran for about ten years. Today Auto-Tune is a standard line in any production, and singers who sing in tune are valued more, not less.
Home studios in the 2000s did the same to recording: fewer studios, and orders of magnitude more musicians releasing tracks.
The pattern is the same: technology makes the craft cheaper, the entry barrier drops, there's more music, and the people with taste and their own voice only stand out more against that background. Evolution can't be stopped, and there's no need to: it has never removed the musician, it removed the monopoly on the instrument.
Slop is not a property of the tool
"AI slop" has become a search query of its own, and I understand where it comes from: there's a lot of faceless generated music out there - open any streaming service and it's full of it, and yes, that needs to be fought. But slop doesn't come from the model; it comes from the absence of an idea and of selection. Someone pressed the button once, took the first result, uploaded it. That track has no author, and you can hear it.
The same tool in the hands of someone with an idea gives something else. They write their own lyric instead of asking the model for "something about love". They run ten versions and throw away nine. They hear that the second verse sags and fix that part, instead of regenerating everything. They know how the chorus should sound because they heard it in their head before generating. The electric guitar in the 1950s was a toy to some and a new language to others - and the difference was not in the guitar.
The formula is simple: AI is a tool. To someone with talent it's a faster road from idea to result and a way to try what there was never time or money for. To someone without an idea it honestly returns a result without an idea.
Who it will actually replace
To not sound like a sales pitch, here's the unpleasant part too. Pressure is already on stock libraries and the part of the market where the client just needs "a track in three days": a jingle for a local car dealer, a bed under a corporate video. Generation does that work faster and cheaper, and that demand is moving to models - our tens of thousands of generations a day show it already.
Not under pressure is everything that rests on personality: an artist with an audience, a live show, a writer with a recognizable voice, a producer with taste. These people have more tools than any generation before them. An arranger who used to write three demos a week now offers a client twenty versions in an evening and spends time on what they're actually paid for - choosing and polishing.
What we're going to do about it
At GPTunneL we invest heavily in every direction of generative models: text, images, video, sound. Music is one of the most important of them. Grom Zvuk runs on our own hardware, and a dedicated team of music-model specialists works on it: we build this in-house rather than reselling someone else's API. We'll keep developing the model: more accurate singing of the lyric, more control over the score, editing individual parts of a track.
Checking all this on your own idea is simple: open Grom Zvuk, describe the style, give it your lyric and listen to what comes out. At launch a track costs $0.04 - with a 50% discount. And you'll test the article's main point along the way: the model will hand you a song, but whether it's a song is still yours to decide.



