What happens in the minute after you press Create.
You type one line, and about a minute later there is a song with a singer, a mix and a cover. It feels like magic and it is really a chain of decisions, each of which you can influence from the prompt. Knowing the chain is the difference between hoping and steering. This is the non-technical tour of the Octa engine, Octa Initium, with a note at each stage on what you can control.
01Reading the brief
The first thing the engine does is turn your sentence into a plan: genre, tempo range, mood, instruments, voice type, language and subject. Anything you did not specify gets filled in from the genre’s conventions. This is why “a rap song” gives you the most average rap song imaginable and “boom-bap about my grandmother’s kitchen, jazzy piano sample, laid-back male flow” gives you something with a face. You control this stage entirely.
02Planning the structure
Next comes the skeleton: intro, verses, chorus placement, bridge, outro, and the length. Genre conventions decide the default (pop wants the chorus early; folk tolerates a long verse). Your lyrics override the default when you use [Verse] and [Chorus] tags, which is why tagged lyrics come out better arranged. Words like “short”, “hook first” or “no intro” also land here.
03Writing the words
If you did not supply lyrics, the engine writes them from the subject and mood in your brief, in the language you wrote the brief in. The quality of these words is proportional to the concreteness of your subject: a name, a place and one object beat an abstract feeling. If you did supply lyrics, they are used exactly, which is why singable lines matter (see the nine rules).
Every vague word in the prompt becomes an average choice in the song.
04Composing and arranging
Melody, chords, drums and instrumentation are generated together against the plan. Tempo words (“slow”, “140 BPM”), instrument names (“Rhodes piano”, “sliding 808s”) and production words (“lo-fi”, “polished”, “raw”) steer this stage. This is also where “instrumental” removes the singer entirely.
05Performing the vocal
The vocal is a performance, not a text-to-speech read: phrasing, breaths, harmonies and ad-libs are generated to fit the melody. Voice choice (female, male, or your cloned MY VOICE model) and character words (“raspy”, “whispered”, “belting”) shape it. When a line comes out rushed, it is almost always a syllable-count problem in the words rather than a performance problem.
06Mix, master and cover
Levels, EQ, compression and stereo width are set automatically, and the result is mastered to a streaming-ready loudness at 48 kHz stereo. The cover art is painted from the same brief, which is why a song about the sea gets a blue cover. You can replace the art later; you cannot remix the stems inside the engine, but you can split them afterwards with the stems tool.
07When something goes wrong
Most failures are refunded automatically: if the engine cannot finish, the credit comes back within seconds. The most common “bad result” is not a failure at all; it is a plan built from a vague brief. Fix the brief, change one thing at a time, and use the section editor for the one bar that bothers you instead of regenerating the whole song.
08Questions
No. It composes new material from your brief; it does not stitch samples of other recordings. Covers of protected commercial recordings are screened and stopped.
Generation is not deterministic: the same brief can be planned and composed differently each run. Lock in what matters (tempo, key instruments, your own lyrics) to reduce the spread.
After generation, the stems tool separates vocals from the instrumental. Stem splits are included in Standard and Premium plans.
Now you know the chain. Pull it.
Write a brief that names genre, mood, subject and one detail, and hear the difference in a minute.
Your first song is free · No card required