AISongMakerLab AI music generator playbooks

One Prompt, Eight AI Music Generators: How to Run the Comparison Yourself

One prompt card fanning out into eight different audio waveform ribbons

Put one identical prompt through every AI music generator you can log into, score each result on six fixed dimensions, and keep the tool that wins the two dimensions you care about most. Published side-by-side write-ups cannot settle this for you, because they tend to use a different prompt per tool, on a different day, judged by a different person. What you need is a controlled test you can re-run whenever a model changes. Below is the whole kit: a neutral prompt, seven variables to hold still, a six-dimension scorecard, and a grid you fill in during one sitting of about ninety minutes.

Published 1 October 2026 · about 8 minutes · no results published here, only the method and the templates to run it

Why one prompt is the only fair test

An AI music generator is a black box: you hand it words, it hands back audio. If you hand each tool different words, you are no longer comparing tools, you are comparing prompts. That is the single biggest reason two people walk away from the same eight tools with opposite rankings.

The trade-off is honesty: a neutral prompt is deliberately average. It will not show off any tool's specialty. That is fine. You are buying comparability, not a demo reel.

The seven variables you hold still

Change any one of these between tools and the comparison stops meaning anything. Print this checklist and keep it next to you while you run the test.

Seven identical slider knobs locked at the same position on a control panel
  1. The prompt text. Character for character, including commas. Save it in a plain text file and paste it, never retype it.
  2. The number of takes. Three per tool. One take is a sample, not a result, because generation carries randomness.
  3. The vocal setting. Instrumental mode for every tool, or lyrics on for every tool. Never mixed.
  4. The target length. Ask for the same duration everywhere, and choose one every tool can actually reach.
  5. The listening chain. Same headphones, same playback level, same room. Speaker versus headphone alone can flip a timbre judgement.
  6. The session length. Cap at twenty minutes, then take a break. Ear fatigue silently drags every later score down.
  7. The file naming. Export as WAV or the highest quality available, then rename files to neutral codes before you listen.

Variables five and six are the ones people skip, and they are the ones that quietly decide the outcome. A tool scored last, after forty minutes of listening, is being scored against tired ears.

The neutral prompt

This prompt is built to be mid-tempo, mid-density, and clearly structured, so that any deviation you hear is the generator's choice rather than your instruction's. Copy it into a text file and paste it untouched into every tool.

Mid-tempo indie pop instrumental, 104 BPM, warm and steady, no vocals,
electric piano + clean guitar + soft drums + subtle pad,
structure: 8-bar intro, verse at 0:15, chorus at 0:40, verse at 1:10, resolved ending by 1:50,
exclude: vocals, exclude: distortion, exclude: key changes, exclude: abrupt cutoff

If a tool has a separate style box and a lyrics box, put everything before the word structure in the style box and leave the lyrics box empty. Instrumental-only tools get the same text minus the vocal line.

What each line actually pins down

LineWhat it locksIf you drop it
Genre plus BPMTempo and style, the two most reliable dialsEvery tool drifts to its own default tempo, and you compare drift instead of quality
Mood adjectiveEmotional colour, which is where tools differ mostResults sound technically fine and emotionally unrelated
Instrument listArrangement densityOne tool fills the mix, another leaves it bare, and the difference is your prompt, not the tool
Structure with timestampsWhether the model follows instructions at allYou cannot measure prompt adherence, the dimension that predicts everything else
Resolve-the-ending instructionUsability of the last five secondsYou discover the cut-off problem only after you have picked a winner
Exclusion linesWhat must not appearA random vocal or a key change ruins one take and you blame the wrong cause

The six-dimension scorecard

Score every take from one to five on each dimension. Resist the urge to add dimensions; six is already the practical limit for reliable listening.

Abstract six-axis radar shape with small sound wave icons near each axis point
DimensionWhat you listen for1 point5 points
Prompt adherenceDid it deliver the genre, tempo, and instruments asked forWrong genre entirelyEvery named element present
Structure controlDo the sections land near the timestamps you asked forOne shapeless loopSections arrive within a few seconds of the plan
Mix qualityBalance, no clipping, no muddy low endAudible clipping or buried melodyClean, balanced, ready to use
Ending usabilityDoes it resolve, or does it stop deadHard cut mid-phraseResolves cleanly and could be faded
Consistency across takesHow close the three takes sound to each otherThree unrelated tracksThree clear variations of one idea
Time to usable resultHow many of the three takes you would actually keepNoneAll three

Consistency is the dimension most people forget, and it is the one that matters most if you release more than one track. A tool that produces one brilliant take and two unusable ones costs you more evenings than a tool that produces three decent ones.

Scoring rules that keep you honest

  1. Rename before you listen. Save each export as a random code and keep the mapping in a separate file until scoring is finished.
  2. Shuffle the order. Never listen to tool one, two, three in the same order you generated them.
  3. Score immediately. Write the number while the take is still playing. Memory smooths out detail within seconds.
  4. Score takes, not tools. Nine takes for three tools means nine rows. Average afterwards, never during.
  5. Take the median of three, not the best of three. The best-of-three rule is how a lucky take turns into a wrong decision.
  6. Stop when tired. If you cannot tell two takes apart, the session is over. Finish another day and keep the old scores.

The grid you fill in

One row per tool, one column per dimension. Keep it as a plain text file so you can reuse it every time a model updates.

tool | adhere | structure | mix | ending | consistency | keep-rate | total
A    |   _    |    _      |  _  |   _    |     _       |    _      |  _
B    |   _    |    _      |  _  |   _    |     _       |    _      |  _
C    |   _    |    _      |  _  |   _    |     _       |    _      |  _
D    |   _    |    _      |  _  |   _    |     _       |    _      |  _
E    |   _    |    _      |  _  |   _    |     _       |    _      |  _
F    |   _    |    _      |  _  |   _    |     _       |    _      |  _
G    |   _    |    _      |  _  |   _    |     _       |    _      |  _
H    |   _    |    _      |  _  |   _    |     _       |    _      |  _
weighting (pick two must-win dimensions): ______________________

Eight rows is the useful ceiling. Past that, listening fatigue does more damage than the extra data does good. If you have access to more tools, run them as a second session on another day with the same prompt.

Turning the grid into a decision

A single total column will mislead you, because it quietly assumes all six dimensions matter equally to you. They do not. Use this step-by-step rule instead:

  1. Pick your two must-win dimensions before you look at the totals. A producer exporting stems cares about structure and mix; someone making background beds cares about consistency and ending usability.
  2. Double those two scores. Add the rest once. Recompute the total.
  3. Break ties with consistency. If two tools are within two points, take the more consistent one. It saves time on every future track.
  4. Check the runner-up on your weakest dimension. If your winner scores two or lower anywhere, keep the runner-up as a backup for that specific job.
  5. Write the date on the grid. Model behaviour changes, and an undated ranking is a rumour by next quarter.

The output of this exercise is not a single winner. It is two tools with clearly different strengths, which is a far more useful thing to know than a ranking with one name on top.

Confounds that fake a winner

What you noticeLikely confoundFix
One tool sounds dramatically betterIt was scored first, on fresh earsShuffle order and re-score the top three
One tool nails the genre every timeIts default style happens to match your promptRe-run with a genre it is unlikely to default to
Everything sounds worse as the evening goes onEar fatigue, not tool qualitySplit into two sessions under twenty minutes each
Two takes from one tool sound unrelatedRandomness, not inconsistency in qualityJudge the median take, not the outlier
A tool wins on headphones and loses on speakersMix translation differencesScore on one listening chain only, and write down which
Longer tracks sound more impressiveLength bias: more time to buildAsk every tool for the same duration

When a tool wins everywhere except one thing

This is the normal outcome, and it is not a reason to keep searching. Decide with a workaround budget: if the single weak dimension can be fixed in two minutes of external work, the tool is still your winner. An abrupt ending can be fixed with a fade or a re-generated tail. A muddy low end can be fixed with a high-pass filter. A tool that ignores your structure line cannot be fixed cheaply, because structure is the thing you would have to rebuild by hand. Judge the gap by the cost of the workaround, not by the score.

Common questions

Do I really need eight generators?

No. Three tools run properly beats eight tools run casually. Eight is the ceiling where the method still works, not the requirement. If you only have two logins, run the same protocol on two and re-run it whenever you get access to a third, keeping the old grids for comparison.

Should the test prompt include lyrics?

Run it twice if vocals matter to you: once instrumental, once with the same short lyric block pasted into every tool. Mixing the two modes inside one test is the fastest way to produce a ranking you cannot interpret, because vocal quality and instrumental quality are separate skills.

Why three takes and not one?

Generation includes randomness, so one take is a single draw from a distribution. Three takes let you judge consistency, which is the dimension that predicts how much time the tool will cost you over a whole project. One take tells you nothing about variance.

What if a generator ignores my structure line?

Score it low on structure control and say so in your notes, then move the structure instruction into whichever field that tool actually respects, and rerun that tool only. A tool that ignores timestamps is not automatically worse, but it does change how you will have to write prompts for it.

How often should I re-run the whole test?

Whenever a generator announces a model change, and otherwise once every three months. Keep every grid file with its date. A stack of dated grids is more informative than any single ranking, because it shows you the direction each tool is moving.

Bottom line: paste the neutral prompt above into every generator you have, three takes each, six dimensions, median of three, and double the two dimensions you care about. Ninety minutes of that gives you a ranking you can defend, and a template you can re-run whenever the models move.