One Prompt, Eight AI Music Generators: How to Run the Comparison Yourself

Put one identical prompt through every AI music generator you can log into, score each result on six fixed dimensions, and keep the tool that wins the two dimensions you care about most. Published side-by-side write-ups cannot settle this for you, because they tend to use a different prompt per tool, on a different day, judged by a different person. What you need is a controlled test you can re-run whenever a model changes. Below is the whole kit: a neutral prompt, seven variables to hold still, a six-dimension scorecard, and a grid you fill in during one sitting of about ninety minutes.
Why one prompt is the only fair test
An AI music generator is a black box: you hand it words, it hands back audio. If you hand each tool different words, you are no longer comparing tools, you are comparing prompts. That is the single biggest reason two people walk away from the same eight tools with opposite rankings.
- Same input, different output means the tool differs. That is the only conclusion a controlled test can support.
- Different input, different output means nothing. A lush result from a long prompt and a flat result from a short one tells you about prompt length, not about the engine.
- One prompt scales. When a generator ships a new model version, you re-run the same prompt and the scorecard tells you what moved, in your own numbers.
The trade-off is honesty: a neutral prompt is deliberately average. It will not show off any tool's specialty. That is fine. You are buying comparability, not a demo reel.
The seven variables you hold still
Change any one of these between tools and the comparison stops meaning anything. Print this checklist and keep it next to you while you run the test.

- The prompt text. Character for character, including commas. Save it in a plain text file and paste it, never retype it.
- The number of takes. Three per tool. One take is a sample, not a result, because generation carries randomness.
- The vocal setting. Instrumental mode for every tool, or lyrics on for every tool. Never mixed.
- The target length. Ask for the same duration everywhere, and choose one every tool can actually reach.
- The listening chain. Same headphones, same playback level, same room. Speaker versus headphone alone can flip a timbre judgement.
- The session length. Cap at twenty minutes, then take a break. Ear fatigue silently drags every later score down.
- The file naming. Export as WAV or the highest quality available, then rename files to neutral codes before you listen.
Variables five and six are the ones people skip, and they are the ones that quietly decide the outcome. A tool scored last, after forty minutes of listening, is being scored against tired ears.
The neutral prompt
This prompt is built to be mid-tempo, mid-density, and clearly structured, so that any deviation you hear is the generator's choice rather than your instruction's. Copy it into a text file and paste it untouched into every tool.
Mid-tempo indie pop instrumental, 104 BPM, warm and steady, no vocals, electric piano + clean guitar + soft drums + subtle pad, structure: 8-bar intro, verse at 0:15, chorus at 0:40, verse at 1:10, resolved ending by 1:50, exclude: vocals, exclude: distortion, exclude: key changes, exclude: abrupt cutoff
If a tool has a separate style box and a lyrics box, put everything before the word structure in the style box and leave the lyrics box empty. Instrumental-only tools get the same text minus the vocal line.
What each line actually pins down
| Line | What it locks | If you drop it |
|---|---|---|
| Genre plus BPM | Tempo and style, the two most reliable dials | Every tool drifts to its own default tempo, and you compare drift instead of quality |
| Mood adjective | Emotional colour, which is where tools differ most | Results sound technically fine and emotionally unrelated |
| Instrument list | Arrangement density | One tool fills the mix, another leaves it bare, and the difference is your prompt, not the tool |
| Structure with timestamps | Whether the model follows instructions at all | You cannot measure prompt adherence, the dimension that predicts everything else |
| Resolve-the-ending instruction | Usability of the last five seconds | You discover the cut-off problem only after you have picked a winner |
| Exclusion lines | What must not appear | A random vocal or a key change ruins one take and you blame the wrong cause |
The six-dimension scorecard
Score every take from one to five on each dimension. Resist the urge to add dimensions; six is already the practical limit for reliable listening.

| Dimension | What you listen for | 1 point | 5 points |
|---|---|---|---|
| Prompt adherence | Did it deliver the genre, tempo, and instruments asked for | Wrong genre entirely | Every named element present |
| Structure control | Do the sections land near the timestamps you asked for | One shapeless loop | Sections arrive within a few seconds of the plan |
| Mix quality | Balance, no clipping, no muddy low end | Audible clipping or buried melody | Clean, balanced, ready to use |
| Ending usability | Does it resolve, or does it stop dead | Hard cut mid-phrase | Resolves cleanly and could be faded |
| Consistency across takes | How close the three takes sound to each other | Three unrelated tracks | Three clear variations of one idea |
| Time to usable result | How many of the three takes you would actually keep | None | All three |
Consistency is the dimension most people forget, and it is the one that matters most if you release more than one track. A tool that produces one brilliant take and two unusable ones costs you more evenings than a tool that produces three decent ones.
Scoring rules that keep you honest
- Rename before you listen. Save each export as a random code and keep the mapping in a separate file until scoring is finished.
- Shuffle the order. Never listen to tool one, two, three in the same order you generated them.
- Score immediately. Write the number while the take is still playing. Memory smooths out detail within seconds.
- Score takes, not tools. Nine takes for three tools means nine rows. Average afterwards, never during.
- Take the median of three, not the best of three. The best-of-three rule is how a lucky take turns into a wrong decision.
- Stop when tired. If you cannot tell two takes apart, the session is over. Finish another day and keep the old scores.
The grid you fill in
One row per tool, one column per dimension. Keep it as a plain text file so you can reuse it every time a model updates.
tool | adhere | structure | mix | ending | consistency | keep-rate | total A | _ | _ | _ | _ | _ | _ | _ B | _ | _ | _ | _ | _ | _ | _ C | _ | _ | _ | _ | _ | _ | _ D | _ | _ | _ | _ | _ | _ | _ E | _ | _ | _ | _ | _ | _ | _ F | _ | _ | _ | _ | _ | _ | _ G | _ | _ | _ | _ | _ | _ | _ H | _ | _ | _ | _ | _ | _ | _ weighting (pick two must-win dimensions): ______________________
Eight rows is the useful ceiling. Past that, listening fatigue does more damage than the extra data does good. If you have access to more tools, run them as a second session on another day with the same prompt.
Turning the grid into a decision
A single total column will mislead you, because it quietly assumes all six dimensions matter equally to you. They do not. Use this step-by-step rule instead:
- Pick your two must-win dimensions before you look at the totals. A producer exporting stems cares about structure and mix; someone making background beds cares about consistency and ending usability.
- Double those two scores. Add the rest once. Recompute the total.
- Break ties with consistency. If two tools are within two points, take the more consistent one. It saves time on every future track.
- Check the runner-up on your weakest dimension. If your winner scores two or lower anywhere, keep the runner-up as a backup for that specific job.
- Write the date on the grid. Model behaviour changes, and an undated ranking is a rumour by next quarter.
The output of this exercise is not a single winner. It is two tools with clearly different strengths, which is a far more useful thing to know than a ranking with one name on top.
Confounds that fake a winner
| What you notice | Likely confound | Fix |
|---|---|---|
| One tool sounds dramatically better | It was scored first, on fresh ears | Shuffle order and re-score the top three |
| One tool nails the genre every time | Its default style happens to match your prompt | Re-run with a genre it is unlikely to default to |
| Everything sounds worse as the evening goes on | Ear fatigue, not tool quality | Split into two sessions under twenty minutes each |
| Two takes from one tool sound unrelated | Randomness, not inconsistency in quality | Judge the median take, not the outlier |
| A tool wins on headphones and loses on speakers | Mix translation differences | Score on one listening chain only, and write down which |
| Longer tracks sound more impressive | Length bias: more time to build | Ask every tool for the same duration |
When a tool wins everywhere except one thing
This is the normal outcome, and it is not a reason to keep searching. Decide with a workaround budget: if the single weak dimension can be fixed in two minutes of external work, the tool is still your winner. An abrupt ending can be fixed with a fade or a re-generated tail. A muddy low end can be fixed with a high-pass filter. A tool that ignores your structure line cannot be fixed cheaply, because structure is the thing you would have to rebuild by hand. Judge the gap by the cost of the workaround, not by the score.
Common questions
Do I really need eight generators?
No. Three tools run properly beats eight tools run casually. Eight is the ceiling where the method still works, not the requirement. If you only have two logins, run the same protocol on two and re-run it whenever you get access to a third, keeping the old grids for comparison.
Should the test prompt include lyrics?
Run it twice if vocals matter to you: once instrumental, once with the same short lyric block pasted into every tool. Mixing the two modes inside one test is the fastest way to produce a ranking you cannot interpret, because vocal quality and instrumental quality are separate skills.
Why three takes and not one?
Generation includes randomness, so one take is a single draw from a distribution. Three takes let you judge consistency, which is the dimension that predicts how much time the tool will cost you over a whole project. One take tells you nothing about variance.
What if a generator ignores my structure line?
Score it low on structure control and say so in your notes, then move the structure instruction into whichever field that tool actually respects, and rerun that tool only. A tool that ignores timestamps is not automatically worse, but it does change how you will have to write prompts for it.
How often should I re-run the whole test?
Whenever a generator announces a model change, and otherwise once every three months. Keep every grid file with its date. A stack of dated grids is more informative than any single ranking, because it shows you the direction each tool is moving.
Bottom line: paste the neutral prompt above into every generator you have, three takes each, six dimensions, median of three, and double the two dimensions you care about. Ninety minutes of that gives you a ranking you can defend, and a template you can re-run whenever the models move.