MiniMax released Music 2.6, and instead of leading with benchmarks, they told four stories. It's the right framing for a music model upgrade: the improvements don't live in spec sheets — they live in "this time, someone used it to do something they couldn't do before." A flamenco dancer who needs silence between notes. An indie game dev who needs boss-fight bass that hits. A cafe owner who needs four hours of background music that's good enough to notice but chill enough to ignore. A daughter who wants to remake her mom's favorite song.

The three technical breakthroughs enabling these stories: Cover mode (melodic extraction + genre transfer), completely reworked low-end frequency response, and intentional vocal/melodic imperfection. Each one solves a specific limitation that made previous AI music feel like "AI music."

What's New

Cover Mode: Same Melody, Everything Else Is New

Cover is the headline feature and it's genuinely new. Upload any track. Music 2.6 extracts the melodic skeleton — the core progression that makes the song recognizable from the first bar — then lets you rebuild everything around it:

This isn't remix or mashup. It's constraint-aware generation: the model knows exactly which elements must stay (melodic identity) and which are free to change (everything else). Previous AI music generation was either fully prompted (no melodic constraint) or simple style transfer (same song, slightly different sound). Cover operates in the middle ground that's actually useful.

The use case that sells it: take a song someone already loves and make it personal. Birthday surprises. Wedding receptions. Memorial tributes. The emotional weight comes from recognition plus novelty — "that's our song, but I've never heard it like this."

Architecture

Redesigned Low End

Music 2.6 completely reworked its low-frequency response. Previous versions had bass that sounded acceptable on laptop speakers but fell apart on decent headphones, car systems, or gaming rigs. The sub-bass turned to mush — boomy, undefined, no chest impact.

The fix is technical but the result is physical: bass that you feel before you hear. Drums that are tight instead of boomy. Sub frequencies that maintain definition across different playback systems.

This matters disproportionately for game scoring (boss fights need bass that hits), electronic music (drop sections lose all impact without proper sub), and any genre where the rhythm section carries physical weight (hip-hop, metal, EDM).

For the indie game developer use case: you can now prompt "start oppressive, build toward power, erupt into invincibility" and get a boss-fight soundtrack that actually sells the power fantasy through audio. No sample library. No composer. No five-figure invoice.

Intentional Imperfection

The subtlest improvement might be the most important: Music 2.6 models vocal and melodic imperfection intentionally. In the right genres — lo-fi, indie folk, indie jazz — perfection is the problem. A voice that hits every note precisely sounds generated. A voice that's slightly casual, slightly off-beat, slightly breathy sounds lived-in.

For the cafe playlist use case: Music 2.6 can generate four hours of music that threads the needle between "good enough to notice" and "chill enough that nobody minds." The prompt is as simple as "late-night feel, urban, slightly tipsy, not too bright." The warmth comes from rolled-off highs, imperfect vocal timing, and bass/drums that carry equal weight to the melody.

This is a deliberate modeling choice. Previous AI music optimized for technical perfection — every note in tune, every beat on time. Music 2.6 understands that different genres have different relationships with perfection, and optimizes for authenticity within each genre's aesthetic.

Benchmarks

The Flamenco Problem: Modeling Silence

The most technically impressive demo is the flamenco use case. Real flamenco's difficulty for AI isn't the notes — it's the silence between them. The guitar accelerates into rapid strumming, then everything halts. Two beats of pure stillness. Then it explodes back. Handclaps don't land on the beat — they orbit around it.

Music 2.6 models "the architecture of tension" — not just what's played, but when things stop and how they restart. This requires understanding music as temporal drama, not just note sequences. The prompt "flamenco, dramatic pauses, the silence matters more than the guitar" actually produces the right result.

What It Means for Developers

For game audio: Full boss-fight soundtrack in an afternoon. Proper low-end means no more supplementing AI music with manual bass passes. Dramatic direction prompts ("build tension, release, loop") actually work now.

For content platforms: Cover mode enables user-generated content at scale — "remake this song in [genre]" is an incredibly engaging feature for music apps, social platforms, and interactive experiences.

For background music needs: Four-hour generation runs that maintain quality and variety. Cafe playlists, podcast intros, YouTube background music — all promptable without manual curation.

API availability: Through MiniMax's standard API platform, same token-based pricing as their other models.

Developer Impact

Bottom Line

Music 2.6's three advances solve specific, named problems: Cover mode lets you preserve what matters (the melody) while changing everything else. Redesigned low-end makes AI music physically impactful instead of just sonically present. Intentional imperfection makes generated music feel human in genres where perfection is the uncanny valley. MiniMax continues to be the most complete multimodal company in China — text (M3), video (Hailuo 2.3), speech (Speech 2.8), and now music that's genuinely production-ready.

Resources


MiniMax Music 2.6 is available via API at minimax.io. Cover mode, enhanced bass, and imperfect vocals all accessible through the standard API.