MiniMax upgraded Music to version 1.5, and the headline change is simple: generation length extended from 90 seconds to 4 minutes. That's not just a number — 4 minutes is full pop song structure. A 90-second clip can cover one verse and a chorus. A 4-minute track can carry intro, verse, chorus, bridge, verse, chorus, outro — the complete arc of a commercially usable piece of music.

The pricing is ¥0.12 per track. Suno v4.5 costs approximately ¥0.30 per generation, is paywalled for longer durations, and has no API. MiniMax Music-1.5 has an API, is priced at roughly 40% of Suno, and is available immediately through the international platform at minimax.io/audio/music and the China platform at minimaxi.com/audio/music.

What Changed in 1.5

What's New

Duration: 90 seconds → 4 minutes. This is the core upgrade. 90-second AI music is a demo; 4-minute AI music is a deliverable. The difference matters for every practical use case: advertisement scoring, short-form video BGM, podcast intros, corporate event music, Vlog background tracks. A 4-minute track requires no stitching, no looping hack, no length extension workaround.

Structure awareness: Music-1.5 generates with compositional structure intact — the model produces pieces with distinct sections rather than continuous undifferentiated audio. Tested outputs show consistent intro/verse/chorus/bridge segmentation with appropriate energy transitions between sections.

Chinese language handling: This is the most differentiated capability. Suno and Udio generate technically competent music, but their training skews heavily toward English and Western structures. Chinese lyrics in these models frequently sound misaligned with the musical phrasing — characters fall on wrong beats, tonal patterns conflict with melodic contour, rhyme schemes don't match genre conventions. Music-1.5's Chinese output is substantially more natural, with tonal alignment, idiomatic rhyme density (tested with end-rhyme schemes sustained across entire verses), and genre-appropriate pronunciation (including hip-hop flow timing, which is particularly demanding for Mandarin).

Interface modes:

Genre Coverage

Genre Coverage

Four genres tested in depth:

Hip-hop / rap — Chinese rap with dense end-rhyme patterns. The tested output maintained a single rhyme sound ("-u" vowel) across an entire verse, with natural breathing cadence, beat alignment on syllable stress, and a distinct hook that repeats differently from the verse structure. The model also applied production elements correctly: lo-fi sampled intro, stuttering hi-hats under rap verses, harder drums lifting the chorus.

Classical Chinese (古风) — Traditional instrumentation generated correctly: guzheng leading melody, pipa providing rhythmic counterpoint. The two instruments maintain distinct timbral roles without clashing. Vocal style shifted appropriately toward light operatic delivery without going full Peking Opera — closer to Chinese TV drama soundtrack aesthetic, which is the commercially useful register.

Jazz — Brass leads in correctly, drums and bass occupy supporting roles. Chord progression has jazz voice leading. Chorus doesn't force a hard dynamic break — it resolves naturally from the verse. The hardest thing to get right in AI jazz generation is swing timing; the tested output handles it adequately.

Rock — Distorted guitar intro, tight drum/bass sync, vocal rhythm aligned with kick pattern. High notes don't crack, low notes don't thin out. Backing vocals added on chorus naturally. The overall feel matches Chinese indie rock aesthetics (think early Joyside or New Pants) rather than defaulting to Western rock templates.

Pricing and Competitive Position

Pricing

The market comparison:

Platform Duration Price per track API
MiniMax Music-1.5 4 min ~¥0.12 Yes
Suno v4.5 8 min (paid only) ~¥0.30 No
Udio 15 min varies Limited

Suno's 8-minute and Udio's 15-minute maximums are longer on paper, but real use cases rarely need more than 4 minutes. The effective comparison for most production workflows is: ¥0.12 with API vs ¥0.30 without API. For volume production — content creators, marketing teams, game studios generating ambient variants — the API matters as much as the price.

The API availability is the structural differentiator. AI music production at scale requires programmatic generation: batch creation from structured briefs, automated style matching, integration into content pipelines. Suno's consumer-only model is a constraint that limits it to direct-to-creator use cases. MiniMax's API makes Music-1.5 viable as infrastructure.

The Multimodal Stack Context

MiniMax has now assembled a complete audio creation pipeline in-house:

These three capabilities compound. A content creator prompt like "generate a lo-fi hip-hop track with a matching looping video" can be fulfilled by MiniMax components without third-party services. The stack is not just feature completeness — it's that each component is trained in awareness of the others, enabling coherent style transfer across modalities.

The Chinese music optimization isn't incidental. Chinese is phonetically demanding for AI music: four tones, word-final consonant patterns, and rhythmic conventions from traditional folk and opera that don't translate from Western training data. MiniMax's Chinese corpus and training methodology give it a baseline advantage in this segment that Western models can't easily replicate.

What It Means for Developers and Creators

The developer-relevant capabilities:

Programmatic generation via API — build music generation into your product, batch-create content variants, integrate with existing content workflows. No manual Suno login, no screen scraping.

Advanced mode lyric scaffolding — feed in your hook and structure, let the model fill in transitions and production. Useful for artists who can write lyrics but not produce, or for brands that need specific messaging set to music.

Chinese content workflows — BGM for Douyin, Xiaohongshu, short drama productions, corporate video, podcast intros. The Chinese language advantage is practically significant here: Western AI music under Chinese vocals consistently sounds wrong in ways audiences notice even if they can't articulate why. Music-1.5 doesn't have this problem.

Volume economics — at ¥0.12 per complete track, A/B testing background music for ads becomes financially trivial. The calculation that made custom music production prohibitive (¥hundreds to ¥thousands per track) no longer applies.

Bottom Line

Music-1.5 completes MiniMax's audio stack. Speech-02 handles voice; Music-1.5 handles composition; Hailuo handles visualization. Combined with API access and Chinese-language optimization, MiniMax has built the most complete AI audio infrastructure for the Chinese market, and the competitive position on cost and capability is strong against both Western (Suno, Udio) and domestic alternatives.

The 4-minute duration is not just a feature — it's a threshold. Below 4 minutes, AI music is a demo. At 4 minutes with structural integrity and Chinese-optimized generation, it's a production tool.

Resources


MiniMax Music-1.5 is available at minimax.io/audio/music. Generation price: approximately ¥0.12 per 4-minute track. API access available at platform.minimaxi.com.