MiniMax has released 'MiniMax-Music3,' a music generation AI, for free, following its video generation AI, allowing users to generate songs with Japanese vocals up to 5 minutes in length locally.



MiniMax, a Chinese AI development company, released its music generation AI, ' MiniMax-Music3, ' on August 14, 2026 (Japan time). MiniMax-Music3 can generate songs with vocals based on text instructions and supports Japanese vocal generation. The model is available free of charge and can be run locally. In addition, AI skills have been released to generate long lyrics and structured prompts that are easier for the AI model to recognize, based on short user input.

GitHub - MiniMax-AI/MiniMax-Music3 · GitHub

https://github.com/MiniMax-AI/MiniMax-Music3




MiniMax-Music3 can generate songs up to 5 minutes in length and output the result as a 32kHz 16-bit stereo music file. The model is available at the following link.

MiniMaxAI/MiniMax-Music3 · Hugging Face
https://huggingface.co/MiniMaxAI/MiniMax-Music3



Many examples of songs generated with MiniMax-Music3 are available at the following link.

Mini Max Music 3
https://minimax-ai.github.io/music3-demo/



A free demo app is also available, so let's try generating a song with Japanese vocals. First, click the link below to open the demo app.

MiniMax Music 3 Studio - a Hugging Face Space by MiniMaxAI
https://huggingface.co/spaces/MiniMaxAI/MiniMax-Music3



Enter 'Rock music with Japanese vocals. Guitar, bass, and drums. Includes phrases like 'hot summer,' 'beware of heatstroke,' and 'drink plenty of water,'' and click 'Write lyrics & review.'



Since 'Japanese lyrics according to the instructions' and 'structured prompts containing the necessary information in accordance with MiniMax-Music3's notation rules' have been generated, click 'Generate' to generate the song.



After waiting for a while, music and video files were output. MiniMax-Music3 is a model with the function of 'outputting music files based on prompts,' while the functions of the demo app, 'generating lyrics and structured prompts based on user input' and 'outputting music as a video,' are handled by a separate system.



The output video is shown below. The Japanese vocals were generated with high quality. This time, the maximum duration was set to 60 seconds, so the song is cut off midway.

Music with Japanese vocals generated by the music generation AI 'MiniMax-Music3' - YouTube


The demo application uses ZeroGPU , Hugging Face's free inference service, and in this case, it reached the free usage limit after generating just one 1-minute song. MiniMax-Music3 already allows local generation using ComfyUI, and it can generate an unlimited number of songs locally.




Additionally, an AI skill has been released to enable the function of 'generating lyrics and structured prompts based on user input.'

MiniMax-Music3/skills at main · MiniMax-AI/MiniMax-Music3 · GitHub
https://github.com/MiniMax-AI/MiniMax-Music3/tree/main/skills



Furthermore, MiniMax has announced plans to release open models such as 'Image generation and editing model based on MiniMax H3' and 'Video upscaling model' in the future.

The development team behind the video generation AI 'MiniMax H3' has confirmed that they are preparing an 'image generation and editing model,' a 'video upscaling model,' and a 'low-step version of MiniMax H3' - GIGAZINE



in AI,   Video,   Review, Posted by log1o_hf