Speech, voice and audio

    Text-to-music

    Also called AI music generation, Music generation.

    Text-to-music generates an original composition from a written description of genre, instrumentation, mood and tempo — with or without sung lyrics.

    The prompt works like a brief to a session musician rather than a search query. Genre, instrumentation, tempo, era and mood steer usefully; naming an artist to imitate is both less reliable and the thing most providers restrict. Some models also take lyrics as a separate input and sing them, which is a different job from generating an instrumental bed.

    For video the practical draw is fit and clearance in one step. A track generated to your brief is the right length and mood without hunting a library, and it does not carry someone else's rights — subject to the provider's terms, which are worth reading before anything commercial goes out.

    Structure is the weak point. Models produce convincing texture and a plausible groove more easily than a piece that develops — an intro that earns a chorus, a drop that lands where your edit needs it. Generating a longer piece and cutting to picture usually beats trying to prompt the arrangement exactly.

    In practice

    • Brief instrumentation, tempo and mood; avoid naming artists.
    • Generate longer than you need and cut to the edit rather than prompting for exact timing.
    • Instrumental beds sit under voiceover far better than anything with a lead vocal.

    The mistake to avoid

    Prompting for a track that is a specific artist's sound. Results are unreliable and provider terms generally prohibit it.

    Where you will run into it

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.