
Audio Jingle
FreeGenerate jingles, voiceovers, and sound effects effortlessly.
Free · Opens the source repo
What Audio Jingle does
The Audio Jingle skill provides a streamlined way to generate audio content, including jingles, voiceovers, and sound effects, by leveraging various AI models. It operates in three distinct modes based on the audioKind specified in the project metadata: music, speech, or sound effects. Each mode routes requests to the appropriate models, such as Suno V5 for music, MiniMax TTS for speech, and ElevenLabs for sound effects, ensuring that users can create high-quality audio tailored to their needs.
To use the skill, developers must define specific parameters like genre, tempo, instrumentation for music, or script and voice for speech. The skill emphasizes the importance of precise input, particularly for TTS, which requires exact voice IDs rather than descriptive phrases. This attention to detail allows for accurate audio generation that meets user expectations. The output is a single MP3 or WAV file saved directly to the project folder, making it easy to incorporate into existing workflows.
This skill is particularly useful for developers and designers looking to enhance their projects with custom audio elements without needing extensive audio production knowledge. By automating the audio generation process, users can focus on other aspects of their project while still achieving professional-quality results. The skill is designed to be straightforward, with a clear workflow that guides users through planning, composing, and dispatching their audio requests.
However, users should be aware that the skill is not suitable for generating longer or complex audio compositions, as it is optimized for shorter segments. The emphasis on specific parameters means that users must have a clear idea of what they want before using the skill, which may limit its use for more experimental or freeform audio projects.
When to use it
Use this skill when you need to create jingles, voiceovers, or sound effects quickly and efficiently for your projects.
When not to use it
This skill may not be suitable for generating lengthy or complex audio compositions, as it is designed for shorter segments.
What you can build with it
Creating a Jingle for a Marketing Campaign
Generate a catchy jingle to promote your product by specifying genre and instrumentation.
Producing a Voiceover for a Video
Use the speech mode to create a professional voiceover by providing a script and selecting a voice.
Designing Sound Effects for a Game
Quickly generate sound effects by defining texture and duration for immersive gameplay.
How to install Audio Jingle
View source1. Install with the skills CLI
npx skills add nexu-io/open-design/audio-jingle --agent claude-code2. Or install it manually
Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.
Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs
Inside SKILL.md
Written by nexu-ioAudio Jingle Skill
Three sub-modes. The active project's audioKind decides which one
runs:
audioKind | Models we route to | Plan focus |
|---|---|---|
music | Suno V5 (default), Udio, Lyria 2 | genre + tempo + instrumentation |
speech | MiniMax TTS (default), Fish, ElevenLabs V3 | script + voice + pacing |
sfx | ElevenLabs SFX (default), AudioCraft | texture + impact + duration |
Resource map
audio-jingle/
├── SKILL.md
└── example.html
Workflow
Step 0 — Read the project metadata
audioKind, audioModel, audioDuration (seconds), and (for speech)
voice. Branch by known values and use them verbatim. Missing metadata is not
an instruction to ask: infer a safe default when possible, and emit a
clarifying form only when the missing answer would materially change the
requested output or prevent generation.
Important: voice is provider-specific. For minimax-tts, --voice
must be a valid MiniMax voice_id (for example male-qn-qingse), not
a natural-language description. If you only have a prose voice brief
("warm female narrator", "neutral Mandarin"), keep that in your plan
but omit --voice so the daemon's default voice id applies, or ask the
user to choose a specific id.
Step 1 — Plan
Music
- Genre + reference artists (1-2)
- Tempo (BPM) + key
- Instrumentation (3-5 instruments max)
- Vocals: yes / no / hummed / choir
- Mood arc (intro → chorus → outro)
Speech
- Script (final, not draft — TTS runs verbatim)
- Voice target + pacing
For MiniMax this means a real
voice_id, not prose in--voice - Pronunciation hints for proper nouns / acronyms
SFX
- Texture (impact / whoosh / ambience / foley)
- Duration + envelope (sharp attack vs. gentle swell)
- Layering note (single hit vs. stacked)
State the plan in 2-3 sentences before dispatching.
Step 2 — Compose the prompt
Use the format the upstream model prefers. Bind audioDuration to the
API parameter directly; never put "make it 30 seconds" in prose.
Step 3 — Dispatch via the media contract
Use the unified dispatcher — do not call provider APIs by hand:
"$OD_NODE_BIN" "$OD_BIN" media generate \
--project "$OD_PROJECT_ID" \
--surface audio \
--audio-kind "<music|speech|sfx>" \
--model "<audioModel from metadata>" \
--duration <audioDuration seconds> \
[--voice "<provider voice id (speech only)>"] \
--output "<short-slug>-<duration>s.mp3" \
--prompt "<assembled prompt from Step 2 — for speech, the literal script>"
The command prints one line of JSON: {"file": {"name": "...", ...}}.
The bytes land in the project; the FileViewer renders the audio
transport controls automatically.
Step 4 — Hand off
Reply with: plan summary, the filename returned by the dispatcher, and one sentence on what to try if the user wants a variation (e.g. "swap tempo from 92 to 108 BPM" rather than "make it different").
Hard rules
- TTS runs your script literally. Proof it before dispatching — even one stray comma changes the cadence.
- MiniMax TTS rejects free-form voice prose in
--voice. Use a real MiniMaxvoice_id(for examplemale-qn-qingse) or omit the flag and let the daemon's default voice apply. - Music: under 30s = single section; 30–90s = intro + body; 90s+ = full arc. Don't try to fit a 3-act song into 15 seconds.
- SFX: prefer one well-described layer over a paragraph of "make it cool" — generators reward specific texture words.
- Save the file every turn. The audio viewer shows transport controls the moment the file lands.
Frequently asked questions about Audio Jingle
Similar skills
Spring Boot Testing
Master testing techniques for Spring Boot 4 applications.
GitHub Issues
Manage GitHub issues efficiently with MCP tools.
Geofeed Tuner
Optimize your IP geolocation feeds in CSV format.
Batch Files
Master Windows batch scripting for automation and task management.
Adobe Illustrator Scripting
Automate your Illustrator workflows with ExtendScript.
Plugin Structure
Create and organize Claude Code plugins effectively.
