The audio contract: ears only, attention partial
▾
Radio, podcasts, and streaming audio share one defining condition: there is no screen. The listener is driving, cooking, or running. Nothing can be shown, nothing can be tapped, and the message competes with whatever the listener is actually doing. Every craft rule in this article falls out of that condition.
The honest concession up front: this channel gives you no CTR, no same-day creative signal, and mostly delayed, indirect response. What it gives you instead is frequency, intimacy (especially in podcasts, where the voice is already trusted), and a low-cost way to stay in memory between purchase moments. Build the creative for memory, then build a decision loop from directional signals (covered below), and do not pretend either one is a click stream.
One idea per spot
▾
A :30 audio spot is roughly 70 to 80 spoken words at a natural pace. A :60 is roughly double. That is enough room for exactly one idea, delivered clearly, with the brand attached. The most common failure in audio scripts is the second idea: a feature list, a secondary offer, a bonus claim. Each addition taxes a listener who is not looking at anything and cannot re-read.
The working test: after one listen, can someone repeat back the brand and the one thing it does? If the answer requires a second listen, the spot is overwritten. Cut until one pass is enough. The same discipline behind a good hook applies; the difference is that in audio the “hook” is the first sentence out of the speaker, and it has to earn the next sentence with words and sound alone. The opening-line craft in writing hooks translates directly; the visuals do not.
Brand early, brand repeated
▾
In a feed ad, the brand is on screen even when nobody says it. In audio, if the name is not spoken, it does not exist. Say the brand name early, ideally inside the first ten seconds, and repeat it at least once more before the close. A spot that saves the name for the final second is betting the entire buy on complete, attentive listens, and audio listening is rarely either.
Repetition feels awkward on the page and normal in the ear. Scripts read aloud always tolerate more brand mentions than scripts read silently. Read every draft out loud before judging it; audio copy that was never spoken during writing almost always runs long and lands flat.
Mnemonics: give the ear something to own
▾
Audio branding compounds through repeatable sound: a sonic logo, a jingle fragment, a consistent voice, a signature sound effect, even a repeated turn of phrase. These are the audio equivalent of a visual identity, and they are what lets exposure number twelve build on exposure number three instead of starting over.
The working rule: pick one mnemonic and keep it across flights. A brand that changes voice, music, and tagline every quarter is buying reach without accumulating recognition. If the brand has a sonic logo from TV work, use it here; the channels reinforce each other. If it has none, a consistent voice and a repeated closing line are the cheapest starting assets. The craft overlaps with sound design, with one inversion: there, audio supports the picture; here, audio is the entire ad.
CTAs that survive audio-only
▾
The listener cannot tap. Any CTA that assumes a screen fails at the moment of delivery: “tap below,” “swipe up,” “link in bio,” and “click the button” are all instructions to a person holding a steering wheel. This is not a style preference; audio platforms commonly reject spots whose CTA language references actions the format cannot perform.
What survives, in rough order of reliability:
- —A memorable URL, short enough to hold in memory until the listener has a screen. If it needs spelling out, it is too long.
- —“Search [brand]”, which outsources the memory problem to a search engine and tolerates imperfect recall.
- —A promo code, ideally the show's or host's name, which doubles as your attribution mechanism.
One CTA per spot, stated near the close and ideally echoed once. The layered-CTA thinking in CTA architecture still applies, compressed: in audio, the primary action is always deferred, so clarity and memorability beat urgency. If the spot runs with a companion banner in a streaming app, the banner must match the audio claim; a mismatch between what was said and what is shown is both a trust problem and a rejection risk.
Host-read vs produced spots
▾
A host-read is the show's own voice delivering your message, usually from talking points rather than a locked script. Its strength is borrowed trust: the recommendation arrives inside a relationship the listener already has. Its costs are control and consistency: the claim discipline lives in your talking points, delivery varies by episode, and an endorsement relationship carries disclosure obligations under FTC endorsement rules. Material connections must be disclosed clearly; that is a legal floor, not a style choice.
A produced spot is your script, your voice talent, your mix, identical on every airing. It scales across shows and stations, protects claim accuracy, and can carry your mnemonic consistently. What it gives up is the intimacy premium: it is audibly an ad arriving from outside the show.
One boundary applies to both: do not mimic the host or DJ, and do not imply the platform or station endorses the product when it does not. A produced spot written to impersonate a host-read, fake sirens or alarm sounds, and notification-style SFX designed to make a listener think their own device is alerting them are all classic rejection territory, and they deserve to be. The same false-urgency instinct that gets feed ads flagged gets audio ads pulled.
Working guidance on choosing: use host-reads where the show's audience is the target and trust is the sell; use produced spots for scale, regulated claims, and brand-building flights where consistency matters more than intimacy. Many programs run both, with the produced spot carrying the mnemonic and the host-read carrying the offer.
The script is the moderation surface
▾
Feed platforms moderate what is shown. Audio platforms moderate what is said. Every claim, disclaimer, comparison, and instruction lives in the script and the ear alone, which means the script review is the compliance review. Regulated-category disclaimers must be spoken, at a pace a human can follow, not rattled off as legal static. Stereotype-driven writing, accent humor and gender stereotypes in particular, is a leading cause of audio ad rejection, independent of how the same joke might fare visually.
The practical upside: because the script is the whole ad, the whole ad can be reviewed before anything is recorded. A script read-through catches most audio failures (overwriting, buried brand, screen-dependent CTA, undeliverable disclaimer) at the cheapest possible stage.
Measurement: directional, and say so
▾
Audio metrics are reach, frequency, recall, and promo-code redemption. Treat all of them as directional. Working practice for building a decision loop anyway:
- —Promo-code redemption per show or station is the cleanest per-placement signal. It undercounts (listeners convert without the code), so read it as a floor and compare placements against each other, not against an absolute target.
- —Vanity URLs per flight give the same relative read with less friction than a code.
- —Recall surveys and pre/post branded-search volume indicate whether the spot registered at all, on a slower clock.
What the channel does not produce: CTR, hook rate, completion curves, saves, shares, or anything from the feed-video vocabulary. Importing that jargon into an audio report is not just sloppy; it invents data. State the metrics the channel actually has, state that they are directional, and make creative decisions by comparing spots and placements against each other over time.
Reviewing an audio spot with The Ad Bench
▾
Paste the script, or upload the produced spot; the analyzer classifies the medium automatically and scores audio creative against a dedicated radio/podcast/audio rubric rather than feed-video criteria, with an explicit format override available. The rubric's primary driver is the script itself, because the script is where audio platforms accept or reject.
Concretely, the review checks for one clear idea in the :30 or :60, the brand name early and repeated, a memorable audio mnemonic or sonic logo, and a CTA that survives audio-only (a memorable URL, “search X,” or a promo code). It flags the policy territory above: visual-only CTA language, fake alarms and notification SFX, host or DJ mimicry and implied platform endorsement, and stereotype-driven writing. Where a companion banner exists, it checks that the banner matches the audio claim. On measurement, the review speaks the channel's own language: reach, frequency, recall, and promo-code redemption, stated as directional, never dressed up as CTR or feed jargon.
What the score can tell you: whether the script's structure, branding, CTA, and policy exposure are sound before you book the buy or record the talent. What it cannot tell you: whether the campaign will perform. See the scoring rubric for how the score is built.
The Ad Bench methodology: The checks and rejection-risk flags in this article reflect The Ad Bench's current audio scoring rubric and creative-review approach. They are designed to identify creative risk before media spend, not predict or guarantee campaign performance.
Compliance note: Platform policies and advertising rules change. This guide is educational, not legal advice. Confirm current requirements in applicable platform documentation and consult qualified counsel for endorsement disclosures, regulated claims, and sensitive categories.
Sources
▾
- The Ad Bench methodology, audio format rubric. Reviewed August 24, 2026.
Reviewed: 2026-08-24 · Last updated: 2026-08-24 · Next review due: 2027-02-20
Read to the end to earn a star.