Listeners on headphones notice breaths. The sharp inhale before a sentence, repeated three hundred times in an hour, is one of the small things that separates a show that sounds produced from one that sounds like a phone call. So podcasters remove them, and then a strange thing happens: the episode starts to sound wrong in a way that is hard to name. The host seems to never stop to think. The guest's answers land too fast. It is subtly exhausting to listen to, and most people who hear it blame the breath removal itself. They are half right.
What Actually Goes Wrong
A breath does not sit on its own. It sits inside a pause, and the pause is doing work. It is the gap where a listener finishes absorbing one thought before the next one arrives, and in conversation it carries meaning: a long pause before an answer says "I am thinking about this", a short one says "obviously". When an editor, human or software, cuts the breath out of the timeline, the pause loses the breath's duration too. A 600 millisecond pause becomes 200 milliseconds. Do that three hundred times and you have removed two minutes of thinking time from an hour of conversation, and every sentence now starts a fraction of a second before the listener was ready for it.
That is the robotic sound. It is not that the breaths are missing; nobody misses a breath. It is that the timing of speech has been compressed in a way no human speaker produces, and the ear, which is extraordinarily good at speech timing, flags it as wrong without being able to say why.
How Many Breaths Are We Talking About?
More than you think, and the number is worth knowing because it decides whether this is a job for a person or a machine. A 2025 study in the journal Digital Biomarkers measured breathing from the speech of 1,005 people and put the rate while talking at about 14 breaths a minute, against about 18 at rest. Talking slows your breathing down, but not by much, and every one of those breaths has to happen somewhere in the stream of words.
Where they happen is the interesting part. Speech scientists have known since the 1990s that breathing during speech is lopsided: the exhale is long and slow, because it is carrying the words, and the inhale is short and fast, squeezed into the pause between one phrase and the next. That is why a breath on a podcast is so audible. A normal resting inhale takes a couple of seconds and makes almost no noise; a speech inhale has to pull the same air through the same nose or mouth in a fraction of the time, so it is quick, sharp and, on a microphone ten centimetres away, about as loud as a quiet word.
Do the sum for a typical show. Two hosts, 45 minutes, each talking about half the time: 14 breaths a minute across 22.5 minutes is roughly 315 breaths per host, around 630 in the episode, one every four seconds or so. If you have ever wondered why breath removal feels endless, that is why. It is endless. It is also perfectly repetitive, which is the other half of the argument: six hundred identical decisions is the definition of a job you should not be doing by hand.
Quieten the Breath, Keep the Pause
The fix is to stop treating breath removal as a cut and treat it as a level change. The breath stays exactly where it is in time, but its level is taken down to the room tone of the recording, so what the listener hears is the pause they would have heard if the speaker had simply inhaled silently. The sentence starts when it was always going to start. Nothing about the rhythm changes.
Three details decide whether this sounds natural or merely different. First, the fade in and out of the quietened region has to be gentle and short enough that it does not clip the tail of the previous word or the onset of the next one; a breath that ends a few milliseconds before a "p" or a "t" is close enough that a careless fade eats the consonant. Second, what fills the gap matters. Digital silence sounds like a hole; the ear hears the room disappear and reappear. The gap has to be filled with the recording's own room tone, at the level the room actually sits at, so the background is continuous. Third, the breath itself has to be identified correctly: a voiced exhale, a laugh, or the "mm" a guest makes while thinking is not a breath and must not be touched.
The test that never lies
Play the processed file alongside the original and count seconds. If the processed version is shorter, breaths were cut. If it is the same length and the breaths are gone, they were quietened. Length is the whole difference.
Three Rules for a Natural Breath-Free Episode
- Never let the pause shrink. Whatever tool you use, the processed file must be the same length as the original to the millisecond. If it is not, the tool is cutting, and the rushed feeling will follow.
- Keep the deliberate ones. A sigh before a hard answer, the sharp intake when a guest is surprised, the audible breath a comedian uses as a beat: these are performance, not noise. Any breath tool worth using lists what it found and lets you keep individual breaths with a click.
- Fill with room, not silence. If the tool replaces the breath with pure digital silence, you will hear the room vanish for half a second three hundred times an hour. Insist on room-tone fill at the recording's real noise floor.
There is also a prevention side worth a sentence. Loud breaths are usually a mic-technique problem, not a lung problem. A mic aimed straight at the mouth from a few centimetres away picks up every inhale at close to speech level; angling it slightly off-axis and backing off to a fist's distance drops the breath level relative to the voice by a surprising amount, and a hydrated speaker produces far fewer of the clicks and smacks that get mistaken for breaths in the first place. Our guide to mouth clicks, plosives and sibilance covers the rest.
How Auto Clean Up Handles Breaths
This is the approach built into Auto Clean Up, and it was rebuilt this year after a professional narrator told us, accurately, that the previous version was clipping the ends of words. Each breath is found on a full-band envelope of the recording (breath is mostly hiss above 4 kHz, which cheaper detection throws away), checked against a voicing test so a spoken "mm" is left alone, and then quietened in place with a release that waits for the previous word's tail to finish decaying. The gap is filled with the steadiest second of your own room tone, matched to the recording's noise floor, so the background never drops out. The pause keeps its exact length. Every breath is listed in a panel with a play button for the quietened and the original version, and one click keeps any of them. There is a "how far to go" setting: obvious breaths only, or all of them.
↗ Try the tool
Auto Clean Up
Upload an episode and hear the difference: breaths quietened, pauses untouched, and a list of every one so you can keep the ones that matter.
Open Auto Clean Up →The Three Breaths You Will Actually Hear
Not all breaths are the same problem, and a tool that treats them all the same is the one that produces the robotic result. Listening back to a few hours of raw podcast audio, they sort into three kinds.
- The nose breath. Soft, low, usually between sentences. On a good mic at a sensible distance it sits only a little above the room tone, and quietening it by a few decibels makes it vanish without any risk to the words either side. This is most of the six hundred.
- The gasp. The sharp mouth inhale before a long sentence, or after a laugh, or when a guest is about to disagree with you. It is the one listeners notice, it is the loudest, and it is the one that sits closest to the next word, which is why a clumsy cut clips the first consonant of what follows. It needs quietening with a gentle fade, not a razor.
- The expressive breath. The sigh before a hard answer, the sharp intake when a guest hears something surprising, the breath that is half a laugh. These are not noise; they are the moment. A tool should find them, list them and let you keep them with a click, because nobody but you knows that the pause before "I have never told anyone this" is the best thing in the episode.
A Ten-Minute Test Before You Trust Any Tool
Run this once on a real episode and you will know whether a breath tool is cutting or quietening, and whether it is safe to leave alone. It takes ten minutes and a pair of headphones.
- Note the length of the original file to the second. Process it. Compare. Same length means quietened; shorter means cut, and every pause in the episode just shrank.
- Find three places where a breath came right before a word starting with a hard consonant (a "p", "t" or "k") and listen to the processed version on headphones. If the consonant is softened or missing, the fades are too tight.
- Find the most dramatic pause in the episode, the one where somebody thought before answering. Is it still the same length, and does it still feel like thinking? If it feels like a jump cut, the tool is removing time, not breath.
- Turn the volume up on a quietened gap. It should sound like the same room the voice is in, not like a hole. A gap that drops to digital silence is the second most common cause of the robotic sound after cut pauses.
- Count how many breaths it reported and spot-check five at random. A tool that cannot tell you how many it touched cannot be checked, and a tool that cannot be checked should not be running unattended on your show.
When to Leave Breaths In
Not every show wants them gone. A slow, intimate interview show can sound sterile without the breath, because breath is one of the cues that tells a listener there is a body in the room. Narrative shows sometimes want it for the same reason. A reasonable middle path for a conversational show is to quieten only the obvious ones, the loud inhales that jump out on headphones, and leave the small ones as texture. The rushed, robotic sound comes from cutting, not from removing; once the pauses are safe, how many breaths you take out is a taste decision, and you can hear both versions before you choose.
The rest of the post-recording workflow, in the order that saves the most time, is in How to Clean Up Podcast Audio Without Spending Hours in the Editor.
Frequently Asked Questions
Should I remove breaths at all?
For most talk shows, quietening the obvious ones is a clear improvement: the sharp inhales before sentences are the sound that separates a produced show from a phone call, and nobody has ever missed one. For intimate or narrative shows, keep more of them. Either way the rule is the same: quieten, never cut, and keep the expressive ones on purpose.
Why does my episode sound rushed after breath removal?
Because the breaths were cut rather than quietened, and every pause lost the length of the breath it held. Six hundred pauses each shortened by a third of a second is more than three minutes of thinking time gone from a 45-minute episode, and the ear notices the compression even when it cannot name it. Compare file lengths; if the processed file is shorter, that is your answer.
Do listeners really notice breaths?
There is no published study that counts breath complaints, so be wary of anyone quoting one. What is documented is the wider effect: acoustically effortful speech is judged more harshly and remembered less well, and a stream of sharp inhales on headphones is effort. The safe reading is that listeners rarely complain about breaths and quietly prefer the show without them.
Sources and Further Reading
- Abrol, A., Das, S., Nallanthighal, V. S., Ouweltjes, O., Grossekathofer, U., and Harma, A. (2025). Measuring Respiration Rate from Speech. Digital Biomarkers, 9(1), N = 1,005: about 14 breaths per minute while speaking against about 18 at rest.
- Winkworth, A. L., Davis, P. J., Adams, R. D., and Ellis, E. (1995). Breathing Patterns During Spontaneous Speech. Journal of Speech and Hearing Research, 38(1): the short-inhale, long-exhale pattern of speech breathing and inhalation landing in the pauses between phrases.
- Peelle, J. E. (2018). Listening Effort. Ear and Hearing, 39(2): the cognitive cost of effortful listening.
- VoiceEditSuite, How to Clean Up Podcast Audio Without Spending Hours in the Editor and Does Bad Audio Make People Stop Listening?.
Corrections
Breathing figures are quoted from the published studies named above; the per-episode arithmetic is a worked estimate for a two-host conversation. Spot an error? Tell us through the contact page and we will correct it with a note.
