How to Clean Up Podcast Audio Without Spending Hours in the Editor, VoiceEditSuite
← Guides

Podcasting

How to Clean Up Podcast Audio Without Spending Hours in the Editor

September 25, 2026·12 min read
A laptop on a sunlit desk showing a blurred audio waveform, headphones resting beside it, a mug half full

Ask a podcaster who has been at it for a year what they would change, and the answer is rarely the microphone. It is the evening after the recording. A 45-minute conversation with a guest reliably becomes two to three hours in the editor, and it is the part of the job that quietly kills shows: not the recording, not the promotion, the Tuesday night spent scrubbing a waveform for breaths and false starts. This guide is about getting that evening back without lowering the standard of what you publish.

The first thing to understand is that editing time is not one thing. It is four or five different jobs that happen to use the same software, and they are not equally deserving of your attention. Some of them are creative decisions about what the episode is. Most of them are mechanical: finding every breath, every fridge hum, every place the level dips because the guest leaned back. Mechanical work is exactly the work a well-built tool does better than a tired human at eleven at night, because it does not get bored and it does not miss the one at 38 minutes.

Where the Time Actually Goes

Timed across a typical interview episode edited by hand, the hours split roughly like this. The numbers will move with your show, but the shape is consistent: the mechanical passes dwarf the creative ones.

Illustrative breakdown of a manual edit. The only slice that needs the host's judgement is the twenty minutes of content decisions.

Look at the biggest slice. Breath removal is the most tedious task in podcast editing and the one with the least creative content. Nobody has ever made an editorial decision about a breath; you either take it out or you do not, and there are three hundred of them in an hour of conversation. The same is true of false starts. When a guest says "so the thing about, sorry, let me start that again, the thing about pricing is", the decision is obvious. Finding it is the work.

The Arithmetic Nobody Does Before Starting a Show

Here is the sum that would stop a lot of shows before episode one, or at least send them shopping for a better workflow. Start with how long episodes are. Across 112,207 active podcasts on one hosting platform in August 2026, the most common episode runs 20 to 40 minutes (31 percent), a fifth run 40 to 60 minutes, and 13 percent run past the hour. A third of everything published is longer than 40 minutes. Now add the rule of thumb every production studio quotes: three to five minutes of editing per minute of audio for a polished result, and about two hours for a basic pass on an hour-long recording. Then add cadence: on the same platform, 37 percent of shows publish every eight to fourteen days and 34 percent every three to seven.

A third of all episodes run past 40 minutes. Source: Buzzsprout platform stats, August 2026.

Multiply it out for a weekly 45-minute show and you get somewhere between 117 and 195 hours of editing a year. That is three to five full working weeks, every year, spent on a task most hosts describe as the part of podcasting they like least. It is also, not coincidentally, the number that explains why the medium has 4.72 million feeds and only about 478,000 that have published anything in the last 90 days: shows do not die of bad ideas, they die on Tuesday nights.

The breath number in the chart above is not a guess either. A 2025 study of 1,005 people, published in the journal Digital Biomarkers, measured breathing rates from speech and found people breathe about 14 times a minute while talking, against about 18 at rest. In a 45-minute conversation where two hosts each talk for roughly half the time, that is around 315 breaths per host and about 630 in the episode, one every four seconds or so. Every one of them is a small, identical, boring decision, which is why that slice of the pie is the biggest and why a human doing it at eleven at night misses the one at 38 minutes.

The Order of Operations That Saves the Most Time

Doing the passes in the wrong order costs time twice. If you normalise loudness first and then remove noise, the noise floor moves and you have to normalise again. If you remove breaths before you take out the room noise, the quietened gaps sit at a different level from the rest and the seams show. The order below is the one that never has to be repeated.

  1. Fix the room before you record, not after. Ten minutes on the room is worth an hour of noise removal. A free measured room check tells you whether the problem is echo, hum, or the fridge, and each one has a cheap fix.
  2. Get the takes right, then clean. Run the raw file through Auto Clean Up first. It drops the false starts and fluffed retakes, quietens every breath while keeping the pause it sat in, and tightens dead air. You are now editing a file that is already ninety percent finished.
  3. Make your content decisions on text. The cleaned file comes back transcribed. Cutting a tangent by deleting a sentence of text is faster than finding it on a waveform, and you are making the one kind of decision the machine should not make for you.
  4. Take the room out of what is left. If the recording still carries hiss, hum or a boxy room, Studio Rescue strips it while leaving the voice alone, and shows you the noise floor before and after so you know it worked.
  5. Set the loudness last. One pass through Loudness Normalizer at the podcast preset (-16 LUFS, true peak under -1 dBTP) and the file matches every other show in the listener's queue. Doing this last means nothing you do afterwards can un-do it.

Followed in that order, the mechanical passes take a few minutes of processing time while you make a coffee. The twenty minutes of content decisions are still yours. The difference is that they are now the whole job rather than a small slice buried inside three hours of clicking.

A Worked Example: One Episode, Two Ways

Take a real-shaped episode: a 47-minute interview, host on a decent USB mic in a spare room, guest on laptop earbuds in a kitchen with a fridge. Two false starts from the guest, one from the host, a tangent about a holiday that goes nowhere, and the usual six hundred breaths. Here is the same evening done both ways.

  • By hand, the old way. Import, listen through once at normal speed to find the problems (47 minutes). Go back and cut the three false starts and the tangent (20 minutes, mostly scrubbing to find the exact spots). Breaths: zoom in, find, fade, repeat, six hundred times, and stop halfway because it is midnight (55 minutes and counting). Noise: try a plugin on the guest track, overdo it, back it off, listen again (35 minutes). Levels: bring the guest up, notice the fridge came up with them, compromise (25 minutes). Export, check the file, notice it is quieter than last week's, export again (15 minutes). Total: a little over three hours, and the breaths in the second half are still in.
  • In the order above. Room check done last week, so the host track is already dry. Upload both tracks to the cleanup: five minutes of processing while the kettle boils, and back come both files transcribed, false starts dropped, every breath quietened with its pause intact, and a list of each decision. Read the transcript and delete the holiday tangent as text: fifteen minutes, because you are reading, not scrubbing. Noise removal on the guest track only, with the noise floor shown before and after: two minutes. Loudness pass at the podcast preset: one minute. Total: under half an hour, and the breath at 38 minutes is gone, because a machine counted to six hundred without getting bored.

The second version is not faster because it skips steps. It does every step the first version does, and one more (it actually finishes the breaths). It is faster because the human only touches the twenty minutes of decisions that need a human, and the searching, which is what the other two and a half hours always were, is done by something that does not mind.

Breaths: Quieten, Do Not Cut

There is a specific mistake worth naming because it is the reason many podcasters distrust automatic breath removal. Early tools cut the breath out of the timeline. The breath disappears, but so does the time it occupied, so every pause in the conversation gets shorter and the whole episode starts to feel rushed and slightly unnatural, like the host never stops to think. Listeners cannot say why it sounds off; they just stop listening.

The right approach is to quieten the breath in place. The pause keeps its exact length, the gap is filled with the room tone from your own recording rather than digital silence, and the rhythm of the conversation is untouched. A good implementation also lists every breath it found so you can keep one deliberately, because a sigh before a hard question is sometimes the best moment in the episode. We wrote up the details in Removing Breaths From a Podcast Without Making It Sound Robotic.

↗ Try the tool

Auto Clean Up

Auto Clean Up takes a raw episode and returns it transcribed, with breaths quietened and pauses kept, false starts dropped, and every decision listed so you can reverse any of them with one click.

Open Auto Clean Up →

What Not to Automate

Be honest about the two things a tool should never decide. The first is what the episode is about. Which tangent stays, which anecdote goes, whether the ten-minute detour at the halfway point is the best part of the show or the reason people drop off: those are yours. The second is your guest's dignity. A tool can drop a stumble. Whether to drop the moment a guest got emotional, or the pause where they thought hard before answering, is an editorial choice, and the tool's job is to make it easy for you to keep those moments, not to remove them for you.

One-voice-per-file gives the cleanest result

If you record each host and guest on separate tracks, run each track through the cleanup on its own and then mix. Breath quietening and noise removal both work best when there is one voice in the file, and a guest recorded in a kitchen can be rescued without touching your studio track.

What "Clean" Means, in Numbers

"Clean" is a feeling until you write down what it measures, and once you write it down, you can check a file in a minute instead of listening to it three times. These are the targets a finished episode should hit, and every one of them is a number a tool can report rather than an opinion you have to form at midnight:

  • Noise floor at -60 dB or lower in the gaps between sentences. That is the audiobook industry's bar, and it is the difference between a pause that sounds like silence and a pause that sounds like a fridge.
  • Both voices within about 2 dB of each other across the episode, so nobody reaches for the volume when the guest speaks.
  • Integrated loudness at -16 LUFS, plus or minus one, with true peaks no higher than -1 dBFS. That is Apple's published recommendation and the target most podcast platforms normalise to; hit it and your show plays at the same volume as the one the listener just left.
  • The same file length before and after breath work. If the processed file is shorter, breaths were cut and the pauses shrank; if it is the same length and the breaths are gone, they were quietened. Length is the whole test.
  • Zero false starts and zero dead air over a couple of seconds, unless you kept a pause on purpose, in which case it should still be there.

Measure, do not listen

Your ears at the end of an edit are the least reliable instrument in the room. A tool that shows you the noise floor before and after, the integrated loudness of the finished file, and the count of breaths it quietened is telling you things you can check; a tool that just says "enhanced" is asking you to trust it.

A Realistic Target

For a 45-minute interview episode recorded in a reasonable room, a realistic post-recording workflow is: five minutes of processing for cleanup and transcription, fifteen to twenty minutes of reading the transcript and cutting what does not belong, two minutes for noise removal if it is needed, and one minute for loudness. Under half an hour, with the standard of the published file higher than it was after three hours by hand, because the machine does not miss the breath at 38 minutes.

That is the whole argument. Not that editing does not matter, but that most of the hours were never editing in the first place. They were searching, and searching is what software is for. Everything the toolkit does for an episode, in order, is laid out on the For Podcasters page.

Frequently Asked Questions

Does cleanup shorten the episode?

Dropping false starts and tightening long dead air shortens it slightly, which is the point. Breath quietening does not: every pause keeps its exact length. If you want no change to the timing at all, keep the retake removal off and run breaths only.

Can I clean a file that already has two people mixed together?

Yes. It works on a mixed file. Separate tracks give a cleaner result because each voice can be treated on its own, but a single stereo or mono export from a remote-recording app is fine.

What happens to my audio?

It is processed and then deleted, never used to train an AI model, never shared. The full promise is on the Your Audio page.

Can it take out the ums and uhs?

Yes, as a switch that is off by default. Turn on "Take out my ums and uhs" and every um, uh and erm it can hear becomes a cut with a short crossfade, counted and listed with its time so any one can go back. It leaves the sounds that mean something alone: mhm, uh-huh, hmm, ah and oh all stay, because a podcast host uses those on purpose.

How long does the processing itself take?

Transcription and cleanup on a 45-minute file takes a few minutes; noise removal and loudness take under a minute each. The honest total for an episode, including your own reading of the transcript, is twenty to thirty minutes, which is the number in the section above and the one to compare against your current Tuesday.

Sources and Further Reading

  • Buzzsprout, Podcast Stats, August 2026: episode length distribution and publishing frequency across 112,207 active podcasts, updated monthly.
  • Abrol, A., Das, S., Nallanthighal, V. S., Ouweltjes, O., Grossekathofer, U., and Harma, A. (2025). Measuring Respiration Rate from Speech. Digital Biomarkers, 9(1), N = 1,005: the roughly 14 breaths per minute while speaking against roughly 18 at rest. The per-episode arithmetic is ours.
  • Podcast Index counts (4.72 million feeds, about 478,000 active in 90 days) as reported by Libsyn, 2026.
  • Apple, Audio requirements for Apple Podcasts: the -16 LKFS loudness recommendation and the -1 dB true-peak ceiling.
  • The three-to-five-minutes-per-minute editing figure is the rule of thumb quoted across podcast production studios for a polished edit; it is a working estimate, not a survey.
  • VoiceEditSuite, Removing Breaths From a Podcast Without Making It Sound Robotic and Does Bad Audio Make People Stop Listening?: the companion pieces on breaths and on why sound decides whether people stay.

Corrections

Platform figures are as published by Buzzsprout for August 2026 and change monthly; the breathing-rate study is quoted from the published paper. Spot an error or a newer number? Tell us through the contact page and we will correct it with a note.

Keep reading

Niches

Meditation and Wellness App Voice Over: A Growing, Different Kind of Read

8 min read

Niches

Movie Trailer Voice Over: What the Work Actually Involves and How to Break In

9 min read

Business

Audio Description Is Expanding to Every TV Market by 2035. The Narration Work Is Expanding With It

13 min read