logo
← All posts

August 3, 2026

Five Steps From a Recorded Talk to a Published Audio Episode

You filmed a talk and only the audio is worth publishing. Five steps turn a 1 GB video into an episode uploaded as a YouTube draft, with nothing downloaded.

You filmed your conference talk on a phone propped against a water jug, and the video is not the point. Fifty-two minutes of a person at a lectern is not something anyone watches. The audio is a different story: it is the talk three people emailed you about afterwards, and you said you would send them something.

So what you have is a 1.2 GB MP4, and what you need is an audio episode with the six minutes of room noise trimmed off the front, a level that survives a phone speaker on a train, and a size that uploads before you lose interest. That is four jobs before you even get to publishing, and none of the tools that do them wants a video file.

What has changed is that you no longer download between them. Each result hands straight to the next tool in the same tab: one upload at the start, and no download at the end, because the last step publishes for you.

The Chain at a Glance

  1. Pull the audio out of the video and leave a gigabyte of picture behind
  2. Cut the dead air at the front and the packing-up at the back
  3. Raise the level so it holds up on a phone speaker
  4. Compress to something that uploads in a reasonable time
  5. Publish it as a draft you finish in YouTube Studio

Step 1: Leave the Picture Behind

Everything after this point is audio work, and audio is roughly a twentieth of what you are holding.

Drop the MP4 in and you get three settings. Quality (bitrate) defaults to Recommended · 192 kbps, which is the right pick. The temptation is to grab Low · 96 kbps because you are compressing later anyway, but that is backwards: step 4 re-encodes this file, and a 96 kbps source compressed a second time sounds like a phone call. Leave Sample rate on Same as source, since your camera almost certainly recorded at 48 kHz. Then set Audio channels to Mono (smaller, voice/podcast). One person, one microphone, and stereo just stores the same signal twice.

A 52-minute talk at those settings lands somewhere around 70 MB. That number matters more than it looks like it does, and step 5 is where you find out why.

The next tool is not in the shortlist. Open Other compatible tools on the result screen and choose Trim audio.

MP4 to MP3 →

Step 2: Cut the Room Noise Off the Front

Every recording of a live talk has the same two useless sections: chairs scraping while someone asks whether it is on, and applause fading into a lapel mic being unclipped.

You get a waveform with two handles. Drag them or type exact timestamps into Start and End, and the Length readout tracks what you have selected. Use Preview selection to check the in-point before you commit, because the difference between cutting on the applause and cutting on your first word is about two seconds and completely audible. Then turn on Fade in and Fade out. A hard cut into a room that already has air in it produces a click at the top of the file, and a fade of a fraction of a second removes it.

Forty-six minutes of actual talk is a realistic result from fifty-two.

Now Boost audio volume is sitting right there under Send to another tool. Pick it and the trimmed file travels across.

Trim audio →

Step 3: Make It Loud Enough to Survive a Train

A phone against a water jug records quietly. On headphones that is fine. On a phone speaker next to traffic it is unlistenable, which is the most common reason people abandon an episode halfway through.

The Gain slider runs from −20 dB to +30 dB, and each file shows its measured peak beside it, so you can see how much headroom you have rather than guessing. Leave Auto-limit to 0 dB on: peaks are held at 0 dB and the loud parts stop rather than distorting. Switch it off and you get pure linear gain, with a warning naming the exact peak at which the output will clip. Press and hold Compare to A/B against the original, on a quiet passage rather than a loud one, because the quiet passages are the ones you are fixing.

Compress audio is in the suggested list. Send it on.

Boost audio volume →

Step 4: Get the Size Down Before You Upload It

Four presets, each labelled with what it costs you per minute rather than with a bitrate you have to do arithmetic on. Light is about 1.4 MB per minute, Standard about 960 KB, Strong about 720 KB and described as fine for voice, and Extreme about 480 KB and genuinely voice-grade.

For one person talking, Strong is the honest answer. Forty-six minutes at roughly 720 KB per minute is around 33 MB, down from the 70 you started with, and speech carries none of the detail that Light would be preserving. The Force mono switch is here too, and will most likely read (already mono) because you made that call in step 1. Each row reports the percentage saved, so you check the result before committing to an upload rather than after.

The last tool is behind the disclosure again. Open Other compatible tools and choose MP3 to YouTube.

Compress audio →

Step 5: Upload It as a Draft

YouTube will not take a bare MP3, so this step builds a video around it and uploads the result. Add a Cover image, or pull one From MP3 tags if your file already has artwork. Without one you get a plain background, which works but looks like an accident. Leave Video size on Landscape 1920 × 1080 unless you are posting to Shorts, and set Image fit to Fill for a photo or Fit for a logo, where letterboxing against your Background colour reads as deliberate rather than cropped.

Then Connect YouTube, write a Video title, and set Visibility to Private. It uploads as a draft either way, so nothing reaches your subscribers before you have listened to it, and you finish the description and publish from YouTube Studio yourself. The audio has to be under 100 MB, which is the constraint step 4 was quietly working towards. Keep the tab open while it renders and uploads.

MP3 to YouTube →

Why This Order

Three of the five positions are load-bearing, and one of them contradicts what the interface suggests.

Extract before anything else, because every tool after step 1 processes what you hand it. Trim the video first instead and you have a shorter video that still needs converting, having spent the processing on a picture you were about to discard.

Trim before you boost. The six minutes at the front are the noisiest part of the recording, and the gain you set is a judgement about the material you are keeping. Boost first and you spend the pass amplifying chair scrapes, then delete them.

Boost before you compress, and this is the one the download screen will happily let you get wrong. Compress audio suggests Boost volume as a next step, because the suggestions run both ways and for plenty of jobs that order is right. For this one it is not. Compression throws away detail to hit a size, and raising the gain afterwards amplifies the artifacts it introduced along with everything else. A compressed file has also had its peaks reshaped already, so the limiter has less to work with. Set the level on the full-quality file, then compress something that is already where you want it.

Compress before you upload for a reason unrelated to quality. The publishing step caps the audio at 100 MB, and your extracted file was already at 70. Trimming pulled it down, but a longer talk, or one extracted in stereo, would hit that wall.

What You End Up With

A 33 MB episode, forty-six minutes long, level enough to hear on a train, sitting in YouTube Studio as a private draft with your cover art on it. One file uploaded at the beginning, nothing downloaded at any point, and the recording never left your browser until the finished audio went to YouTube.

Ready to start? Open MP4 to MP3 →