Remote Podcast Interviews: From Zoom and Riverside to a Clean MP3
The interview went great. Then the files arrived: your side is a WAV from your audio interface, the guest's side is an M4A that Zoom produced, there is a phone-recorded backup "just in case," and every one of them is at a different volume, a different format, and — worst of all — a slightly different start time. Remote podcast production is not hard because of the conversation; it is hard because of the pile of mismatched files the conversation leaves behind. Here is the workflow that turns that pile into one broadcast-ready episode.
Know what each platform actually hands you
The cleanup starts before the call, with knowing what you will receive:
- Zoom: records M4A audio (and MP4 video). With "record a separate audio file for each participant" enabled in settings, you get one M4A per speaker — enable it, always. Cloud recordings compress harder than local ones.
- Riverside, SquadCast, Zencastr: record locally on each participant's machine and upload progressively, so you get uncompressed or lightly compressed WAV per track at 44.1 or 48 kHz — the best-case scenario.
- Phone voice-memo backups: M4A on iPhone, M4A or AMR on Android, mono, recorded from across a desk. A last resort that has saved many episodes.
The takeaway: your session will almost always mix M4A and WAV sources at mixed sample rates. That is normal, and the workflow below assumes it.
Step 1: collect everything, then standardize the format
First, get every file into one folder and rename immediately — ep42-host.wav, ep42-guest-zoom.m4a, ep42-guest-phone.m4a. Guests delete their local recordings within days; collect while the call is still warm.
Then convert the M4A files so everything speaks the same language. If you edit in a DAW that imports M4A natively you can skip ahead, but many editors — and almost all quick-and-simple tools — behave better with uniform inputs. An M4A to MP3 converter handles the Zoom and phone files in seconds; if a guest sends a voice-memo file and you are not sure what is inside it, the voice memo conversion FAQ covers the common cases. Convert at a generous bitrate (192 kbps or higher) — this is an intermediate file, and you do not want to stack a harsh encode on top of Zoom's own compression. Match sample rates too: if your WAV is 48 kHz and the M4A decodes at 44.1 kHz, pick one rate for the session or your tracks will slowly slide against each other.
Step 2: align the tracks
Drop all tracks into your editor and line them up. Two tricks make this fast:
- Record a sync marker. At the top of every call, have everyone count down and clap into their own mic. That clap is a visible spike in every waveform — align the spikes and you are done.
- Check the end, not just the start. Consumer recording apps occasionally drift over a long session. If voices are aligned at minute 1 but the guest answers questions slightly early by minute 50, one file's clock drifted; a tiny time-stretch on that track (most DAWs offer this) fixes it.
Once aligned, mute the backup tracks — do not delete them. You will want the phone recording the moment you discover Zoom dropped four seconds of the best answer.
Step 3: level the voices
Remote tracks arrive at wildly different volumes: your interface was gain-staged properly, the guest's laptop mic was not. Normalize each speaker's track so nobody dominates. Loudness normalization — not simple peak normalization — is what you want, because a quiet speaker with one loud laugh defeats peak-based leveling entirely; the loudness explainer covers the difference. If you are working outside a DAW, run each exported voice track through a loudness normalizer before assembly, targeting the same LUFS value for every speaker.
Step 4: edit, then assemble
Make your content edit with the tracks separate — cutting crosstalk and cleaning up interruptions is only possible per-speaker. Then mix down. If your episode structure is simple (intro, interview, outro as separate finished files), you can assemble the final sequence with an MP3 merger instead of a DAW timeline: normalize the segments to the same loudness first, then join them in order. Speech-only interviews should ship mono — the guest was effectively mono from the moment Zoom recorded them, and mono halves your file size; the mono vs stereo FAQ explains when stereo actually earns its bits.
Export settings for interview episodes
Once the assembled episode is mixed down, resist the urge to export at the intermediate bitrate you used during cleanup. Interview audio is speech, and speech is forgiving: 96 kbps mono is transparent for a two-person conversation, and 128 kbps stereo is only worth it if you have stereo music beds in the intro. A one-hour interview lands around 43 MB at 96 kbps mono versus 58 MB at 128 kbps stereo — a meaningful difference across a season of episodes for both your hosting bill and your listeners' data plans. Keep the pre-mixdown session files (or at least the aligned, converted tracks) for a few weeks after publishing; corrections and clip requests always arrive after the episode is live, never before.
Prevention: the five-minute pre-call setup
Every hour of cleanup traces back to a skipped minute of setup. Before the next interview:
- Ask guests to wear wired headphones (kills echo-cancellation artifacts) and sit in a soft-furnished room, not a kitchen.
- Enable per-participant recording in Zoom, or use a local-recording platform for any interview you care about.
- Do the countdown clap.
- Have guests start a phone voice memo as backup — and actually collect it afterward.
The remote-formats problem never fully goes away; guests will always send you something strange. But with a standard collect-convert-align-normalize-merge pipeline, "strange" costs you ten minutes instead of an evening — and the listener hears one conversation, in one room, at one volume, which is the entire illusion remote podcasting is trying to pull off.