VALORAE Arc
Craft

Turning a horizontal podcast into vertical video without losing the conversation

Framing decisions for converting two-camera horizontal podcast footage into 9:16, including when to cut, when to split the frame, and why static crops underperform.

Short answer

Convert horizontal podcast footage to vertical by cutting between tracked single-speaker frames rather than cropping the wide shot. Cut on the exchange, not on a timer. Split-screen works for genuine back-and-forth and fails for monologue. A static centre crop is the cheapest option and it is why most repurposed podcast clips look like leftovers.

A horizontal two-shot cropped to 9:16 gives you two people at the edges and an empty table in the middle. Everyone has seen this clip and everyone scrolls past it.

The conversion is a framing job, not a crop job.

Cut, do not crop

The default should be full-frame single-speaker shots, cut between as the conversation moves.

Vertical is a tall narrow frame. It suits one face. Filling it with one person, properly framed, gives you a shot that looks made for the format rather than salvaged from another one.

Which means the edit is now a decision about where to cut, and that is the whole craft.

Cut on the exchange

The rule that produces good rhythm: cut when the conversational energy moves, not when a timer says to.

Cut to a speaker as they begin an important line, not a beat late. Cut to the listener when their reaction is the content. Hold on someone through a pause when the pause is doing work.

Automatic reframing tools cut on audio detection. That gets you a cut every time someone speaks, including on a two-word interjection, which produces a jumpy edit that fights the conversation instead of following it.

When split-screen works

Stacked frames, one speaker above the other, work in specific cases:

  • Rapid back-and-forth where cutting would be too busy
  • Disagreement, where seeing both reactions is the point
  • A punchline landing, where the reaction is half the joke

They fail for:

  • Any sustained single-speaker passage
  • Clips where one person is clearly the subject
  • Anything where half the frame shows someone nodding politely

Used as a default because it avoids making cut decisions, split-screen halves your visual real estate for no benefit.

Framing within the shot

Eyes in the upper third. Standard portrait framing, and it still applies. Centring the face vertically leaves dead space above the head.

Room in the direction they face. If someone is angled left, give space on the left. Framing them tight against the edge they are looking toward feels wrong even to viewers who cannot articulate why.

Leave the caption zone. Captions occupy the middle band. Frame so nothing important sits behind them.

Respect the platform safe zones. Roughly the top 10 percent and bottom 25 percent of the frame will be covered by interface on at least one platform. Nothing load-bearing goes there.

Resolution and source

Cropping into a wide shot to isolate one speaker throws away pixels. Whether that is acceptable depends on the source.

Recording at 4K and delivering 1080p vertical gives comfortable room to punch in on individual speakers without visible softness. Recording at 1080p and cropping to a single speaker means upscaling, which is visible.

If the podcast is new and the setup is being decided, this is the single most useful production choice: shoot larger than you deliver, so the vertical edit has room to work.

Motion

Some movement helps a static two-person setup read as designed rather than salvaged.

Slow push in over a long shot. A slight reframe as someone leans forward. Both are subtle enough not to be noticed as effects and they keep the frame from feeling frozen.

Avoid: constant zoom oscillation, shake effects, and fast whip transitions between speakers. These read as a template rather than an edit, and in a feed full of the same templates they make the clip look generic.

A workable default

  • Single-speaker full-frame shots as the base
  • Cuts driven by the exchange, roughly every three to six seconds
  • Split-screen reserved for genuine back-and-forth
  • Eyes in the upper third, space in the direction of gaze
  • Nothing important in the top 10 or bottom 25 percent
  • A slow push on any shot held longer than about six seconds
  • Source recorded larger than the delivery format

That set of decisions is most of the difference between a clip that looks native to vertical and one that looks like a widescreen video wearing a costume.

Frequently asked questions

Is auto-reframe good enough?

Automatic reframing tools handle the mechanical part well and they cut on audio detection, which means they cut on who is speaking rather than on what matters. They are a good first pass that still needs a human deciding where the cuts belong.

Should I show both people at once?

Only when both are participating. Split-screen during a two-minute monologue wastes half the frame on someone listening. Use it for rapid exchange, disagreement and reactions, and cut to single frames for anything sustained.

What if I only have one wide camera?

Punch in digitally on the speaker, provided the source resolution supports it. Shooting in 4K for a 1080p vertical output gives enough room to crop into individual speakers without visible softness. This is the main reason to record at higher resolution than you deliver.

How often should I cut?

On the exchange rather than on a clock. A useful check is that no single framing sits still long enough for the viewer to notice it is still. In practice that tends to mean every three to six seconds, driven by the conversation.

Want your podcast turned into clips that hold?

We cut, caption and distribute short-form for podcasts and founders. Bring one episode and we will show you what comes out of it.

Book a call

Keep reading