You are mid-episode and your guest describes a dish, a building, a play, a piece of footage. Everyone knows the thing being described. Nobody can see it. On a video podcast, that is a cutaway you will hunt for later. On an audio show, it is a link you meant to put in the show notes and forgot.
The Media Finder agent handles both cases by surfacing the visual while you are still talking about it.
What the Agent Does
Media Finder is one of six live agents in Podmod's Agents menu. Its role in the interface: Surfaces images and videos related to what you said.
It produces two kinds of cards during a recording.
Images. Mention a restaurant on a food show and you get plating photographs. Mention a building, a product, a person, or a place, and you get visual references for it.
Images Noma plating, Copenhagen
Video. Mention a specific moment and you get clips, with titles, rather than a search you have to run yourself.
Video Curry Finals last-second three Steph Curry game-winner, NBA Finals Top clutch threes of the decade Warriors Finals highlights reel
Both appear silently in the panel beside your transcript, on the host side only. Nothing plays. Your guest sees nothing.
Why This Matters More Than It Used To
Podcasting became a visual medium while a lot of workflows stayed audio-only.
Edison Research now finds that YouTube is the most-used podcast service among weekly consumers, with video reshaping how audiences find and consume shows. The same research puts monthly podcast consumption at an all-time high of 58% of Americans age 12 and older.
The practical consequence for a working show is that visuals are no longer a bonus deliverable. They are part of the episode: cutaways in the video edit, thumbnails, clip B-roll, and images on the episode page. Sourcing all of that after the fact, from a finished recording, is one of the least enjoyable jobs in podcasting.
Turning It On
Open the Agents menu in the control bar and toggle Media Finder. Like every agent, it can be switched on or off before or during a take.
Mid-session toggling is worth using here. Some segments are visual and some are not. A conversation about a physical place, a product, or a piece of footage benefits. Twenty minutes of abstract discussion produces cards you will not use.
Three Ways to Use the Cards
During the recording, as a prompt. This is the underrated one, and it applies even to audio-only shows. Seeing the thing you are describing makes you describe it better. Hosts get vaguer the further they get from a concrete memory of something. A photograph on screen pulls the description back toward specifics.
During the recording, as a live cutaway. On a video show where you are sharing a screen, a surfaced image or clip is something you can put up while the conversation is on it, rather than promising the audience you will link it later.
After the recording, as a source list. Every card is saved with the session and appears in Session Viewer, anchored to the transcript line that triggered it. This is where most of the value lands for most shows.
The Post-Recording Workflow
Open the session in Session Viewer. The Transcript tab shows the conversation with agent cards attached to the lines that produced them.
For a video edit, this gives you a shot list. Scroll the transcript, and every point where a visual was surfaced is a candidate cutaway with the reference already found and timestamped to the exact moment it belongs. Compare that to the alternative, which is watching your own episode back with a notepad.
For the episode page, the image and video cards are your embeds and links. Podcast discoverability increasingly depends on the episode page carrying real content: a title, a summary, chapters, and a transcript that turns a conversation into indexable text. Visual references belong in that mix.
For clips, the cards tell you which moments had something to look at. A clip built around a moment with a visual reference generally performs better than a clip of two people talking.
Pairing It With Markers
If you are recording in Podmod Studio, you have three marker hotkeys during a take:
- M for a generic marker
- C for "clip this"
- X for "edit out"
The combination worth building into your habits: when a Media Finder card appears for something you know you want in the video edit, press C. Now you have a timestamped flag and a found visual on the same moment. When you sit down to edit, those two things together are most of the work already done.
Which Shows Get the Most From It
Food and travel. Places and dishes are the entire subject and almost nobody can picture them from a description alone.
Sports. Specific plays, specific games. The footage exists and finding it mid-conversation rather than mid-edit is a large time saving.
Tech and product shows. Screenshots, hardware, interfaces. Reviews are hard to follow without seeing the thing.
Culture, film, and music. Clips, covers, performances, stills.
History. Photographs, maps, documents. A visual reference does more for a historical episode than another paragraph of description.
Where it helps least: abstract advice shows, therapy and coaching formats, and pure interview shows about ideas rather than things. Some hosts on these formats find the cards distracting, which is a fair reason to leave the agent off.
Practical Notes
Verify before you publish. A surfaced image is a reference, not a cleared asset. Check the source and the rights before you put anything in a published video or on an episode page. The agent finds material fast; it does not resolve licensing for you.
Fewer agents means readable cards. Media Finder running alongside five other agents on a fast conversation produces a panel you cannot scan. If visuals are the priority for an episode, run Media Finder plus one or two others rather than the full set.
It responds to specificity. "That restaurant in Copenhagen" produces weaker results than naming it. This is a useful discipline anyway, since specific references are also better for listeners and better for search. Hosts who run Media Finder for a few episodes tend to get noticeably more precise in how they reference things, because the feedback is immediate: name the thing and a card appears, gesture at it vaguely and nothing does.
Do not stop to look properly. The temptation the first few times is to actually examine a card mid-conversation. Resist it. A glance is enough to register that the reference exists and was found. Everything is saved with the session, so proper evaluation belongs afterward, in Session Viewer, when you are not also hosting.
Nothing is audible. No cues, no pings, no interruption to the conversation. This holds for every agent in Podmod.
A Realistic Session
A food show recording an interview with a chef:
- Load the chef's bio and their restaurant's press coverage into the Context panel before recording.
- Turn on Media Finder, Prep Desk, and Producer Coach.
- Record. When the chef describes a signature dish, an image card appears. Press C to flag the moment.
- Keep talking. Do not stop to look at anything properly.
- Afterward, open Session Viewer. Your C markers are the clip list. The image cards attached to those timestamps are the visuals.
- Generate show notes and chapters from the Episode Assets tab.
- Build the episode page with the images already sourced, and cut two vertical clips from the flagged moments.
The time saved is not dramatic on any single step. It is that the sourcing, flagging, and structuring all happened during the conversation instead of forming a second job afterward.
Turning Cards Into an Episode Page
The episode page is where surfaced media pays off for audio-only shows, and it is worth being deliberate about.
A page with a player and two sentences of description gives search engines almost nothing to work with. A page with a real summary, chapter timestamps, a transcript, and the visual references from the conversation gives them a great deal. The transcript alone converts a half-hour conversation into thousands of indexable, quotable words, and every reference you cite becomes something an AI-driven search result can attribute back to you.
Media Finder cards contribute the visual half of that. Work through the session transcript in Session Viewer, pull the images and clips that were surfaced, and place them next to the section of the show notes where they belong. Because each card is anchored to the transcript line that triggered it, you already know where each one goes.
Do this once and you will notice the second-order effect: it makes you reference specific, nameable things more often during recordings, because you know the reference will become an asset rather than a loose end.
Try It on a Visual Episode
Pick an episode where you know you will be describing things people cannot see. Turn on Media Finder, press C when something looks worth using, and see what your session looks like when you open it afterward.
Start at app.podmod.ai.