Your recordings had a second transcript all along.
Video frames carry a second layer of evidence - error messages, field labels, URLs, and button states. Request visual analysis when you need that layer; the transcript-first experience is already usable while it runs.
The most important noun in the sentence is rarely in the transcript.
Walk through the hero frame: the user said “it.” The screen showed “Payment declined - code 402.” The UI state is error. The URL reveals which environment they were in. Each layer is extracted, indexed, and queryable on its own - the value is in how they align.
Speech-only tools save the words. Not the thing.
When a customer says “this is broken,” Fireflies saves the words. Zoom AI saves the words. Otter saves the words. None of them know what this was. The most important noun in the sentence is not in the transcript.
Visual evidence can resolve it when the user requests that analysis. Once verified, it enriches existing artifacts with the actual on-screen target without replacing their IDs or review state.
Six signals from every frame, structured and queryable.
On-screen text
Buttons, labels, menus, URLs, form fields, headings.
Error states
Toasts, banners, red form validation, HTTP errors visible in DevTools.
UI elements
Primary CTAs, inputs, navigation, modals, tables, menus.
App context
Which product, page, environment, build is visible.
Selection & focus
Which element the user was interacting with at the moment.
Change detection
What appeared, what disappeared, frame to frame.
Requested visual evidence aligns speech with what was on screen.
- 01
Use the core transcript
Permanent timestamped speech and transcript artifacts stay intact.
- 02
Sample keyframes
The requested visual generation reads relevant frames for OCR and UI classification.
- 03
Bind to time
Visual findings are aligned to the available transcript timing without replacing the core transcript index.
- 04
Resolve deixis
“This,” “that,” “here” - resolved to the nearest on-screen target.
Five capabilities that literally cannot exist without visual grounding.
- Bug artifacts that carry the actual error screen, not a generic summary.
- Q&A answers that include the frame, not just the quote.
- Requirement extraction that references the specific screen being discussed.
- Search queries like “show me every moment a 500 error appeared on screen.”
- Evidence cards with embedded frame thumbnails - proof a reader can see in one glance.
The OCR is not optional.
Many tools claim “visual features” and mean “we store the video alongside the transcript.” That is not visual intelligence - that’s a video file.
Citesvue’s visual layer is structured data: every frame produces a queryable record of what text appeared, where, in what state. You can search it, filter on it, and cite it the same way you cite a transcript segment.
A UAT bug report - assembled from one frame and one sentence.
A client walks through checkout. At 14:23.4 the screen shows Payment declined - code 402. At 14:24.1 the client says “see, it doesn’t tell me why.” After visual evidence is requested and ready, the bug artifact can join error text from the frame with the user quote, timestamp, URL on screen, and a click-to-jump link to that exact second of the recording.
Visual intelligence comes with rules.
Frames, OCR text, and UI state inherit the same workspace-scoped access controls as the transcript itself, and deleting a recording removes every derived frame and extraction with it. Automated redaction and region masking are on the roadmap, described here as roadmap because they have not shipped.
What enterprise buyers ask about visual intelligence.
- English at GA. Arabic, French, Spanish, German and others on the roadmap. Reach out if your team needs a specific language scoped in earlier.
- Not yet. Region and application masking rules are on the roadmap and described on this page as roadmap. Today the control is what you choose to record, plus deletion: removing a recording removes every derived frame and extraction with it.
- Confidence scores are exposed per-frame. Low-confidence text is flagged for review rather than silently used in artifacts or answers.
- Yes. After an eligible recording has ready visual evidence, captured frames and their extracted text are viewable on the recording itself. The frames cited by an artifact are included in exports and pushes only when that visual evidence exists.
- No. Recap, transcript, transcript-derived artifacts, and transcript Q&A are available first. Visual evidence runs only when requested, with its own allowance and progress, while those features remain usable.
- Not yet. Frames and OCR text are available in the UI and in exports; a public API is on the roadmap.
Your next recording could be
your most
valuable asset.
Or it could sit in a Drive folder nobody opens again. The difference is whether it has citations attached.
- SetupOne drag-and-drop upload, or send the notetaker. No plugins.
- First insightCited Q&A on a 60-min recording in under 6 minutes.
- Cancel anytimeFull data export, full right to erasure.