Capture

Upload a recording

Uploading is the universal path into Citesvue: it works on every plan, with recordings from any tool. This guide covers formats and limits, what processing does, what you get back, and how failures behave.

Updated
30 August 2026
Read
6 min
For
Everyone; no paid plan required
The steps

Six steps, one of them optional.

From file to finished recording. Total hands-on time: about a minute.

  1. STEP 01

    Choose the session type

    From the dashboard, choose what the recording is: a UAT session, meeting, sales call, demo, or customer research. This label helps organise the library; Citesvue probes the actual file to determine whether it contains audio, video, or both.

  2. STEP 02

    Drop the file

    Drag the file onto the upload area or browse for it. Video (MP4, MOV, WebM, MKV) and audio (MP3, M4A, WAV) both work, up to 4 hours per recording. The size ceiling depends on your plan.

  3. STEP 03

    Optionally file it into a project

    Assign the recording to a project (a client, product, or sprint) so project-scoped questions can search it together with its documents. You can also move it later.

  4. STEP 04

    Let it upload, then leave

    The file uploads directly with a progress bar. Once you see "Upload complete - processing continues in the background" you can close the page; nothing depends on your browser staying open.

  5. STEP 05

    Watch the stages if you like

    The transcript-first processing view narrates what is happening: probing the media, preparing audio, transcribing speech, creating the recap and transcript artifacts, and making them searchable. Visual analysis is not part of this wait.

  6. STEP 06

    Open the finished recording

    As soon as the transcript-first pass completes, Recap, Transcript, Artifacts, and transcript Q&A are usable. If the file has video, choose Analyze visual evidence later to add Deep Dive, scenes, OCR, visual findings, and frame citations.

The Citesvue new content screen with upload source options
Steps 1 and 2. Choose what you are adding before you pick the file.
An upload configured with a title, project, and recording type
Step 3. Title, project, and recording type set. The type is what changes the result: a UAT session is mined for defects, a sales call for objections.

Core processing shows its transcript-first stages as it progresses. You can leave the page and return later; optional visual analysis is a separate action after core processing is ready.

What you get back

Four views of one recording.

Everything below is generated from your file and linked back to it, moment by moment.

Recap

A short summary with key outcomes. Each outcome seeks the player to the moment behind it, so the summary is a map of the evidence rather than a replacement for it.

Transcript

Full text with speakers separated and timestamps throughout, readable as a stream or grouped by speaker.

Deep Dive

An optional visual record: what was on screen at any moment, scene by scene, with every keyframe clickable to jump the player. It appears with progress when requested and becomes usable after the visual generation is verified. Audio-only files do not offer it.

Artifacts

Transcript-derived bugs, requirements, decisions, action items, risks, questions, and insights are available before visual analysis. A completed visual pass enriches those same findings without replacing their IDs or review state, and adds genuinely new visual-only findings when needed.

A recording recap with a meeting summary and key outcomes
The recap. Summary first, then the decisions and action items that were actually said - each one seeks the player to the moment behind it.
A recording transcript with speaker labels and timestamps
The transcript, split by speaker. Every timestamp is a control: select one and the player jumps there.
The Deep Dive tab showing on-screen frames from a recording
After you request and receive ready visual evidence for an eligible video, Deep Dive answers what was on screen at a given moment.
Limits

Three numbers worth knowing.

All limits are published on the pricing page and enforced exactly as published.

  • Visual-evidence hours

    Transcription does not consume this allowance. Before optional visual analysis starts, Citesvue shows the monthly allowance, used and reserved time, this recording’s duration, estimated remaining time, and reset date.

  • Stored recordings

    Plans include a recording library size. Deleting a recording frees its slot immediately, along with everything derived from it.

  • Size per file

    Per-recording ceilings by plan: Free 0.5 GB, Pro 4 GB, Team 4 GB, Business 4 GB, Enterprise 4 GB.

Full allowance tables live in the plans and usage guide and on pricing. If an upload is refused for quota, the message names the limit involved and the recording is not lost.

When it goes wrong

Four failure modes, all recoverable.

  • The upload fails partway

    Usually a dropped connection. Re-drop the file: uploads are safe to retry and never leave a half-recording in your library.

  • The file is too large

    The uploader tells you before any waiting: reduce the resolution or trim the recording, or upgrade for a higher per-file ceiling.

  • Transcript processing reports a failure

    The recording page shows the failed transcript state without hiding playback. A silent video does not pretend a transcript exists and can still be submitted for optional visual evidence.

  • It stops moving

    Progress is watched continuously; a stalled run restarts itself within minutes, and the page offers a manual restart as well.

More states and their fixes: troubleshooting guide.

Common questions

Uploading, answered.

  • MP4, MOV, WebM, MKV video and MP3, M4A, WAV audio, up to 4 hours per recording. Video can be submitted for optional visual analysis when you need citable on-screen evidence; audio-only files use the transcript-first flow.
  • It depends on plan: 0.5 GB on Free and 4 GB on paid plans today. The uploader checks before the upload starts, so an oversize file fails fast rather than at the end.
  • Only while the file itself uploads. Once you see the upload-complete message, processing runs server-side; the activity centre and a browser notification tell you when it is done, even days later.
  • Only video duration that you explicitly submit to Analyze visual evidence. Uploading, transcribing, reading recaps and transcripts, asking transcript questions, and exporting transcript content do not consume visual-evidence hours.
  • Yes. Any recording file works regardless of where it was made - the meeting assistant is a convenience, never a requirement. Download the recording from your meeting tool and upload it like any other file.
  • The media and everything derived from it - transcript, findings, search index entries, and any requested visual scenes, OCR, or frames - are removed together, and the library slot frees immediately.