VoxFrame FAQ

Product information

Frequently asked questions

Clear answers about VoxFrame, pricing, supported features and secure customer delivery.

EN NL DE FR

Start here

Getting started and release

What is VoxFrame for?

VoxFrame is a Windows video player built around subtitles. It plays local video and legal IPTV/live streams, looks for existing human subtitles first, and can translate, synchronize or transcribe when a usable subtitle is missing.

It also includes subtitle quality checks, repair and review tools, export, LAN Watch and separate Cast support.

Can I download VoxFrame publicly yet?

Not yet. Public customer downloads open after the installer is signed and the final release checks have passed.

Until then, test builds use the protected account and in-app update routes and should not be shared.

Which files are supported?

Local video: MP4, MKV, AVI, MOV, WebM, M4V and TS. Subtitle files: SRT, VTT, WebVTT, ASS and SSA.

For IPTV, VoxFrame can open legal M3U/M3U8 playlists and supported direct stream URLs. VoxFrame does not bypass DRM or access restrictions.

Human subtitles first

Subtitle routes

How does Subtitle Assistant choose a subtitle?

VoxFrame first checks embedded tracks and local sidecar files. It then searches OpenSubtitles in the selected language. If that language is unavailable, it can translate a suitable human English subtitle. Speech-to-text generation is the last fallback.

The selected source, language and match confidence remain visible so you can review what was chosen.

Does live transcript work with local videos?

No. Live transcript / translate is for IPTV/live streams. A local video uses the normal subtitle workflows: search, generate AI subtitles, synchronize with fast STT, translate, repair and export.

Synchronized captions

Live IPTV

Which live subtitle setup should I choose?
  • Local best quality - recommended: the strongest tested local setup. Audio stays on the PC, a capable PC/GPU is recommended and synchronized playback starts at least 35 seconds behind live.
  • Local lighter setup - weaker PC/GPU: uses the selected local speech model with a lighter live setup and at least a 30-second synchronization buffer.
  • Own OpenAI API - direct: sends live audio directly to OpenAI with your saved key and uses at least a 30-second synchronization buffer. It appears when VoxFrame Cloud is off and a key is stored.
  • Cloud synced quality - recommended: the first of exactly two Cloud profiles; it processes 10-second speech chunks through VoxFrame Cloud.
  • Cloud longer dialogue context: the second Cloud profile; it uses 12-second chunks for more context but does not identify speakers.

Cloud profiles appear only when the account is entitled to use VoxFrame Cloud. Local profiles require an installed local STT engine and model.

Why is live playback delayed?

The delay is intentional: video, audio and captions share one buffered timeline. Local best starts at least 35 seconds behind live; Local lighter and Own OpenAI API - direct start at least 30 seconds behind. Cloud playback starts at least 60 seconds behind live. Transcription and optional translation time are added to those buffers.

The 0 / +1 / +2 / +3 / +5 seconds caption delay setting only moves captions further back; it does not remove the synchronization buffer.

Does a live stream create an SRT file?

No. Live captions are kept in memory for the active session and shown as an overlay. VoxFrame does not create or retain a live SRT. Temporary audio and buffer files are removed after stopping.

An active LAN Watch session can temporarily carry captions as WebVTT, but that is not a normal downloadable live subtitle file.

What does the music note mean?

The symbol means speech recognition returned an explicit music label. VoxFrame normalizes short labels such as “music” or “[background music]” to this symbol. It does not identify the song or performer.

Does Cloud longer dialogue context identify speakers?

No. The longer-dialogue Cloud profile gives transcription a slightly longer audio block, but it does not perform speaker diarization and does not name or label speakers.

Can I change target language while live?

Yes. New chunks use the newly selected target language without restarting the stream. Captions that were already shown are not translated again.

Choose your route

Privacy and billing

Do I need Premium Cloud?

No. VoxFrame remains usable with local functions and your own provider credentials. Premium Cloud is optional. A direct qualifying Cloud purchase also unlocks Player Pro automatically, so you do not need to buy the Player License first.

What is the difference between local, own API and Cloud?
  • Local: speech-to-text runs on this PC; performance depends on the installed model and hardware.
  • Own API: supported requests go directly to the provider using your key, and that provider bills your account.
  • VoxFrame Cloud: supported requests use the metered VoxFrame server route without exposing a VoxFrame provider key in the app.
Where does my audio go?

With a local live profile, audio processing stays on the PC. With Own OpenAI API - direct, audio chunks go directly to OpenAI and use your OpenAI billing; VoxFrame Cloud credit is not used. With a Cloud profile, audio chunks are sent over HTTPS to VoxFrame for OpenAI processing.

Cloud transcription usage is booked after a successful transcription. Live translation is a separate processing action.

Every screen

LAN Watch and Cast

What is the difference between LAN Watch and Chromecast?

LAN Watch starts a temporary HLS web player and shows a link and QR code for a browser on the same local network. It is the primary browser/phone/tablet/TV-browser route.

Chromecast, Google TV and DLNA are separate device routes. Availability and subtitle compatibility depend on the actual device, network and media codec; support for every TV or format is not guaranteed.

When something goes wrong

Troubleshooting and support

Why is a line sometimes missing or why do captions stop?

Speech recognition is not word-perfect. Silence, music, unclear speech, an overloaded PC, a slow provider or an unstable stream can produce an empty or late chunk. A chunk that has already missed its playback window may be skipped to prevent permanent lag.

Local and own-key inputs are read at real-time speed to avoid a startup flood. An isolated local transcription failure can leave a short gap while the session continues; three consecutive local provider failures stop the live workflow. The direct own-key view keeps short rapid cues visible for at least about 1.4 seconds and can retain up to three closely spaced cues; a real pause of more than about 0.8 seconds breaks that rolling history. Known short fake-credit phrases are filtered, but VoxFrame cannot reconstruct speech the provider did not transcribe.

What do Listening, No speech yet and Failed mean?
  • Listening: the live session is active and waiting for usable speech.
  • No speech yet: chunks were checked, but no clear speech was found; this is not automatically an error.
  • Degraded: part of the workflow is temporarily limited; source captions can continue when translation falls behind.
  • Reconnecting: VoxFrame is making a bounded retry.
  • Failed: a technical problem stopped the live workflow.
What should I send to support?

Send the app version, the action you started, the selected local/own-key/Cloud profile, spoken and target languages, the exact visible status and a sanitized support log from Settings > About > Export safe support log.

Never send license keys, provider keys, access tokens, payment details or unrelated personal files.