Guide baseline: v0.9.11
Video and Audio
Create draft transcripts and captions, then check them against the full recording and its educational purpose.
Transcription and caption outputs
The transcription route uses the Whisper path and requires its model and media dependencies, including FFmpeg. It can generate WebVTT and SRT text. Review timing, speaker identification, technical terms and meaningful non-speech audio. Automatic captions, including Aelira’s output, need verification and correction; the source being automatic does not by itself decide adequacy.
W3C guidance on accurate captions and required audio information
Audio description in this release
POST /education/multimedia/scan exposes generate_audio_descriptions and generate_spoken_descriptions as query options. Video keyframes can be described through the authorized workspace vision provider. Optional spoken output requires the TTS processor, and media assembly requires FFmpeg. Missing providers, frames or TTS can leave descriptions empty or spoken output unavailable.
Delivery and quality boundaries
The processor has companion audio and video generation paths, but the stored scan response exposes description counts and transcription/caption data, not a guaranteed downloadable described-video deliverable. Check the actual returned artifacts before planning distribution. Generated companion files remain unverified: the route sets manual_review_required and verification_passed=false. A source score does not verify caption accuracy or whether audio description is appropriate. Enabled flashing detection samples the first 30 seconds and cannot clear a whole recording.
Supported formats
MP4, MOV, AVI, MKV, WebM, MP3, WAV, M4A and OGG are accepted subject to content validation and installed codecs. MAX_FILE_SIZE_VIDEO defaults to 500 MiB and is configurable. Check the final caption track in the target player and review the complete recording before publishing.
Authenticated API usage
Use your deployment URL and an authorized API key. The localhost URL below is for a locally running API.
curl --fail --silent --show-error \
-H "Authorization: Bearer $AELIRA_API_KEY" \
-F "[email protected]" \
"http://localhost:8000/education/multimedia/transcribe?generate_captions=true&whisper_model=base"The request returns a scan_id while work runs in the queue. Poll GET /education/scans/{scan_id}/progress, then GET /education/scans/{scan_id}. Completed scan.result.structure.transcription_output contains captions.webvtt and captions.srt when generated; transcription segments are limited to the first ten in this stored preview, while full_text contains the joined transcript. No output_format form field is used. The transcribe route disables audio-description generation.
Evidence and review
This guide describes released Core source; availability in a hosted deployment depends on its release and configuration. Review output before publishing it.
Read the v0.9.11 source and limits