Skip to content

Transcription pipeline

worker/README.md says how to run a worker. This page records what the pipeline does and what it measured, so capacity and quality decisions have numbers.

Stage Tool Where What it produces
Enqueue a worker’s claim on an empty queue (200 at a time), or the admin panel التفريغ الآلي ahead of time API one transcription-jobs row per eligible material (src/transcription/eligible.ts); gaps before caption redos, newest first
Claim POST /api/pipeline/claim worker → API up to N jobs, atomically; stale running jobs handed back first; a worker whose last three jobs all failed within 15 minutes is refused (blocked)
Download urllib worker the mp3 from files.kalelm.com in a temp dir
ASR cohere-transcribe (CohereLabs 2B Arabic/English, Silero VAD, segment timing) worker GPU cues with start times; the detector cuts on every pause, so raw cues average 5–8 words with many of one or two
Cue shape build_cues in worker/worker.py, from the word timings (the library’s own builder is subtitle-sized and cuts at every comma) worker never under CUE_MIN_WORDS (10); past that a cue ends at a sentence mark, a pause over CUE_MAX_GAP (1.5 s), or the soft ceilings CUE_MAX_CHARS 200 / CUE_MAX_SECONDS 15; CUE_HARD_SECONDS 40 cuts regardless; a short tail joins the cue before it. A re-queued done job replaces its earlier output, so changed settings apply by re-queueing
Upload POST /api/pipeline/complete worker → API transcript created (source: cohere), material flagged, Meilisearch updated by the hooks

A result is written to results/job-<id>.json before upload and retried three times; a job that fails three claims stays failed until re-queued from the panel.

Published, has audioUrl, no youtubeId (videos have caption tracks already), language ar or en, no transcript yet. On 2026-09-10: 30,974 materials, 12,696 known audio hours across the two thirds with a recorded duration, so roughly 19,000 hours in all, averaging 36 minutes each.

All on the 59-minute fatawa episode unless noted; Apple M1 Max GPU via Metal.

What Result
ASR, default model 3 min 09 s total, of which model load 1 min 30 s first time, 11 s cached; 18.8× real time
ASR + word-level alignment 3 min 34 s; every word aligned by CTC, zero fallbacks; word boundaries land on the right audio 14 of 15 times, start edges ~80 ms tight
ASR on CPU (M1 Max) 0.4× real time — impractical for the backlog
Shadda / tanween recovery 71 % / 55 % against a noisy reference
Where Throughput 19,000 hours takes
This Mac, full pipeline ~10× ~7 weeks
This Mac 20–45× measured on real jobs 2–4 weeks
Rented RTX 3090/4090 50–100× ~1 week; $20–60 at spot prices

After the backlog, new content is a few recordings a day; the panel’s local worker button or a Mac mini covers it. Nothing has to press enqueue: a worker that finds the queue empty pulls the next eligible materials in itself, so a published lesson is transcribed as soon as a worker is idle. The pause switch is the only thing that stops it.

  • Panel shows backlog, queue, running jobs with stage and heartbeat age, every worker that has ever reported (running or last seen, jobs in 24 h), last failures with a re-queue button, pause and the caption-redo toggle. A start/stop card for a worker on the backend host itself appears only where one can exist (PIPELINE_WORKER_PYTHON set on a GPU box), never on the web host.
  • Pause stops workers taking new jobs; they finish what they hold.
  • Stopped workers (SIGTERM, a restart) release their job at once through /api/pipeline/release, attempt not counted.
  • Stale running jobs (no heartbeat for 10 min; workers heartbeat every 60 s from their own thread) return to the queue on the next claim. A native crash under the model kills a worker with nothing in its log; this is how its job comes back.
  • Never overwrites: if a transcript appeared by another route while a job ran, the job is marked skipped.
  • Stats per job (audio seconds, RTFx, stage timings, cue and word counts) are stored on the job row.

An LLM pass (Ollama qwen3:8b, later Gemini) corrected spelling and punctuation after ASR, guarded so a rewritten line kept its original text. It changed ~4 % of words, cost ~6 minutes per hour of audio, and its failures (no key, provider overloaded) produced transcripts that were silently uncorrected. Removed as not worth its cost. The CAMeL hamza/shadda pass followed it out the same day (2026-09-10): with the ASR output used as-is, the worker needs nothing beyond the ASR library.