Transcription pipeline
worker/README.md says how to run a worker. This page records what the pipeline does and what it measured, so capacity and quality decisions have numbers.
Stages
Section titled “Stages”| Stage | Tool | Where | What it produces |
|---|---|---|---|
| Enqueue | a worker’s claim on an empty queue (200 at a time), or the admin panel التفريغ الآلي ahead of time | API | one transcription-jobs row per eligible material (src/transcription/eligible.ts); gaps before caption redos, newest first |
| Claim | POST /api/pipeline/claim |
worker → API | up to N jobs, atomically; stale running jobs handed back first; a worker whose last three jobs all failed within 15 minutes is refused (blocked) |
| Download | urllib |
worker | the mp3 from files.kalelm.com in a temp dir |
| ASR | cohere-transcribe (CohereLabs 2B Arabic/English, Silero VAD, segment timing) |
worker GPU | cues with start times; the detector cuts on every pause, so raw cues average 5–8 words with many of one or two |
| Cue shape | build_cues in worker/worker.py, from the word timings (the library’s own builder is subtitle-sized and cuts at every comma) |
worker | never under CUE_MIN_WORDS (10); past that a cue ends at a sentence mark, a pause over CUE_MAX_GAP (1.5 s), or the soft ceilings CUE_MAX_CHARS 200 / CUE_MAX_SECONDS 15; CUE_HARD_SECONDS 40 cuts regardless; a short tail joins the cue before it. A re-queued done job replaces its earlier output, so changed settings apply by re-queueing |
| Upload | POST /api/pipeline/complete |
worker → API | transcript created (source: cohere), material flagged, Meilisearch updated by the hooks |
A result is written to results/job-<id>.json before upload and retried three times; a job that fails three claims stays failed until re-queued from the panel.
Eligibility
Section titled “Eligibility”Published, has audioUrl, no youtubeId (videos have caption tracks already), language ar or en, no transcript yet. On 2026-09-10: 30,974 materials, 12,696 known audio hours across the two thirds with a recorded duration, so roughly 19,000 hours in all, averaging 36 minutes each.
Measurements
Section titled “Measurements”All on the 59-minute fatawa episode unless noted; Apple M1 Max GPU via Metal.
| What | Result |
|---|---|
| ASR, default model | 3 min 09 s total, of which model load 1 min 30 s first time, 11 s cached; 18.8× real time |
| ASR + word-level alignment | 3 min 34 s; every word aligned by CTC, zero fallbacks; word boundaries land on the right audio 14 of 15 times, start edges ~80 ms tight |
| ASR on CPU (M1 Max) | 0.4× real time — impractical for the backlog |
| Shadda / tanween recovery | 71 % / 55 % against a noisy reference |
Sizing the backlog
Section titled “Sizing the backlog”| Where | Throughput | 19,000 hours takes |
|---|---|---|
| This Mac, full pipeline | ~10× | ~7 weeks |
| This Mac | 20–45× measured on real jobs | 2–4 weeks |
| Rented RTX 3090/4090 | 50–100× | ~1 week; $20–60 at spot prices |
After the backlog, new content is a few recordings a day; the panel’s local worker button or a Mac mini covers it. Nothing has to press enqueue: a worker that finds the queue empty pulls the next eligible materials in itself, so a published lesson is transcribed as soon as a worker is idle. The pause switch is the only thing that stops it.
Operating it
Section titled “Operating it”- Panel shows backlog, queue, running jobs with stage and heartbeat age, every worker that has ever reported (running or last seen, jobs in 24 h), last failures with a re-queue button, pause and the caption-redo toggle. A start/stop card for a worker on the backend host itself appears only where one can exist (
PIPELINE_WORKER_PYTHONset on a GPU box), never on the web host. - Pause stops workers taking new jobs; they finish what they hold.
- Stopped workers (SIGTERM, a restart) release their job at once through
/api/pipeline/release, attempt not counted. - Stale running jobs (no heartbeat for 10 min; workers heartbeat every 60 s from their own thread) return to the queue on the next claim. A native crash under the model kills a worker with nothing in its log; this is how its job comes back.
- Never overwrites: if a transcript appeared by another route while a job ran, the job is marked
skipped. - Stats per job (audio seconds, RTFx, stage timings, cue and word counts) are stored on the job row.
Post-edit, removed 2026-09-10
Section titled “Post-edit, removed 2026-09-10”An LLM pass (Ollama qwen3:8b, later Gemini) corrected spelling and punctuation after ASR, guarded so a rewritten line kept its original text. It changed ~4 % of words, cost ~6 minutes per hour of audio, and its failures (no key, provider overloaded) produced transcripts that were silently uncorrected. Removed as not worth its cost. The CAMeL hamza/shadda pass followed it out the same day (2026-09-10): with the ASR output used as-is, the worker needs nothing beyond the ASR library.