Run a worker on a new machine
This page is for a Claude session (or a person) sitting at a machine that has never run a worker. Follow it top to bottom; each step says how to check it worked. Nothing here needs access to the servers: a worker only makes outbound HTTPS calls to the API.
What a worker is
Section titled “What a worker is”A Python process in worker/ that asks the backend for a job, downloads the mp3, runs Cohere ASR on the GPU, builds cues of ten words or more from the word timings, and uploads them. The backend hands each job to exactly one worker, so any number of machines can run one. Details and measurements: Transcription pipeline.
0. Is this machine worth it?
Section titled “0. Is this machine worth it?”| Machine | Speed | Verdict |
|---|---|---|
| Apple silicon Mac (M1 or later, 16 GB+) | 12 to 28× real time | yes |
| Linux box with an NVIDIA GPU (8 GB VRAM+) | similar or better | yes |
| Any machine on CPU only | 0.4× real time | no: one hour of audio takes two and a half |
| Windows | untested | use WSL2 with the Linux steps, or skip |
Memory: ASR needs about 5 GB.
1. Get the code
Section titled “1. Get the code”git clone https://github.com/mohanad-a/kalelm-backend.git kalelmcd kalelm/workerThe repository is private; the owner’s GitHub account (mohanad-a) has to grant access or clone it for you. Only worker/ is needed on this machine.
2. Install
Section titled “2. Install”macOS:
brew install ffmpeg uvUbuntu:
sudo apt install -y ffmpegcurl -LsSf https://astral.sh/uv/install.sh | shThen, in worker/:
uv venv --python 3.12 && source .venv/bin/activateuv pip install -r requirements.txt --torch-backend=autoCheck: python test_worker.py prints ok.
3. Two things only the owner can give you
Section titled “3. Two things only the owner can give you”Ask for them; do not guess and do not put them in a file that is committed.
- The API secret.
CRON_SECRETfrom the backend’s environment (Hangar → sitekalelm-api→ env, or ask the owner). A wrong secret shows up asclaim failed: HTTP Error 403. - A Hugging Face login that has accepted the model’s terms. The ASR model
CohereLabs/cohere-transcribe-arabic-07-2026is gated. Runhf auth loginand paste a token from an account that accepted it (the owner’s is namedmacunder usercoder0x9). Check:hf auth whoaminames the account.
Also confirm the API address. As of 2026-09-10 the test deployment answers at https://kalelm-api.mohanad.xyz; the docs index page’s Hosts table has the current one.
5. First run: one job
Section titled “5. First run: one job”KALELM_API=https://kalelm-api.mohanad.xyz CRON_SECRET=… WORKER_NAME=<this-machine> DEVICE=mps \python worker.py --onceIf you keep it running with nohup rather than launchd or systemd (step 6, which restart it themselves), run run.sh with WORKER_PYTHON pointing at the venv’s python: it restarts the worker after a native crash, which the model stack does now and then with nothing in the log.
DEVICE is mps on a Mac, cuda on Linux with NVIDIA, and can be left unset to auto-detect. WORKER_NAME must be unique per machine; the panel lists workers by it.
Expected output, within a couple of minutes:
14:21:56 worker <name> -> https://kalelm-api.mohanad.xyz14:21:56 claimed job 5: '…'14:23:11 job 5 material 49014: done (81 cues, 73.4s, RTFx 12.0)RTFx is how many times faster than real time it ran. If instead you see:
claim failed: HTTP Error 403— wrong secret, or a request without the worker’s own user agent (the code sets one; do not replaceurllibcalls with something that sends a different one, Cloudflare blocks Python’s default).nothing to do (paused)— the queue is paused in the panel; the owner unpauses it.nothing to do— the queue is empty and nothing is eligible; that is fine.- an out-of-memory error during
asr— setDEVICE=cputo prove the rest works, then find a machine with more memory.
6. Keep it running
Section titled “6. Keep it running”macOS, ~/Library/LaunchAgents/com.kalelm.worker.plist (fill in the paths, the secret and the name; PATH must include where ffmpeg and ollama live):
<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd"><plist version="1.0"><dict> <key>Label</key><string>com.kalelm.worker</string> <key>ProgramArguments</key> <array><string>/ABS/PATH/kalelm/worker/.venv/bin/python</string><string>worker.py</string></array> <key>WorkingDirectory</key><string>/ABS/PATH/kalelm/worker</string> <key>EnvironmentVariables</key><dict> <key>KALELM_API</key><string>https://kalelm-api.mohanad.xyz</string> <key>CRON_SECRET</key><string>…</string> <key>WORKER_NAME</key><string>mac-mini</string> <key>DEVICE</key><string>mps</string> <key>PYTHONUNBUFFERED</key><string>1</string> <key>PATH</key><string>/opt/homebrew/bin:/usr/bin:/bin</string> </dict> <key>RunAtLoad</key><true/> <key>KeepAlive</key><true/> <key>StandardOutPath</key><string>/ABS/PATH/kalelm/worker/worker.log</string> <key>StandardErrorPath</key><string>/ABS/PATH/kalelm/worker/worker.log</string></dict></plist>launchctl load ~/Library/LaunchAgents/com.kalelm.worker.plist # starts now and at every loginlaunchctl unload ~/Library/LaunchAgents/com.kalelm.worker.plist # stopsLinux, /etc/systemd/system/kalelm-worker.service:
[Unit]Description=Kalelm transcription workerAfter=network-online.target
[Service]WorkingDirectory=/opt/kalelm/workerEnvironment=KALELM_API=https://kalelm-api.mohanad.xyzEnvironment=CRON_SECRET=…Environment=WORKER_NAME=gpu-boxEnvironment=PYTHONUNBUFFERED=1ExecStart=/opt/kalelm/worker/.venv/bin/python worker.pyRestart=alwaysRestartSec=10
[Install]WantedBy=multi-user.targetsudo systemctl enable --now kalelm-workerjournalctl -u kalelm-worker -fKaggle (free GPU, no machine needed)
Section titled “Kaggle (free GPU, no machine needed)”Kaggle gives every account 30 GPU-hours a week for nothing, with sessions capped at 12 hours. scripts/kaggle/kalelm-worker.ipynb is the worker as a notebook: it copies worker.py and secrets.env from the private Kaggle dataset kalelm-worker, installs ffmpeg and cohere-transcribe, and runs for MAX_MINUTES (230, under four hours), stopping between jobs. Accelerator GPU T4 x2 (machine_shape: NvidiaTeslaT4 in scripts/kaggle/kernel-metadata.json; a push without it lands on a P100). The torch that cohere-transcribe installs has no kernels for the P100’s generation (no kernel image is available for execution on the device), so on a P100 the notebook reinstalls the same torch from the CUDA 12.6 wheel index, two extra minutes. On a T4 expect 10 to 40× real time, roughly 100 archive-hours per run. Measured 2026-09-10: setup 2.5 minutes, first job done 4 minutes after start.
Weekly hours are per account, however many GPUs or sessions: parallel runs spend the quota faster, they do not add to it. Two Kaggle limits shape the design. Its scheduler refuses GPU notebooks, so the daily run is started from outside through the API (root’s crontab on the search VPS, infra/search-vps/kaggle-run.sh, 03:00 UTC). And runs started through the API never receive the notebook’s Add-ons → Secrets, so CRON_SECRET and HF_TOKEN travel in secrets.env inside the private dataset instead; the notebook falls back to Kaggle’s Secrets only when that file is absent. If the dataset ever leaks, rotate CRON_SECRET in Hangar and the Hugging Face token.
Once, in a browser on kaggle.com: Settings → Phone verification (notebooks get internet only on verified accounts) and Settings → API → Create access token, kept as KAGGLE_API_TOKEN on the pushing machine and in /srv/kalelm/.env on the VPS.
From the repository, whenever worker.py or the notebook changes, or a run is wanted now:
uv tool install kaggleKAGGLE_API_TOKEN=… CRON_SECRET=… scripts/kaggle/push.shIt uploads a new dataset version (worker, secrets), saves the notebook, which starts a run, and copies the notebook to the VPS for the cron. HF_TOKEN defaults to the token hf auth login stored. A T4 x2 session has two GPUs, so the notebook runs two workers, kaggle-<container>-gpu0 and -gpu1; Kaggle allows two such sessions at once, and the cron starts both, so up to four workers run until the week’s 30 GPU-hours are spent (about four days). The workers appear in the panel under those names; the notebook’s log on Kaggle shows claimed job … and done (…) lines.
Kaggle allows two GPU sessions at once and a second copy of the same worker name only confuses the panel: cancel a running session (the Active events panel, bottom right, then Stop Session) before starting another by hand. A run whose environment is broken stops itself after three failed jobs, and the server refuses further claims from a name whose last three jobs failed; press retry failed in the panel afterwards. Nothing on Kaggle’s free tier can incur a bill; when the weekly quota is spent, the push fails until it resets.
7. Check from the panel
Section titled “7. Check from the panel”The owner opens /admin/transcription on the backend. The Workers card lists this machine by name, running or last seen, with jobs done in the last 24 hours. The Active workers table shows the job it holds and the stage.
8. Stopping, and what happens to the job it holds
Section titled “8. Stopping, and what happens to the job it holds”Stop the service (launchctl unload / systemctl stop) or the process (pkill -f 'wor[k]er.py'). The job it was on goes back to the queue after 10 minutes without a heartbeat, or at once when the owner re-queues it from the panel. Nothing is lost: results are written to results/job-<id>.json before upload and deleted once the server accepts them. A file left there means the upload failed three times; fix the cause, then re-queue the job in the panel rather than re-running ASR.
Things to know
Section titled “Things to know”- The queue feeds itself. A worker that finds it empty pulls the next eligible materials in, so there is no button to press after publishing.
- Every worker shares one secret. Rotating it (Hangar env) means updating every machine.
- A job that fails three times stays failed until someone re-queues it; the panel shows the error text.
DEVICE,IDLE_SECONDS(default 30),RESULTS_DIRand theCUE_*cue-shape values are the only other knobs; see the top ofworker/worker.py.