Skip to content

Run a worker on a new machine

This page is for a Claude session (or a person) sitting at a machine that has never run a worker. Follow it top to bottom; each step says how to check it worked. Nothing here needs access to the servers: a worker only makes outbound HTTPS calls to the API.

A Python process in worker/ that asks the backend for a job, downloads the mp3, runs Cohere ASR on the GPU, builds cues of ten words or more from the word timings, and uploads them. The backend hands each job to exactly one worker, so any number of machines can run one. Details and measurements: Transcription pipeline.

Machine Speed Verdict
Apple silicon Mac (M1 or later, 16 GB+) 12 to 28× real time yes
Linux box with an NVIDIA GPU (8 GB VRAM+) similar or better yes
Any machine on CPU only 0.4× real time no: one hour of audio takes two and a half
Windows untested use WSL2 with the Linux steps, or skip

Memory: ASR needs about 5 GB.

Terminal window
git clone https://github.com/mohanad-a/kalelm-backend.git kalelm
cd kalelm/worker

The repository is private; the owner’s GitHub account (mohanad-a) has to grant access or clone it for you. Only worker/ is needed on this machine.

macOS:

Terminal window
brew install ffmpeg uv

Ubuntu:

Terminal window
sudo apt install -y ffmpeg
curl -LsSf https://astral.sh/uv/install.sh | sh

Then, in worker/:

Terminal window
uv venv --python 3.12 && source .venv/bin/activate
uv pip install -r requirements.txt --torch-backend=auto

Check: python test_worker.py prints ok.

Ask for them; do not guess and do not put them in a file that is committed.

  1. The API secret. CRON_SECRET from the backend’s environment (Hangar → site kalelm-api → env, or ask the owner). A wrong secret shows up as claim failed: HTTP Error 403.
  2. A Hugging Face login that has accepted the model’s terms. The ASR model CohereLabs/cohere-transcribe-arabic-07-2026 is gated. Run hf auth login and paste a token from an account that accepted it (the owner’s is named mac under user coder0x9). Check: hf auth whoami names the account.

Also confirm the API address. As of 2026-09-10 the test deployment answers at https://kalelm-api.mohanad.xyz; the docs index page’s Hosts table has the current one.

Terminal window
KALELM_API=https://kalelm-api.mohanad.xyz CRON_SECRET=… WORKER_NAME=<this-machine> DEVICE=mps \
python worker.py --once

If you keep it running with nohup rather than launchd or systemd (step 6, which restart it themselves), run run.sh with WORKER_PYTHON pointing at the venv’s python: it restarts the worker after a native crash, which the model stack does now and then with nothing in the log.

DEVICE is mps on a Mac, cuda on Linux with NVIDIA, and can be left unset to auto-detect. WORKER_NAME must be unique per machine; the panel lists workers by it.

Expected output, within a couple of minutes:

14:21:56 worker <name> -> https://kalelm-api.mohanad.xyz
14:21:56 claimed job 5: '…'
14:23:11 job 5 material 49014: done (81 cues, 73.4s, RTFx 12.0)

RTFx is how many times faster than real time it ran. If instead you see:

  • claim failed: HTTP Error 403 — wrong secret, or a request without the worker’s own user agent (the code sets one; do not replace urllib calls with something that sends a different one, Cloudflare blocks Python’s default).
  • nothing to do (paused) — the queue is paused in the panel; the owner unpauses it.
  • nothing to do — the queue is empty and nothing is eligible; that is fine.
  • an out-of-memory error during asr — set DEVICE=cpu to prove the rest works, then find a machine with more memory.

macOS, ~/Library/LaunchAgents/com.kalelm.worker.plist (fill in the paths, the secret and the name; PATH must include where ffmpeg and ollama live):

<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0"><dict>
<key>Label</key><string>com.kalelm.worker</string>
<key>ProgramArguments</key>
<array><string>/ABS/PATH/kalelm/worker/.venv/bin/python</string><string>worker.py</string></array>
<key>WorkingDirectory</key><string>/ABS/PATH/kalelm/worker</string>
<key>EnvironmentVariables</key><dict>
<key>KALELM_API</key><string>https://kalelm-api.mohanad.xyz</string>
<key>CRON_SECRET</key><string>…</string>
<key>WORKER_NAME</key><string>mac-mini</string>
<key>DEVICE</key><string>mps</string>
<key>PYTHONUNBUFFERED</key><string>1</string>
<key>PATH</key><string>/opt/homebrew/bin:/usr/bin:/bin</string>
</dict>
<key>RunAtLoad</key><true/>
<key>KeepAlive</key><true/>
<key>StandardOutPath</key><string>/ABS/PATH/kalelm/worker/worker.log</string>
<key>StandardErrorPath</key><string>/ABS/PATH/kalelm/worker/worker.log</string>
</dict></plist>
Terminal window
launchctl load ~/Library/LaunchAgents/com.kalelm.worker.plist # starts now and at every login
launchctl unload ~/Library/LaunchAgents/com.kalelm.worker.plist # stops

Linux, /etc/systemd/system/kalelm-worker.service:

[Unit]
Description=Kalelm transcription worker
After=network-online.target
[Service]
WorkingDirectory=/opt/kalelm/worker
Environment=KALELM_API=https://kalelm-api.mohanad.xyz
Environment=CRON_SECRET=…
Environment=WORKER_NAME=gpu-box
Environment=PYTHONUNBUFFERED=1
ExecStart=/opt/kalelm/worker/.venv/bin/python worker.py
Restart=always
RestartSec=10
[Install]
WantedBy=multi-user.target
Terminal window
sudo systemctl enable --now kalelm-worker
journalctl -u kalelm-worker -f

Kaggle gives every account 30 GPU-hours a week for nothing, with sessions capped at 12 hours. scripts/kaggle/kalelm-worker.ipynb is the worker as a notebook: it copies worker.py and secrets.env from the private Kaggle dataset kalelm-worker, installs ffmpeg and cohere-transcribe, and runs for MAX_MINUTES (230, under four hours), stopping between jobs. Accelerator GPU T4 x2 (machine_shape: NvidiaTeslaT4 in scripts/kaggle/kernel-metadata.json; a push without it lands on a P100). The torch that cohere-transcribe installs has no kernels for the P100’s generation (no kernel image is available for execution on the device), so on a P100 the notebook reinstalls the same torch from the CUDA 12.6 wheel index, two extra minutes. On a T4 expect 10 to 40× real time, roughly 100 archive-hours per run. Measured 2026-09-10: setup 2.5 minutes, first job done 4 minutes after start.

Weekly hours are per account, however many GPUs or sessions: parallel runs spend the quota faster, they do not add to it. Two Kaggle limits shape the design. Its scheduler refuses GPU notebooks, so the daily run is started from outside through the API (root’s crontab on the search VPS, infra/search-vps/kaggle-run.sh, 03:00 UTC). And runs started through the API never receive the notebook’s Add-ons → Secrets, so CRON_SECRET and HF_TOKEN travel in secrets.env inside the private dataset instead; the notebook falls back to Kaggle’s Secrets only when that file is absent. If the dataset ever leaks, rotate CRON_SECRET in Hangar and the Hugging Face token.

Once, in a browser on kaggle.com: Settings → Phone verification (notebooks get internet only on verified accounts) and Settings → API → Create access token, kept as KAGGLE_API_TOKEN on the pushing machine and in /srv/kalelm/.env on the VPS.

From the repository, whenever worker.py or the notebook changes, or a run is wanted now:

Terminal window
uv tool install kaggle
KAGGLE_API_TOKEN=… CRON_SECRET=… scripts/kaggle/push.sh

It uploads a new dataset version (worker, secrets), saves the notebook, which starts a run, and copies the notebook to the VPS for the cron. HF_TOKEN defaults to the token hf auth login stored. A T4 x2 session has two GPUs, so the notebook runs two workers, kaggle-<container>-gpu0 and -gpu1; Kaggle allows two such sessions at once, and the cron starts both, so up to four workers run until the week’s 30 GPU-hours are spent (about four days). The workers appear in the panel under those names; the notebook’s log on Kaggle shows claimed job … and done (…) lines.

Kaggle allows two GPU sessions at once and a second copy of the same worker name only confuses the panel: cancel a running session (the Active events panel, bottom right, then Stop Session) before starting another by hand. A run whose environment is broken stops itself after three failed jobs, and the server refuses further claims from a name whose last three jobs failed; press retry failed in the panel afterwards. Nothing on Kaggle’s free tier can incur a bill; when the weekly quota is spent, the push fails until it resets.

The owner opens /admin/transcription on the backend. The Workers card lists this machine by name, running or last seen, with jobs done in the last 24 hours. The Active workers table shows the job it holds and the stage.

8. Stopping, and what happens to the job it holds

Section titled “8. Stopping, and what happens to the job it holds”

Stop the service (launchctl unload / systemctl stop) or the process (pkill -f 'wor[k]er.py'). The job it was on goes back to the queue after 10 minutes without a heartbeat, or at once when the owner re-queues it from the panel. Nothing is lost: results are written to results/job-<id>.json before upload and deleted once the server accepts them. A file left there means the upload failed three times; fix the cause, then re-queue the job in the panel rather than re-running ASR.

  • The queue feeds itself. A worker that finds it empty pulls the next eligible materials in, so there is no button to press after publishing.
  • Every worker shares one secret. Rotating it (Hangar env) means updating every machine.
  • A job that fails three times stays failed until someone re-queues it; the panel shows the error text.
  • DEVICE, IDLE_SECONDS (default 30), RESULTS_DIR and the CUE_* cue-shape values are the only other knobs; see the top of worker/worker.py.