Skip to content

Decision log

Newest first. Each entry: what was decided, what it was weighed against, and the measurement.

2026-09-12 — One account hub on Better Auth, not accounts in Payload

Section titled “2026-09-12 — One account hub on Better Auth, not accounts in Payload”

Readers needed a login for the learning platform, and the owner wanted it the way Wasl does it: one account for several sites, with each site’s data for the reader visible in one place. A second Payload auth collection was built first and thrown away the same day: it would tie every future site to this backend, and every req.user check in the archive would have had to learn that a reader is not staff. The hub is its own small Astro app (hub/) on Better Auth with its OAuth provider plugin, so a site signs readers in with standard OpenID Connect and needs nothing from the hub’s database. Weighed against Auth.js (no provider side), Keycloak, Ory, Authentik and Zitadel (each a service of its own, on a box with 3.7 GB), Logto and Clerk (hosted or heavier). Better Auth’s tables live in a second database on the same Postgres. Its opaque access tokens are looked up by hash for the hub’s data API, since the plugin does not put the client id in a JWT unless a resource is registered.

2026-09-12 — Certificates as PDF from SVG, without a browser on the host

Section titled “2026-09-12 — Certificates as PDF from SVG, without a browser on the host”

The Hangar host runs Node only, and Arabic in a PDF needs shaping and right-to-left ordering that pdf-lib and pdfkit do not do. The certificate is drawn as SVG, which resvg (a native binary, prebuilt for the host’s Linux and the owner’s Mac) sets with the site’s own fonts, shaping and bidi included; the bitmap goes onto an A4 page with pdf-lib. A 2480-pixel-wide raster prints cleanly; text is not selectable, which a certificate does not need. Weighed against printing HTML through Chromium (no browser on the host, 300 MB) and a PDF service (another dependency for a page a reader downloads once).

2026-09-12 — Claude writes questions and grades written answers; matching stands in without it

Section titled “2026-09-12 — Claude writes questions and grades written answers; matching stands in without it”

Editors have thousands of transcribed lessons and no time to write tests. POST /learn/ai gives Claude (claude-opus-5, structured JSON output) the level’s transcripts and gets back questions with explanations to append and edit; the same key grades a written answer by meaning with a sentence of feedback. Both are additions to an editor’s or a grader’s work, never the only path: without the key the buttons say so, and written answers are matched against the model answer with tashkeel and punctuation stripped. Ollama on the search VPS was not an option: the box has 3.7 GB and no text model; the owner’s Mac is not always on.

2026-09-12 — Tests graded on the backend, progress at the hub

Section titled “2026-09-12 — Tests graded on the backend, progress at the hub”

The right answers never leave the API: answer is a staff-only field and POST /learn/attempt grades. A certificate is a row with a code, so the certificate page is also the verification page. What a reader has heard is the browser’s record first (works signed out, like the library) and mirrored to the hub as a per-site kind when signed in; the hub stores whatever a site sends, keyed by kind, and only knows title, href and when. Weighed against keeping progress in the backend (a third store, and the hub’s dashboard would be empty) and against grading in the browser (the answers would be in the page).

2026-09-11 — Frontend features are folders, discovered by glob

Section titled “2026-09-11 — Frontend features are folders, discovered by glob”

Every optional part of the website (dark mode, the player bar, the personal library, podcast feeds, related items, transcript export, offline) is a folder under src/features/ that the site finds by glob: routes from its pages/, components at declared extension points from its ext.ts, a browser script, its own strings. Features never import each other and core never imports a feature; they talk through named browser events. Weighed against a plugin registry (one more file to edit per feature, and one more place to forget) and against a monorepo of packages (build tooling for a site with one deployable). The folder is the unit: delete it and the feature is gone with nothing dangling, which the gate checks. Details: Frontend features.

2026-09-10 — Cohere stays the ASR engine; Audar-ASR-V1 tested and removed

Section titled “2026-09-10 — Cohere stays the ASR engine; Audar-ASR-V1 tested and removed”

Audar-ASR-V1 (a Qwen3-ASR fine-tune; Flash 0.8B under an open license, Turbo 1.7B under a community license that caps commercial use) was wired in as a second worker engine for a day. It returns text without timings and reads 30 s at a time, so cue starts had to be spread evenly over each span. Measured on the owner’s Mac against the migrated YouTube captions (WER is agreement, not truth; the GPU was shared with a production worker):

Material Band Length Cohere WER Audar Flash WER Audar Turbo WER Cohere speed Flash speed Turbo speed
27399 short 6 min 3.7% 12.1% 6.9% 14× 8× 6×
43472 medium 13 min 1.4% 12.8% 5.6% 31× 11× 9×
7108 long 52 min 6.8% 18.3% – 41× 11× –
12422 very_long 60 min 6.7% 21.5% – 48× 13× –

Flash dropped a six-word phrase from the five-minute sample outright; Turbo kept every word but is still two to four times Cohere’s error rate at a fifth of the speed. Removed the same day (commit 333c6c1 has the engine); the audar value stays in enum_transcripts_source because Postgres cannot drop one, and nothing writes it.

Surveyed the same day for something freer or faster than Cohere, none adopted:

Option Why not
Meta OmniASR-LLM-7B (Apache 2.0) second on the Open Universal Arabic ASR leaderboard behind Cohere (28.3 vs 25.9 WER) and 8× slower (RTFx 66 vs 525)
Qwen3-ASR-1.7B + ForcedAligner (Apache 2.0) what Audar is built on; real word timings via the aligner, but Audar’s tuned version already lost to Cohere here
Whisper large-v3 (local or Groq’s free 8 h/day) 36.9 WER on the Arabic leaderboard; Groq’s free quota would take seven years for the backlog
Voxtral Transcribe 2 (Mistral) best batch model is API-only, $0.003/min ≈ $3,700 for the archive; the open 4B realtime model has no timestamps
Gemini 3.5 Transcribe (Aug 2026 preview) $0.003/min ≈ $3,700; free tier is rate-limited and lets Google train on the audio
Cloud free tiers (Google 60 min/month, AWS 60 min/month, Azure 5 h/month) rounding error against 20,700 hours

The free, fast route is more GPUs running the worker we have: Kaggle gives 30 GPU-hours a week (P100/T4) and Colab about 15–30, and scripts/kaggle/asr_eval.py already runs there. At the Mac’s 39× that is roughly 1,200 archive-hours a week per account. Renting one H100 for the Cohere model’s 525× would do the whole backlog in about 40 GPU-hours, under $100.

2026-09-10 — Postgres self-hosted on the search VPS, not Neon

Section titled “2026-09-10 — Postgres self-hosted on the search VPS, not Neon”

Neon’s free tier caps egress at 5 GB/month; the test site used 5.06 GB in its first day, mostly crawlers walking the sitemap and listing pages, and the compute never idled because Hangar’s health probe hits the API every few seconds. Options were Neon’s paid plan (~$19/month), or Postgres in the existing Compose stack on the box we already pay for. Chose the latter: no caps, TLS with a self-signed certificate, port open only to the Hangar IP, pg_stat_statements on so the next surprise is visible. Nightly dumps go to the Hangar box.

2026-09-10 — Copy the search database instead of re-embedding

Section titled “2026-09-10 — Copy the search database instead of re-embedding”

Building the semantic indexes means embedding ~1 M sentences and ~360 k passages with bge-m3, about 10 GPU-hours here, then Meilisearch building vector graphs, which ran at ~14 docs/s on the Mac and would take over a day on 2 vCPUs. The local instance already held the finished 25 GB database. Copying it took most of a day only because of a bad link and three script bugs, but it needs no GPU and no graph build; the transfer itself is ~35 minutes at 100 Mbit/s. Same Meilisearch version on both ends is mandatory.

2026-09-09 — Search box on its own VPS; site and API on Hangar

Section titled “2026-09-09 — Search box on its own VPS; site and API on Hangar”

Hangar hosts Node apps well and gives immutable releases with rollback, but provides no database and no long-running non-Node services. Meilisearch plus an embedder plus Postgres fit naturally in one Compose stack on a plain box. The apps reach it over HTTPS; the firewall admits only the Hangar server.

2026-09-09 — Ollama with bge-m3 stays the embedder

Section titled “2026-09-09 — Ollama with bge-m3 stays the embedder”

The index already carries bge-m3 vectors and the code calls Ollama’s API. Alternatives were Hugging Face TEI (3–5× faster on CPU, small code change, same vectors), Meilisearch’s built-in embedder (would re-embed everything), or a paid API (different vectors, metered). Query-time embedding on a 2-vCPU box costs a few hundred milliseconds, acceptable. TEI is the upgrade if latency matters.

2026-09-09 — Search falls back to keyword; vectors can be switched off

Section titled “2026-09-09 — Search falls back to keyword; vectors can be switched off”

A half-copied vector index crashed Meilisearch on every semantic query, taking keyword search down for seconds each time and rendering the site’s search page as a 404. Now: any failure on the vector path answers with keyword results marked degraded, and SEARCH_VECTORS=off forces that for all modes without a deploy. Cost: a request that silently degrades; the response says so and the UI can show it.

2026-09-09 — Cache at the edge, plus small in-process caches

Section titled “2026-09-09 — Cache at the edge, plus small in-process caches”

Public pages are identical for every visitor. Cloudflare cache rule for the frontend host honouring s-maxage=600, stale-while-revalidate=3600; facets and the browse tree cached in the API for five minutes; the sitemap rebuilt hourly. Rate limit of 60 requests per 10 s per IP. Rejected: a Redis layer (a second service for what a Map does with one process).

2026-09-09 — Transcription via a queue in Payload and workers anywhere

Section titled “2026-09-09 — Transcription via a queue in Payload and workers anywhere”

Options were Payload’s job queue running Python on the backend host (needs the GPU on that host), or a transcription-jobs collection claimed by workers over HTTPS with the cron secret. The second lets the GPU be a Mac at home, a Mac mini, or a rented spot instance, and survives worker preemption via heartbeats and a 30-minute stale timeout. Atomic claims use FOR UPDATE SKIP LOCKED.

2026-09-09 — Cohere ASR default model; the tashkeel fine-tune rejected

Section titled “2026-09-09 — Cohere ASR default model; the tashkeel fine-tune rejected”

Measured on a 59-minute fatawa episode on Apple GPU: default model 3 min 09 s (18.8× real time), 6,447 words. The NAMAA-Space/Cohere-Speech-Tashkeel-2B fine-tune took 35 min 39 s, generated 3× the tokens, hit the token limit on 14 segments and looped on 10, and diverged from the plain transcript on 48.9 % of words. Diacritics come from a CPU pass instead (next entry).

2026-09-09 — Hamza and shadda/tanween from CAMeL Tools, no full tashkeel

Section titled “2026-09-09 — Hamza and shadda/tanween from CAMeL Tools, no full tashkeel”

The owner wanted hamzas and only the necessary marks, and a light CPU path for thousands of files. CAMeL’s MLE disambiguator restores hamza, ta-marbuta and alef-maqsura at 90.3 % on 1,344 reference words, recovers 71 % of shadda and 55 % of tanween, at ~3,000 words/s on one core. Alternatives: a 350 M-parameter LLM diacritizer (slow, full vowels, unwanted) and Mishkal (weaker, no hamza).

2026-09-11 — A series is the unit for lessons; series pages by kind

Section titled “2026-09-11 — A series is the unit for lessons; series pages by kind”

19,810 of 19,822 lessons belong to a series, so the flat lesson list was the wrong first page. A series now has a page (/series/<slug>/) and a kind (derived by migration 20260911_190000, editable in the admin): a book explanation (33; its lessons carry book chapters and are grouped under them, chapters in book order or by first lesson), a program (26; dated episodes, several sheikhs, paged newest first by month), or a series (843; a numbered list, oldest first). Chapters are navigation inside a book explanation, not a global filter: the same chapter is explained in several series by several sheikhs, and the chapter page is that cross-series view. Each sheikh has their own series, whatever the data shares: the unit a reader sees is (series, sheikh). The lessons landing page (/lessons/ with no parameters) is a compact grid of the sheikhs with their counts; a sheikh’s page lists their series by kind with a book explanation’s chapters folded under it; a series page is read for one sheikh (?sheikh=<slug>, else the one who taught most of it) with a switch to the others, and its groups are folds, one open at a time. The flat, filterable archive of every lesson is the other tab (/lessons/?view=all); any filter, sort or page parameter answers it. The landing and the sheikh page are fed by GET /series-index (one row per series and sheikh).

It corrected ~4 % of words (spelling, punctuation) for ~6 minutes per hour of audio, and when its provider was missing or overloaded the job still completed with uncorrected text. Not worth its cost; the CAMeL hamza/shadda pass was removed the same day for the same reason. The cue shape problem it was thought to help with is solved where it arose: the worker builds cues from the word timings with a ten-word floor.

2026-09-09 — LLM post-edit optional, guarded, off by default for bulk (superseded above)

Section titled “2026-09-09 — LLM post-edit optional, guarded, off by default for bulk (superseded above)”

qwen3:8b through Ollama changed 3.8 % of words on a 20-chunk sample, mostly ta-marbuta and punctuation, with two harmful edits. Kept as an option behind a guard that rejects any line whose word content changes, at ~6 minutes per audio hour on a GPU. Recommended off for the 19,000-hour backlog and on for new content.

2026-09-09 — Meilisearch, not Elasticsearch

Section titled “2026-09-09 — Meilisearch, not Elasticsearch”

Pre-dates this log; recorded here because the previous project on the search VPS ran Elasticsearch. Meilisearch runs in ~200 MB idle, has Arabic normalization built in, and supports hybrid search with user-provided vectors; Elasticsearch needs 2–4 GB just to start. The box has 3.7 GB.