Learned the hard way
Each of these was found by breaking something. None is derivable from the code.
Deploying
Section titled “Deploying”-
Hangar env
PUTreplaces the whole set.GETit, merge, thenPUT. A partialPUTdeletes every other key and the app restarts without them. -
Cloudflare cuts uploads at 100 s with a 524. Hangar usually finishes the deploy anyway. Check
GET /sites/<site>/releasesbefore re-uploading. -
The backend bundle must include
.next/staticandpublic. Without them the admin renders blank with 404s on every chunk.scripts/deploy-hangar.shdoes the copying. -
Bundles are built on a Mac for an x86_64 host.
package.json→pnpm.supportedArchitecturespulls the Linux binaries (sharp, swc); the deploy script strips the Darwin ones. The frontend getsnpm ci --os=linux --cpu=x64in a separate directory, because installing Linux binaries in place removes the Mac build tools. -
next buildneeds no database. DummyDATABASE_URLandPAYLOAD_SECRETat build time; the real ones come from Hangar env at runtime. -
Never build in the working tree, and never edit a running shell script. A production build in
.nextdeleted the dev server’s output from under it (every route 500); builds now go to.next-build. And bash reads scripts incrementally, so editingdeploy-hangar.shwhile it ran produced a syntax error mid-deploy. -
Uploads through Cloudflare die at 100 s. The deploy script now rsyncs the bundle to the Hangar box (resumable, retried) and hands it to Hangar locally with
--resolve admin.mohanad.xyz:443:127.0.0.1. -
A CASE of string literals cannot be assigned to an enum column. Cast it:
(case … end)::enum_transcription_jobs_status. The reclaim query 500’d every claim until it was.
Database
Section titled “Database”- Neon free tier caps egress at 5 GB/month. Crawlers on the sitemap and listing pages burned it in a day. Postgres now lives on the search VPS;
infra/search-vps/top-queries.sqlshows who is hitting it. - Self-signed Postgres cert:
sslmode=requirerefuses it, andPG_CA_CERTalone still failed withERR_TLS_CERT_ALTNAME_INVALID … Host: localhost: node-postgres hands no server name to TLS for an IP host, so Node checked the certificate against “localhost”.src/utilities/pgSsl.tsverifies the identity against the host inDATABASE_URLitself, the certificate carriesIP:49.12.78.95as a subject-alt-name, and Hangar holds it base64-encoded (multi-line env values do not survive). Since 2026-09-10 the pool runs verify-full;pg_stat_sslon the VPS shows every Hangar connection on TLS. payload migrate:creatediffs against a stale drizzle snapshot (the last generated migration’s JSON), not the live database. It proposed dropping tables. Hand-scope generated migrations; seesrc/migrations/20260909_043541_pipeline.ts.- A hook that calls the Local API without
reqopens a second transaction. Creating a transcript deadlocked on the material’s foreign-key lock, and the cue reindex read an uncommitted row and indexed nothing. Passreqthrough every read and write inside hooks (src/search/sync.ts). materials_textsandmaterials_relsscans equal to material reads means a query walking all materials with relations. The sitemap did that per request; it is cached hourly now.
Search
Section titled “Search”- Meilisearch only opens a database written by the same version. Pin the image tag (
v1.53.1) ininfra/search-vps/docker-compose.yml; a mismatch crash-loops with exit 0 and no useful log. - Ollama on the 4 GB search box dies on embedding inputs past ~130 words (“llama-server process no longer running”, then a 30 s reload) when the request leaves bge-m3 at its 8,192-token default context: the KV cache for that window does not fit beside Meilisearch and Postgres.
embed()sendsnum_ctx: 1024, the container’s cap went from 1.5 GB to 2.2 GB (Meilisearch’s down to 1.4 GB, it is mostly reclaimable cache) andOLLAMA_NUM_PARALLEL=1; the longest input is a 125-word passage, now 3 s. Before the fix every passage push failed silently and new transcripts had sentence vectors only.POST /api/pipeline/reindex-vectorsre-pushes vectors for recently transcribed materials. - A truncated vector index crashes Meilisearch on every semantic query, taking keyword search down for seconds each time.
SEARCH_VECTORS=offin Hangar env serves keyword results for all modes while an index is being moved or rebuilt. - Moving
data.msis faster than re-embedding: 25 GB copied versus ~10 GPU-hours plus a day of vector-graph building on 2 vCPUs. Copy with local Meilisearch stopped for the final pass, verify withmd5sumbefore swapping. - The Docker default log driver copies container stdout to disk. Streaming a 23 GB tar through
docker runfilled the Docker VM and took local Postgres down. Use--log-driver nonefor streams. - Busybox
tail -c +Nfails past 2 GiB with “Bad address”. Useddwith a block offset. sshinside awhile readloop eats the loop’s stdin. Usessh -n. Two scripts silently processed one file each because of this.- A shell that mentions a process name in its own command line matches
pkill -ffor that name. Use a self-excluding pattern like'wor[k]er.py'. - Never truncate a remote file based on a size you did not actually read. A failed
statover ssh became “0” and wiped 5 GB. The copy script now refuses to act on a non-numeric size.
Worker and pipeline
Section titled “Worker and pipeline”-
Cohere ASR on CPU is 0.4× real time; on Apple GPU 20–45×. Backlog is ~19,000 hours of audio. A rented consumer GPU does it in about a week.
-
The ASR library’s cue builder cuts at every comma (
،is in its sentence-ending set) and caps cues at 80 characters, so it produced one- and two-word cues a quarter of the time.worker/worker.pybuilds cues from the word timings itself, with a ten-word floor. -
Cloudflare answers Python’s default user agent with a 403 before the request reaches the API, so a worker saw
claim failed: HTTP Error 403while the same call from curl worked. The worker sendsUser-Agent: kalelm-worker/1 (<name>); keep a real user agent on anything else that talks to the API through the edge. -
The ASR library’s default audio decoder (torchcodec) segfaults on some mp3s. Two crashes in one afternoon, both in
libavformatunderSingleStreamDecoder::getFramesPlayedInRangeAudio, each killing the worker with no Python error. The worker passesaudio_backend="ffmpeg";worker/run.shrestarts it if anything else of the kind happens, and a job whose worker goes silent returns to the queue after 10 minutes. -
Kaggle’s Active events panel shows three entries; older sessions hide below them. A run from the first afternoon’s notebook (no torchvision fix, a worker without the failure guard) kept running for two hours under the same worker name as the good run and failed forty jobs in seconds each. Now every Kaggle run names itself
kaggle-<container>, and the claim endpoint refuses a worker whose last three jobs all failed within fifteen minutes. When cancelling sessions, expand the panel and check every entry. -
Meilisearch sizes indexing memory from the host, not its container. With
MEILI_MAX_INDEXING_MEMORYat 1200Mb inside a 1400m container, four Kaggle workers uploading transcripts had the kernel kill it at 1.36 GB nineteen times in a day, and every search during a restart was a 500. It is 600Mb with one indexing thread now; the gate checks the number stays at most half the limit. -
Ollama on this Mac has two servers:
localhostresolves to an empty IPv6 one; models live on127.0.0.1:11434. -
Payload builds collections, access rules and custom endpoints at startup.
next devhot-reloads page code but not the config; an access-control or endpoint change looks unapplied until the dev server restarts. The gate failed three checks that way. -
A self-signed Postgres certificate needs an IP subject-alt-name (
-addext "subjectAltName=IP:…") or Node’sverify-fullrejects it by hostname; a CN alone is not enough. -
A field validator that rejects legacy data blocks unrelated edits. The file-URL check accepts an unchanged
previousValue, otherwise the 42 migrated rows with off-origin URLs could never be saved again.
The account hub and the learning platform
Section titled “The account hub and the learning platform”- Astro’s origin check (
security.checkOrigin) refuses the OAuth token endpoint. It rejects every form-encoded POST without a matchingOrigin, and a site’s server-to-server code exchange has none:openid-clientreported “unexpected response content-type” (a 403 HTML page). The hub turns the check off and does its own inhub/src/middleware.tsfor its forms; Better Auth checks its own routes. - The OAuth provider plugin registers web clients only with https callbacks, and refuses
http://localhostandhttp://127.0.0.1for them. A plain-http site is registered as a native client (/admin/sitesdoes this by the URL), and its callback must be127.0.0.1, notlocalhost. - The plugin has no
trustedClientsin 1.7.4: sites are rows inoauthClient, created throughcreateOAuthClientby a signed-in admin (clientPrivileges). WithoutclientPrivilegesany signed-in user could register a client. - Access tokens are opaque, not JWTs, so the hub’s data API cannot read the client id from the token. It hashes the bearer (SHA-256, base64url, no padding — the plugin’s
defaultHasher) and looks it up inoauthAccessToken, which also checks expiry and revocation. astro devdoes not put.envonprocess.env; Better Auth read no base URL and the plugin crashed onnew URL(undefined). The hub reads env throughenv()inhub/src/lib/auth.ts:process.envfirst (Hangar),import.meta.envafter (dev).- The site prefetches links on hover, and a prefetch of
/learn/loginstarts a sign-in flow. Each start rotates the state cookie; the prefetch and the real click raced, and the callback failed with “unexpected state response parameter value”. The sign-in links carrydata-astro-prefetch="false"and the route answers a prefetch (Sec-Purpose) with 204. auth.api.oauth2Consentcannot finish a consent: on accept the plugin resumes the authorization, which needs the raw request (request not found). The consent page posts to the plugin throughauth.handlerwith a realRequest, forwarding the browser’s cookies, and follows the redirect it answers with.- resvg loads TTF only, and fontTools cannot merge Reem Kufi’s subsets. The site’s woff2 files were converted with fontTools; the IBM Plex subsets merged into one TTF each, Reem Kufi’s variable-font tables refused, so its Arabic and Latin subsets are two files under one family name and resvg falls back between them per glyph. The fonts live in
public/certificate/, which the deploy copies beside the server; a font read fromsrc/would not be in the standalone bundle. payload migrate:createand drizzle’s push both stop on an interactive question (“isenum_courses_statuscreated or renamed?”) that piped input cannot answer. The learn migration was written by hand in Payload’s naming and checked by creating a course through the Local API; the slug field’s helper column isgenerate_slug, notslug_lock.
Repo hygiene
Section titled “Repo hygiene”pnpm verifyis the gate and a Stop hook runs it. It needs local Postgres, the dev backend on 3001, the frontend on 4322, Meilisearch on 7701 and Ollama. If local Meilisearch is stopped the gate fails for that reason alone.- ESLint’s config crashes on every file.
tsc --noEmitand the gate are the working checks. docs/is excluded from the backendtsconfig; Astro’s virtual modules would otherwise fail the type check.- Moving a Payload admin component means regenerating the import map (
pnpm generate:importmap, also part ofbuild) and restartingnext dev; stale.next*/types/validator.tsfiles can also failtscafter a route file is deleted — remove them, a build recreates them.
Frontend
Section titled “Frontend”offsetTopis not “offset inside the scroll box”. It is measured from the nearest positioned ancestor, and the transcript list has none, so the auto-scroll used page coordinates and landed anywhere but on the spoken line. Measure withgetBoundingClientRect()against the list’s own box (MaterialDetail.astro,bring()).- A Payload endpoint edit is not picked up by the running
next dev. The search API kept answering with the old result shape until the dev server was restarted (preview_start payload-admin); the same rule as config changes. - A page kept for offline is nothing without its assets, and a script’s imports are not in the HTML. The worker used to
cache.addthe offline page and a saved material’s page on their own; after the next deploy their scripts were gone from the origin and never in the cache, so offline they rendered as text with no player, no list, no menus.keepDocumentinsw.js.tsreads the HTML for its stylesheet, script and font addresses, keeps them, then reads each script for its ownimports and keeps those. The worker is also stamped per deploy (BUILD), so the offline page is re-kept with the live build; and a kept recording is answered from the cache whether online or not, sincenavigator.onLinesays online on a wifi with no internet. caches.match()searches every cache. The record of a kept recording is a JSON response stored under the page’s own path in the records cache, so the offline fallback, looking the page up across all caches, handed the reader that JSON instead of the page. Pages are matched in the page cache only; the gate refusescaches.match(in the worker.- A page served from the worker’s cache runs the build it was cached with. The library page kept offline by an earlier worker asked that worker for the list; the next worker no longer answered, and “available offline” read empty over a cache full of recordings. A new worker keeps answering the old messages, and on activation drops every merely-visited page from the page cache and re-keeps the offline page and the kept recordings’ pages with the live build.
- A service worker cannot download an hour of audio. Chrome stops a worker whose event has run for about five minutes, whatever
waitUntilpromised; a 27 MB save stalled at 80% and was never finished or reported. The download runs in the page now (features/pwa/client.ts), where module state survives the client router’s swaps; a full reload ends it, which is at least visible. - Cache Storage is best-effort until the site asks for persistence. Three recordings kept during one afternoon of testing were gone by the evening while two older ones stayed, with no code path that deletes them; the browser is allowed to evict. The save button now calls
navigator.storage.persist()first, and the worker’s list falls back to the audio cache itself, so a recording whose record is missing still shows (by file name) and still plays. - A deploy must keep the previous build’s hashed assets. Cloudflare holds a material page for an hour; that HTML names the chunks of the build it was rendered by, and Hangar serves only the current release, so after a deploy every edge-cached page loaded a 404 for its scripts: no player, no continue, no menus, until the cache expired.
deploy-hangar.sh frontendnow copies the live release’s_astrointo the new one (each release holds its predecessors’, so the set carries forward) and fails the deploy if the previous build’s first chunk stops answering. The clean fix is purging the edge on deploy, which needs a Cloudflare API token with cache-purge rights in the deploy environment. lib/api.tsis imported by client scripts. The browse tree’s script imports it forAPIand a few helpers; an import of the Node-only loopback fetcher in it shippednode:httpsto the browser and the tree died at load with “Agent is not a constructor”. The server installs its fetcher from the middleware (setApiFetcher), the module itself stays plain; the gate refuses a Node import there.- The frontend must call the API over the loopback, not its public hostname. Both run on the Hangar box behind the same Caddy; the public name goes out through Cloudflare and back in, 0.5–1.2 s per call against 30 ms on the loopback, and a lesson page made five calls before it could start (3.9 s to first byte).
frontend/src/lib/apiFetch.tssends requests toAPI_PIN(127.0.0.1, set by the deploy script at build) with the hostname kept for SNI and the certificate check; the gate checks both sides. - One inline SVG per transcript cue was 450 KB of a 1 MB page. Anything repeated per row is drawn by CSS once (
.tr__link::before), and the gate counts<svg>inside the transcript list. - A service worker must not proxy the mp3 while online. Passing
fetch(req)through forfiles.kalelm.comgives the<audio>element an opaque response: Chrome cannot readContent-Rangeon it, treats every range request as the whole file, and the bar sat at 0:00 for twenty seconds while an hour of audio came down. The worker only answers media requests whennavigator.onLineis false, from the cache (features/pwa/pages/sw.js.ts); the gate checks the guard is there. - Under the client router, a click listener sees the link already handled. The router’s own listener runs first and calls
preventDefault, sonavProgressnever started its bar and every navigation looked dead. Start it fromastro:before-preparation. - Astro scoped styles stop at the component. Rows rendered by a child component (
PlaylistItems) do not match the parent’s.plist__items …rules; the styles live with the markup that owns them.
Hiding a sheikh hides their material, in four places at once
Section titled “Hiding a sheikh hides their material, in four places at once”The «إخفاء الشيخ ومواده من الموقع» switch on a sheikh (2026-09-12) is read by four
readers that must agree: the Materials REST access rule ('sheikh.hidden': { not_equals: true },
which keeps materials with no sheikh), every SQL reader through PUBLIC_MATERIAL() in
src/utilities/publicMaterial.ts (facets, tree, series summary, series index), and the
search and related endpoints through hiddenSheikhFilter() (sheikhId NOT IN […], all four
indexes carry sheikhId, no reindex needed; read on every search, one indexed lookup). The gate takes one hidden sheikh’s published material and asserts it is
absent from each. Add a fifth reader and the gate’s list in the data section is where it goes.
A sheikh, series or chapter with nothing published is not shown either (2026-09-12): the directories list only what the facets count, their pages 404, the sitemap and the tree’s sheikh search skip them. That rule lives in the frontend and the tree query, not in the Sheikhs access rule, because Payload cannot filter a collection by the absence of related rows.
Two things learned on the way:
- Once the access rule joins the sheikh, Payload reports
totalDocs: 0for a page past the end (valid pages count correctly). The archive redirects an overrun to its first page. ?limit=0on/api/materialsreturns every document. Against 45,000 materials the dev server dies withMap maximum size exceeded. Uselimit=1to read a count. Andcurltreats[]in a URL as a glob: pass-gor the request never leaves the shell.
A persisted component wears the current page’s stylesheet
Section titled “A persisted component wears the current page’s stylesheet”The player bar (and the save strip) persist across pages through transition:persist, but
their CSS comes from whichever page is showing. After a deploy that adds a rule to the bar,
a reader who navigates from a fresh page into an edge-cached old page sees the new markup
without the new rule: the caption toggle rendered as a loose 42px ring under the bar
(2026-09-12). Every page is cached ten minutes at the edge (material pages were an hour). Since
2026-09-13 the service worker closes the window for anyone who has visited before: it
fetches every page under ?r=<release> (the worker’s build stamp changes with each deploy),
which is a new edge cache key, and middleware.ts strips the parameter before any page
renders. A first-time visitor, or a browser without the worker, can still meet a cached page
for up to ten minutes after a deploy; a Cloudflare purge on deploy would close that too, and
needs a token the deploy does not have. The owner’s own view is fresh from the second page
load after a deploy: the first load installs the new worker, every navigation after it is
keyed by the new release.