Skip to content

Learned the hard way

Each of these was found by breaking something. None is derivable from the code.

  • Hangar env PUT replaces the whole set. GET it, merge, then PUT. A partial PUT deletes every other key and the app restarts without them.

  • Cloudflare cuts uploads at 100 s with a 524. Hangar usually finishes the deploy anyway. Check GET /sites/<site>/releases before re-uploading.

  • The backend bundle must include .next/static and public. Without them the admin renders blank with 404s on every chunk. scripts/deploy-hangar.sh does the copying.

  • Bundles are built on a Mac for an x86_64 host. package.json → pnpm.supportedArchitectures pulls the Linux binaries (sharp, swc); the deploy script strips the Darwin ones. The frontend gets npm ci --os=linux --cpu=x64 in a separate directory, because installing Linux binaries in place removes the Mac build tools.

  • next build needs no database. Dummy DATABASE_URL and PAYLOAD_SECRET at build time; the real ones come from Hangar env at runtime.

  • Never build in the working tree, and never edit a running shell script. A production build in .next deleted the dev server’s output from under it (every route 500); builds now go to .next-build. And bash reads scripts incrementally, so editing deploy-hangar.sh while it ran produced a syntax error mid-deploy.

  • Uploads through Cloudflare die at 100 s. The deploy script now rsyncs the bundle to the Hangar box (resumable, retried) and hands it to Hangar locally with --resolve admin.mohanad.xyz:443:127.0.0.1.

  • A CASE of string literals cannot be assigned to an enum column. Cast it: (case … end)::enum_transcription_jobs_status. The reclaim query 500’d every claim until it was.

  • Neon free tier caps egress at 5 GB/month. Crawlers on the sitemap and listing pages burned it in a day. Postgres now lives on the search VPS; infra/search-vps/top-queries.sql shows who is hitting it.
  • Self-signed Postgres cert: sslmode=require refuses it, and PG_CA_CERT alone still failed with ERR_TLS_CERT_ALTNAME_INVALID … Host: localhost: node-postgres hands no server name to TLS for an IP host, so Node checked the certificate against “localhost”. src/utilities/pgSsl.ts verifies the identity against the host in DATABASE_URL itself, the certificate carries IP:49.12.78.95 as a subject-alt-name, and Hangar holds it base64-encoded (multi-line env values do not survive). Since 2026-09-10 the pool runs verify-full; pg_stat_ssl on the VPS shows every Hangar connection on TLS.
  • payload migrate:create diffs against a stale drizzle snapshot (the last generated migration’s JSON), not the live database. It proposed dropping tables. Hand-scope generated migrations; see src/migrations/20260909_043541_pipeline.ts.
  • A hook that calls the Local API without req opens a second transaction. Creating a transcript deadlocked on the material’s foreign-key lock, and the cue reindex read an uncommitted row and indexed nothing. Pass req through every read and write inside hooks (src/search/sync.ts).
  • materials_texts and materials_rels scans equal to material reads means a query walking all materials with relations. The sitemap did that per request; it is cached hourly now.
  • Meilisearch only opens a database written by the same version. Pin the image tag (v1.53.1) in infra/search-vps/docker-compose.yml; a mismatch crash-loops with exit 0 and no useful log.
  • Ollama on the 4 GB search box dies on embedding inputs past ~130 words (“llama-server process no longer running”, then a 30 s reload) when the request leaves bge-m3 at its 8,192-token default context: the KV cache for that window does not fit beside Meilisearch and Postgres. embed() sends num_ctx: 1024, the container’s cap went from 1.5 GB to 2.2 GB (Meilisearch’s down to 1.4 GB, it is mostly reclaimable cache) and OLLAMA_NUM_PARALLEL=1; the longest input is a 125-word passage, now 3 s. Before the fix every passage push failed silently and new transcripts had sentence vectors only. POST /api/pipeline/reindex-vectors re-pushes vectors for recently transcribed materials.
  • A truncated vector index crashes Meilisearch on every semantic query, taking keyword search down for seconds each time. SEARCH_VECTORS=off in Hangar env serves keyword results for all modes while an index is being moved or rebuilt.
  • Moving data.ms is faster than re-embedding: 25 GB copied versus ~10 GPU-hours plus a day of vector-graph building on 2 vCPUs. Copy with local Meilisearch stopped for the final pass, verify with md5sum before swapping.
  • The Docker default log driver copies container stdout to disk. Streaming a 23 GB tar through docker run filled the Docker VM and took local Postgres down. Use --log-driver none for streams.
  • Busybox tail -c +N fails past 2 GiB with “Bad address”. Use dd with a block offset.
  • ssh inside a while read loop eats the loop’s stdin. Use ssh -n. Two scripts silently processed one file each because of this.
  • A shell that mentions a process name in its own command line matches pkill -f for that name. Use a self-excluding pattern like 'wor[k]er.py'.
  • Never truncate a remote file based on a size you did not actually read. A failed stat over ssh became “0” and wiped 5 GB. The copy script now refuses to act on a non-numeric size.
  • Cohere ASR on CPU is 0.4× real time; on Apple GPU 20–45×. Backlog is ~19,000 hours of audio. A rented consumer GPU does it in about a week.

  • The ASR library’s cue builder cuts at every comma (، is in its sentence-ending set) and caps cues at 80 characters, so it produced one- and two-word cues a quarter of the time. worker/worker.py builds cues from the word timings itself, with a ten-word floor.

  • Cloudflare answers Python’s default user agent with a 403 before the request reaches the API, so a worker saw claim failed: HTTP Error 403 while the same call from curl worked. The worker sends User-Agent: kalelm-worker/1 (<name>); keep a real user agent on anything else that talks to the API through the edge.

  • The ASR library’s default audio decoder (torchcodec) segfaults on some mp3s. Two crashes in one afternoon, both in libavformat under SingleStreamDecoder::getFramesPlayedInRangeAudio, each killing the worker with no Python error. The worker passes audio_backend="ffmpeg"; worker/run.sh restarts it if anything else of the kind happens, and a job whose worker goes silent returns to the queue after 10 minutes.

  • Kaggle’s Active events panel shows three entries; older sessions hide below them. A run from the first afternoon’s notebook (no torchvision fix, a worker without the failure guard) kept running for two hours under the same worker name as the good run and failed forty jobs in seconds each. Now every Kaggle run names itself kaggle-<container>, and the claim endpoint refuses a worker whose last three jobs all failed within fifteen minutes. When cancelling sessions, expand the panel and check every entry.

  • Meilisearch sizes indexing memory from the host, not its container. With MEILI_MAX_INDEXING_MEMORY at 1200Mb inside a 1400m container, four Kaggle workers uploading transcripts had the kernel kill it at 1.36 GB nineteen times in a day, and every search during a restart was a 500. It is 600Mb with one indexing thread now; the gate checks the number stays at most half the limit.

  • Ollama on this Mac has two servers: localhost resolves to an empty IPv6 one; models live on 127.0.0.1:11434.

  • Payload builds collections, access rules and custom endpoints at startup. next dev hot-reloads page code but not the config; an access-control or endpoint change looks unapplied until the dev server restarts. The gate failed three checks that way.

  • A self-signed Postgres certificate needs an IP subject-alt-name (-addext "subjectAltName=IP:…") or Node’s verify-full rejects it by hostname; a CN alone is not enough.

  • A field validator that rejects legacy data blocks unrelated edits. The file-URL check accepts an unchanged previousValue, otherwise the 42 migrated rows with off-origin URLs could never be saved again.

  • Astro’s origin check (security.checkOrigin) refuses the OAuth token endpoint. It rejects every form-encoded POST without a matching Origin, and a site’s server-to-server code exchange has none: openid-client reported “unexpected response content-type” (a 403 HTML page). The hub turns the check off and does its own in hub/src/middleware.ts for its forms; Better Auth checks its own routes.
  • The OAuth provider plugin registers web clients only with https callbacks, and refuses http://localhost and http://127.0.0.1 for them. A plain-http site is registered as a native client (/admin/sites does this by the URL), and its callback must be 127.0.0.1, not localhost.
  • The plugin has no trustedClients in 1.7.4: sites are rows in oauthClient, created through createOAuthClient by a signed-in admin (clientPrivileges). Without clientPrivileges any signed-in user could register a client.
  • Access tokens are opaque, not JWTs, so the hub’s data API cannot read the client id from the token. It hashes the bearer (SHA-256, base64url, no padding — the plugin’s defaultHasher) and looks it up in oauthAccessToken, which also checks expiry and revocation.
  • astro dev does not put .env on process.env; Better Auth read no base URL and the plugin crashed on new URL(undefined). The hub reads env through env() in hub/src/lib/auth.ts: process.env first (Hangar), import.meta.env after (dev).
  • The site prefetches links on hover, and a prefetch of /learn/login starts a sign-in flow. Each start rotates the state cookie; the prefetch and the real click raced, and the callback failed with “unexpected state response parameter value”. The sign-in links carry data-astro-prefetch="false" and the route answers a prefetch (Sec-Purpose) with 204.
  • auth.api.oauth2Consent cannot finish a consent: on accept the plugin resumes the authorization, which needs the raw request (request not found). The consent page posts to the plugin through auth.handler with a real Request, forwarding the browser’s cookies, and follows the redirect it answers with.
  • resvg loads TTF only, and fontTools cannot merge Reem Kufi’s subsets. The site’s woff2 files were converted with fontTools; the IBM Plex subsets merged into one TTF each, Reem Kufi’s variable-font tables refused, so its Arabic and Latin subsets are two files under one family name and resvg falls back between them per glyph. The fonts live in public/certificate/, which the deploy copies beside the server; a font read from src/ would not be in the standalone bundle.
  • payload migrate:create and drizzle’s push both stop on an interactive question (“is enum_courses_status created or renamed?”) that piped input cannot answer. The learn migration was written by hand in Payload’s naming and checked by creating a course through the Local API; the slug field’s helper column is generate_slug, not slug_lock.
  • pnpm verify is the gate and a Stop hook runs it. It needs local Postgres, the dev backend on 3001, the frontend on 4322, Meilisearch on 7701 and Ollama. If local Meilisearch is stopped the gate fails for that reason alone.
  • ESLint’s config crashes on every file. tsc --noEmit and the gate are the working checks.
  • docs/ is excluded from the backend tsconfig; Astro’s virtual modules would otherwise fail the type check.
  • Moving a Payload admin component means regenerating the import map (pnpm generate:importmap, also part of build) and restarting next dev; stale .next*/types/validator.ts files can also fail tsc after a route file is deleted — remove them, a build recreates them.
  • offsetTop is not “offset inside the scroll box”. It is measured from the nearest positioned ancestor, and the transcript list has none, so the auto-scroll used page coordinates and landed anywhere but on the spoken line. Measure with getBoundingClientRect() against the list’s own box (MaterialDetail.astro, bring()).
  • A Payload endpoint edit is not picked up by the running next dev. The search API kept answering with the old result shape until the dev server was restarted (preview_start payload-admin); the same rule as config changes.
  • A page kept for offline is nothing without its assets, and a script’s imports are not in the HTML. The worker used to cache.add the offline page and a saved material’s page on their own; after the next deploy their scripts were gone from the origin and never in the cache, so offline they rendered as text with no player, no list, no menus. keepDocument in sw.js.ts reads the HTML for its stylesheet, script and font addresses, keeps them, then reads each script for its own imports and keeps those. The worker is also stamped per deploy (BUILD), so the offline page is re-kept with the live build; and a kept recording is answered from the cache whether online or not, since navigator.onLine says online on a wifi with no internet.
  • caches.match() searches every cache. The record of a kept recording is a JSON response stored under the page’s own path in the records cache, so the offline fallback, looking the page up across all caches, handed the reader that JSON instead of the page. Pages are matched in the page cache only; the gate refuses caches.match( in the worker.
  • A page served from the worker’s cache runs the build it was cached with. The library page kept offline by an earlier worker asked that worker for the list; the next worker no longer answered, and “available offline” read empty over a cache full of recordings. A new worker keeps answering the old messages, and on activation drops every merely-visited page from the page cache and re-keeps the offline page and the kept recordings’ pages with the live build.
  • A service worker cannot download an hour of audio. Chrome stops a worker whose event has run for about five minutes, whatever waitUntil promised; a 27 MB save stalled at 80% and was never finished or reported. The download runs in the page now (features/pwa/client.ts), where module state survives the client router’s swaps; a full reload ends it, which is at least visible.
  • Cache Storage is best-effort until the site asks for persistence. Three recordings kept during one afternoon of testing were gone by the evening while two older ones stayed, with no code path that deletes them; the browser is allowed to evict. The save button now calls navigator.storage.persist() first, and the worker’s list falls back to the audio cache itself, so a recording whose record is missing still shows (by file name) and still plays.
  • A deploy must keep the previous build’s hashed assets. Cloudflare holds a material page for an hour; that HTML names the chunks of the build it was rendered by, and Hangar serves only the current release, so after a deploy every edge-cached page loaded a 404 for its scripts: no player, no continue, no menus, until the cache expired. deploy-hangar.sh frontend now copies the live release’s _astro into the new one (each release holds its predecessors’, so the set carries forward) and fails the deploy if the previous build’s first chunk stops answering. The clean fix is purging the edge on deploy, which needs a Cloudflare API token with cache-purge rights in the deploy environment.
  • lib/api.ts is imported by client scripts. The browse tree’s script imports it for API and a few helpers; an import of the Node-only loopback fetcher in it shipped node:https to the browser and the tree died at load with “Agent is not a constructor”. The server installs its fetcher from the middleware (setApiFetcher), the module itself stays plain; the gate refuses a Node import there.
  • The frontend must call the API over the loopback, not its public hostname. Both run on the Hangar box behind the same Caddy; the public name goes out through Cloudflare and back in, 0.5–1.2 s per call against 30 ms on the loopback, and a lesson page made five calls before it could start (3.9 s to first byte). frontend/src/lib/apiFetch.ts sends requests to API_PIN (127.0.0.1, set by the deploy script at build) with the hostname kept for SNI and the certificate check; the gate checks both sides.
  • One inline SVG per transcript cue was 450 KB of a 1 MB page. Anything repeated per row is drawn by CSS once (.tr__link::before), and the gate counts <svg> inside the transcript list.
  • A service worker must not proxy the mp3 while online. Passing fetch(req) through for files.kalelm.com gives the <audio> element an opaque response: Chrome cannot read Content-Range on it, treats every range request as the whole file, and the bar sat at 0:00 for twenty seconds while an hour of audio came down. The worker only answers media requests when navigator.onLine is false, from the cache (features/pwa/pages/sw.js.ts); the gate checks the guard is there.
  • Under the client router, a click listener sees the link already handled. The router’s own listener runs first and calls preventDefault, so navProgress never started its bar and every navigation looked dead. Start it from astro:before-preparation.
  • Astro scoped styles stop at the component. Rows rendered by a child component (PlaylistItems) do not match the parent’s .plist__items … rules; the styles live with the markup that owns them.

Hiding a sheikh hides their material, in four places at once

Section titled “Hiding a sheikh hides their material, in four places at once”

The «إخفاء الشيخ ومواده من الموقع» switch on a sheikh (2026-09-12) is read by four readers that must agree: the Materials REST access rule ('sheikh.hidden': { not_equals: true }, which keeps materials with no sheikh), every SQL reader through PUBLIC_MATERIAL() in src/utilities/publicMaterial.ts (facets, tree, series summary, series index), and the search and related endpoints through hiddenSheikhFilter() (sheikhId NOT IN […], all four indexes carry sheikhId, no reindex needed; read on every search, one indexed lookup). The gate takes one hidden sheikh’s published material and asserts it is absent from each. Add a fifth reader and the gate’s list in the data section is where it goes.

A sheikh, series or chapter with nothing published is not shown either (2026-09-12): the directories list only what the facets count, their pages 404, the sitemap and the tree’s sheikh search skip them. That rule lives in the frontend and the tree query, not in the Sheikhs access rule, because Payload cannot filter a collection by the absence of related rows.

Two things learned on the way:

  • Once the access rule joins the sheikh, Payload reports totalDocs: 0 for a page past the end (valid pages count correctly). The archive redirects an overrun to its first page.
  • ?limit=0 on /api/materials returns every document. Against 45,000 materials the dev server dies with Map maximum size exceeded. Use limit=1 to read a count. And curl treats [] in a URL as a glob: pass -g or the request never leaves the shell.

A persisted component wears the current page’s stylesheet

Section titled “A persisted component wears the current page’s stylesheet”

The player bar (and the save strip) persist across pages through transition:persist, but their CSS comes from whichever page is showing. After a deploy that adds a rule to the bar, a reader who navigates from a fresh page into an edge-cached old page sees the new markup without the new rule: the caption toggle rendered as a loose 42px ring under the bar (2026-09-12). Every page is cached ten minutes at the edge (material pages were an hour). Since 2026-09-13 the service worker closes the window for anyone who has visited before: it fetches every page under ?r=<release> (the worker’s build stamp changes with each deploy), which is a new edge cache key, and middleware.ts strips the parameter before any page renders. A first-time visitor, or a browser without the worker, can still meet a cached page for up to ten minutes after a deploy; a Cloudflare purge on deploy would close that too, and needs a token the deploy does not have. The owner’s own view is fresh from the second page load after a deploy: the first load installs the new worker, every navigation after it is keyed by the new release.