feedc0de 196286783a
Test and deploy / test (push) Successful in 11s
Test and deploy / deploy (push) Failing after 14s
Add authenticated PDF editor for Paperless scans
2026-10-10 15:41:26 +02:00
2026-10-04 16:40:38 +02:00

Paperless-ngx on the homelab

One Paperless-ngx instance, PostgreSQL, and Valkey in the paperless-ngx namespace. Documents, scanner staging, and the incoming folder use the retained CephFS PVC. PostgreSQL and Paperless's search index use separate Ceph block PVCs. The app is served at https://paper.brunner.ninja through Traefik. Back up the three PVCs and the paperless-secret-key and paperless-postgres Secrets together.

Install

Run ./install.sh from any directory. It creates the namespace and random, persistent application, PostgreSQL, and OIDC secrets on first install, creates the Authentik applications, installs the pinned PostgreSQL Helm chart with the official PostgreSQL image, checks the manifests with a server dry run, applies them, then waits for the deployments. Open the URL and sign in with the brunner.ninja Auth button. Drop scanned files in the web UI or into /usr/src/paperless/consume on the PVC. An existing SQLite installation needs a database migration before switching app.yaml to PostgreSQL; the homelab instance was migrated on 2026-10-05.

Change Paperless settings in app.yaml and PostgreSQL settings in postgresql-values.yaml, then run ./test.sh and ./install.sh. Both images and the chart version are pinned. The CephFS inbox is checked every five seconds because network filesystems may not deliver file notifications reliably.

Paperless prefers the 24-core arschrock node for OCR. This is a scheduler preference, so it can run on an Odroid when arschrock is unavailable. It has 24 task workers with one OCR thread each and no Kubernetes memory limit. This can be heavy on an Odroid during failover; reduce PAPERLESS_TASK_WORKERS in app.yaml if that happens. The control-plane taint is tolerated only by Paperless, not Valkey. A node failure may briefly interrupt OCR while Kubernetes reschedules the single Paperless pod.

The 20 Gi CephFS PVC uses rook-cephfs-ec, whose data pool is erasure coded with three data and two coding chunks across hosts. CephFS metadata uses three replicas. PostgreSQL and Paperless's runtime data use rook-ceph-block-ec. The claims can be expanded if scans or the index outgrow them; erasure coding does not replace backups.

Scanner

The ScanservJS web UI is at https://scanner.brunner.ninja. Point its DNS CNAME at the same target as paper.brunner.ninja. Traefik requires Authentik login and access to the Paperless Users group. The scanner's eSCL address is set once as AIRSCAN_DEVICES in app.yaml; update it there if the printer gets a new IP. The UI pulls scans directly from the HP OfficeJet Pro 8020 at 192.168.4.189:8080, without HP cloud services or a running laptop.

Load the automatic document feeder, choose the HP scanner and its feeder source, select Auto batch mode, then press Scan. Review the staged result in the Files view, select its row, then choose Action selected → Send to Paperless. ScanservJS may save a multipage JPG batch as one ZIP row; the action converts its ordered images into one multipage PDF before placing it in the existing Paperless consume directory. PDF and single-image results also work. Paperless checks the directory every five seconds and handles OCR. The scan remains in staging until you send or delete it. ScanservJS does not offer reliable individual-page deletion or retry before sending: rescan the batch, or edit/split the PDF after import in Paperless. The OfficeJet Pro 8020's feeder is single-sided. This workflow starts scans from the web UI rather than a button on the printer panel.

ScanservJS stores staged scans in the scanner-staging directory on the same retained CephFS claim. The deployment uses the pinned upstream image; its config.local.js action is in app.yaml and has a local handoff test in test.sh. The Authentik proxy application is created by setup-authentik.sh along with Paperless's OIDC application.

If an ADF scan immediately says Document feeder out of documents and 0 pages scanned, remove and reseat the stack in the feeder. That response comes from the OfficeJet before the first page is captured; no scan file is created. ScanservJS currently displays the raw scanner error.

To merge separate Paperless documents imported from JPG scans, enable Try to include archive version in merge for non-PDF files in the merge dialog. Paperless's default merge tries to open the original JPG as a PDF and skips it. Enable Delete original documents after successful merge if you want only the merged entry in the active document list; the source entries go to Paperless's trash. A single feeder batch sent through the scanner action already becomes one document, without a later merge.

Rotate or crop an existing document

Paperless's PDF Editor works only when the selected file version is a PDF. For a scanned JPG, download its archived PDF from the document's download menu; Paperless creates that PDF during OCR. Open https://scanner.brunner.ninja/pdf/ to use the self-hosted Stirling PDF editor, protected by the same Authentik gate as ScanservJS. Its Rotate and Crop tools let you turn pages and visually select the area to keep. Download the edited PDF, then return to the original Paperless document and choose Versions → Upload new version. The document ID, title, tags, correspondent, document type, permissions, and custom fields stay with that Paperless entry; the previous file remains available in version history. Once the new PDF version has imported, Paperless's own PDF Editor can rotate or rearrange it too.

Paperless has no native crop tool or custom button that sends a document to another editor. Cropping a PDF changes its visible page boundary without resampling the scan image. For sensitive content outside the crop, use a redaction tool rather than relying on crop alone. Stirling PDF uses the pinned upstream Ultra-Lite image and temporary container storage; no document PVC or cloud account is involved.

CI/CD

./test.sh checks syntax, the deployment manifests, and the scanner PDF handoff. ./deploy.sh updates the existing workloads, services, ConfigMap, and ingresses using app.yaml and pdf-editor.yaml. The GitLab and Gitea workflows call these same scripts. Bootstrap once with ./create-ci-kubeconfig.sh, then save its single output line as the protected CI variable / Gitea Actions secret KUBE_CONFIG_BASE64. Do not commit it. The deployer has named-object read and patch permissions and cannot read Secrets. Run ./install.sh manually for first installation or changes to storage.yaml, postgresql-values.yaml, or Authentik setup.

The Git remote is gitea@brunner.ninja:feedc0de/paperless-ngx-deployment.git. No credentials or generated kubeconfig are committed. To publish changes: git add . && git commit -m 'Deploy Paperless-ngx' && git push origin main.

Authentik login

./setup-authentik.sh creates the OAuth2/OpenID Connect provider, application, and Paperless Users and Paperless Admins groups in Authentik, adds feedc0de to both groups, and stores the client secret in the paperless-oidc Kubernetes Secret. It preserves that secret on later runs. ./install.sh invokes it automatically. To grant others access, add them to Paperless Users in Authentik's web UI; use Paperless Admins only for people who should see and administer all documents. Paperless syncs those Authentik group claims at login, including administrator status. Existing Paperless users can link their Authentik account from My Profile. Local login stays enabled as a recovery option. The callback is https://paper.brunner.ninja/accounts/oidc/authentik/login/callback/ per Authentik's Paperless guide.

The sign-in button says brunner.ninja Auth; the internal OIDC provider ID remains authentik so existing account links keep working.

The first-visit create-user form is Paperless's bootstrap screen. Since the initial local feedc0de account was created, first sign in at /accounts/login/ with its local password, then choose My Profile → Connect new social account and authenticate with Authentik. Starting with the Authentik button before linking opens a new-account registration screen and conflicts with the existing username. After linking, use the Authentik button for future logins; the Paperless Admins claim grants administrator status.

Recovery

The PVCs have retained storage. Do not delete them on uninstall. Restore the PostgreSQL PVC, document PVC, runtime PVC, and both application and database Secrets together. The former SQLite database remains on the CephFS data directory as a migration fallback; it is no longer used by Paperless. For a consistent backup, use PostgreSQL's backup tools and copy the document storage at a coordinated point in time, or use Paperless's document exporter.

S
Description
No description provided
Readme GPL-3.0
165 KiB
Languages
Shell 65.4%
Python 34.6%