# Feature List (backend) Structured log of shipped backend features, updated by the `n`/`next` workflow (see [AGENTS.md](../AGENTS.md) Part B) whenever a task in [plans/next-enhancements.md](../plans/next-enhancements.md) is marked `[DONE]`. Split out 2026-07-08 from root `docs/feature-list.md`'s backend sections — this file is the sole home for backend feature history going forward. ## Format ``` ##
- **** — shipped ``` --- ## Existing Features (pre-kit) Backfilled 2026-07-08 during adoption of this kit — these predate the `e`/`n` workflow and have no task numbers; see `git log` for real dates/history. ### Backend — Next.js API Gateway - Upload/parse/documents CRUD routes, GPU status endpoint, vLLM proxy, manual-label review tool. ### Backend — OCR Pipeline & Accuracy - PaddleOCR + vLLM classification pipeline with DB layout caching, table column-shift correction, date normalization, and an accuracy regression harness (`pfm-web-app/scripts/accuracy-check.mts`) — **95.10% overall as of 2026-07-08** (target 95% met; see task 2.3 below for the investigation and `CLAUDE.md`). ### Backend — Postgres Data Layer - Schema/init in `pfm-web-app/src/db/init.ts`, served via the canonical root `docker-compose.yml` stack. ### DevOps — Docker & Dev Tunnel - Root `docker-compose.yml` (dev, hot-reload) and `docker-compose.demo.yml` (production-mode override); `start-dev-tunnel.ps1` syncs the host LAN IP into the Flutter app config and starts an ngrok tunnel. *(New features shipped via `n`/`next` go below, organized the same way, with task numbers.)* ## Backend — Next.js API Gateway - **1.4** Enforced real 401 auth on `/api/v1/documents/*` (list, PUT-by-id, upload) — the actual production API surface, already fully supported by the Flutter client (real login + `Authorization: Bearer` on every request). Previously none of these three routes rejected a missing/invalid token; upload only optionally read it. Added the pre-existing `getAccountFromAuthHeader()` helper (`utils/auth.ts`) + a 401 guard to all three; `OPTIONS` (CORS preflight) untouched. The original task 1.3 (auth on the *classic* routes) was cancelled instead — those routes are dev-only web UI surface with no login flow, going away in production. Verified via `curl`: 401 with no token, success with a real token from `/api/v1/auth/login` — shipped 2026-07-08. ## Backend — OCR Pipeline & Accuracy - **2.1** Built the Product/SKU scan classifier's model artifacts: `models/dinov2_index.pkl` (118/118 reference photos indexed across 16 SKU classes) and `models/produk-pfm-classifier-26n-100e-2026-07-08.pt` (+ `.onnx` export) — a YOLO classifier fine-tuned 100 epochs, 83.3% top-1 / 90% top-5 validation accuracy on the current (thin, 2-16 photos/class) dataset. Built via a one-off `docker run` from a freshly-rebuilt `pipeline-api` image (bare-metal training isn't viable on Windows — `paddlepaddle-gpu`'s wheel index is Linux-only). `pipeline-api` restarted and confirmed loading both models from logs. Also fixed `scripts/install-pipeline.sh`, which was missing `ultralytics`/`torch` — shipped 2026-07-08. - **2.1 (verification pass)** Ran a full browser walkthrough of `/scan-pfm` (classification, top-5, OCR expiry extraction + crop, SKU-master matching, Visual/Spotting Grid, Raw Response — all confirmed working with real data). Found and fixed a real bug: "Save Ground Truth" was returning success but silently writing into the `pfm-web-app` container's ephemeral filesystem instead of the host, because `/sources` wasn't a bind-mounted path in root `docker-compose.yml`. Added `./backend/sources:/sources` to the `pfm-web-app` service, recovered an orphaned entry via `docker cp`, and re-verified the save now persists to `backend/sources/product_manual_labels.json` on the host (confirmed the DO-flow's `manual_labels.json` save was fixed by the same change too) — shipped 2026-07-08. - **2.3** Ran the accuracy regression harness and discovered `sources/accuracy_report.md` was badly stale (claimed 75.04%; real current baseline is **95.10% overall, already at/above the 95% target** — added a staleness banner to that file). Root-caused every remaining mismatch by pulling raw OCR text from Postgres (`documents.layout_parsing_result`): the worst field, `plat` (67.6%), is almost entirely the license-plate region being classified as an image/seal by the layout model rather than OCR'd as text — not fixable in `parser.ts`. Found and fixed one genuine parser logic bug along the way: the "global pattern scanning fallback" could duplicate an already-correctly-extracted `noDO` value into a still-missing `noSO` field; fixed by excluding already-assigned values from that fallback's candidate pool (`pfm-web-app/src/utils/parser.ts`). Doesn't change the aggregate score (a wrong value and "Not Found" score the same) but stops a fabricated-looking wrong number from silently reaching the database. All 48 parser unit tests still pass — shipped 2026-07-08. ## Backend — Postgres Data Layer - **3.1** Wrapped the `ocr_items` delete-then-reinsert in `/api/parse` and `/api/v1/documents/[id]` PUT inside a DB transaction (`withTransaction` helper, `pfm-web-app/src/db/index.ts`) — a mid-loop insert failure now rolls back to the previous item set instead of leaving a document with a correct header but partial/missing items — shipped 2026-07-08. (Renumbered from root's `7.1` when this file split from root `docs/feature-list.md`.) - **3.2** Hashed `accounts.password` with `bcryptjs` (pure-JS, no native compile step — the `pfm-web-app` Docker image has no build toolchain). `db/init.ts` hashes the seed and idempotently migrates any pre-existing plaintext rows on every startup; `api/v1/auth/login/route.ts` now compares with `bcrypt.compareSync` and cleanly rejects missing credentials with a 401 instead of risking a raw-query edge case. Verified via `psql` (hash format) and `curl` (correct login succeeds, wrong/missing password returns 401) — shipped 2026-07-08.