4.6 KiB
4.6 KiB
Fix: Optimized getAgents DB Queries for Production Scale
Trigger: Solve agents page timeouts/errors on demoplace production server Date: 2026-07-23 18:30 WIB Affected files:
src/lib/actions/agents.tssrc/app/(dashboard)/agents/columns.tsxbackend/routes/dashboard/summary.jssrc/proxy.ts
Solution
-
Diagnosed Root Cause:
- The Agents Inventory page on the production domain (
https://demoplace.my.id/agents) was failing with "An unexpected response was received from the server." (HTTP 500/504). - Remote backend logs showed no active errors, but database queries on
flowstimed out or hung. - Identified that
getAgentsserver action performed collection-wide aggregations anddistinctqueries on theflowscollection to calculate the last seen dates and active status. - On the production database, the
flowscollection holds over 11.9 million documents and lacks a general index starting withtimestampfor those queries. This resulted in full collection scans and sorts, triggering timeouts.
- The Agents Inventory page on the production domain (
-
Implemented Indexed Per-Agent Queries:
- Refactored
getAgentsto perform fast, individual queries per agent. - Utilized the existing composite index
{ agent_uuid: 1, timestamp: -1 }on theflowscollection. - Checked active status using
findOne({ agent_uuid, timestamp: { $gte: twentyFourHoursAgo } }, { projection: { _id: 1 } }). - Found flow last seen date using
findOne({ agent_uuid }, { projection: { timestamp: 1 }, sort: { timestamp: -1 } }).
- Refactored
-
Data Size Formatting:
- Updated the
Data Sizecolumn renderer insrc/app/(dashboard)/agents/columns.tsxto format dynamically:sizeMB >= 1024 * 1024formats asTBsizeMB >= 1024formats asGB- Otherwise formats as
MB.
- Wrote unit tests in
test/test-data-size-format.jsand successfully verified them.
- Updated the
-
Pruned Cumulative Database Telemetry:
- Diagnosed that the Overview Dashboard on demoplace displayed corrupted bandwidth totals (e.g.
16.27 TB) compared to Netify Portal (153 MB) because the database contained a mixture of historical cumulative telemetry and newly ingested incremental 5-minute deltas. - Executed a migration script
scripts/prune-production-cumulative.json the production MongoDB to delete the older cumulative summary documents from before the PM2 reload (pre-18:50WIB), resolving the TB/MB discrepancy.
- Diagnosed that the Overview Dashboard on demoplace displayed corrupted bandwidth totals (e.g.
-
Overview Flows Summation & Alignment:
- Resolved the issue where the Flows count KPI card displayed real-time concurrent flows (the latest 5-minute snapshot, e.g.,
126) instead of aggregating them over the selected time range (e.g., 24 hours). - Refactored
backend/routes/dashboard/summary.jsto count the actual number of documents in theFlowcollection matching the filter. - This ensures that both the Overview Dashboard Flows card and the
/flowslist page display identical, consistent counts (e.g.,30,759flows).
- Resolved the issue where the Flows count KPI card displayed real-time concurrent flows (the latest 5-minute snapshot, e.g.,
-
Overview Threats Fallback & Alignment:
- Resolved the discrepancy where the threats page showed
1threat, but the Overview Dashboard showed0threats. - Identified that the
/threatsAPI endpoint falls back to counting cybersecurity-related events from theEventcollection when there are no real threats in theThreatcollection. - Refactored
backend/routes/dashboard/summary.jsto implement the same fallback logic for the dashboard's "Threats" card count when the primaryThreatcount is0. - Both pages now consistently display
1threat.
- Resolved the discrepancy where the threats page showed
-
Next.js 16 Middleware Verification:
- Verified that Next.js 16 deprecates the
middleware.tsnaming convention in favor ofproxy.ts(exporting aproxyfunction). - Confirmed that
src/proxy.tsis fully active and automatically redirects unauthenticated users to/login(while logged-in users with a valid token cookie are bypassed to the dashboard directly).
- Verified that Next.js 16 deprecates the
-
Verification:
- Ran queries directly on the production database via SSH; response time dropped from hanging (>30s) to 192ms total.
- Executed local tests using
npx tsx test/test-actions-agents.js, verifying logic correctness. - Compiled Next.js locally (
npm run build) successfully with zero errors. - Deployed changes to production using
node scripts/deploy-sftp.js. - Verified that the
https://demoplace.my.id/agentsdashboard loaded successfully, showing formatted Data Sizes (e.g.7.48 GB) and correct real-time aggregate bandwidth (e.g.,2.02 MB). - Confirmed that Overview Dashboard displays matching flows (
30,797) and threats (1) in full alignment with their respective list pages.