- Refactor server initialization to separate concerns and improve error handling. - Implement centralized environment validation using Zod. - Introduce database, Redis, and queue lifecycle management. - Add health check endpoints for liveness and readiness. - Enhance error handling middleware for better response structure. - Implement rate limiting for API endpoints. - Add request ID middleware for traceability. - Create Sequelize CLI configuration and baseline migration for schema management. - Establish CI workflow with Gitea for testing and syntax checks. - Document foundational changes and migration strategy in PHASE_0_FOUNDATION_STABILIZATION.md. - Add Docker Compose configuration for local development and testing. - Implement unit and integration tests for critical functionality.
10 KiB
ZUMRI Phase 0 Foundation Stabilization
Objective
Stabilize the existing modular monolith so later identity and commerce work can build on deterministic startup, explicit schema management, observable health, baseline security, testability, and controlled shutdown. This phase does not add e-commerce domain behavior or intentionally redesign existing modules.
Starting Problems
The API listener started before asynchronous database authentication, runtime sequelize.sync() was the schema strategy, /health returned before its database check and hardcoded other dependencies as healthy, Redis clients connected during imports, and there was no environment validation, global error/404 handling, request correlation, security headers, rate limiting, graceful shutdown, migration system, or tests. Bull Board was public. Docker targeted Node 20 while the architecture targets Node 22. See Documentation/CURRENT_BACKEND_STATUS.md for the full baseline audit.
Changes Implemented
- Aligned package and container runtime to Node 22.
- Added centralized Zod environment validation with safe error messages.
- Separated Express construction (
app.js) from dependency initialization and listening (server.js). - Added database, Redis, queue, cron, API, and worker lifecycle handling.
- Removed runtime
sequelize.sync()and introduced a Sequelize CLI baseline migration. - Added liveness/readiness endpoints, request IDs, Helmet, body limits, general/sensitive rate limits, centralized errors, and centralized 404 behavior.
- Protected Bull Board with the existing JWT middleware and
admin/superadminaccount guard; registered activity, document, and log queues. - Kept
/apiand added/api/v1as a backward-compatible alias. - Added Jest/Supertest baseline tests and a portable JavaScript syntax-check command.
- Added Node 22 Docker hardening, development Compose, an Nginx example, and Gitea Actions CI.
- Restored the existing S3 utility interface through a feature-gated S3 client.
- Removed unsafe JWT/refresh-secret fallbacks and moved
nodemonto development dependencies. - Pinned
uuidto the CommonJS-compatible v11 line after tests exposed that v13 could not be loaded by this CommonJS application.
Application Startup Lifecycle
The API sequence is now:
- Load
.env. - Validate critical configuration without displaying values.
- Authenticate Sequelize (no schema mutation).
- connect to and ping Redis.
- Load the Express app and queue resources.
- Start cron only when
RUN_CRON=true. - Start the HTTP listener.
- On SIGTERM/SIGINT or fatal process error, stop accepting traffic, stop cron, close queues, Redis, and Sequelize, with a timeout guard.
Tests can import app.js without opening a TCP port. Startup failures prevent the listener from opening.
Environment Variables
Required for API/worker startup: NODE_ENV, PORT, DB_HOST, DB_PORT, DB_NAME, DB_USER, DB_PASSWORD, JWT_SECRET, REFRESH_TOKEN_SECRET, REDIS_HOST, REDIS_PORT, and FRONTEND_URL. JWT secrets must each be at least 32 characters. REDIS_PASSWORD is optional at schema level for deployments without Redis authentication.
Runtime controls: TRUST_PROXY (numeric trusted proxy hop count; keep 0 when directly exposed), JSON_BODY_LIMIT, API_RATE_LIMIT_WINDOW_MS, API_RATE_LIMIT_MAX, SENSITIVE_RATE_LIMIT_WINDOW_MS, SENSITIVE_RATE_LIMIT_MAX, RUN_CRON, CACHE, and SHUTDOWN_TIMEOUT_MS.
Optional mail variables are required as a complete group only when ENABLE_MAIL=true: MAIL_HOST, MAIL_PORT, MAIL_USER, MAIL_PASS, MAIL_FROM; MAIL_SECURE is optional. Optional S3 variables are required as a complete group only when ENABLE_S3=true: AWS_REGION, AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_S3_BUCKET_NAME. Docs credentials remain optional. See .env.sample; never commit real values.
Database Migration Strategy
Sequelize CLI uses .sequelizerc, app/config/sequelize-cli.config.js, and migrations/20260903000000-current-schema-baseline.js. Commands:
npm run db:migrate:status
npm run db:migrate
npm run db:migrate:undo
The baseline represents all currently registered models and creates only tables whose names are absent. It does not drop, alter, or validate columns during up. For an existing deployment: take a backup, compare its schema to the migration/model definitions, test against a restored copy, resolve drift explicitly, then run the migration so Sequelize records it. Do not blindly run the baseline undo in an existing environment; its standard down removes baseline tables. No migration was executed during this phase.
Future schema changes require new forward migrations. Production no longer calls sequelize.sync().
Health Endpoints
GET /health/live: process-only liveness; no dependencies queried.GET /health/ready: checks MySQL and Redis concurrently; returns 200/readyor 503/not_ready, exposing onlyok/errorstates.GET /health: backward-compatible alias to liveness so existing monitors are not broken.
Container orchestration should normally use liveness for process restart and readiness for traffic admission.
Security Middleware
Helmet is enabled; CSP is disabled globally for compatibility with the existing Bull Board/static documentation, while the board itself is authorization-protected. The API disables X-Powered-By, limits JSON and URL-encoded bodies to 1 MB by default, retains current credentialed CORS behavior, and accepts X-Request-ID only in a constrained safe format. Unknown routes and uncaught request errors use a standard response. Production 500 responses hide internal details.
JWT helpers have no public fallback secrets. Startup rejects missing/weak secrets. Morgan does not log Authorization, Cookie, or request bodies. Error logs contain request ID plus error name/message, not arbitrary error objects.
Rate Limiting
Both /api and /api/v1 use the configurable general limiter. The entire authentication router uses the stricter limiter, as do password reset requests/submissions and docs login. Health endpoints are outside limiters. TRUST_PROXY is an explicit numeric hop count rather than universal trust; set it to the exact Nginx hop count in deployment.
Bull Board Security
/admin/queues now runs through the existing authentication middleware and accepts only actual User model administrator values: admin and superadmin. It has no hardcoded secondary credentials. Activity, document, and log queues are registered.
Logging / Request IDs
Every request receives req.id and X-Request-ID; a safe incoming ID may be preserved. Development retains Morgan dev; production uses a concise method/path/status/timing line with request ID and no auth headers. The existing log queue remains in place. Phase 1 must remove remaining OTP/reset/session logging inside legacy authentication flows.
Graceful Shutdown
The API handles SIGTERM, SIGINT, unhandled rejections, and uncaught exceptions. It closes the HTTP listener, cron tasks, API-owned queues, shared Redis, and Sequelize. Workers initialize required dependencies before accepting jobs and close worker instances, Redis, and Sequelize on the same signals/fatal conditions. SHUTDOWN_TIMEOUT_MS protects against indefinitely stuck shutdown.
Cron Deployment Model
Cron is disabled unless RUN_CRON=true; tests do not start it. Until a distributed scheduler lock is added, enable it on exactly one API/scheduler instance. The cron launcher returns a stopper used during graceful shutdown.
Testing
npm run check:syntax
npm test -- --runInBand
npm run test:unit
npm run test:integration
Tests mock external infrastructure. Coverage includes liveness and readiness success/failure, /health compatibility, centralized 404/error responses, environment validation and optional features, authentication-required rejection, Bull Board rejection, and request IDs. No real MySQL, Redis, email, or S3 is required by the baseline suite.
Docker
The existing image now uses node:22-slim, keeps system Chromium/Puppeteer support, installs production dependencies, copies files as the unprivileged node user, and includes a /health/live healthcheck. Build and configuration are still environment-driven.
Local Development
Copy .env.sample to an ignored .env, replace all placeholder credentials/secrets, then run migrations explicitly before starting the API. compose.yaml provides API, worker, MySQL 8.4, and Redis 7.4 using the same application image and named data volumes. It does not auto-run migrations. Only the API port is published; MySQL/Redis remain internal.
CI
.gitea/workflows/ci.yml uses checkout/setup-node actions, Node 22, npm ci, syntax checks, and Jest. It performs no deployment and requires no production credentials. Runner action mirroring/network policy remains an installation-specific Gitea concern.
API Versioning Strategy
The existing router is mounted at both /api and /api/v1. Existing frontend calls remain valid, while new consumers should adopt /api/v1. A later compatibility window can deprecate /api; no route was mass-renamed in Phase 0.
Known Remaining Issues
- Authentication contains in-memory OTP/refresh-session behavior and needs the dedicated Phase 1 security/session design; no Phase 1 feature was implemented here.
- Ownership/IDOR and role/account naming inconsistencies remain in legacy controllers/routes.
- Permission routes remain imported but unmounted and role linkage/cache behavior needs repair.
- Existing authentication utilities may still log OTP/reset/session material; remove and test during Phase 1.
- Docs authentication still uses a weak boolean cookie and should receive a server-authenticated session design.
- Activity/log queues need standardized retry, retention, idempotency, and sensitive-data sanitation.
- S3 is now correctly constructed only when enabled, but object authorization/content validation and lifecycle remain later work.
- The dependency audit still reports transitive vulnerabilities; forced/major upgrades were intentionally avoided.
- Migration baseline schema drift must be reviewed against any deployed database before first use.
- A distributed cron lock is not yet present.
Phase 1 Prerequisites
The foundation is ready to begin Phase 1 once the baseline migration has been reviewed/tested against a copy of the deployment database and deployment secrets are configured. Phase 1 should focus on authentication/session durability, secure OTP/reset behavior, account status, role/permission integration, and ownership authorization without starting commerce modules.