Multi-channel notification engine built from scratch as an architecture case study: runtime-swappable queue backends and delivery channels, with end-to-end distributed tracing.
- NestJS
- TypeScript
- PostgreSQL
- TypeORM
- BullMQ
- +7
What I did
- Designed and built an asynchronous notification engine end-to-end in NestJS, with queue-based processing, automatic retries and dead-letter handling.
- Implemented a queue abstraction layer (Adapter pattern, swappable BullMQ/SQS) and a delivery abstraction layer (Strategy pattern, email/Slack/SMS), keeping business logic decoupled from any concrete implementation.
- Instrumented the system with OpenTelemetry, exporting traces to Grafana Tempo, covering both asynchronous processing paths.
- Ran a structured self-audit against my own codebase and resolved the critical findings: API-key authentication, rate limiting, idempotency against duplicate sends, and a real npm dependency vulnerability.
Results
Swappable queues without touching business logic
BullMQ (Redis) and AWS SQS sit behind a single IQueue interface. The active backend is chosen with one environment variable; when SQS is active, the app makes no attempt to connect to Redis at all.
One environment variable to switch queue backends
Security self-audit and remediation
Found that the main endpoint operated as an unauthenticated open relay (could be used to send arbitrary HTML through my Resend account). Implemented API-key auth, rate limiting and idempotency to prevent duplicate sends on retry.
Critical audit findings resolved
Distributed tracing across both async paths
Centralized the sending logic into a single service method shared by the BullMQ worker and the SQS consumer, guaranteeing identical OpenTelemetry traces regardless of which backend processed the notification.
End-to-end traces in Grafana Tempo for both backends
Technologies
- NestJS
- TypeScript
- PostgreSQL
- TypeORM
- BullMQ
- AWS SQS
- Redis
- Docker
- OpenTelemetry
- Grafana Tempo
- Resend
- Slack API
