Skip to main content

Monitoring and Logging

Logging​

Logging is configured in service application*.yml files rather than a centralized log aggregation stack.

Examples in the repository:

  • org.springframework.cloud.gateway: INFO
  • com.bookstore.analytics: INFO
  • com.bookstore.notification: INFO / DEBUG depending on profile
  • org.springframework.web: DEBUG in payment-service dev config

In production (EKS), logs are available via kubectl logs on individual pods. There is no centralized log collector (ELK, Loki, CloudWatch agent) configured in the current infrastructure.

Metrics (Prometheus)​

Most backend services expose Micrometer metrics through Spring Actuator:

management:
endpoints:
web:
exposure:
include: health,prometheus

Services with Prometheus exposure​

ServiceHealthPrometheusCustom metrics
auth-serviceYesYesBusinessMetrics counters
user-serviceYesYesBusinessMetrics counters
book-serviceYesYesBusinessMetrics counters
order-serviceYesYesBusinessMetrics counters
payment-serviceYesYesBusinessMetrics counters
notification-serviceYesYesBusinessMetrics counters
analytics-serviceYesYesBusinessMetrics counters
api-gatewayPartialNot configured in yml—

Production Prometheus stack​

The bookstore-infra repository deploys a full monitoring stack in the monitoring namespace, managed by separate Argo CD Applications:

ComponentManifest pathPurpose
Prometheusk8s/monitoring/prometheus/Scrapes annotated services in bookstore namespace
Grafanak8s/monitoring/grafana/Dashboards
Alertmanagerk8s/monitoring/alertmanager/Alert routing

Prometheus uses Kubernetes service discovery to find endpoints in the bookstore namespace. Services must carry these annotations to be scraped:

prometheus.io/scrape: "true"
prometheus.io/path: /actuator/prometheus
prometheus.io/port: "8083"

Prometheus stores time-series data on a 10 Gi EBS-backed PVC (gp2 storage class).

Prometheus is configured to forward alerts to Alertmanager at alertmanager.monitoring.svc.cluster.local:9093.

Health endpoints​

Actuator health endpoints are available at /actuator/health on services that expose Actuator. These are used for basic liveness checks but are not wired to Kubernetes liveness/readiness probes in all manifests.

Event observability (Kafka)​

Domain events provide an async audit trail:

EventProducerConsumers
payment-successpayment-serviceorder-service
payment-failedpayment-serviceanalytics-service
order-createdorder-servicenotification-service, analytics-service

The analytics-service deduplicates consumed events in a processed_events table.

Current observability posture​

CapabilityStatus
Service logs (stdout)Present
Actuator healthPresent on most services
Prometheus metrics (application)Present on most services
Prometheus server (production)Deployed via bookstore-infra
Grafana dashboardsDeployed via bookstore-infra
AlertmanagerDeployed via bookstore-infra
Distributed tracingNot implemented
Centralized log aggregationNot implemented