Actuator and Observability: Health, Metrics and Monitoring
Actuator is the dashboard Spring Boot installs in your application. Once the program runs, you cannot see "how busy am I, are my dependencies fine, how much heap is left, which build is deployed". Actuator packages that internal state into a set of HTTP addresses (called endpoints); you or your ops system asks with a browser or curl, and it answers. It draws no charts and fires no alerts — it just answers honestly.
Five terms, one line each:
- Endpoint: a fixed URL such as
/actuator/health; request it and you get JSON back - Exposure: which endpoints are actually reachable over HTTP; Boot ships only
healthandinfoby default - HealthIndicator: a tiny self-check function meaning "the part I own is UP or DOWN"; the container aggregates all of them into one status
- Micrometer: the metrics facade (what SLF4J is to logging) — your code only instruments, and swapping Prometheus for Datadog changes no business code
- Probe: the two doors Kubernetes knocks on periodically —
livenessasks "is the process alive",readinessasks "can it take traffic now"
think of it as a car dashboard plus the OBD diagnostic port. The tacho and the warning light are metrics and health — the driver can read them any time. The mechanic plugging into the OBD port to read fault codes is loggers, threaddump, heapdump. A dashboard is fine for anyone to see, but an OBD port lets a knowledgeable person rewrite your engine control unit — which is why only the workshop may touch it (internal network plus authentication), never a loose cable on the roadside.
health probes are like taking your blood pressure before a shift handover. Measuring yourself to confirm you can start work is readiness. "You look unwell today, so we are rushing you to resuscitation" is liveness — its criterion must be your own body (the process), never "the ambulance is late today" (a database wobble). Get that backwards and one network hiccup sends every nurse on the ward to emergency, emptying a department that was still functioning.

That tree is the map of this article: one branch is safe to publish, one is internal-only, one stays closed by default. Section 3 walks through each entry.
After this article you should be able to answer three questions:
- Why does
/actuator/healthsay DOWN without telling you which part broke — and how do you make it tell you? - What exactly is dangerous about
include: "*", and which lines belong in a production config? - Can you spot "the endpoints are getting slow" ten minutes early from a metric? Which one?
After launch, the sentence you dread most is "the system feels a bit slow". Behind it are three concrete questions: what is happening now, what changed, and why. The industry answers them with the "three pillars of observability":
| Pillar | Question it answers | Data shape | Common tools |
|---|---|---|---|
| Logging | what exactly happened at a moment | discrete event text | Logback + ELK / Loki |
| Metrics | what is the overall trend and magnitude | aggregatable time series | Micrometer + Prometheus + Grafana |
| Tracing | which hops a request took and where it slowed | causal request chains | OpenTelemetry / SkyWalking |
Logs are great at "pinpointing one incident", but to compute "the error rate over the last hour" you would have to count millions of log lines — far too slow. Metrics precompute "aggregatable numbers": QPS, latency percentiles, error rate, connection pool usage — time series by nature, ideal for dashboards and alerts. Tracing then stitches together cross-service call chains.
Spring Boot's Actuator is precisely the entry point for the latter two (especially metrics): it exposes the app's internal state through a set of HTTP endpoints. This article goes from "enabling Actuator" to "wiring up Prometheus and Grafana".

Actuator is not the whole monitoring system — it is the monitored side. It produces health and metrics; the storage, visualization and alerting are handled by Prometheus and Grafana.
Add one starter and it works with no other change:
<dependency> <groupId>org.springframework.boot</groupId> <artifactId>spring-boot-starter-actuator</artifactId></dependency>By default only health and info are exposed; the rest are not visible over HTTP — a safe default. To use more, list them explicitly:
management: endpoints: web: exposure: include: health,info,metrics,prometheus # only what you need endpoint: health: show-details: when_authorized # details only for authorized users server: port: 9090 # isolate management from business port- Use a whitelist instead of
; *never writeinclude: ""in production* management.server.port=9090moves management endpoints to a separate port, so the business port can face the internet while management stays internalshow-detailscontrols whether/healthreveals component details; tighten it when publicly exposed
Warning: include: "*" plus a public-facing app hands out env (possibly with DB passwords), heapdump (the whole heap, downloadable) and shutdown (remote kill). Several real incidents started exactly here.
Run the kernel lab first to see how this gate actually opens and closes:
Now push the gate fully open and watch what an attacker collects within thirty seconds:
You do not have to memorise the difference between those two configs — flip the switches and see it. Here is this article's exposure sandbox: on one side the "just show me everything" config people shout for in chat, on the other the pre-launch shape:
GET :9090/actuator -> _links: [self, health, info, metrics, prometheus]GET :9090/actuator/env -> 404 Not FoundPrometheus scrape OK: 15s interval#recommended shape: four endpoints, separate port
One more thing worth confirming: which file these settings live in, and who overrides whom, decides why "I changed the yml locally but production ignored it".
And the two endings side by side, as one picture:

Actuator has many endpoints, but a dozen matter in practice. Organized by purpose and risk:
| Endpoint | Purpose | Risk |
|---|---|---|
health | aggregate health for probes and load balancers | low (public OK, tighten details) |
info | build info, Git commit, custom properties | low |
metrics | runtime metrics (drill into individual meters) | medium |
prometheus | metrics in Prometheus text format | medium |
loggers | read/change log levels at runtime | medium |
beans | all beans and their dependencies | high |
mappings | all URL to handler mappings | high |
conditions | auto-configuration condition report | high |
threaddump | thread snapshot for deadlocks | high |
heapdump | downloadable heap dump (very large) | extreme |
env / configprops | env vars and config props (may hold secrets) | extreme |
httpexchanges | recent HTTP request records | high |
beans, mappings and conditions are development-time diagnostic gems — "why was this bean not registered", "which method is this URL mapped to" — answered at a glance. But they expose internal structure, so restrict access in production.
Endpoint names are not the thing to memorise — the question you are asking comes first. Here are six everyday questions, each matched with the endpoint to open before anything else. Match one wrongly and it will describe the price you just paid:
/actuator/health aggregates several built-in indicators (DB, Redis, disk, ping) into structured JSON:
{ "status": "UP", "components": { "db": { "status": "UP", "details": { "database": "MySQL", "validationQuery": "isValid()" } }, "diskSpace": { "status": "UP", "details": { "free": 107374182400, "total": 536870912000 } }, "redis": { "status": "UP", "details": { "version": "7.2.4" } }, "ping": { "status": "UP" } }}- The top-level
statusis an aggregate: any DOWN component makes the whole thing DOWN - HTTP mapping:
UPbecomes 200,DOWNbecomes 503 - The
dbandredisindicators are auto-registered by their starters; a configured datasource is enough
For a third-party service that a health check should not blanket-cover, write a custom HealthIndicator:
@Component("smsProvider")public class SmsProviderHealthIndicator implements HealthIndicator { private static final Logger log = LoggerFactory.getLogger(SmsProviderHealthIndicator.class); private final SmsClient smsClient; public SmsProviderHealthIndicator(SmsClient smsClient) { this.smsClient = smsClient; } @Override public Health health() { try { long latency = smsClient.ping(); return Health.up() .withDetail("latencyMs", latency) .withDetail("endpoint", smsClient.getEndpoint()) .build(); } catch (Exception ex) { log.warn("SMS provider probe failed", ex); return Health.down(ex).withDetail("endpoint", smsClient.getEndpoint()).build(); } }}In Kubernetes, health maps to probes, and liveness and readiness have completely different meanings. You must group them, or you invite the disaster of "a DB hiccup removes every instance":
management: endpoint: health: group: liveness: include: ping # only "is the process alive"; keep db out readiness: include: db,redis # "can it take traffic"; dependencies allowedTrap: putting db into liveness means a brief DB blip makes Kubernetes decide the container is "dead" and restart the process — a network glitch that would recover in tens of seconds escalates into a rolling restart of every instance. liveness should judge the process alone; dependency health belongs to readiness.
The aggregation deserves a live look: with three indicators and one DOWN, how is the top-level status voted out, and under what condition does it bother showing the details?
The act health lab showed you the outcome, but the word overall hides a real vote. Spread that voting rule into a single-step debugger and walk it line by line — click Next and watch how the severity ranking flips the whole verdict at step 5:
List<HealthIndicator> indicators = List.of(ping, db, redis); // 1 registration orderStatus overall = Status.UNKNOWN; // 2 starting point: no verdict yetfor (HealthIndicator ind : indicators) { // 3 call the roll Status cur = ind.health().getStatus(); // 4 this round's answer if (severity(cur) > severity(overall)) overall = cur; // 5 the graver one wins}return overall; // 6 the vote's outcome| indicators | 3 |
| order | ping → db → redis |
HealthEndpoint.readCompositeHealthIndicator.healthThen zoom out to the rollout itself — what each probe answers at startup, during a wobble, on recovery and at shutdown:

Watch the graceful-shutdown chain next to it and you will see why draining traffic must come before killing the process:
Spring Boot uses Micrometer as its metrics facade (like SLF4J for logging): business code depends only on MeterRegistry, and the backend can be swapped to Prometheus, Datadog and so on. Know the four fundamental meters:
| Meter | Semantics | Example |
|---|---|---|
| Counter | monotonically increasing total | orders placed, error count |
| Gauge | instantaneous value, up or down | current connections, queue length |
| Timer | timing plus call count | endpoint latency, method duration |
| DistributionSummary | distribution of values | order amount, payload size |
Built-in metrics by source, the common ones:
| Prefix | Source | Why it matters |
|---|---|---|
jvm.* | JVM (memory, GC, threads) | detect leaks and frequent GC |
http.server.requests | web requests | QPS, P99 latency, error rate |
hikaricp.connections.* | DB connection pool | pool exhaustion (a slow-query omen) |
tomcat.threads.* | Tomcat thread pool | thread starvation |
system.cpu.usage | OS | CPU saturation |
hikaricp.connections.pending (threads waiting for a connection) staying above zero usually precedes "endpoints getting slow" and is one of the most valuable early-warning metrics.
Those numbers do not materialise on a dashboard — they are pulled out of /actuator/metrics one question at a time:
Technical metrics tell you "the system is healthy"; business metrics tell you "the business is healthy". Inject MeterRegistry to instrument:
@Servicepublic class OrderMetrics { private final Counter orderCreated; private final Counter orderFailed; private final Timer payLatency; public OrderMetrics(MeterRegistry registry) { this.orderCreated = Counter.builder("beeorder.order.created") .description("total orders created") .register(registry); this.orderFailed = Counter.builder("beeorder.order.failed") .description("total orders failed") .register(registry); this.payLatency = Timer.builder("beeorder.pay.latency") .description("payment callback latency") .publishPercentileHistogram() .register(registry); } public void markCreated() { orderCreated.increment(); } public void markFailed() { orderFailed.increment(); } public void recordPay(Runnable task) { payLatency.record(task); // times the call and counts it }}- Counter is monotonic, perfect for "cumulative count"; one
incrementcall does it - Timer records both count and duration, and
publishPercentileHistogramemits percentiles - Prefix names with your business domain (like
beeorder.) to avoid mixing with built-ins
Method-level timing is also available via @Timed (requires a registered TimedAspect):
@Timed(value = "beeorder.product.detail", description = "product detail query latency")public ProductVO detail(Long productId) { return productRepository.findById(productId).map(ProductVO::from).orElseThrow(NotFoundException::new);}Key point: business metrics are the starting point for alerts. Expose "payment success rate = successes / callbacks" as a metric and you can write "alert when success rate drops below 99%" — instead of waiting for user complaints.
First turn metrics into a format Prometheus can read — add a registry dependency:
<dependency> <groupId>io.micrometer</groupId> <artifactId>micrometer-registry-prometheus</artifactId></dependency>Hitting /actuator/prometheus yields text like this (excerpt):
# HELP http_server_requests_seconds Timer for HTTP server requests# TYPE http_server_requests_seconds histogramhttp_server_requests_seconds_count{method="POST",uri="/api/orders",status="201"} 1284.0http_server_requests_seconds_sum{method="POST",uri="/api/orders",status="201"} 41.52http_server_requests_seconds_bucket{method="POST",uri="/api/orders",status="201",le="0.1"} 990.0http_server_requests_seconds_bucket{method="POST",uri="/api/orders",status="201",le="0.5"} 1270.0# HELP beeorder_order_created_total Orders created successfully# TYPE beeorder_order_created_total counterbeeorder_order_created_total 13245.0Prometheus scrapes this endpoint on a schedule and stores it as time series:
scrape_configs: - job_name: beeorder metrics_path: /actuator/prometheus scrape_interval: 15s static_configs: - targets: ['beeorder-app:9090']Grafana turns the time series into dashboards, and alert rules notify you automatically:
groups: - name: beeorder-alerts rules: - alert: HighErrorRate expr: | sum(rate(http_server_requests_seconds_count{status="500"}[5m])) / sum(rate(http_server_requests_seconds_count[5m])) > 0.01 for: 5m labels: { severity: critical } annotations: summary: "5xx error rate exceeds 1% for 5 minutes"rate(...[5m])gives the per-second rate over 5 minutes, better reflecting "now" than a raw counterfor: 5mmeans "alert only if the breach persists 5 minutes", avoiding flapping false alarms- A dashboard needs at least three panels: traffic and error rate, P99 latency, and pool/thread water levels

A production endpoint errors out, but the log level is INFO and hides the detail. Restarting to change config is too costly — use the loggers endpoint to change it online:
# 1) inspect the current level of a loggercurl http://localhost:9090/actuator/loggers/com.beeorder.order# 2) temporarily switch the order package to DEBUGcurl -X POST http://localhost:9090/actuator/loggers/com.beeorder.order \ -H "Content-Type: application/json" \ -d '{"configuredLevel":"DEBUG"}'# 3) revert to INFO once donecurl -X POST http://localhost:9090/actuator/loggers/com.beeorder.order \ -H "Content-Type: application/json" \ -d '{"configuredLevel":"INFO"}'- The change takes effect immediately, with no restart, scoped precisely to a package or class
- It affects only the current instance; with multiple instances, change each or push via a config center
- Always revert afterwards, or DEBUG logs will quickly fill the disk
Tip: combined with the httpexchanges endpoint you can review recent requests (path, status, duration), often locating "which request broke" faster than grepping the log.
What does this endpoint actually change? The level field on a Logback logger object — run the whole pipeline once and it is obvious:
/actuator/info shows build info, the Git commit and custom properties, answering "which version is actually running". First make Maven produce build metadata, then expose it:
<plugin> <groupId>org.springframework.boot</groupId> <artifactId>spring-boot-maven-plugin</artifactId> <executions> <execution> <goals> <goal>build-info</goal> <!-- generates META-INF/build-info.properties --> </goals> </execution> </executions></plugin>management: info: git: mode: full # requires spring-boot-starter-actuator + the git-commit-id plugin env: enabled: trueNow hitting /actuator/info shows the version, build time and Git commit id — "is production new or old" is no longer a guessing game.
Attention: Git info requires the extra `git-commit-id-maven-plugin` and a `.git` directory at build time; with a shallow clone in CI the Git info may be empty, so configure the plugin to fetch it.
Actuator is useful but double-edged by default. Check each item before launch:
| Item | Action |
|---|---|
| Endpoint whitelist | include only health,info,metrics,prometheus |
| Isolate management port | separate management.server.port, internal only |
| Disable sensitive endpoints | turn off env, configprops, heapdump (or enabled: false) |
| Enforce auth | protect /actuator/** with Spring Security, allow only health |
| Disable shutdown | management.endpoint.shutdown.enabled: false (default; do not turn it on) |
| Tighten health details | show-details: never publicly, when_authorized internally |
You do not have to memorise the checklist — tick the lines that matter and let the generator assemble a yml skeleton you can take straight to production:
server:
port: 8080
spring:
application:
name: demo-service
management:
server:
port: 9090 # 管理端口与业务端口隔离
endpoints:
web:
exposure:
include: health,info,metrics,prometheus # 白名单,绝不写 *
endpoint:
health:
show-details: when_authorized
group:
liveness: { include: ping }
readiness: { include: db,redis,diskSpace }
the generator hands you a skeleton, not a clearance certificate — the value of show-details, whether your gateway blocks the /actuator path, and whether the network policy admits only the internal plane are three things it cannot know. Confirm each one yourself.
the heapdump endpoint downloads the entire heap, which may contain DB URLs, keys and user data. Disable it unless in a controlled internal debugging window.
putting db into the liveness group is the classic incident of "a DB hiccup removes instances from the load balancer, or even gets them restarted by Kubernetes". Remember the grouping rule: liveness asks only "is the process alive"; readiness asks "are dependencies fine and can it take traffic".
management.endpoints.web.exposure.include: "*" plus a 0.0.0.0 bind puts /actuator/env, /actuator/heapdump and /actuator/shutdown on the public internet. Countless apps in the real world have leaked config or been killed with one request this way.
Actuator turns the app from a black box into something that answers questions — health answers "is it alive", metrics answers "how fast, how busy, how stable", info answers "which version is running", and beans/mappings/conditions answer "why is it wired this way". Production rollout needs just three steps: whitelist exposure + management port isolation + Prometheus/Grafana for visualization and alerting. Remember the bottom line: observability presupposes security — do not let a debugging entry point become an attack surface.
With the tools known and the traps marked, string together the order in which your hands should reach for things on the night it breaks — every cell uses one endpoint from this article, and no cell may be skipped:

And the summary that goes with it: observing endpoints make the story tellable, diagnosing endpoints make it findable, and hardening decides who may even ask — miss any of the three and the four-capability picture in Section 0 stays a capability table instead of a triage workflow.
Every row below can be copied and searched as it stands. Beginners stall in the same four places: the endpoint 404s before you have even begun, health says DOWN without naming who, /env leaks a secret, and the management port gets confused with the business port.
| Error text (fragment) | Real cause | 30-second rescue | Read more in | |
|---|---|---|---|---|
Whitelabel Error Page ... type=Not Found, status=404 on /actuator or /actuator/health | Two candidates: spring-boot-starter-actuator was never added; or it is on the classpath but this endpoint is not in the management.endpoints.web.exposure.include whitelist (Boot ships only health and info) | First `mvn dependency:tree \ | grep actuator to confirm the dependency is really there, then curl localhost:8080/actuator and read which keys the index page _links` actually lists — the index page is the one entry point that cannot lie to you | Section 2 of this article · the exposure sandbox |
/actuator/health returns {"status":"DOWN"}, but the body carries no components | management.endpoint.health.show-details defaults to never; the aggregation rule lets any single DOWN component drag the whole status down — so you see the verdict but not the culprit | Set show-details: always temporarily (internal network only), or scope it with management.endpoint.health.group.<name>.show-details=always so just one group opens up; then read components to find the culprit | Section 4 of this article · the health aggregation experiment | |
In GET /actuator/env, spring.datasource.password prints as ******, but your own app.oss.accessKey and app.wx.secret2 come out in plain text | Sanitising is a regex match on the property name (by default covering words like password, secret, key, token); pick a different name and nothing matches. Objects bound through @ConfigurationProperties are dumped whole in configprops | Extend the sanitised key list: management.endpoint.env.keys-to-sanitize=password,secret,key,token,certificate,accessKey (on Boot 3.x, management.endpoint.env.show-values=never plus keys-to-sanitize), and switch configprops off as well | #21 Configuration in Full | |
management.server.port=9090 is configured, yet the Kubernetes probe still gets 404 or cannot connect at all | The probe URL still points at the business port 8080; once the management port changes, /actuator/** listens on 9090 only | Rewrite livenessProbe/readinessProbe from port: http(8080) to port: management(9090), and confirm the Service's network policy lets the cluster reach 9090 | Section 4 of this article · the probe animation | |
Both 8080 and 9090 answer locally, but the moment production Nginx forwards, /actuator is on the public internet | Nothing isolates them: either both ports share one origin, or the gateway forwards by path without ever blocking /actuator | Deny that path explicitly at the gateway; more robust, leave the management port out of the public port list of the Ingress/Service altogether | Section 10 · the hardening checklist | |
POST /actuator/loggers/com.example.order answers 415 Unsupported Media Type or 400 | No Content-Type: application/json, or the body field was written as level (the real field is configuredLevel) | curl -X POST -H "Content-Type: application/json" -d '{"configuredLevel":"DEBUG"}'; sending {"configuredLevel":null} means "inherit from the parent again" | Section 8 | |
The startup log contains Cannot find factory with name "Prometheus", or there simply is no /actuator/prometheus | Only the actuator starter was added, without io.micrometer:micrometer-registry-prometheus; or the registry is there but the endpoint was never included | Add the dependency → restart → put prometheus in include → curl :9090/actuator/prometheus should return # HELP jvm_... text | Section 7 | |
The dependency is clearly there, but /actuator/shutdown keeps 404ing, and ops says "we need to restart it remotely" | By default the shutdown endpoint does not even create its bean; you need management.endpoint.shutdown.enabled=true and it in the include list — both switches or nothing is listening | Think it through first: in containers you restart with SIGTERM plus graceful shutdown, not with an HTTP call. If you really must open it, put authentication in front of it | Section 4 · the probe animation · #46 going live | |
The heapdump download is several hundred MB, MAT cannot open it, or the machine goes OOM | Taking a heap snapshot triggers a stop-the-world pause; on a large heap that single step can halt the service for ten-plus seconds — which is an incident by itself | Run it only inside a controlled window, or switch to jcmd <pid> GC.heap_dump; in production this endpoint stays disabled | Section 10 · the hardening checklist |
The rows above can all be rescued from the error text alone. One class of incident is less kind: the error does not name the culprit — the clue hides in the numbers inside the parentheses. This stack is an excerpt thrown on a probe thread during a real Pod-cycling incident — do not read the conclusion yet, point out the frame you believe is the culprit:
At 01:40 the on-call phone erupts: instances of pay-api have been restarted twice within ten minutes, each time preceded by a probe timeout. A restart fixes it at once, and it recurs a while later. This is the stack excerpt the operators recovered from the killed instance — thrown on a probe thread. Do not read the conclusion yet, point out the frame you believe is the culprit.
Goal: put together a minimal monitoring project that runs entirely on your laptop and whose config is already a production skeleton — and see for yourself the 200-versus-404 difference a whitelist makes, the components block inside /health, and one line of curl switching a package to DEBUG.
Before you type anything, run the five labs in this kernel console, in order — each one is a conclusion from this article and every echo is computed by the kernel itself:
five labs walk you through expose, tighten, observe and probe end to end; the three tiers below then move that flow into a project of your own.
Step one, pom.xml:
<?xml version="1.0" encoding="UTF-8"?><project xmlns="http://maven.apache.org/POM/4.0.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 https://maven.apache.org/xsd/maven-4.0.0.xsd"> <modelVersion>4.0.0</modelVersion> <parent> <groupId>org.springframework.boot</groupId> <artifactId>spring-boot-starter-parent</artifactId> <version>3.3.4</version> <relativePath/> </parent> <groupId>com.example</groupId> <artifactId>monitor-lab</artifactId> <version>0.0.1-SNAPSHOT</version> <properties> <java.version>17</java.version> </properties> <dependencies> <dependency> <groupId>org.springframework.boot</groupId> <artifactId>spring-boot-starter-web</artifactId> </dependency> <!-- Actuator itself --> <dependency> <groupId>org.springframework.boot</groupId> <artifactId>spring-boot-starter-actuator</artifactId> </dependency> <!-- the registry that emits the Prometheus text format --> <dependency> <groupId>io.micrometer</groupId> <artifactId>micrometer-registry-prometheus</artifactId> </dependency> <!-- gives us a real db health indicator to aggregate --> <dependency> <groupId>org.springframework.boot</groupId> <artifactId>spring-boot-starter-jdbc</artifactId> </dependency> <dependency> <groupId>com.h2database</groupId> <artifactId>h2</artifactId> <scope>runtime</scope> </dependency> </dependencies> <build> <plugins> <plugin> <groupId>org.springframework.boot</groupId> <artifactId>spring-boot-maven-plugin</artifactId> <executions> <execution> <goals> <goal>build-info</goal> <!-- so /actuator/info has something to say --> </goals> </execution> </executions> </plugin> </plugins> </build></project>Step two, src/main/resources/application.yml:
server: port: 8080 # business port: open to the internetspring: application: name: monitor-lab datasource: url: jdbc:h2:mem:lab;DB_CLOSE_DELAY=-1 driver-class-name: org.h2.Drivermanagement: server: port: 9090 # management port: reachable from the internal network only endpoints: web: base-path: /actuator exposure: include: health,info,metrics,prometheus # a whitelist, never a * endpoint: health: show-details: when_authorized group: liveness: include: ping readiness: include: db,diskSpace info: env: enabled: truelogging: level: com.example.monitorlab: INFOStep three, one custom health indicator and one order-count instrument (package com.example.monitorlab):
package com.example.monitorlab;import java.util.concurrent.atomic.AtomicLong;import org.springframework.boot.actuate.health.Health;import org.springframework.boot.actuate.health.HealthIndicator;import org.springframework.stereotype.Component;@Component("inventoryCache")public class InventoryCacheHealthIndicator implements HealthIndicator { private final AtomicLong lastSyncOkAt = new AtomicLong(System.currentTimeMillis()); @Override public Health health() { long staleMs = System.currentTimeMillis() - lastSyncOkAt.get(); if (staleMs > 60_000) { return Health.down().withDetail("staleMs", staleMs).withDetail("reason", "inventory cache has not been synced for over 60s").build(); } return Health.up().withDetail("staleMs", staleMs).build(); } public void markSynced() { lastSyncOkAt.set(System.currentTimeMillis()); }}Step four, start the app and run these curls in order (watch which ones are 8080 and which are 9090):
# ① the index page: confirm the exposure surface is exactly four endpointscurl -s http://localhost:9090/actuator# ② the health summary (unauthorized here, so no components)curl -s http://localhost:9090/actuator/health# ③ the grouped probe addressescurl -s http://localhost:9090/actuator/health/livenesscurl -s http://localhost:9090/actuator/health/readiness# ④ a management endpoint on the business port -> should be 404curl -i -s http://localhost:8080/actuator/health | head -1# ⑤ a sensitive endpoint should be 404curl -i -s http://localhost:9090/actuator/env | head -1# ⑥ change the log level at runtime, then observe immediatelycurl -s -X POST http://localhost:9090/actuator/loggers/com.example.monitorlab \ -H "Content-Type: application/json" -d '{"configuredLevel":"DEBUG"}'curl -s http://localhost:9090/actuator/loggers/com.example.monitorlabcurl -s -X POST http://localhost:9090/actuator/loggers/com.example.monitorlab \ -H "Content-Type: application/json" -d '{"configuredLevel":null}'# ⑦ drilling into metricscurl -s http://localhost:9090/actuator/metricscurl -s "http://localhost:9090/actuator/metrics/hikaricp.connections.pending"curl -s http://localhost:9090/actuator/prometheus | grep -E "^jvm_memory_used_bytes|^http_server_requests" | head -3Expected output (the default shape of ② when you are unauthorized, the two groups from ③, and the status codes of ④ and ⑤):
{ "status": "UP" }{ "status": "UP", "group": "liveness" }{ "status": "UP", "group": "readiness", "components": { "db": { "status": "UP" }, "diskSpace": { "status": "UP" } } }HTTP/1.1 404HTTP/1.1 404Acceptance checklist: ① you can explain why there is no db inside liveness; ② change show-details to always, call ② again, and the custom inventoryCache indicator now shows up in components; ③ you can say what causes each of the two 404s in ④ and ⑤ (the first is port isolation, the second is the whitelist).
One line at a time — and the conclusion flips:
- Replace
includewith"*"and restart. What you will observe: the_linksobject fromcurl :9090/actuatorsuddenly has a dozen or more keys; insideenvthe value ofspring.datasource.passwordis masked, but theapp.my.access-key-alias=plain-secretyou just added prints verbatim — that is the actual boundary of the sanitising rules. Put the config back when you are done. - Comment out the
management.server.portline. What you will observe: ④ turns into 200, because the management endpoints are back on the business port — and if your gateway then forwards by path,/actuatoris public. This one is what makes "port isolation" click. - Add
dbto thelivenessgroup, then break startup with a wrong URL (jdbc:h2:mem:lab;INIT=RUNSCRIPT FROM 'nonexistent.sql'). What you will observe:/actuator/health/livenessgoes straight to DOWN — a rehearsal of "Kubernetes repeatedly restarts an instance that only needed to be pulled out of rotation". - Change the threshold in
InventoryCacheHealthIndicatorfrom60_000to-1. What you will observe: the top-levelstatusturns DOWN and the HTTP code becomes 503, but by default you still cannot see why — so you are pushed into changingshow-details, which is exactly the second row of Section 12. - After step ⑥, hit an endpoint a dozen or more times. What you will observe: DEBUG logs pour out; set the level back to
null(inherit from the parent) and it goes quiet at once. The value and the risk of "debugging without a restart", in one minute.
Hint: after item 3, go back to the probe animation in Section 4 — the two timelines should line up exactly.
Build yourself an "on-call self-service panel": one browser screen that shows the key conclusions, with no Grafana installed.
Requirements:
- A custom endpoint
@Endpoint(id = "opsboard")whose@ReadOperationreturns one aggregated JSON: JVM heap usage, the GC count delta, pending connections in the pool, the 5xx ratio over the last 5 minutes, the row count oft_order, the running Git commit, and uptime - All of it comes from an injected
MeterRegistry,DataSourceandHealthContributor; no new database tables allowed - A
@TransactionalEventListener(AFTER_COMMIT)or an interceptor records the business success rate, as one tile of the panel - The panel endpoint is not exposed by default; it can only be opened through the
includelist of an internal profile, and your documentation says plainly how to open it and when to close it again - A
README.mdwith three curls that perform the standard sequence: "spot the anomaly → locate the component → raise the log level temporarily"
Acceptance checklist: ① every field in the output of curl :9090/actuator/opsboard is traceable to a metric name you can point at; ② make the database unavailable on purpose (a wrong URL, or cut the network with Testcontainers) — the matching tile must turn red, /health/readiness must be DOWN and /health/liveness must still be UP; ③ after 60 seconds of load, the 5xx ratio on the panel is within 5% of the figure you compute by hand from http_server_requests_seconds_count{status="500"}; ④ with the internal profile switched off, the endpoint returns 404.
which two endpoints does Actuator expose over the Web by default, and why that particular pair?
how is the top-level status of /actuator/health computed? With the three values of show-details, name exactly what each one hides and what it reveals.
what question does liveness answer, and what question does readiness answer? What specific incident does db in the wrong group cause?
what is the semantic difference between Counter, Gauge and Timer? Which one fits "the latency distribution of payment callbacks", and why?
where do the sanitising rules on /actuator/env stop working? Give two hardening measures that follow directly from that.
a whitelist opens the doors, a separate port closes them again, details only on authorization; liveness judges only itself, readiness asks about the dependencies; metrics come in, dashboards come out — and never leave the OBD port on the roadside.
if you can answer the five questions above without scrolling back, you have what it takes to put Actuator on a live service. The shape never changes — instrument once with Micrometer, publish four endpoints, move them onto their own port, split liveness from readiness, let Prometheus and Grafana do the watching, and keep the table in Section 12 where you can paste straight from it at 3 a.m. The one line to remember: observability is only worth having if the entry point was designed the day before the incident, not improvised during it.