Actuator and Observability: Health, Metrics and Monitoring

bee2026-10-0860 min read0 views
After enabling Actuator: custom health checks, exposing metrics to Prometheus, changing log levels at runtime and custom endpoints — making "what is happening in production" answerable.
1 / 147
Section
0. The 30-second version
2 / 147

Actuator is the dashboard Spring Boot installs in your application. Once the program runs, you cannot see "how busy am I, are my dependencies fine, how much heap is left, which build is deployed". Actuator packages that internal state into a set of HTTP addresses (called endpoints); you or your ops system asks with a browser or curl, and it answers. It draws no charts and fires no alerts — it just answers honestly.

3 / 147

Five terms, one line each:

4 / 147
  • Endpoint: a fixed URL such as /actuator/health; request it and you get JSON back
  • Exposure: which endpoints are actually reachable over HTTP; Boot ships only health and info by default
  • HealthIndicator: a tiny self-check function meaning "the part I own is UP or DOWN"; the container aggregates all of them into one status
  • Micrometer: the metrics facade (what SLF4J is to logging) — your code only instruments, and swapping Prometheus for Datadog changes no business code
  • Probe: the two doors Kubernetes knocks on periodically — liveness asks "is the process alive", readiness asks "can it take traffic now"
5 / 147
类比|Analogy

think of it as a car dashboard plus the OBD diagnostic port. The tacho and the warning light are metrics and health — the driver can read them any time. The mechanic plugging into the OBD port to read fault codes is loggers, threaddump, heapdump. A dashboard is fine for anyone to see, but an OBD port lets a knowledgeable person rewrite your engine control unit — which is why only the workshop may touch it (internal network plus authentication), never a loose cable on the roadside.

6 / 147
类比|Analogy

health probes are like taking your blood pressure before a shift handover. Measuring yourself to confirm you can start work is readiness. "You look unwell today, so we are rushing you to resuscitation" is liveness — its criterion must be your own body (the process), never "the ambulance is late today" (a database wobble). Get that backwards and one network hiccup sends every nurse on the ward to emergency, emptying a department that was still functioning.

7 / 147
Diagram
Figure · The Actuator endpoint family
Figure · The Actuator endpoint family
8 / 147

That tree is the map of this article: one branch is safe to publish, one is internal-only, one stays closed by default. Section 3 walks through each entry.

9 / 147

After this article you should be able to answer three questions:

10 / 147
  • Why does /actuator/health say DOWN without telling you which part broke — and how do you make it tell you?
  • What exactly is dangerous about include: "*", and which lines belong in a production config?
  • Can you spot "the endpoints are getting slow" ten minutes early from a metric? Which one?
11 / 147
Section
1. The three pillars of observability: answering "how is production?"
12 / 147

After launch, the sentence you dread most is "the system feels a bit slow". Behind it are three concrete questions: what is happening now, what changed, and why. The industry answers them with the "three pillars of observability":

13 / 147
Table
PillarQuestion it answersData shapeCommon tools
Loggingwhat exactly happened at a momentdiscrete event textLogback + ELK / Loki
Metricswhat is the overall trend and magnitudeaggregatable time seriesMicrometer + Prometheus + Grafana
Tracingwhich hops a request took and where it slowedcausal request chainsOpenTelemetry / SkyWalking
14 / 147

Logs are great at "pinpointing one incident", but to compute "the error rate over the last hour" you would have to count millions of log lines — far too slow. Metrics precompute "aggregatable numbers": QPS, latency percentiles, error rate, connection pool usage — time series by nature, ideal for dashboards and alerts. Tracing then stitches together cross-service call chains.

15 / 147

Spring Boot's Actuator is precisely the entry point for the latter two (especially metrics): it exposes the app's internal state through a set of HTTP endpoints. This article goes from "enabling Actuator" to "wiring up Prometheus and Grafana".

16 / 147
Diagram
Figure 1 · Four capabilities of Actuator
Figure 1 · Four capabilities of Actuator
17 / 147
Note

Actuator is not the whole monitoring system — it is the monitored side. It produces health and metrics; the storage, visualization and alerting are handled by Prometheus and Grafana.

18 / 147
Section
2. Quick start: expose first, then tighten
19 / 147

Add one starter and it works with no other change:

20 / 147
xml
<dependency>    <groupId>org.springframework.boot</groupId>    <artifactId>spring-boot-starter-actuator</artifactId></dependency>
21 / 147

By default only health and info are exposed; the rest are not visible over HTTP — a safe default. To use more, list them explicitly:

22 / 147
Code
Codeyaml
management:  endpoints:    web:      exposure:        include: health,info,metrics,prometheus   # only what you need  endpoint:    health:      show-details: when_authorized              # details only for authorized users  server:    port: 9090                                   # isolate management from business port
Notes
  • Use a whitelist instead of ; *never write include: "" in production*
  • management.server.port=9090 moves management endpoints to a separate port, so the business port can face the internet while management stays internal
  • show-details controls whether /health reveals component details; tighten it when publicly exposed

Warning: include: "*" plus a public-facing app hands out env (possibly with DB passwords), heapdump (the whole heap, downloadable) and shutdown (remote kill). Several real incidents started exactly here.

23 / 147

Run the kernel lab first to see how this gate actually opens and closes:

24 / 147
Kernel lab
TeaVMExposure control: who can see which endpointsidle
Pick "Exposure": notice that the /actuator index lists only what you put in include; everything else 404s
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
25 / 147

Now push the gate fully open and watch what an attacker collects within thirty seconds:

26 / 147
Kernel lab
TeaVMFull exposure: from env to shutdownidle
Switch to "Everything open": /env, /heapdump, /loggers and /shutdown are demonstrated one by one
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
27 / 147

You do not have to memorise the difference between those two configs — flip the switches and see it. Here is this article's exposure sandbox: on one side the "just show me everything" config people shout for in chat, on the other the pre-launch shape:

28 / 147
Sandbox
SandboxExposure switch: include=* versus a whitelist
Result
GET :9090/actuator -> _links: [self, health, info, metrics, prometheus]
GET :9090/actuator/env -> 404 Not Found
Prometheus scrape OK: 15s interval
#recommended shape: four endpoints, separate port
This is the target config of Tier 1 below: full observability, not one sensitive endpoint exposed
29 / 147

One more thing worth confirming: which file these settings live in, and who overrides whom, decides why "I changed the yml locally but production ignored it".

30 / 147
Kernel lab
TeaVMWhich source overrides this management propertyidle
Pick "Who wins": command line > JVM args > environment variables > application-{profile}.yml > application.yml
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
31 / 147

And the two endings side by side, as one picture:

32 / 147
Diagram
Figure · Convenient on the LAN, catastrophic on the internet
Figure · Convenient on the LAN, catastrophic on the internet
33 / 147
Section
3. The core endpoints, one by one
34 / 147

Actuator has many endpoints, but a dozen matter in practice. Organized by purpose and risk:

35 / 147
Table
EndpointPurposeRisk
healthaggregate health for probes and load balancerslow (public OK, tighten details)
infobuild info, Git commit, custom propertieslow
metricsruntime metrics (drill into individual meters)medium
prometheusmetrics in Prometheus text formatmedium
loggersread/change log levels at runtimemedium
beansall beans and their dependencieshigh
mappingsall URL to handler mappingshigh
conditionsauto-configuration condition reporthigh
threaddumpthread snapshot for deadlockshigh
heapdumpdownloadable heap dump (very large)extreme
env / configpropsenv vars and config props (may hold secrets)extreme
httpexchangesrecent HTTP request recordshigh
36 / 147
Key point

beans, mappings and conditions are development-time diagnostic gems — "why was this bean not registered", "which method is this URL mapped to" — answered at a glance. But they expose internal structure, so restrict access in production.

37 / 147

Endpoint names are not the thing to memorise — the question you are asking comes first. Here are six everyday questions, each matched with the endpoint to open before anything else. Match one wrongly and it will describe the price you just paid:

38 / 147
Match
MatchSix questions, matched with the first endpoint to openMatched 0/6 · Missed 0
Both columns are shuffled; the criterion is whether the question is about seeing state, finding a cause, or taking action — the wrong endpoint costs you half an hour in front of a 404 or sensitive data
Pick a card on the left first
39 / 147
Section
4. Health: from built-in to custom
40 / 147

/actuator/health aggregates several built-in indicators (DB, Redis, disk, ping) into structured JSON:

41 / 147
Code
Codejson
{  "status": "UP",  "components": {    "db": { "status": "UP", "details": { "database": "MySQL", "validationQuery": "isValid()" } },    "diskSpace": { "status": "UP", "details": { "free": 107374182400, "total": 536870912000 } },    "redis": { "status": "UP", "details": { "version": "7.2.4" } },    "ping": { "status": "UP" }  }}
Notes
  • The top-level status is an aggregate: any DOWN component makes the whole thing DOWN
  • HTTP mapping: UP becomes 200, DOWN becomes 503
  • The db and redis indicators are auto-registered by their starters; a configured datasource is enough
42 / 147

For a third-party service that a health check should not blanket-cover, write a custom HealthIndicator:

43 / 147
java
@Component("smsProvider")public class SmsProviderHealthIndicator implements HealthIndicator {    private static final Logger log = LoggerFactory.getLogger(SmsProviderHealthIndicator.class);    private final SmsClient smsClient;    public SmsProviderHealthIndicator(SmsClient smsClient) {        this.smsClient = smsClient;    }    @Override    public Health health() {        try {            long latency = smsClient.ping();            return Health.up()                    .withDetail("latencyMs", latency)                    .withDetail("endpoint", smsClient.getEndpoint())                    .build();        } catch (Exception ex) {            log.warn("SMS provider probe failed", ex);            return Health.down(ex).withDetail("endpoint", smsClient.getEndpoint()).build();        }    }}
44 / 147

In Kubernetes, health maps to probes, and liveness and readiness have completely different meanings. You must group them, or you invite the disaster of "a DB hiccup removes every instance":

45 / 147
Code
Codeyaml
management:  endpoint:    health:      group:        liveness:          include: ping            # only "is the process alive"; keep db out        readiness:          include: db,redis        # "can it take traffic"; dependencies allowed
Notes

Trap: putting db into liveness means a brief DB blip makes Kubernetes decide the container is "dead" and restart the process — a network glitch that would recover in tens of seconds escalates into a rolling restart of every instance. liveness should judge the process alone; dependency health belongs to readiness.

46 / 147

The aggregation deserves a live look: with three indicators and one DOWN, how is the top-level status voted out, and under what condition does it bother showing the details?

47 / 147
Kernel lab
TeaVMHealth aggregation: who drags the whole thing to DOWNidle
Switch to "Health aggregation": watch how show-details changes whether components are visible at all
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
48 / 147

The act health lab showed you the outcome, but the word overall hides a real vote. Spread that voting rule into a single-step debugger and walk it line by line — click Next and watch how the severity ranking flips the whole verdict at step 5:

49 / 147
Stepper
StepperStep through it: how three indicators vote the overall status1 / 7
Click Next; from step 5 watch the severity ranking and how one DOWN rewrites the whole ballot
Code under debug
1List<HealthIndicator> indicators = List.of(ping, db, redis); // 1 registration order
2Status overall = Status.UNKNOWN; // 2 starting point: no verdict yet
3for (HealthIndicator ind : indicators) { // 3 call the roll
4 Status cur = ind.health().getStatus(); // 4 this round's answer
5 if (severity(cur) > severity(overall)) overall = cur; // 5 the graver one wins
6}
7return overall; // 6 the vote's outcome
Variables now
indicators3
orderping → db → redis
Call stack
1HealthEndpoint.read
2CompositeHealthIndicator.health
1The three indicators queue up in registration order. The order does not change the final tally, only whom you see first in the details block.
50 / 147

Then zoom out to the rollout itself — what each probe answers at startup, during a wobble, on recovery and at shutdown:

51 / 147
Animation
Animation · Health probes across a rollout
Animation · Health probes across a rollout
52 / 147

Watch the graceful-shutdown chain next to it and you will see why draining traffic must come before killing the process:

53 / 147
Kernel lab
TeaVMReadiness → drain → in-flight requests complete → exitidle
Start on "Readiness" to see when it turns true, then switch to "Drain in-flight" to watch running requests finish
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
54 / 147
Section
5. Metrics: the Micrometer facade
55 / 147

Spring Boot uses Micrometer as its metrics facade (like SLF4J for logging): business code depends only on MeterRegistry, and the backend can be swapped to Prometheus, Datadog and so on. Know the four fundamental meters:

56 / 147
Table
MeterSemanticsExample
Countermonotonically increasing totalorders placed, error count
Gaugeinstantaneous value, up or downcurrent connections, queue length
Timertiming plus call countendpoint latency, method duration
DistributionSummarydistribution of valuesorder amount, payload size
57 / 147

Built-in metrics by source, the common ones:

58 / 147
Table
PrefixSourceWhy it matters
jvm.*JVM (memory, GC, threads)detect leaks and frequent GC
http.server.requestsweb requestsQPS, P99 latency, error rate
hikaricp.connections.*DB connection poolpool exhaustion (a slow-query omen)
tomcat.threads.*Tomcat thread poolthread starvation
system.cpu.usageOSCPU saturation
59 / 147
Tip

hikaricp.connections.pending (threads waiting for a connection) staying above zero usually precedes "endpoints getting slow" and is one of the most valuable early-warning metrics.

60 / 147

Those numbers do not materialise on a dashboard — they are pulled out of /actuator/metrics one question at a time:

61 / 147
Kernel lab
TeaVMHow metrics are fetched: from the name list to drill-down tagsidle
Switch to "Metrics": first GET /actuator/metrics for the roster, then drill in with ?name= to read availableTags
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
62 / 147
Section
6. Custom business metrics: make business numbers observable
63 / 147

Technical metrics tell you "the system is healthy"; business metrics tell you "the business is healthy". Inject MeterRegistry to instrument:

64 / 147
Code
Codejava
@Servicepublic class OrderMetrics {    private final Counter orderCreated;    private final Counter orderFailed;    private final Timer payLatency;    public OrderMetrics(MeterRegistry registry) {        this.orderCreated = Counter.builder("beeorder.order.created")                .description("total orders created")                .register(registry);        this.orderFailed = Counter.builder("beeorder.order.failed")                .description("total orders failed")                .register(registry);        this.payLatency = Timer.builder("beeorder.pay.latency")                .description("payment callback latency")                .publishPercentileHistogram()                .register(registry);    }    public void markCreated() { orderCreated.increment(); }    public void markFailed() { orderFailed.increment(); }    public void recordPay(Runnable task) {        payLatency.record(task);   // times the call and counts it    }}
Notes
  • Counter is monotonic, perfect for "cumulative count"; one increment call does it
  • Timer records both count and duration, and publishPercentileHistogram emits percentiles
  • Prefix names with your business domain (like beeorder.) to avoid mixing with built-ins
65 / 147

Method-level timing is also available via @Timed (requires a registered TimedAspect):

66 / 147
Code
Codejava
@Timed(value = "beeorder.product.detail", description = "product detail query latency")public ProductVO detail(Long productId) {    return productRepository.findById(productId).map(ProductVO::from).orElseThrow(NotFoundException::new);}
Notes

Key point: business metrics are the starting point for alerts. Expose "payment success rate = successes / callbacks" as a metric and you can write "alert when success rate drops below 99%" — instead of waiting for user complaints.

67 / 147
Section
7. Wiring up Prometheus + Grafana
68 / 147

First turn metrics into a format Prometheus can read — add a registry dependency:

69 / 147
xml
<dependency>    <groupId>io.micrometer</groupId>    <artifactId>micrometer-registry-prometheus</artifactId></dependency>
70 / 147

Hitting /actuator/prometheus yields text like this (excerpt):

71 / 147
text
# HELP http_server_requests_seconds Timer for HTTP server requests# TYPE http_server_requests_seconds histogramhttp_server_requests_seconds_count{method="POST",uri="/api/orders",status="201"} 1284.0http_server_requests_seconds_sum{method="POST",uri="/api/orders",status="201"} 41.52http_server_requests_seconds_bucket{method="POST",uri="/api/orders",status="201",le="0.1"} 990.0http_server_requests_seconds_bucket{method="POST",uri="/api/orders",status="201",le="0.5"} 1270.0# HELP beeorder_order_created_total Orders created successfully# TYPE beeorder_order_created_total counterbeeorder_order_created_total 13245.0
72 / 147

Prometheus scrapes this endpoint on a schedule and stores it as time series:

73 / 147
yaml
scrape_configs:  - job_name: beeorder    metrics_path: /actuator/prometheus    scrape_interval: 15s    static_configs:      - targets: ['beeorder-app:9090']
74 / 147

Grafana turns the time series into dashboards, and alert rules notify you automatically:

75 / 147
Code
Codeyaml
groups:  - name: beeorder-alerts    rules:      - alert: HighErrorRate        expr: |          sum(rate(http_server_requests_seconds_count{status="500"}[5m]))          / sum(rate(http_server_requests_seconds_count[5m])) > 0.01        for: 5m        labels: { severity: critical }        annotations:          summary: "5xx error rate exceeds 1% for 5 minutes"
Notes
  • rate(...[5m]) gives the per-second rate over 5 minutes, better reflecting "now" than a raw counter
  • for: 5m means "alert only if the breach persists 5 minutes", avoiding flapping false alarms
  • A dashboard needs at least three panels: traffic and error rate, P99 latency, and pool/thread water levels
76 / 147
Animation
Animation · Metrics from app to dashboard
Animation · Metrics from app to dashboard
77 / 147
Section
8. Changing log levels at runtime: debugging without a restart
78 / 147

A production endpoint errors out, but the log level is INFO and hides the detail. Restarting to change config is too costly — use the loggers endpoint to change it online:

79 / 147
Code
Codebash
# 1) inspect the current level of a loggercurl http://localhost:9090/actuator/loggers/com.beeorder.order# 2) temporarily switch the order package to DEBUGcurl -X POST http://localhost:9090/actuator/loggers/com.beeorder.order \     -H "Content-Type: application/json" \     -d '{"configuredLevel":"DEBUG"}'# 3) revert to INFO once donecurl -X POST http://localhost:9090/actuator/loggers/com.beeorder.order \     -H "Content-Type: application/json" \     -d '{"configuredLevel":"INFO"}'
Notes
  • The change takes effect immediately, with no restart, scoped precisely to a package or class
  • It affects only the current instance; with multiple instances, change each or push via a config center
  • Always revert afterwards, or DEBUG logs will quickly fill the disk

Tip: combined with the httpexchanges endpoint you can review recent requests (path, status, duration), often locating "which request broke" faster than grepping the log.

80 / 147

What does this endpoint actually change? The level field on a Logback logger object — run the whole pipeline once and it is obvious:

81 / 147
Kernel lab
TeaVMThe full chain behind changing a level via /actuator/loggersidle
Pick "Level filtering" to see INFO blocked before the Appender, then "Async appender" to understand why DEBUG flooding the disk slows requests
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
82 / 147
Section
9. The info endpoint and build information
83 / 147

/actuator/info shows build info, the Git commit and custom properties, answering "which version is actually running". First make Maven produce build metadata, then expose it:

84 / 147
xml
<plugin>    <groupId>org.springframework.boot</groupId>    <artifactId>spring-boot-maven-plugin</artifactId>    <executions>        <execution>            <goals>                <goal>build-info</goal>   <!-- generates META-INF/build-info.properties -->            </goals>        </execution>    </executions></plugin>
85 / 147
yaml
management:  info:    git:      mode: full        # requires spring-boot-starter-actuator + the git-commit-id plugin    env:      enabled: true
86 / 147

Now hitting /actuator/info shows the version, build time and Git commit id — "is production new or old" is no longer a guessing game.

87 / 147

Attention: Git info requires the extra `git-commit-id-maven-plugin` and a `.git` directory at build time; with a shallow clone in CI the Git info may be empty, so configure the plugin to fetch it.

88 / 147
Kernel lab
89 / 147
Section
10. A hardening checklist: do not leave a back door on the internet
90 / 147

Actuator is useful but double-edged by default. Check each item before launch:

91 / 147
Table
ItemAction
Endpoint whitelistinclude only health,info,metrics,prometheus
Isolate management portseparate management.server.port, internal only
Disable sensitive endpointsturn off env, configprops, heapdump (or enabled: false)
Enforce authprotect /actuator/** with Spring Security, allow only health
Disable shutdownmanagement.endpoint.shutdown.enabled: false (default; do not turn it on)
Tighten health detailsshow-details: never publicly, when_authorized internally
92 / 147

You do not have to memorise the checklist — tick the lines that matter and let the generator assemble a yml skeleton you can take straight to production:

93 / 147
Generator
GeneratorTurn the hardening checklist into a few management linesapplication.yml1 / 3
Tick management alone for the whitelist, the 9090 management port and the liveness/readiness groups; add logging for levels and rolling policy, add profiles for per-environment blocks — then check every line against the table above
Output
server:
  port: 8080

spring:
  application:
    name: demo-service

management:
  server:
    port: 9090                            # 管理端口与业务端口隔离
  endpoints:
    web:
      exposure:
        include: health,info,metrics,prometheus   # 白名单,绝不写 *
  endpoint:
    health:
      show-details: when_authorized
      group:
        liveness: { include: ping }
        readiness: { include: db,redis,diskSpace }
Why each choice matters
managementA whitelist plus a separate management port is the floor; * is the number-one leak incident cause.
94 / 147
Note

the generator hands you a skeleton, not a clearance certificate — the value of show-details, whether your gateway blocks the /actuator path, and whether the network policy admits only the internal plane are three things it cannot know. Confirm each one yourself.

95 / 147
Warning

the heapdump endpoint downloads the entire heap, which may contain DB URLs, keys and user data. Disable it unless in a controlled internal debugging window.

96 / 147
Section
11. Two traps + decision + summary
97 / 147
Trap

putting db into the liveness group is the classic incident of "a DB hiccup removes instances from the load balancer, or even gets them restarted by Kubernetes". Remember the grouping rule: liveness asks only "is the process alive"; readiness asks "are dependencies fine and can it take traffic".

98 / 147
Trap

management.endpoints.web.exposure.include: "*" plus a 0.0.0.0 bind puts /actuator/env, /actuator/heapdump and /actuator/shutdown on the public internet. Countless apps in the real world have leaked config or been killed with one request this way.

99 / 147
Decision
Decisionshould the health check include the database?
100 / 147
Summary

Actuator turns the app from a black box into something that answers questions — health answers "is it alive", metrics answers "how fast, how busy, how stable", info answers "which version is running", and beans/mappings/conditions answer "why is it wired this way". Production rollout needs just three steps: whitelist exposure + management port isolation + Prometheus/Grafana for visualization and alerting. Remember the bottom line: observability presupposes security — do not let a debugging entry point become an attack surface.

101 / 147

With the tools known and the traps marked, string together the order in which your hands should reach for things on the night it breaks — every cell uses one endpoint from this article, and no cell may be skipped:

102 / 147
Animation
Animation · An incident triage loop
Animation · An incident triage loop
103 / 147

And the summary that goes with it: observing endpoints make the story tellable, diagnosing endpoints make it findable, and hardening decides who may even ask — miss any of the three and the four-capability picture in Section 0 stays a capability table instead of a triage workflow.

104 / 147
Section
105 / 147

Every row below can be copied and searched as it stands. Beginners stall in the same four places: the endpoint 404s before you have even begun, health says DOWN without naming who, /env leaks a secret, and the management port gets confused with the business port.

106 / 147
Table
Error text (fragment)Real cause30-second rescueRead more in
Whitelabel Error Page ... type=Not Found, status=404 on /actuator or /actuator/healthTwo candidates: spring-boot-starter-actuator was never added; or it is on the classpath but this endpoint is not in the management.endpoints.web.exposure.include whitelist (Boot ships only health and info)First `mvn dependency:tree \grep actuator to confirm the dependency is really there, then curl localhost:8080/actuator and read which keys the index page _links` actually lists — the index page is the one entry point that cannot lie to youSection 2 of this article · the exposure sandbox
/actuator/health returns {"status":"DOWN"}, but the body carries no componentsmanagement.endpoint.health.show-details defaults to never; the aggregation rule lets any single DOWN component drag the whole status down — so you see the verdict but not the culpritSet show-details: always temporarily (internal network only), or scope it with management.endpoint.health.group.<name>.show-details=always so just one group opens up; then read components to find the culpritSection 4 of this article · the health aggregation experiment
In GET /actuator/env, spring.datasource.password prints as ******, but your own app.oss.accessKey and app.wx.secret2 come out in plain textSanitising is a regex match on the property name (by default covering words like password, secret, key, token); pick a different name and nothing matches. Objects bound through @ConfigurationProperties are dumped whole in configpropsExtend the sanitised key list: management.endpoint.env.keys-to-sanitize=password,secret,key,token,certificate,accessKey (on Boot 3.x, management.endpoint.env.show-values=never plus keys-to-sanitize), and switch configprops off as well#21 Configuration in Full
management.server.port=9090 is configured, yet the Kubernetes probe still gets 404 or cannot connect at allThe probe URL still points at the business port 8080; once the management port changes, /actuator/** listens on 9090 onlyRewrite livenessProbe/readinessProbe from port: http(8080) to port: management(9090), and confirm the Service's network policy lets the cluster reach 9090Section 4 of this article · the probe animation
Both 8080 and 9090 answer locally, but the moment production Nginx forwards, /actuator is on the public internetNothing isolates them: either both ports share one origin, or the gateway forwards by path without ever blocking /actuatorDeny that path explicitly at the gateway; more robust, leave the management port out of the public port list of the Ingress/Service altogetherSection 10 · the hardening checklist
POST /actuator/loggers/com.example.order answers 415 Unsupported Media Type or 400No Content-Type: application/json, or the body field was written as level (the real field is configuredLevel)curl -X POST -H "Content-Type: application/json" -d '{"configuredLevel":"DEBUG"}'; sending {"configuredLevel":null} means "inherit from the parent again"Section 8
The startup log contains Cannot find factory with name "Prometheus", or there simply is no /actuator/prometheusOnly the actuator starter was added, without io.micrometer:micrometer-registry-prometheus; or the registry is there but the endpoint was never includedAdd the dependency → restart → put prometheus in include → curl :9090/actuator/prometheus should return # HELP jvm_... textSection 7
The dependency is clearly there, but /actuator/shutdown keeps 404ing, and ops says "we need to restart it remotely"By default the shutdown endpoint does not even create its bean; you need management.endpoint.shutdown.enabled=true and it in the include list — both switches or nothing is listeningThink it through first: in containers you restart with SIGTERM plus graceful shutdown, not with an HTTP call. If you really must open it, put authentication in front of itSection 4 · the probe animation · #46 going live
The heapdump download is several hundred MB, MAT cannot open it, or the machine goes OOMTaking a heap snapshot triggers a stop-the-world pause; on a large heap that single step can halt the service for ten-plus seconds — which is an incident by itselfRun it only inside a controlled window, or switch to jcmd <pid> GC.heap_dump; in production this endpoint stays disabledSection 10 · the hardening checklist
107 / 147

The rows above can all be rescued from the error text alone. One class of incident is less kind: the error does not name the culprit — the clue hides in the numbers inside the parentheses. This stack is an excerpt thrown on a probe thread during a real Pod-cycling incident — do not read the conclusion yet, point out the frame you believe is the culprit:

108 / 147
Triage
Error triageSQLTransientConnectionException: the probe could not get a connection
Probes wobble, Pods get cycled: the culprit hides on line 38 of a health indicator

At 01:40 the on-call phone erupts: instances of pay-api have been restarted twice within ten minutes, each time preceded by a probe timeout. A restart fixes it at once, and it recurs a while later. This is the stack excerpt the operators recovered from the killed instance — thrown on a probe thread. Do not read the conclusion yet, point out the frame you believe is the culprit.

java.sql.SQLTransientConnectionException: beeHikari - Connection is not available, request timed out after 1000ms (total=20, active=20, idle=0, waiting=31)
at com.zaxxer.hikari.pool.HikariPool.createTimeoutException(HikariPool.java:692)
at com.zaxxer.hikari.pool.HikariPool.getConnection(HikariPool.java:189)
at com.zaxxer.hikari.HikariDataSource.getConnection(HikariDataSource.java:100)
at org.springframework.jdbc.datasource.DataSourceUtils.fetchConnection(DataSourceUtils.java:160)
at org.springframework.jdbc.core.JdbcTemplate.queryForObject(JdbcTemplate.java:802)
at com.bee.pay.ReportHealthIndicator.health(ReportHealthIndicator.java:38)
at org.springframework.boot.actuate.health.HealthEndpoint.health(HealthEndpoint.java:53)
at org.springframework.boot.actuate.endpoint.web.servlet.AbstractWebMvcEndpointHandlerMapping$OperationHandler.handle(AbstractWebMvcEndpointHandlerMapping.java:367)
at org.springframework.web.servlet.DispatcherServlet.doDispatch(DispatcherServlet.java:1089)
Click the frame you blame — guessing is allowed
No pressure: guess the exception first, then which line actually made the call.
109 / 147
Quiz
Check yourselfProduction needs exactly two things from Actuator: Prometheus scrapes the metrics and Kubernetes runs its probes. All troubleshooting happens on the internal network. Which config fits best?
Pick one — you get feedback right away
110 / 147
Quiz
Check yourself/actuator/health reports status=DOWN and you want to know which part broke inside five minutes. What is the fastest correct move?
Pick one — you get feedback right away
111 / 147
Section
13. Hands-on practice
112 / 147
Section
Tier 1 · Follow along
113 / 147

Goal: put together a minimal monitoring project that runs entirely on your laptop and whose config is already a production skeleton — and see for yourself the 200-versus-404 difference a whitelist makes, the components block inside /health, and one line of curl switching a package to DEBUG.

114 / 147

Before you type anything, run the five labs in this kernel console, in order — each one is a conclusion from this article and every echo is computed by the kernel itself:

115 / 147
Console
116 / 147
Tip

five labs walk you through expose, tighten, observe and probe end to end; the three tiers below then move that flow into a project of your own.

117 / 147

Step one, pom.xml:

118 / 147
xml
<?xml version="1.0" encoding="UTF-8"?><project xmlns="http://maven.apache.org/POM/4.0.0"         xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"         xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 https://maven.apache.org/xsd/maven-4.0.0.xsd">    <modelVersion>4.0.0</modelVersion>    <parent>        <groupId>org.springframework.boot</groupId>        <artifactId>spring-boot-starter-parent</artifactId>        <version>3.3.4</version>        <relativePath/>    </parent>    <groupId>com.example</groupId>    <artifactId>monitor-lab</artifactId>    <version>0.0.1-SNAPSHOT</version>    <properties>        <java.version>17</java.version>    </properties>    <dependencies>        <dependency>            <groupId>org.springframework.boot</groupId>            <artifactId>spring-boot-starter-web</artifactId>        </dependency>        <!-- Actuator itself -->        <dependency>            <groupId>org.springframework.boot</groupId>            <artifactId>spring-boot-starter-actuator</artifactId>        </dependency>        <!-- the registry that emits the Prometheus text format -->        <dependency>            <groupId>io.micrometer</groupId>            <artifactId>micrometer-registry-prometheus</artifactId>        </dependency>        <!-- gives us a real db health indicator to aggregate -->        <dependency>            <groupId>org.springframework.boot</groupId>            <artifactId>spring-boot-starter-jdbc</artifactId>        </dependency>        <dependency>            <groupId>com.h2database</groupId>            <artifactId>h2</artifactId>            <scope>runtime</scope>        </dependency>    </dependencies>    <build>        <plugins>            <plugin>                <groupId>org.springframework.boot</groupId>                <artifactId>spring-boot-maven-plugin</artifactId>                <executions>                    <execution>                        <goals>                            <goal>build-info</goal>   <!-- so /actuator/info has something to say -->                        </goals>                    </execution>                </executions>            </plugin>        </plugins>    </build></project>
119 / 147

Step two, src/main/resources/application.yml:

120 / 147
yaml
server:  port: 8080                      # business port: open to the internetspring:  application:    name: monitor-lab  datasource:    url: jdbc:h2:mem:lab;DB_CLOSE_DELAY=-1    driver-class-name: org.h2.Drivermanagement:  server:    port: 9090                    # management port: reachable from the internal network only  endpoints:    web:      base-path: /actuator      exposure:        include: health,info,metrics,prometheus   # a whitelist, never a *  endpoint:    health:      show-details: when_authorized      group:        liveness:          include: ping        readiness:          include: db,diskSpace  info:    env:      enabled: truelogging:  level:    com.example.monitorlab: INFO
121 / 147

Step three, one custom health indicator and one order-count instrument (package com.example.monitorlab):

122 / 147
java
package com.example.monitorlab;import java.util.concurrent.atomic.AtomicLong;import org.springframework.boot.actuate.health.Health;import org.springframework.boot.actuate.health.HealthIndicator;import org.springframework.stereotype.Component;@Component("inventoryCache")public class InventoryCacheHealthIndicator implements HealthIndicator {    private final AtomicLong lastSyncOkAt = new AtomicLong(System.currentTimeMillis());    @Override    public Health health() {        long staleMs = System.currentTimeMillis() - lastSyncOkAt.get();        if (staleMs > 60_000) {            return Health.down().withDetail("staleMs", staleMs).withDetail("reason", "inventory cache has not been synced for over 60s").build();        }        return Health.up().withDetail("staleMs", staleMs).build();    }    public void markSynced() {        lastSyncOkAt.set(System.currentTimeMillis());    }}
123 / 147

Step four, start the app and run these curls in order (watch which ones are 8080 and which are 9090):

124 / 147
bash
# ① the index page: confirm the exposure surface is exactly four endpointscurl -s http://localhost:9090/actuator# ② the health summary (unauthorized here, so no components)curl -s http://localhost:9090/actuator/health# ③ the grouped probe addressescurl -s http://localhost:9090/actuator/health/livenesscurl -s http://localhost:9090/actuator/health/readiness# ④ a management endpoint on the business port -> should be 404curl -i -s http://localhost:8080/actuator/health | head -1# ⑤ a sensitive endpoint should be 404curl -i -s http://localhost:9090/actuator/env | head -1# ⑥ change the log level at runtime, then observe immediatelycurl -s -X POST http://localhost:9090/actuator/loggers/com.example.monitorlab \     -H "Content-Type: application/json" -d '{"configuredLevel":"DEBUG"}'curl -s http://localhost:9090/actuator/loggers/com.example.monitorlabcurl -s -X POST http://localhost:9090/actuator/loggers/com.example.monitorlab \     -H "Content-Type: application/json" -d '{"configuredLevel":null}'# ⑦ drilling into metricscurl -s http://localhost:9090/actuator/metricscurl -s "http://localhost:9090/actuator/metrics/hikaricp.connections.pending"curl -s http://localhost:9090/actuator/prometheus | grep -E "^jvm_memory_used_bytes|^http_server_requests" | head -3
125 / 147

Expected output (the default shape of ② when you are unauthorized, the two groups from ③, and the status codes of ④ and ⑤):

126 / 147
json
{ "status": "UP" }
127 / 147
json
{ "status": "UP", "group": "liveness" }
128 / 147
json
{ "status": "UP", "group": "readiness", "components": { "db": { "status": "UP" }, "diskSpace": { "status": "UP" } } }
129 / 147
text
HTTP/1.1 404HTTP/1.1 404
130 / 147

Acceptance checklist: ① you can explain why there is no db inside liveness; ② change show-details to always, call ② again, and the custom inventoryCache indicator now shows up in components; ③ you can say what causes each of the two 404s in ④ and ⑤ (the first is port isolation, the second is the whitelist).

131 / 147
Section
Tier 2 · Change one thing
132 / 147

One line at a time — and the conclusion flips:

133 / 147
  1. Replace include with "*" and restart. What you will observe: the _links object from curl :9090/actuator suddenly has a dozen or more keys; inside env the value of spring.datasource.password is masked, but the app.my.access-key-alias=plain-secret you just added prints verbatim — that is the actual boundary of the sanitising rules. Put the config back when you are done.
  2. Comment out the management.server.port line. What you will observe: ④ turns into 200, because the management endpoints are back on the business port — and if your gateway then forwards by path, /actuator is public. This one is what makes "port isolation" click.
  3. Add db to the liveness group, then break startup with a wrong URL (jdbc:h2:mem:lab;INIT=RUNSCRIPT FROM 'nonexistent.sql'). What you will observe: /actuator/health/liveness goes straight to DOWN — a rehearsal of "Kubernetes repeatedly restarts an instance that only needed to be pulled out of rotation".
  4. Change the threshold in InventoryCacheHealthIndicator from 60_000 to -1. What you will observe: the top-level status turns DOWN and the HTTP code becomes 503, but by default you still cannot see why — so you are pushed into changing show-details, which is exactly the second row of Section 12.
  5. After step ⑥, hit an endpoint a dozen or more times. What you will observe: DEBUG logs pour out; set the level back to null (inherit from the parent) and it goes quiet at once. The value and the risk of "debugging without a restart", in one minute.
134 / 147

Hint: after item 3, go back to the probe animation in Section 4 — the two timelines should line up exactly.

135 / 147
Section
Tier 3 · Build one yourself
136 / 147

Build yourself an "on-call self-service panel": one browser screen that shows the key conclusions, with no Grafana installed.

137 / 147

Requirements:

138 / 147
  • A custom endpoint @Endpoint(id = "opsboard") whose @ReadOperation returns one aggregated JSON: JVM heap usage, the GC count delta, pending connections in the pool, the 5xx ratio over the last 5 minutes, the row count of t_order, the running Git commit, and uptime
  • All of it comes from an injected MeterRegistry, DataSource and HealthContributor; no new database tables allowed
  • A @TransactionalEventListener(AFTER_COMMIT) or an interceptor records the business success rate, as one tile of the panel
  • The panel endpoint is not exposed by default; it can only be opened through the include list of an internal profile, and your documentation says plainly how to open it and when to close it again
  • A README.md with three curls that perform the standard sequence: "spot the anomaly → locate the component → raise the log level temporarily"
139 / 147

Acceptance checklist: ① every field in the output of curl :9090/actuator/opsboard is traceable to a metric name you can point at; ② make the database unavailable on purpose (a wrong URL, or cut the network with Testcontainers) — the matching tile must turn red, /health/readiness must be DOWN and /health/liveness must still be UP; ③ after 60 seconds of load, the 5xx ratio on the panel is within 5% of the figure you compute by hand from http_server_requests_seconds_count{status="500"}; ④ with the internal profile switched off, the endpoint returns 404.

140 / 147
Section
14. Self-check
141 / 147
Self-check

which two endpoints does Actuator expose over the Web by default, and why that particular pair?

142 / 147
Self-check

how is the top-level status of /actuator/health computed? With the three values of show-details, name exactly what each one hides and what it reveals.

143 / 147
Self-check

what question does liveness answer, and what question does readiness answer? What specific incident does db in the wrong group cause?

144 / 147
Self-check

what is the semantic difference between Counter, Gauge and Timer? Which one fits "the latency distribution of payment callbacks", and why?

145 / 147
Self-check

where do the sanitising rules on /actuator/env stop working? Give two hardening measures that follow directly from that.

146 / 147
Mnemonic

a whitelist opens the doors, a separate port closes them again, details only on authorization; liveness judges only itself, readiness asks about the dependencies; metrics come in, dashboards come out — and never leave the OBD port on the roadside.

147 / 147
Summary

if you can answer the five questions above without scrolling back, you have what it takes to put Actuator on a live service. The shape never changes — instrument once with Micrometer, publish four endpoints, move them onto their own port, split liveness from readiness, let Prometheus and Grafana do the watching, and keep the table in Section 12 where you can paste straight from it at 3 a.m. The one line to remember: observability is only worth having if the entry point was designed the day before the incident, not improvised during it.