Docker Deployment: Image Layers and Compose
Plain language first. A container is not "a miniature computer" — it is a single Java process shut into a small room of its own: it has its own view of the filesystem, its own network card, its own memory allowance, but it shares one Linux kernel with the host. Because it virtualizes no hardware and installs no operating system, it starts in seconds and begins at a few dozen MB. An image is that packaged set of files and configuration — the same image expands to exactly the same thing on any machine with Docker installed, which is how it cures "it works on my machine".
Six terms, one line each (used throughout):
- Image: a template built by stacking read-only layers; built once, reused everywhere; it is a deliverable, not something running
- Container: a running instance of an image — the read-only layers plus one writable layer on top; delete the container and the writable layer is gone
- Layer: one
COPY/RUNin a Dockerfile; a layer only cares whether "I and everything below me changed" - Build cache: if a layer and everything below it is unchanged, it is not re-executed — this is the entire art of writing a Dockerfile
- cgroup: the Linux kernel's quota switch;
docker run -m 512mlimits memory through it, and a container-aware JVM reads it - namespace: Linux's view-isolation mechanism, which is why a process inside a container cannot see the host's other processes or network
an image is a standardized shipping container, and a container is the truck body loaded with that box. The steel box has ISO-standard corner castings and fixed dimensions, so any crane in any port can lift it; the truck adds its own route, its own driver's paperwork (env vars), and a small open box at the back for today's loose parcels (the writable layer) — unload the truck and today's parcels are gone. Two consequences follow immediately. The cargo must be packed once, correctly, to be accepted anywhere — that is exactly why COPY pom.xml comes before COPY src: no port inspector unpacks a container to find one crate shifted. And anything that must outlive the truck needs a separate warehouse (a volume), because nobody stores today's parcels inside the returnable steel box.

That tree splits the whole article into four gates: separate image from container → write a Dockerfile that does not waste the cache → how the JVM and the container behave once it runs → orchestration and graceful shutdown. Every later section sits somewhere on that tree; when you get stuck, come back and see which gate you are standing on.
After this article you should be able to answer three questions:
- Why does editing one line of business code make my image build re-download dependencies? In what order should Dockerfile instructions go?
- The container log says
Exit code (137)/Reason: OOMKilled, but there is not a singleOutOfMemoryErrorin the Java log — who killed whom? - I configured graceful shutdown, yet
docker stopstill severs requests. Which step swallowed the signal?
Every backend engineer has heard: "it works on my machine." The real cause is environment drift — the dev box runs JDK 21 with MySQL 8.0, staging runs JDK 17 with MySQL 5.7, and production is something else again. Containerizing means packaging "app + dependencies + runtime" into one immutable artifact so all three environments run the same thing.
| Aspect | Bare metal | Virtual machine | Container (Docker) |
|---|---|---|---|
| Isolation | None | Hardware-level virtualization | Process-level (namespaces + cgroups) |
| Startup | — | Minutes | Seconds to milliseconds |
| Overhead | High (one per host) | High (a full OS each) | Low (shares the host kernel) |
| Image size | — | GBs | MBs to hundreds of MB |
| Consistency | Poor, manual alignment | Decent | Excellent, the image is the environment |
| Elasticity | Slow | Slow | Fast, start and serve |
Containers are light because they virtualize no hardware and install no full OS — they share the host kernel, using namespaces for view isolation and cgroups for limits. That brings fast startup and high density at the cost of weaker isolation than a VM (a kernel vulnerability can be exploited across containers).
containers are not a silver bullet. They solve "environment consistency, dependency isolation, elasticity"; they do nothing for "is the application itself written correctly". Do not expect a Docker image to fix a memory leak.
Separate the four most confused terms:
| Concept | In one sentence | Key point |
|---|---|---|
| Image | A read-only layered template | Build artifact, immutable |
| Container | A running instance of an image | Adds a writable layer on top |
| Registry | Where images are stored | Docker Hub, Harbor, cloud registries |
| Volume | Persistent storage outside the union FS | Data survives container removal |
An image is not one blob but a stack of read-only layers, one per COPY, RUN, ADD. During a build, Docker compares the cache layer by layer: if a layer and everything below it are unchanged, that layer is reused and not re-executed.
This explains a classic optimization: copy pom.xml first.
# Bad: copy all sources first; a one-line change invalidates the dependency layerCOPY . .RUN mvn dependency:go-offlineRUN mvn package -DskipTests# Good: copy only the pom to fetch dependencies; source edits keep the cacheCOPY pom.xml .RUN mvn dependency:go-offline # deps unchanged → this layer always hitsCOPY src ./src # source changes → only later layers invalidateRUN mvn package -DskipTests- In the bad version,
COPY . .changes whenever source changes, so every later layer (dependency download, packaging) reruns - The good version splits "rarely-changing dependency manifest" from "frequently-changing source": a code edit only reruns packaging, cutting build time from minutes to seconds

layers depend upward, so a change invalidates only what is above it. The rule for Dockerfile instruction order is therefore stable things high, volatile things low.
This rule is easy to remember backwards when you only read about it; watching one cache decision once fixes it for good. The experiment below uses the default layer argument to spell out "which layer each instruction creates", "what the cache compares", "what CACHED in the build output means" — watch frame 3 (why COPY pom.xml goes first) and frame 6 (how much layered extraction saves):
The cache decision is not obvious from a picture alone; what people forget is what makes up a layer's fingerprint. The single-stepper below runs a five-line builder stage frame by frame — at each step, watch which fingerprint is recomputed. Step ③ is the heart of it:
# the five-line builder stageCOPY pom.xml .RUN mvn -B dependency:go-offlineCOPY src ./srcRUN mvn -B package -DskipTests| changes | none |
| cache | — |
Keep that fingerprint table in mind for 2.2: if COPY src ./src gets polluted with target/-type junk, every layer changes and the cache never hits again.
The build context is packaged as a whole and handed to the daemon first, so target/, .git/, node_modules/ are uploaded once even though they never make it into the image. Put a .dockerignore at the root:
target/.git/.idea/*.imlsrc/test/- The builder stage's
~/.m2(local Maven repository) is re-downloaded on every rebuild by default; BuildKit can mount and reuse it:RUN --mount=type=cache,target=/root/.m2 mvn -B package - Leaving
.gitinside the builder layer is not just slow, it can ship repository history into the deliverable - Forgetting to exclude
target/causes the classic chain reaction:COPY . .changes every time → every layer above it invalidates → exactly cancelling the 2.1 optimization
A naive version that runs but is not yet production-grade:
# 1. Base image: JRE only, much smaller than a JDKFROM eclipse-temurin:21-jre# 2. Working directory: the relative root for later COPY / RUN / CMDWORKDIR /app# 3. Copy the built jar into the imageCOPY target/app.jar app.jar# 4. Timezone, so container logs are not 8 hours offENV TZ=Asia/Shanghai# 5. Declare the listening port (documentation only, does not publish it)EXPOSE 8080# 6. Startup command: exec form so PID 1 receives SIGTERMENTRYPOINT ["java", "-XX:MaxRAMPercentage=75.0", "-jar", "app.jar"]FROM: everything starts from a base image. Usejrerather thanjdk— runtime only, saving one or two hundred MBWORKDIR: sets the working directory; likecdbut safer — it creates the directory if missingCOPY: copies host files into the image; in-build paths are relative to the build contextENV: sets environment variables;TZaffects log timestamps, and the default UTC often makes them look wrongEXPOSE: metadata only, telling users the container listens on 8080; it does not open the port (publishing is-p)ENTRYPOINT: use the exec array form rather than the shell form so the Java process is PID 1 and receives the SIGTERM thatdocker stopsends — that is what makes graceful shutdown work
The Dockerfile above assumes the host already ran mvn package. But CI machines may lack Maven, and even if they have it, build and runtime environments should not be mixed. The answer is a multi-stage build — two FROM sections in one Dockerfile, the first building, the second taking only the artifact.
# ---------- Stage 1: build (JDK + Maven, thrown away after) ----------FROM maven:3.9-eclipse-temurin-21 AS builderWORKDIR /build# Copy only pom.xml first; the dependency layer always hits when deps are unchangedCOPY pom.xml .RUN mvn -B dependency:go-offline# Then copy sources; a source edit only invalidates later layersCOPY src ./srcRUN mvn -B clean package -DskipTests# ---------- Stage 2: runtime (JRE only, no Maven, no source) ----------FROM eclipse-temurin:21-jreWORKDIR /appCOPY --from=builder /build/target/*.jar app.jarENV TZ=Asia/ShanghaiEXPOSE 8080ENTRYPOINT ["java", "-XX:MaxRAMPercentage=75.0", "-jar", "app.jar"]AS buildernames the stage forCOPY --from=builder- The final image contains stage two only — Maven, the JDK and the source are all discarded, which is where the size drop comes from
COPY --from=buildercarries over just the artifact (the jar)
| Approach | Base images | Rough size | Note |
|---|---|---|---|
| JDK + source image | maven + jdk | ~700MB+ | Build tooling shipped to production — wasteful and unsafe |
| Multi-stage build | jre only | ~200MB | Runtime only, recommended |
Multi-stage builds fix the size, but a cache problem remains: every code change alters app.jar, invalidating the COPY app.jar layer and forcing the dependency layer to rebuild too. Spring Boot's layered extraction splits the fat jar into separate layers that can be copied independently:
<plugin> <groupId>org.springframework.boot</groupId> <artifactId>spring-boot-maven-plugin</artifactId> <configuration> <layers> <enabled>true</enabled> </layers> </configuration></plugin># Build stage: split the jar into layered directoriesRUN java -Djarmode=layertools -jar /build/target/app.jar extract# Runtime stage: copy layer by layer — unchanged deps hit the cacheCOPY --from=builder /build/dependencies/ ./COPY --from=builder /build/spring-boot-loader/ ./COPY --from=builder /build/snapshot-dependencies/ ./COPY --from=builder /build/application/ ./ENTRYPOINT ["java", "org.springframework.boot.loader.launch.JarLauncher"]layertools extractsplits the jar intodependencies/(stable),application/(volatile) and more- Stable layers high, application layer low: when dependencies are unchanged only the
application/layer invalidates, and rebuilds are nearly instant - The entry point becomes
JarLauncher(the executable-jar launcher from the previous article) so it loads from the layered directories
Sections one to four are really three stages of one pipeline: first build a jar that runs, then cut that jar into layers inside an image, finally start the container and publish the port. Beginners break exactly at "which one holds which", so walk the whole road in order:

Step six deserves a note: -p 8080:8080 and /actuator/health/readiness are two different responsibilities — the first is Docker's job (mapping a host port into the container), the second is the application's own (answering "can I take traffic right now"). This article nails the Docker half; the probe half is Article 42.
And the premise of steps one and two is that you really know what the jar you baked into the image looks like. The previous article opened up the executable jar: Main-Class is JarLauncher rather than your startup class, and the dependencies sit in BOOT-INF/lib read by a custom class loader. Confirm that with this experiment before containerizing — especially the war branch: once you deploy to an external Tomcat, COPY target/*.jar simply finds nothing.
| Command | Effect | Common flags |
|---|---|---|
docker build -t app:1.0 . | Build an image | -t tags it; . is the context |
docker run -d app:1.0 | Run a container | -d detached, -p port, -e env, -v mount |
docker ps -a | List containers | -a includes stopped |
docker logs -f <id> | Follow container logs | -f follows |
docker exec -it <id> sh | Shell into the container | -it allocates a terminal |
docker stop <id> | Stop (SIGTERM, graceful) | docker kill is the hard version |
docker rm <id> | Remove a container | -f forces it |
docker images / docker rmi | List / remove images | — |
docker system prune -a | Clean unused images and cache | Use with care; deletes unused resources |
docker stop waits only 10 seconds by default before sending SIGKILL. If your graceful shutdown needs 30 seconds, always docker stop -t 30 <id>; otherwise shutdown is cut short and you are back to severed requests.
Containers run on the bridge network by default, and each container has its own IP — inside a container localhost means the container itself, not the host. Within a custom network, containers reach each other by service name:
# Create a custom networkdocker network create bee-net# Start MySQL named mysql on that networkdocker run -d --name mysql --network bee-net \ -e MYSQL_ROOT_PASSWORD=root -e MYSQL_DATABASE=bee \ -v mysql-data:/var/lib/mysql \ mysql:8.0# Start the app, connecting by service name mysql (not localhost)docker run -d --name bee-app --network bee-net \ -p 8080:8080 \ -e SPRING_PROFILES_ACTIVE=prod \ -e DB_HOST=mysql \ -v "$(pwd)/logs:/app/logs" \ bee-app:1.0.0--network bee-netputs both containers on one network, reachable by name-e DB_HOST=mysql: the app writesmysqlin its JDBC URL and Docker's embedded DNS resolves it-v mysql-data:/var/lib/mysql: a named volume persists data even if the container is removed-v "$(pwd)/logs:/app/logs": a bind mount maps the container log directory to the host for inspection
Hand-typing docker run quickly spirals out of control. Compose describes the whole environment declaratively, and docker compose up brings it all up:
services: mysql: image: mysql:8.0 environment: MYSQL_ROOT_PASSWORD: ${MYSQL_ROOT_PASSWORD:-root} MYSQL_DATABASE: bee command: --character-set-server=utf8mb4 --collation-server=utf8mb4_general_ci volumes: - mysql-data:/var/lib/mysql - ./sql:/docker-entrypoint-initdb.d # init scripts run on first start ports: - "3306:3306" healthcheck: test: ["CMD", "mysqladmin", "ping", "-h", "localhost"] interval: 10s timeout: 5s retries: 10 redis: image: redis:7-alpine command: redis-server --appendonly yes volumes: - redis-data:/data healthcheck: test: ["CMD", "redis-cli", "ping"] interval: 10s retries: 5 app: build: . # build from the Dockerfile here image: bee-app:1.0.0 depends_on: mysql: condition: service_healthy # wait until MySQL is healthy redis: condition: service_healthy environment: TZ: Asia/Shanghai SPRING_PROFILES_ACTIVE: prod DB_HOST: mysql DB_PORT: 3306 DB_NAME: bee REDIS_HOST: redis ports: - "8080:8080"volumes: mysql-data: redis-data:Section by section:
- Service dependencies and order:
depends_onalone guarantees only "start order", not "readiness". Addingcondition: service_healthywaits for the health check — the key to avoiding "the app starts before the database is ready" - Health checks:
healthcheck.testis the liveness command —mysqladmin pingfor MySQL,redis-cli pingfor Redis.service_healthyindepends_onrelies on its result - Injecting DB config via env vars: the app never hard-codes a connection string; it reads
DB_HOST/DB_NAMEinjected by Compose.${MYSQL_ROOT_PASSWORD:-root}means "take the env var, default to root" - Volumes:
mysql-dataandredis-dataare named volumes declared at the top level, so data survives container recreation - Build vs image:
build: .lets Compose build on the spot;image:names it for reuse

Compose's healthcheck tells the orchestrator whether a container can take work — and the Spring Boot app inside has to answer that question itself. Article 42 covered the Actuator health endpoint; here it is through the container lens: a health check is not a ping, it is the worst result after several contributors vote:
- DB up, Redis up →
UP; anything down → the whole chain reportsDOWN, and K8s readiness pulls you out of the load balancer - So never put a business query in the health check: one slow SQL turns "the process is alive" into "the instance is unavailable"
- To see which contributor drags the status down, turn on
management.endpoint.health.show-details=always
Containerizing is more than a Dockerfile; the app needs two adjustments too.
Never hard-code localhost in application.yml; use placeholders with defaults:
spring: datasource: url: jdbc:mysql://${DB_HOST:localhost}:${DB_PORT:3306}/${DB_NAME:bee}?useSSL=false&serverTimezone=Asia/Shanghai username: ${DB_USER:root} password: ${DB_PASSWORD:root} data: redis: host: ${REDIS_HOST:localhost} port: ${REDIS_PORT:6379}${DB_HOST:localhost}means "use env varDB_HOST, elselocalhost"- Locally no env vars are set, so defaults connect to the local machine; in a container Compose injects
DB_HOST=mysql - One codebase, one jar, adapted by environment variables — exactly the third factor of twelve-factor apps: store config in the environment
One container principle: the process writes to stdout; where logs go is the platform's business. So production config should log to the console only, not to files:
logging: file: name: "" # disable file output, keep console only pattern: console: "%d{HH:mm:ss.SSS} %-5level [%X{traceId}] %logger{36} - %msg%n"- Writing files inside a container has a fatal flaw: the file dies with the container; mounting the host needs volumes and multi-instance aggregation gets hard
- Writing to stdout lets
docker logsshow it directly and lets the log driver or ELK/Fluentd collect it uniformly - This connects nicely to the logging article — JSON to stdout for machines, human-readable for development
This is the spot in containerizing that "is fine in normal times and explodes right after a scale-out". When the JVM decides how large a heap to open, it looks by default at the physical memory of the whole machine, not at the quota handed to it. A process running inside a 512m container that sizes its heap as if the host had 32GB will race up to the cgroup ceiling and get erased by the kernel. Since Java 8u191 UseContainerSupport is on by default and the JVM reads the cgroup limit files (v1 memory.limit_in_bytes, v2 memory.max) — but it reads the limit, not "your heap".
The correct form is to declare a percentage rather than hard-code -Xmx:
FROM eclipse-temurin:21-jre-jammyWORKDIR /appCOPY --from=builder /build/target/app.jar app.jar# The heap follows 75% of the container limit; metaspace and thread stacks get their own capsENV JAVA_TOOL_OPTIONS="-XX:MaxRAMPercentage=75.0 \ -XX:InitialRAMPercentage=75.0 \ -XX:MaxMetaspaceSize=256m \ -Xss512k"ENTRYPOINT ["java", "-jar", "app.jar"]JAVA_TOOL_OPTIONSbeats baking it intoENTRYPOINT: ops can override it with-ewithout rebuilding the image- The remaining 25% is not waste — it is the budget for off-heap memory: Metaspace (class metadata), roughly 1MB of stack per thread, CodeCache (the machine code JIT produces) and
DirectByteBuffer(direct memory used by NIO and Netty) - One addition people forget: Tomcat defaults to 200 worker threads, so the stacks alone approach 200MB. That is why high-concurrency containers often push
-Xssdown to512k - To see how much memory the JVM actually thinks it has, just ask it inside the container:
docker exec bee-app java -XX:+PrintFlagsFinal -version | grep -E 'MaxHeapSize|InitialHeapSize'Two failure modes must be kept apart, because their investigation directions are opposite:
OutOfMemoryError: Java heap space | OOMKilled / Exit code (137) | |
|---|---|---|
| Who raised it | The JVM itself (the heap is full) | cgroup / kernel / K8s (total usage over the limit) |
| Is there a Java-side exception | Yes, a full stack trace in the log | No — the log just stops mid-line |
| Can you keep the evidence | Yes, with -XX:+HeapDumpOnOutOfMemoryError for an hprof | Nothing survives |
| What to tune | Hunt the leak, adjust the heap ratio | Do the total sum: heap plus off-heap |
The boundary between the two failure modes comes down to one number: what percentage of the limit the heap takes. Drag -XX:MaxRAMPercentage from 25 to 100 inside a 512MB container and watch heap, off-heap and the total bill split apart at every stop — note the far end, where it becomes equivalent to hard-coding -Xmx512m:
- 384MB heap; Metaspace + 200 thread stacks + CodeCache ≈ 130MB, leaving roughly 40MB of headroom
- Starts fine and GC frequency is acceptable — this is where the 75% in the example above comes from
- Precondition: cap Metaspace and -Xss so off-heap cannot grow without bound
an airline luggage weight limit. "This case may weigh at most 23 kg" is the weight of the whole case — clothes, the lock, the wheels, the power bank inside all count. Pack the clothes up to 23 kg (set -Xmx equal to the container limit) and the counter will refuse it before even opening the lid (OOMKilled); while "the clothes no longer fit in the case" (an in-heap OutOfMemoryError) is something you discover at home, and you can still spread it out to see which item is the culprit. The first is an external execution, the second is an internal error — symptoms, fixes and whether evidence survives are all different.

The left column explains why "just give it more memory" so often only hides the problem: the root cause may be Netty direct memory that was never released, and off-heap memory is not governed by -Xmx at all. This experiment walks "how much does the JVM read → how is the default heap computed → what each failure looks like" in one pass:
| Item | How | Benefit |
|---|---|---|
| Base image | Choose jre-slim / alpine variants | Size from hundreds of MB to ~200MB |
| Multi-stage build | Separate build and runtime | Drop Maven/JDK/source |
| Run as non-root | USER app | Reduces container-escape risk |
.dockerignore | Exclude target/, .git/, node_modules/ | Smaller context, faster builds |
| Combine RUN layers | RUN a && b && c | Fewer layers |
| Layered COPY | layertools + per-layer copy | Better cache hits |
| Pin base image versions | Use :21-jre not :latest | Reproducible builds |
alpine images are small but use musl libc rather than glibc, which some components relying on native libraries (certain JDK features, JNI libraries) do not support. To be safe, prefer a slim variant like eclipse-temurin:21-jre-jammy over blindly going alpine.
Checklist read; now turn it into your own project's behavior. The generator below turns "which blocks a production-grade Dockerfile needs" into checkboxes. Tick only multi-stage + layered first for the minimal skeleton, then add non-root, the memory ratio, the health check and the stop signal one by one, verifying each against the benefits in the table above:
# ---------- 构建阶段:要完整 JDK 与 Maven ----------
FROM maven:17-eclipse-temurin AS build
WORKDIR /src
COPY pom.xml .
RUN mvn -B dependency:go-offline # 先只拷 pom,依赖层可被缓存
COPY src ./src
RUN mvn -B -DskipTests package \
&& java -Djarmode=layertools -jar target/*.jar extract --destination /app
# ---------- 运行阶段:只要 JRE ----------
FROM eclipse-temurin:17-jre
WORKDIR /app
# 变化频率从低到高排列,改代码不会让依赖层缓存失效
COPY --from=build /app/dependencies/ ./
COPY --from=build /app/spring-boot-loader/ ./
COPY --from=build /app/snapshot-jar/ ./
COPY --from=build /app/application/ ./
ENV JAVA_OPTS="-XX:MaxRAMPercentage=75.0 -XX:+ExitOnOutOfMemoryError"
ENV SPRING_PROFILES_ACTIVE="prod"
EXPOSE 8080 9090
# exec 形式:java 就是 PID 1,SIGTERM 能送达(shell 形式做不到)
ENTRYPOINT ["java", "-jar", "app.jar"]
# docker build -t demo-service:0.0.1 .
# docker run --rm -p 8080:8080 -m 512m demo-service:0.0.1healthcheck and stopsignal are the two blocks that "look like nothing until the day they matter" — the first decides whether the orchestrator notices a wedged container, the second decides whether your app gets a chance to drain on docker stop. Read them together with the probe division of labour in Section 13.
The most common misconception: that localhost:3306 inside a container reaches the host's MySQL. Wrong — a container's network namespace is separate, so localhost means the container itself. Either use the host IP (172.17.0.1 on Linux), or — better — put the database in Compose and use service names.
EXPOSE 8080 only declares; it does not make the port reachable from outside. What publishes a port is docker run -p 8080:8080 (or ports in Compose). Forget -p and the container runs fine while your browser cannot connect.
Containers default to UTC, so log timestamps and database times are 8 hours behind Beijing. Fix it with ENV TZ=Asia/Shanghai in the Dockerfile, or -e TZ=Asia/Shanghai at run time, plus serverTimezone=Asia/Shanghai in the JDBC URL.
Some images (especially slim ones) have a small entropy pool, so SecureRandom blocks reading /dev/random — the app appears stuck at startup while Tomcat never binds the port. Linux kernel 5.6+ no longer blocks on /dev/random; on older kernels mount /dev/urandom or pass -Djava.security.egd=file:/dev/./urandom.
Running an x86-only image on an Apple Silicon (M-series) Mac, or pushing a locally built arm64 image to an amd64 server, produces:
standard_init_linux.go:228: exec user process caused: exec format errorThis has nothing to do with Java — Docker failed while trying to execute the ENTRYPOINT binary: the image carries a linux/arm64 java, the host kernel is linux/amd64, the instruction sets differ and the loader refuses. Three self-rescue routes:
- Declare the platform at build time:
docker buildx build --platform linux/amd64 -t bee-app:1.0 . - Declare it at run time (only on machines with qemu emulation — slow and unstable):
docker run --platform linux/amd64 bee-app:1.0 - Produce both architectures in CI:
docker buildx bake, or multiple--platform linux/amd64,linux/arm64values plus a manifest push
One command is enough to investigate — read the image's own architecture rather than your machine's:
docker inspect --format '{{.Os}}/{{.Architecture}}' bee-app:1.0Trap: docker pull fetches the image matching your machine's architecture by default, so "it runs locally but CI reports exec format error" is often not a code problem at all — the two architectures published under one tag simply differ. Pin the platform for production images instead of relying on the default.
Article 36 covered the six actions of a graceful shutdown; the premise is that the JVM actually received that signal. In a container the easiest link to break is exactly this one.

The fork sits at frames 2 and 3:
- exec (JSON array) form:
ENTRYPOINT ["java","-jar","app.jar"]—javaitself is PID 1, SIGTERM reaches the JVM directly and the shutdown hooks run as usual - shell form:
ENTRYPOINT java -jar app.jar— what actually starts is/bin/sh -c "java -jar app.jar", so the shell is what occupies slot 1;/bin/shdoes not forward received signals by default, so none of your graceful-shutdown configuration executes and the process is hard-killed by SIGKILL ten seconds later
The symptom is remarkably fixed: not one shutdown-related line in the log, yet the container did stop cleanly. Three fixes — use the exec form; or keep the shell form and write ENTRYPOINT exec java -jar app.jar; or let tini / dumb-init act as PID 1 to manage signals and reap zombies.
Align both sides too: if Spring's server.shutdown=graceful plus spring.lifecycle.timeout-per-shutdown-phase waits 30 seconds while Docker's grace period is the 10-second default, the last 20 seconds of requests are severed anyway.
server: shutdown: gracefulspring: lifecycle: timeout-per-shutdown-phase: 20sThe verification method is crude but the most effective: run docker stop -t 30 bee-app in one shell and docker logs -f bee-app in another; you only pass if you can see Closing org.springframework.... This experiment walks the whole signal path — focus on the fork at frames 2 and 3:
Container environment variables and conditional wiring are the same idea: one codebase assembling different results per environment. This demo runs the IoC container's condition switches — just like toggling container env vars.
- Flip switches like
@ConditionalOnPropertyand watch which beans are registered - Container env vars (such as
DB_HOST) bind via Spring's@ConfigurationPropertiesand ultimately affect wiring - Grasp this and you grasp the mechanism behind "one image, many environments"
With images, containers and the app itself covered, one layer of verification is left: the command line. This console is wired to the same Java kernel in your browser — start with whoami, then type the lab experiments from this section one by one; every line comes back computed by the kernel itself:
type lab dockerimg mem and lab dockerimg signal back to back — the first answers "how much memory does the JVM see", the second "did SIGTERM reach it"; the main conclusions of 8.3 and 10.6 both live in that output.
All of the above finally lands on one real rolling release: the new container takes over while the old one retires, and users must not feel the hiccup in between. Retiring an old container is not "kill the process" but three things happening in order — stop receiving traffic, finish the in-flight requests, then exit. Docker only owns half of the first step; the rest belongs to the application and the orchestrator.
The two Actuator probe endpoints from Article 42 have completely different semantics, and mixing them up causes "the database wobbles and the whole fleet restarts":
/actuator/health/livenessanswers "does this process still have hope" — it should only judge the process itself; misconfiguring it causes restart loops/actuator/health/readinessanswers "can I take traffic now" — when a dependency is unhealthy this is what should pull you out of the load balancer, not kill the process
- Align the grace period: K8s's
terminationGracePeriodSeconds(30s by default) must exceed Spring'sspring.lifecycle.timeout-per-shutdown-phase, otherwise you are SIGKILLed halfway through waiting - The preStop sleep is not mysticism: endpoint updates propagate with delay, so an old instance may still receive new requests the instant SIGTERM arrives — the convention is to
sleep 2~5before starting to wind down docker stop -tis the same story: locally or under Compose there is no preStop, and the only knob is-t; the trap in Section 5 is exactly this
retiring a container is like a cinema closing for the night. Turning off the lights is "no longer selling tickets" (readiness goes red, stop accepting new requests); the audience already inside must finish this scene before leaving (in-flight requests drain); and the announcement only works if it actually reaches the usher on duty (the signal has to reach PID 1). Ten minutes after closing and the security still clears the hall (SIGKILL) — and the popcorn is on the floor.
Every "error text" below can be pasted into a search engine as-is — no paraphrasing, no abbreviating. The three places beginners get stuck most: the entrypoint script that will not run, the port that will not bind, and the memory limit nobody configured correctly.
| Error text (fragment) | Real cause | 30-second rescue | Read more in | |
|---|---|---|---|---|
Exit code (137), with not one OutOfMemoryError in the Java log | 137 = 128 + 9, i.e. killed by SIGKILL; cgroup killed the process because total usage (heap + metaspace + thread stacks + direct memory) exceeded the limit — not because the heap was too small | First confirm what limit the JVM sees: `docker exec <id> java -XX:+PrintFlagsFinal -version \ | grep MaxHeapSize; swap -Xmx for -XX:MaxRAMPercentage=75.0, then use -XX:NativeMemoryTracking=summary plus jcmd <pid> VM.native_memory` to see off-heap | Section 8.3 · #36 JVM parameters |
exec /entrypoint.sh: no such file or directory, or the container exits instantly with only that line in docker logs | The script was edited on Windows and stored with CRLF line endings; the kernel reads the interpreter as /bin/sh\r, and that file genuinely does not exist. ENTRYPOINT ["sh","entrypoint.sh"] hides it (sh parses the argument), but the exec form or a direct RUN chmod +x blows up | Run file entrypoint.sh on the host (it reports with CRLF line terminators) or `cat -A entrypoint.sh \ | grep -c '\^M\$'; fix with dos2unix, or add *.sh text eol=lf to .gitattributes, or RUN sed -i 's/\r$//' entrypoint.sh` in the Dockerfile | Section 10 · #46 delivery checklist |
OCI runtime create failed: ... exec /entrypoint.sh: no such file or directory | The container-level twin of the row above, reported by runc; the same family of triggers also includes a missing executable bit (COPY preserves source permissions), a relative ENTRYPOINT combined with a changed WORKDIR, and a base image built for another architecture | Look first with docker run --rm --entrypoint sh <image> -c 'ls -l; head -1 entrypoint.sh' — is the file there at all, and is it -rwxr-xr-x; then handle CRLF as above, and add RUN chmod +x for the permission case | Section 10 | |
Bind for 0.0.0.0:8080 failed: port is already allocated | Host port 8080 is already taken (another container, a process that survived the last run, a local dev server). Note the message names the host port — the container's own 8080 is fine | docker ps --format '{{.Names}}\t{{.Ports}}' to find the squatter; in development just move the host side with -p 8081:8080; in production the real fix is to let the orchestrator allocate rather than hard-coding a host port | Sections 6 and 7 | |
java.lang.OutOfMemoryError: Java heap space | The heap genuinely filled up: an unbounded collection, a query returning a million rows, a cache with no ceiling | Reproduce once with -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/tmp, grab the hprof and read the dominator tree in MAT; this class of failure never leaves an exit 137 behind | Section 8.3 · #28 connection pool and memory | |
Received SIGTERM, JVM shutdown is in progress in the log, yet requests are severed at the very moment of shutdown | The signal did reach the JVM (so half of Section 10.6 works), but draining did not finish before SIGKILL: Spring's phase timeout exceeds Docker's grace period, or a long query / long-lived connection is still holding the thread | Align both sides: docker stop -t 30 <id> while keeping spring.lifecycle.timeout-per-shutdown-phase below it (say 20s); and make sure no in-flight request runs a multi-minute SQL | Section 10.6 · Section 13 · #36 graceful shutdown | |
standard_init_linux.go:228: exec user process caused: exec format error | Image architecture and host instruction set mismatch (an arm64 image on an amd64 host, or the reverse); nothing to do with Java | docker inspect --format '{{.Os}}/{{.Architecture}}' <image> to learn the image's identity; pin --platform linux/amd64 at build time | Section 10.5 | |
Editing one line of business code makes mvn dependency:go-offline run again | COPY . . sits before dependency installation, so a source change invalidates that layer and everything above it | Split it: COPY pom.xml . → RUN mvn -B dependency:go-offline → COPY src ./src → RUN mvn package; plus a .dockerignore excluding target/ and .git/ | Section 2.1 · the layer experiment above | |
| The container runs fine but the browser cannot connect at all | Only EXPOSE 8080 (a pure declaration) was written; -p 8080:8080 or Compose's ports is missing | Check the PORTS column of docker ps for 0.0.0.0:8080->8080/tcp; no entry means nothing was published | Section 10.2 | |
Communications link failure connecting to a MySQL on the host | localhost inside a container means the container itself, not the host | Use 172.17.0.1 on Linux, host.docker.internal on Mac/Windows; the proper fix is putting the database in Compose too and reaching it by service name | Section 10.1 · Section 6 | |
Sporadic UnknownHostException: redis on an Alpine base image, while nslookup redis works | musl libc's getaddrinfo handles IPv6 / short TTLs differently from glibc, and slim images hit it more often | Switch back to eclipse-temurin:21-jre-jammy to confirm it is libc-related before committing to alpine | The Section 9 tip · the build experiment above | |
| The app hangs at startup, Tomcat never binds the port, logs stop at a random-number line | A slim image has a small entropy pool, so SecureRandom blocks reading /dev/random | On older kernels add -Djava.security.egd=file:/dev/./urandom; Linux 5.6+ no longer blocks on /dev/random | Section 10.4 |
three rows in this table (137, the swallowed signal, alpine DNS) never appear as a "Java exception" — the symptom is that the log simply stops. When debugging a container, start with docker inspect <id> --format '{{.State.ExitCode}} {{.State.OOMKilled}}'; one glance tells you who did the killing.
Row five of the error table above (OutOfMemoryError: Java heap space) looks a lot like row one (137) and yet is a completely different scene. The stack below is real; don't peek at the answer — click the line you blame, then read the boundary between the two failure modes:
A reporting endpoint exports a month of order details (SELECT * over hundreds of thousands of rows). The pod was OOMKilled and restarted twice at peak hours; this time the log kept a complete Java exception — note how it differs from Exit Code 137.
A warm-up on the caching rule from Section 2:
Then a comprehensive one, tying 8.3, 10.6 and Section 13 together:
"Just raise the limit" is the most expensive form of comfort in container debugging — it removes the symptom while ensuring you never learn who was eating the memory. The sandbox below turns heap ratio and container limit into two switches; flipping one shows the computed default heap, where total usage lands, and which failure you end up with:
-Xmx512m → MaxHeapSize = 512MB (the entire limit)# Metaspace 180MB + 200 thread stacks ≈ 200MB + CodeCache 64MBRSS climbs to 519MB against a 512MB cgroup ceilingReason: OOMKilled · Exit Code 137 · no OutOfMemoryError in the log
the numbers are illustrative; the conclusions are real — the costs outside the heap do not shrink with the limit, yet they grow with concurrency. That is why "the same percentage is safe or not depending on the limit", and why hard-coded -Xmx loses in both directions: small boxes blow up anyway, big boxes waste half the memory.
Goal: starting from an empty directory, produce an image with "multi-stage build + layered COPY + limit awareness + exec-form entry point", then bring it up together with MySQL via Compose. Every step has an expected output; stop and check when it does not match.
Create three files: pom.xml, .dockerignore and src/main/java/com/example/lab/LabApplication.java. The pom.xml needs layered extraction enabled, because that is the precondition for the four-layer COPY later:
<?xml version="1.0" encoding="UTF-8"?><project xmlns="http://maven.apache.org/POM/4.0.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 https://maven.apache.org/xsd/maven-4.0.0.xsd"> <modelVersion>4.0.0</modelVersion> <parent> <groupId>org.springframework.boot</groupId> <artifactId>spring-boot-starter-parent</artifactId> <version>3.3.0</version> <relativePath/> </parent> <groupId>com.example</groupId> <artifactId>docker-lab</artifactId> <version>1.0.0</version> <properties> <java.version>21</java.version> </properties> <dependencies> <dependency> <groupId>org.springframework.boot</groupId> <artifactId>spring-boot-starter-web</artifactId> </dependency> </dependencies> <build> <finalName>app</finalName> <plugins> <plugin> <groupId>org.springframework.boot</groupId> <artifactId>spring-boot-maven-plugin</artifactId> <configuration> <layers> <enabled>true</enabled> </layers> </configuration> </plugin> </plugins> </build></project>.dockerignore — without this line, tens of seconds of context upload are wasted:
target/.git/.idea/*.imlsrc/test/The minimal application, with one endpoint that reports what the JVM believes its heap is:
package com.example.lab;import org.springframework.boot.SpringApplication;import org.springframework.boot.autoconfigure.SpringBootApplication;import org.springframework.web.bind.annotation.GetMapping;import org.springframework.web.bind.annotation.RestController;@SpringBootApplication@RestControllerpublic class LabApplication { @GetMapping("/hello") public String hello() { long heapMb = Runtime.getRuntime().maxMemory() / 1024 / 1024; return "heap max = " + heapMb + "MB"; // step 7 uses this to verify container awareness } public static void main(String[] args) { SpringApplication.run(LabApplication.class, args); }}The Dockerfile — multi-stage, layered, exec entry point, proportional heap:
# ---------- stage one: build ----------FROM maven:3.9-eclipse-temurin-21 AS builderWORKDIR /buildCOPY pom.xml .RUN mvn -B dependency:go-offlineCOPY src ./srcRUN mvn -B clean package -DskipTestsRUN java -Djarmode=layertools -jar target/app.jar extract# ---------- stage two: runtime ----------FROM eclipse-temurin:21-jre-jammyWORKDIR /appCOPY --from=builder /build/dependencies/ ./COPY --from=builder /build/spring-boot-loader/ ./COPY --from=builder /build/snapshot-dependencies/ ./COPY --from=builder /build/application/ ./ENV TZ=Asia/Shanghai \ JAVA_TOOL_OPTIONS="-XX:MaxRAMPercentage=75.0 -XX:MaxMetaspaceSize=256m"EXPOSE 8080ENTRYPOINT ["java", "org.springframework.boot.loader.launch.JarLauncher"]Build and watch the cache:
docker build -t docker-lab:1.0 .Expected output (the key is how many layers print CACHED):
=> [builder 2/6] WORKDIR /build 0.0s=> [builder 3/6] COPY pom.xml . 0.1s=> [builder 4/6] RUN mvn -B dependency:go-offline 0.0s=> CACHED=> [builder 5/6] COPY src ./src 0.2s=> [builder 6/6] RUN mvn -B clean package -DskipTests 41.8s=> [stage-1 3/6] COPY --from=builder /build/dependencies 1.4sRun it with a limit and check your arithmetic against Section 8.3:
docker run -d --name lab -p 8080:8080 -m 512m docker-lab:1.0curl -s localhost:8080/helloExpected response:
heap max = 384MB384 is exactly 75% of 512, proving MaxRAMPercentage took effect. If you read a large fraction of the host's physical memory instead, your base image JDK is too old and UseContainerSupport is off.
Edit one line of business code and rebuild to confirm the dependency layer is still CACHED:
sed -i 's/hello = /hello .. /' src/main/java/com/example/lab/LabApplication.javadocker build -t docker-lab:1.1 .Expected: the build-stage line RUN mvn -B dependency:go-offline still prints CACHED, and only the layers after COPY src rerun. If you see dependencies downloading again, go back to Section 2.1 and re-check the instruction order.
Verify that graceful shutdown really reaches the JVM:
docker stop -t 30 lab &docker logs -f lab 2>&1 | grep -i "closing\|shutdown"You should see a line like Closing org.springframework.boot.web.servlet.context.ServletWebServerApplicationContext.... If not, return to Section 10.6 and inspect the form of your ENTRYPOINT.
Finally, settle it as Compose (docker-compose.yml, next to the Dockerfile):
services: mysql: image: mysql:8.0 environment: MYSQL_ROOT_PASSWORD: root MYSQL_DATABASE: lab healthcheck: test: ["CMD", "mysqladmin", "ping", "-h", "localhost"] interval: 10s retries: 10 volumes: - mysql-data:/var/lib/mysql app: build: . depends_on: mysql: condition: service_healthy deploy: resources: limits: memory: 512M ports: - "8080:8080"volumes: mysql-data:docker compose up -ddocker compose psExpected: app is running and mysql is healthy; the PORTS column for app must contain 0.0.0.0:8080->8080/tcp — if it does not, Section 7 lost its ports.
Acceptance checklist: ① explain why the dependency layer still hits the cache in the second build (the Section 2 rule);② where does the 384MB come from in your curl;③ what does docker inspect lab --format '{{.State.ExitCode}} {{.State.OOMKilled}}' return after your manual docker stop, and why is it not 137 True.
Change exactly one thing and watch the conclusion flip:
- Turn
ENTRYPOINTinto shell form,ENTRYPOINT java -jar app.jar(dropping the JSON array). You will observe:docker stop -t 30 labstill ends after about ten seconds, and the log contains noClosing ...at all — the signal swallowed by the shell as described in 10.6. Then rundocker exec lab ps -efand look at who occupies slot 1: it issh. - Delete the
target/line from.dockerignore, runmvn packagelocally, thendocker build. You will observe:transferring contextat the start of the build grows from a few hundred KB to tens of MB and takes noticeably longer; and any change outsideCOPY src ./srcinvalidates the layers above. That is "the context is a cost too". - Change
-m 512mto-m 256mwithout touchingJAVA_TOOL_OPTIONS. You will observe:curlbecomes connection-refused,docker inspectreportsOOMKilled truewith Exit Code 137, and the last line ofdocker logsis still the normal startup banner — a perfect sample of "external execution, no evidence". Now drop the ratio to 50 and try again: the symptom flips from 137 to an in-heap error the log can record. - Switch the base image to
eclipse-temurin:21-jre-alpine. You will observe: the image does shrink by about 20MB, but if you connect by service name rather than IP you may hit sporadic resolution failures — compare the Section 9 tip with frame 4 of thebuildexperiment.
After finishing item 1, go back to the Section 16 sandbox: set policy to fixed-xmx and limit to 512m; the two conclusions should line up.
Give yourself a "one-command reproducible environment plus a release health check", so that anyone who receives the repository can run it without asking you a single question.
Requirements:
- One
compose.yamlwith three services —mysql(withhealthcheckand init SQL),redis(persisted by a named volume),app(depends_on: condition: service_healthy) — all declaring explicit memory limits - The
appDockerfile must satisfy all of: multi-stage,.dockerignore, four-layerlayertoolsCOPY, exec-form entry point,MaxRAMPercentagerather than-Xmx, non-root user (USER app) - A
Makefilewith at least four targets:make build,make up,make smoke(curl the health endpoint plus/hello, exit non-zero if either fails),make graceful-test(backgrounddocker stop -t 30whiledocker logs -fscrapes the shutdown output; the assertion must seeClosing, otherwise it reports "the signal never reached the JVM" and exits non-zero) - A "memory budget table" in your README: the ratio you picked, the resulting heap size, Metaspace, thread count ×
-Xss, your direct-memory estimate, and why the sum stays under the limit
Acceptance checklist: ① make up && make smoke is all green on a brand-new machine;② if you deliberately make one bean's constructor throw, make smoke must fail rather than hang;③ make graceful-test passes with the exec form and fails after you switch the entry point to shell form (proving the script really catches regressions);④ the final image is under 300MB in docker images, and docker inspect --format '{{.Os}}/{{.Architecture}}' matches the deployment target;⑤ explain "why does docker stop end with ExitCode 143 rather than 137".
without looking above, state the invalidation direction of image layers — does a change affect the layers above it or below? Why must COPY pom.xml precede COPY src?
who raises OOMKilled / Exit Code 137 and who raises OutOfMemoryError: Java heap space? Which one leaves a heap dump, and why?
why does a shell-form ENTRYPOINT break graceful shutdown? What is each of the three fixes for?
inside a container, who does localhost point to? What is the proper way to reach across containers, and what role does EXPOSE play in it?
once depends_on carries condition: service_healthy, who decides "readiness"? What is the half it still lacks (hint: probe semantics in Article 42)?
stable first, volatile last — change a thing and it explodes from there upward; the limit is the bill for the whole case, and the signal only counts once it reaches slot 1.
three takeaways — image layers are a read-only stack, so put stable instructions high and COPY pom.xml early to hit the cache; multi-stage builds separate the build environment from the runtime environment, slimming an image from 700MB to 200MB; inside a container there is no localhost, EXPOSE is not publishing a port, and logs go to stdout. Get the Dockerfile right and the Compose orchestration clean, and you have a reproducible "app + MySQL + Redis" environment.