Capstone 4: Testing, CI/CD and Going Live

bee2026-10-08102 min read0 views
Complete quality and delivery for BeeOrder: a layered test strategy, a GitHub Actions pipeline, production Compose, graceful shutdown and probes, an Nginx reverse proxy, a go-live checklist and a load-testing primer.
1 / 180
Section
0. The 30-second version
2 / 180

The first 45 articles finished BeeOrder: it takes orders, deducts stock, and receives payment callbacks. This article does exactly one thing — turn "it runs on my machine" into "anyone can run it somewhere else, and we can back out when it breaks." Delivery, in substance, is writing three kinds of hidden precondition down as explicit assets: a reproducible build (the same code produces the same artifact on any machine), verifiable quality (the mistakes worth blocking get blocked by a machine, not by somebody remembering), and a reversible runtime (when the new version comes in, the old one drains before it leaves). As long as one of the three still depends on "I remembered," the incident will come from that one.

3 / 180

Six words, one line each (they appear everywhere below):

4 / 180
  • Artifact: the output of one build, never modified afterwards — a jar or an image tag. What you deploy is an artifact, never source
  • Image tag: the artifact's fingerprint. Tag with the commit sha and you can answer the lethal question "which build is production running right now?"
  • Readiness probe: answers "may I send it traffic?" A dependency not ready should pull traffic, not kill the process
  • Liveness probe: answers "is this process salvageable?" It should judge the process alone; wire it wrong and a database blip restarts every instance
  • Graceful shutdown: stop accepting new requests, let in-flight requests finish, then exit
  • Canary release: hand 5% of traffic to the new version first, widen only once it looks healthy
5 / 180
Analogy

closing a restaurant for the night. The correct shutdown order is: flip the sign from OPEN to CLOSED first (readiness goes red, no new requests), then let the staff finish the dishes already on the table and settle the bills (in-flight requests drain), and only then kill the lights and lock the door (the process exits). What happens if you just pull the mains (SIGKILL)? Customers leave with food still in their mouths and no receipt — in system terms, a row where stock was deducted but the order was never written. So graceful shutdown is not a switch; it is an entire closing procedure.

6 / 180
Analogy

a bento box with separate compartments. An image is not one solid block but a stack of read-only layers: rice (dependencies, almost never change), the dish (your code, changes daily), the sauce (configuration, swapped per environment). The payoff of separating them is that replacing the dish does not require steaming fresh rice: change one line of business code and that several-hundred-MB dependency layer still comes straight out of the cache. Do the opposite — stir everything into one lump (COPY . . and only then install dependencies) — and a single character forces you to remake the whole box, turning a 30-second build back into a several-minute one.

7 / 180
Diagram
Figure · Your map for this article: five gates of delivery
Figure · Your map for this article: five gates of delivery
8 / 180

Those five gates map onto the article's spine: quality gates (Sections 2-3), build gates (Sections 4-5), runtime gates (Sections 6-10), exit discipline (Section 8), rollout cadence (Section 12). Every failed release can be pinned to exactly one spot on this tree — which makes it both the starting point of triage and the index of the error table at the end.

9 / 180
Animation
Animation · The full timeline of one rolling release
Animation · The full timeline of one rolling release
10 / 180

This timeline is the single most important picture in the article; every section later dissects one of its frames. Memorise three conclusions first: only the tag CI verified may ship (frame 1), an old instance leaves in two steps — pulled from the pool, then drained — not one (frames 4-5), and a hard kill is the fallback when the grace period expires, never the normal method (frame 7).

11 / 180

After this article you should be able to answer three questions:

12 / 180
  1. "It works on my machine" is missing which three kinds of precondition, and which stage supplies each one?
  2. Why may readiness depend on the database while liveness absolutely may not? What concrete incident does the reversal cause?
  3. After one docker compose up -d, in what order should you confirm what, during those 30 seconds?
13 / 180
Section
1. The last gate before delivery: why "it works on my machine" is not delivery
14 / 180

"It works on my machine" is the most expensive sentence in the history of shipping software. What it really means is that my machine carries a pile of hidden conditions — versions, time zones, memory, configuration — that happen to add up to a working state, and none of those conditions were ever written down. Delivery is precisely the act of making all of those hidden conditions explicit.

15 / 180

Three real disasters, every one of which I have seen with my own eyes:

16 / 180
Table
DisasterSymptomRoot causeThe guardrail delivery needs
JDK mismatchjava -jar throws UnsupportedClassVersionError on line oneBuilt locally on JDK 21, the server only has 17Pin the JDK inside the image; build and run on the same base image
Time zone driftOrder times are 8 hours off what users see; reconciliation failsContainers default to UTC while app and DB each follow their own defaultOne time zone across the stack; set TZ explicitly in the image
Pool copied from localTraffic arrives and you get Connection is not available, request timed outProduction runs the default maximumPoolSize=10Size the pool from DB limits and concurrency, then verify under load
17 / 180

What these three share is not "a wrong parameter" but an environment that was never version-controlled. Spread the hidden preconditions out and there are only three families, each with its own place to be written down and its own verifier:

18 / 180
Table
Hidden preconditionWhere it normally livesWhere it gets written downWho verifies it
Build: JDK version, Maven plugins, the dependency treeYour local ~/.m2 and the SDK dropdown in your IDEThe first FROM line of the Dockerfile + pom.xmlCI, on a clean runner with no shared state
Runtime: time zone, memory limit, pool size, GCYour machine's 32GB and the fact that nobody hits your laptop at 3amCompose environment / JAVA_TOOL_OPTIONS / deploy.resourcesHealth checks + load tests
Process: who reviewed it, did the tests run, was the canary switchedTeam memory and chat historyBranch protection rules + the pipeline's needs gatesA machine, which does not accept "I ran it locally"
19 / 180
Key point

delivery is not "copy the jar up there"; it is "let anyone run a single command on any machine and get the same working system." Any precondition that is not written into code or configuration is a future incident ticket.

20 / 180
Section
2. A layered test strategy: each layer has its own job
21 / 180

More tests is not the goal. The point of testing is to cover a specific risk at a specific cost. Split tests into three layers, give each layer one class of problem, and you avoid the embarrassment of "hundreds of tests and the production incident still slipped through."

22 / 180
Table
LayerKey annotations / toolsWhat it coversDependenciesSpeedSuggested share
UnitJUnit 5 + MockitoService branches, amount math, state transitionsAll mocked, no IOMilliseconds60%
Slice@WebMvcTestController contract: routing, validation, status codes, JSON shapeWeb layer only, Service mockedHundreds of ms25%
Integration@SpringBootTest + TestcontainersSQL, transactions, unique constraints, idempotency, pool behaviorReal MySQL / Redis containersSeconds15%
23 / 180
Note

the ratios are a cost constraint, not dogma. Do not use an integration test to cover an if-else branch — that is spending second-level cost to verify a millisecond-level question. And do not use a unit test to confirm "did the SQL use the index" — a mock can never give you the real answer.

24 / 180
Section
2.1 Unit tests: isolating dependencies in the Service layer with Mockito
25 / 180

Start with the happy path: enough stock, correct amount, order persisted.

26 / 180
Code
Codejava
@ExtendWith(MockitoExtension.class)class OrderServiceTest {    @Mock private OrderMapper orderMapper;    @Mock private OrderItemMapper orderItemMapper;    @Mock private InventoryService inventoryService;    @InjectMocks private OrderService orderService;    @Test    void createOrder_shouldSucceed_whenStockIsEnough() {        CreateOrderCmd cmd = new CreateOrderCmd(1001L, List.of(new ItemCmd(2001L, 2)));        when(inventoryService.deduct(2001L, 2)).thenReturn(true); // deduction succeeds        OrderVO vo = orderService.create("idem-key-001", cmd);        assertThat(vo.orderNo()).startsWith("BO");        assertThat(vo.amount()).isEqualByComparingTo("199.00"); // compare money as BigDecimal        verify(orderMapper, times(1)).insert(any(Order.class));  // persisted exactly once    }}
Notes
  • @Mock fakes the mappers and the inventory service, so the test never touches a database
  • when(...).thenReturn(true) stages the "stock is enough" premise
  • isEqualByComparingTo compares money and sidesteps BigDecimal.equals' scale sensitivity
  • verify(..., times(1)) asserts "persisted once" — exactly the floor idempotency must hold
27 / 180
Section
2.2 Insufficient stock: the failure branch is what prevents losses
28 / 180
Code
Codejava
@Testvoid createOrder_shouldFail_whenStockIsNotEnough() {    CreateOrderCmd cmd = new CreateOrderCmd(1001L, List.of(new ItemCmd(2001L, 2)));    when(inventoryService.deduct(2001L, 2)).thenReturn(false); // deduction fails    BizException ex = assertThrows(BizException.class,            () -> orderService.create("idem-key-002", cmd));    assertThat(ex.getCode()).isEqualTo(ErrorCode.STOCK_NOT_ENOUGH);    verify(orderMapper, never()).insert(any(Order.class)); // the key: never insert}
Notes
  • When deduction returns false, the domain must throw "insufficient stock", not silently return
  • never() asserts that not a single order row was inserted — blocking the "deduction failed but the order got written anyway" corruption
  • The exception flows through the unified error-code system (see #44) so the client can handle it reliably
29 / 180
Section
2.3 Duplicate payment callbacks: idempotency tests matter more than concurrency tests
30 / 180

The most typical behavior of a payment gateway is to call back more than once. When the same trade_no arrives twice, the system may book it only once.

31 / 180
Code
Codejava
@Testvoid handleCallback_shouldBeIdempotent_whenDuplicated() {    PayCallback cb = new PayCallback("trade-9f3a", "BO20260207001", new BigDecimal("199.00"));    boolean first  = paymentService.handleCallback(cb);    boolean second = paymentService.handleCallback(cb); // an identical duplicate    assertThat(first).isTrue();          // the first one books for real    assertThat(second).isFalse();        // the second is stopped by idempotency    verify(orderMapper, times(1)).markPaid(anyString()); // state advances only once}
Notes
  • The first callback moves the order from CREATED to PAID and returns true
  • The second hits the unique key trade_no or the state-machine check and returns false, doing no work
  • times(1) is the soul of this test: the definition of idempotency is "the side effect happens exactly once"
32 / 180
Section
2.4 Slice tests: locking the API contract with `@WebMvcTest`
33 / 180

Once the contract is frozen, the front end can build in parallel. @WebMvcTest loads only the web layer, starts no database, and tests the HTTP face alone.

34 / 180
Code
Codejava
@WebMvcTest(OrderController.class)class OrderControllerTest {    @Autowired private MockMvc mockMvc;    @MockBean private OrderService orderService;    @Test    void createOrder_shouldReturn200_withOrderNo() throws Exception {        when(orderService.create(anyString(), any()))                .thenReturn(new OrderVO("BO20260207001", new BigDecimal("199.00"), "CREATED"));        mockMvc.perform(post("/api/orders")                        .header("Idempotency-Key", "idem-key-001")                        .contentType(MediaType.APPLICATION_JSON)                        .content("{\"userId\":1001,\"items\":[{\"productId\":2001,\"quantity\":2}]}"))                .andExpect(status().isOk())                .andExpect(jsonPath("$.code").value(0))                .andExpect(jsonPath("$.data.orderNo").value("BO20260207001"));    }}
Notes
  • @WebMvcTest scans only the controller; the service is replaced by @MockBean, so startup is fast
  • What is asserted is the response contract: status code, the uniform body's code, and the data shape
  • The front end treats this test as the API spec; the moment a field name drifts, CI turns red
35 / 180
Section
2.5 Integration tests: real MySQL with `@SpringBootTest` + Testcontainers
36 / 180

SQL, unique constraints and transaction rollback can only be tested against a real database. Testcontainers' trick is declaring a throwaway real MySQL in code, so the test machine needs nothing preinstalled.

37 / 180
Code
Codejava
@SpringBootTest@Testcontainersclass OrderRepositoryIT {    @Container    static MySQLContainer<?> mysql = new MySQLContainer<>("mysql:8.0")            .withDatabaseName("beeorder").withUsername("bee").withPassword("bee");    @DynamicPropertySource    static void props(DynamicPropertyRegistry r) {        r.add("spring.datasource.url", mysql::getJdbcUrl);        r.add("spring.datasource.username", mysql::getUsername);        r.add("spring.datasource.password", mysql::getPassword);    }    @Autowired private OrderService orderService;    @Test    void shouldNotOversell_underConcurrency() {        // run concurrent checkouts on a real DB and prove the conditional update blocks oversell    }}
Notes
  • @Testcontainers lets JUnit manage the container's lifecycle: it starts with the class and is destroyed after
  • @DynamicPropertySource feeds the container's random port back into Spring, with zero hardcoded config
  • Complex SQL and the behavior of WHERE available >= ? can only be verified here
38 / 180
Section
2.6 Coverage is not the goal; assertion strength is
39 / 180

The number that fools reviewers most reliably is "82% line coverage". Both tests below make that line count as covered; only the second actually blocks a regression.

40 / 180
Code
Codejava
// Weak: only proves the line ran. It buys coverage and locks nothing.@Testvoid createOrder_any() {    OrderVO vo = orderService.create("k", cmd);    assertThat(vo).isNotNull();}// Strong: nails down the math, the state transition, and how many side effects ran.@Testvoid createOrder_contract() {    OrderVO vo = orderService.create("k", cmd);    assertThat(vo.amount()).isEqualByComparingTo("199.00");    assertThat(vo.status()).isEqualTo("CREATED");    verify(inventoryService, times(1)).deduct(2001L, 2);    verify(notificationService, never()).notify(any());   // never notify before commit}
Notes
  • Coverage measures "was this line executed"; assertions measure "is this behavior nailed". Neither substitutes for the other
  • One practical rule: every new business branch earns at least one verify or assertThrows; a bare isNotNull() does not count
  • That final never() is precisely the guardrail of #45 Section 14 — notify before commit and coverage is still a pristine 100%
41 / 180
Section
3. Test data and concurrency tests: oversell can only be caught concurrently
42 / 180

A unit test can never catch oversell — a mock does not run concurrently. Concurrency problems must be verified with concurrency; this is the one point in this article on which there is no compromise.

43 / 180
Code
Codejava
@Testvoid shouldNeverOversell_when100RequestsCompeteFor5Items() throws Exception {    int threads = 100;    ExecutorService pool = Executors.newFixedThreadPool(threads);    CountDownLatch ready = new CountDownLatch(threads);    CountDownLatch start = new CountDownLatch(1);   // the starting gun    CountDownLatch done  = new CountDownLatch(threads);    AtomicInteger success = new AtomicInteger();    for (int i = 0; i < threads; i++) {        final long userId = 1000L + i;        pool.submit(() -> {            ready.countDown();            try {                start.await();                       // all threads wait here, released together                orderService.create("key-" + userId,                        new CreateOrderCmd(userId, List.of(new ItemCmd(2001L, 1))));                success.incrementAndGet();            } catch (BizException e) {                // insufficient stock, an expected failure            } catch (Exception ignored) {            } finally {                done.countDown();            }        });    }    ready.await();    start.countDown();        // the real burst happens after this line    done.await(30, TimeUnit.SECONDS);    Integer remain = jdbcTemplate.queryForObject(            "SELECT available FROM inventory WHERE product_id = 2001", Integer.class);    assertThat(success.get()).isEqualTo(5);   // exactly 5 orders succeed    assertThat(remain).isZero();              // stock is 0, no oversell}
Notes
  • ready seats all 100 threads; start is the starting gun that compresses them into one instant
  • success asserts exactly 5 orders succeed; one more means oversell
  • Finally the DB is queried for available to prove the books balance (stock must be 0)
44 / 180
Table
MetricPass criterionWhy
Successful ordersExactly 5More means oversell, fewer means false rejections
Remaining stock0It must never go negative
Ledger rows5One per successful order
Failed orders95All return the "insufficient stock" code
45 / 180
Trap

a concurrency test without a start starting gun degrades into serial execution — it "passes" while testing nothing. Whether it works depends on everything really happening at once, not on how large a thread count you typed.

46 / 180
Section
3.1 Three disciplines for test data
47 / 180
Table
DisciplineAnti-patternThe right move
Each test carries its own dataThe whole class shares order_id=1; whichever test runs first mutates the state and the second failsCreate in @BeforeEach, clean in @AfterEach, or simply mint a new ID every time
Never trust the system clock"Close the order after 15 unpaid minutes" turns into sleep(15, MINUTES) in the testInject a Clock and wind a fake time forward in the test
Every external dependency gets a stubThe test really calls the payment gateway, and one CI run books a real charge@MockBean or a WireMock fake gateway — and keep signing the request anyway
48 / 180
Tip

always spin integration databases with Testcontainers. A shared "test environment database" is the number one reason regression suites stop being trustworthy — somebody pokes one row by hand, your test is red the next morning, and nobody can explain why.

49 / 180
Section
4. The CI pipeline: let the machine do the checking
50 / 180

CI's value fits in one sentence: expose mistakes before the merge, not after go-live. Manual checks will always miss; a machine will not, as long as the rules are right.

51 / 180
Diagram
Figure 1 · From commit to production
Figure 1 · From commit to production
52 / 180

Here is a ready-to-use GitHub Actions pipeline with three stages: build-and-test → package-and-push → deploy. The stages are chained with needs, so a failure upstream stops everything downstream — fail fast.

53 / 180
Code
Codeyaml
name: beeorder-cion:  push:    branches: [ main ]    tags: [ 'v*' ]  pull_request:    branches: [ main ]jobs:  build-and-test:    runs-on: ubuntu-latest    steps:      - uses: actions/checkout@v4      - name: Set up JDK 21        uses: actions/setup-java@v4        with:          distribution: temurin          java-version: '21'          cache: maven                     # cache ~/.m2 to speed up dependency downloads      - name: Run tests        run: mvn -B -ntp clean verify       # unit + slice + integration together      - name: Upload test report        if: always()        uses: actions/upload-artifact@v4        with:          name: surefire-reports          path: target/surefire-reports  package:    needs: build-and-test                  # tests fail → no image build    if: github.ref == 'refs/heads/main' || startsWith(github.ref, 'refs/tags/v')    runs-on: ubuntu-latest    permissions:      contents: read      packages: write    steps:      - uses: actions/checkout@v4      - name: Log in to GHCR        uses: docker/login-action@v3        with:          registry: ghcr.io          username: ${{ github.actor }}          password: ${{ secrets.GITHUB_TOKEN }}      - name: Build and push image        uses: docker/build-push-action@v6        with:          context: .          push: true          tags: |            ghcr.io/acme/beeorder:${{ github.sha }}            ghcr.io/acme/beeorder:latest  deploy:    needs: package    if: startsWith(github.ref, 'refs/tags/v')   # only a tag actually ships    runs-on: ubuntu-latest    environment: production    steps:      - name: Deploy over SSH        uses: appleboy/ssh-action@v1        with:          host: ${{ secrets.DEPLOY_HOST }}          username: ${{ secrets.DEPLOY_USER }}          key: ${{ secrets.DEPLOY_KEY }}          script: |            cd /opt/beeorder            export IMAGE_TAG=${{ github.sha }}            docker compose pull app            docker compose up -d app            ./scripts/wait-healthy.sh app   # only success once health passes
Notes
  • cache: maven is the single most important speedup; repeat builds stop re-downloading dependencies
  • needs chains the stages: red tests → no package; a failed package → no deploy, saving a lot of waiting
  • if: startsWith(github.ref, 'refs/tags/v') narrows "ship it" to a tag, so an everyday push can never publish by accident
  • wait-healthy.sh polls container health, preventing "the container is up but the service is not" from being counted as success

Tip: integration tests inside CI really start a MySQL container, so never point test environment variables at the production database. The test connection string always comes from Testcontainers, and production secrets live only in the deployment environment's secrets.

54 / 180
Section
4.1 The image is the artifact: one multi-stage Dockerfile
55 / 180

The pipeline's last step produces an image rather than a jar, because an image seals "application + JDK + time zone + startup flags" together, while a jar only seals the application. This is #37's recipe in its final delivery form:

56 / 180
Code
Codedockerfile
# ---------- stage 1: build (JDK + Maven, thrown away) ----------FROM maven:3.9-eclipse-temurin-21 AS builderWORKDIR /buildCOPY pom.xml .RUN mvn -B dependency:go-offline          # deps unchanged → this layer is always CACHEDCOPY src ./srcRUN mvn -B clean package -DskipTests# split the fat jar into four layers so the dependency layer caches inside the image tooRUN java -Djarmode=layertools -jar target/beeorder.jar extract# ---------- stage 2: runtime (JRE only, no Maven, no sources) ----------FROM eclipse-temurin:21-jre-jammyWORKDIR /appENV TZ=Asia/Shanghai \    JAVA_TOOL_OPTIONS="-XX:MaxRAMPercentage=75.0 -XX:MaxMetaspaceSize=256m -Xss512k"# run as a non-root user to shrink the blast radius of an escapeRUN useradd -r -u 1001 -s /usr/sbin/nologin beeUSER beeCOPY --from=builder /build/dependencies/ ./COPY --from=builder /build/spring-boot-loader/ ./COPY --from=builder /build/snapshot-dependencies/ ./COPY --from=builder /build/application/ ./EXPOSE 8080# exec form: java is PID 1, so SIGTERM actually lands (the premise of Section 8)ENTRYPOINT ["java", "org.springframework.boot.loader.launch.JarLauncher"]
Notes
  • The order of those four COPY lines is the caching strategy: the most stable on top, application/ last
  • JAVA_TOOL_OPTIONS lives in the runtime stage's env so ops can override it with -e without rebuilding
  • After USER bee, remember the health check in Section 7 needs curl to exist in this image at all
  • ENTRYPOINT must be the exec array form, or none of Section 8's graceful-shutdown config will ever execute
57 / 180

Three hygiene problems of the build stage hide in the trade-offs of this section — whether .dockerignore has to be written first, why the builder's ~/.m2 re-downloads every time, and why an alpine/musl base intermittently fails name resolution. Pick build and you get them frame by frame:

58 / 180
Kernel lab
TeaVMMulti-stage builds and image hygieneidle
Choose build: frame 2 explains why .dockerignore must be written first, frame 4 is where that alpine/musl UnknownHostException comes from
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
59 / 180

The order of those four COPY lines follows Docker's layer-cache rule: the most stable layer gets copied first and the most volatile one stays outermost, so a rebuild only invalidates the top few layers. This figure peels the image from bottom to top — click each layer to see what update forces it to rebuild:

60 / 180
Diagram
LayersImage layers: the stable at the bottom, the volatile on top1 / 5
Click the five layers bottom-up: the base image almost never changes, dependencies change only with pom.xml, application changes on every build — the further out, the more often it breaks and the cheaper it is to rebuild
→
→
→
→
Base image eclipse-temurin:21-jre
Never invalidated unless you change the base. It also pins the Java version inside the image — the first line of defence against that UnsupportedClassVersionError row in Section 15.
All clearBottom to top: the most stable gets COPYed first, the ever-changing application always last — that is the entire reason for the four COPY lines above.
61 / 180

The skeleton can also be generated directly: tick the items this version needs and every line comes annotated with the precondition it satisfies (the shutdown signal and memory awareness are precisely what make Sections 8 and 9 work inside a container):

62 / 180
Generator
GeneratorGenerate a delivery Dockerfile for what you needDockerfile5 / 8
Start with just 'multi-stage + JRE' to see the minimal skeleton; layer the 'layered extraction' on and watch how the COPY order follows; finish by adding non-root, memory awareness, a healthcheck and the exec-form stop signal — each one maps to a precondition from this section or Section 8
Output
# ---------- 构建阶段:要完整 JDK 与 Maven ----------
FROM maven:21-eclipse-temurin AS build
WORKDIR /src
COPY pom.xml .
RUN mvn -B dependency:go-offline        # 先只拷 pom,依赖层可被缓存
COPY src ./src
RUN mvn -B -DskipTests package \
    && java -Djarmode=layertools -jar target/*.jar extract --destination /app

# ---------- 运行阶段:只要 JRE ----------
FROM eclipse-temurin:21-jre
WORKDIR /app
# 变化频率从低到高排列,改代码不会让依赖层缓存失效
COPY --from=build /app/dependencies/ ./
COPY --from=build /app/spring-boot-loader/ ./
COPY --from=build /app/snapshot-jar/ ./
COPY --from=build /app/application/ ./

RUN groupadd -r app && useradd -r -g app app && chown -R app:app /app
USER app

ENV JAVA_OPTS="-XX:MaxRAMPercentage=75.0 -XX:+ExitOnOutOfMemoryError"
ENV SPRING_PROFILES_ACTIVE="prod"

EXPOSE 8080 9090
# exec 形式:java 就是 PID 1,SIGTERM 能送达(shell 形式做不到)
ENTRYPOINT ["java", "-jar", "app.jar"]

# docker build -t beeorder:0.0.1 .
# docker run --rm -p 8080:8080 -m 512m beeorder:0.0.1
Why each choice matters
多阶段构建Build needs JDK+Maven, runtime only a JRE — the image drops from ~700MB to ~200MB.
分层解包The Boot plugin writes layers.idx so dependencies and your code land in separate layers — a one-line change stops re-pushing 200MB.
运行镜像用 JRE 而非 JDKNo javac/jar needed in the container; a JRE image is a smaller attack surface.
非 root 用户运行The first barrier on container escape; many PodSecurity policies reject root outright.
-XX:MaxRAMPercentage 而不是 -Xmx 写死The heap follows the container limit; a hard-coded -Xmx is the classic cause of OOMKilled (exit 137).
63 / 180
Section
4.2 The deploy script must wait, not sleep
64 / 180

That ./scripts/wait-healthy.sh app line is the only place in the whole pipeline that separates "the container exists" from "the service works". Do not sleep 30; poll:

65 / 180
Code
Codebash
#!/usr/bin/env bash# scripts/wait-healthy.sh — returns 0 only when readiness is genuinely UPset -euo pipefailSVC="${1:-app}"URL="${HEALTH_URL:-http://127.0.0.1:8080/actuator/health/readiness}"DEADLINE=$((SECONDS + 120))while [ $SECONDS -lt $DEADLINE ]; do  body=$(curl -fsS "$URL" 2>/dev/null || true)  case "$body" in    *'"status":"UP"'*) echo "healthy after ${SECONDS}s"; exit 0 ;;  esac  sleep 3doneecho "not healthy in 120s, dumping logs:" >&2docker compose logs --tail=200 "$SVC" >&2exit 1          # non-zero → pipeline red → later stages never run, i.e. an automatic halt
Notes
  • Use readiness, not health: the first answers "may I route traffic here", the second is an aggregate that may include irrelevant components
  • Print the logs, then exit — that leaves "why it went red" in the build output so nobody has to SSH anywhere
  • The value of exit 1 is that rollback becomes a decision instead of an accident: the stage is red, so the next batch simply never starts
66 / 180
Section
5. How the branch strategy relates to the pipeline
67 / 180

A pipeline is not isolated; it has to mesh with the branch model. BeeOrder uses the simplest trunk-based variant: main protected + PRs must be green + releases triggered by tags.

68 / 180
Table
Branch / eventPipeline triggeredMust passArtifact
push to a feature branchbuild-and-test onlyAll three test layersNone (a quick check)
PR into mainbuild-and-testTests + contract checkNone (a gate)
merge into mainbuild-and-test + packageTests green:sha and :latest images
tag v1.2.0Full pipeline + deployEverythingDeployed to production
69 / 180
Note

main should turn on branch protection, requiring "at least 1 review plus all status checks passing" before merge. The GitLab CI equivalent is rules matching branch/tag plus only: tags — the same idea, just different yml words.

70 / 180
Section
6. Shipping format: executable jar, war, or image?
71 / 180

The choice you really have to make before shipping is "what shape is the deliverable". These three are not alternatives on one axis: jar and war are packaging formats; an image is a runtime environment.

72 / 180
Table
DimensionExecutable fat jarwar (external Tomcat)Container image
Who supplies the Servlet containerThe app itself (embedded Tomcat, packed inside)Tomcat on the hostThe app itself, inside the image
Who guarantees the JDK versionOps, and it drifts from CIWhatever JDK Tomcat runs onThe FROM line of the Dockerfile
Where dependencies liveBOOT-INF/lib, read by a custom class loaderWEB-INF/lib + WEB-INF/lib-providedSame as #37 once split into layers
Start commandjava -jar app.jarDrop into webapps/, Tomcat unpacks itdocker run; the entrypoint is still java
Several instances per machineEasy (change the port)Constrained to one TomcatEasiest (change the container)
Best forMicroservices, K8s, most new projectsA mandated corporate Tomcat baselineWhen the environment must ship too
73 / 180

**A war that wants to boot with java -jar and deploy into an external Tomcat has two hard prerequisites**; miss either and you get a 404:

74 / 180
java
// Prerequisite 1: extend SpringBootServletInitializer so an external container finds an entry pointpublic class BeeOrderApplication extends SpringBootServletInitializer {    @Override    protected SpringApplicationBuilder configure(SpringApplicationBuilder builder) {        return builder.sources(BeeOrderApplication.class);    }}
75 / 180
Code
Codexml
<!-- Prerequisite 2: make the embedded container provided, or it fights the external Tomcat for servlet-api --><dependency>    <groupId>org.springframework.boot</groupId>    <artifactId>spring-boot-starter-tomcat</artifactId>    <scope>provided</scope></dependency>
Notes

Analogy: a tin of stew versus a ready-made meal pouch. A fat jar is the tin — meat, sauce and seasoning all sealed inside, eat straight out of it (java -jar just works), the price being the weight of the tin itself (those tens of MB of embedded Tomcat). A war is the meal pouch: the food really is cooked, but you must pour it into somebody else's pot (an external Tomcat) to eat, and you never had to buy a stove. An image ships the microwave as well — the difference is no longer about the food, it is about delivering the whole kitchen.

76 / 180

This also explains the most common delivery-stage "it packaged fine but will not start": launching a repackaged fat jar with java -cp app.jar com.beeorder.BeeOrderApplication, and getting class-not-found — the jars in BOOT-INF/lib are not on the system classpath, and only JarLauncher's custom class loader knows they exist. The lab below reproduces exactly that:

77 / 180
Kernel lab
TeaVMWhat is actually inside the jar you shipidle
Start with layout to see the BOOT-INF tree, switch to loader to watch LaunchedClassLoader find the dependencies, and finish on war to see how loading changes under an external Tomcat
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
78 / 180
Decision
DecisionBeeOrder is a monolith, the company mandates no Tomcat baseline, and ops handed over two empty Docker-capable machines. Which deliverable should you pick?
79 / 180
Section
7. Composing production with Compose
80 / 180

A development Compose only needs "it starts." A production Compose must answer four questions: how to restart, how to keep data, how to cap resources, and how to collect logs.

81 / 180
Code
Codeyaml
services:  app:    image: ghcr.io/acme/beeorder:${IMAGE_TAG:-latest}    restart: unless-stopped    env_file: [ .env.production ]    environment:      SPRING_PROFILES_ACTIVE: prod      TZ: Asia/Shanghai      JAVA_TOOL_OPTIONS: "-XX:MaxRAMPercentage=75 -XX:+UseG1GC"      SERVER_SHUTDOWN: graceful                      # Section 8: graceful shutdown, injected here      SPRING_LIFECYCLE_TIMEOUT_PER_SHUTDOWN_PHASE: 25s    depends_on:      mysql: { condition: service_healthy }      redis: { condition: service_healthy }    ports:      - "127.0.0.1:8080:8080"          # host-local only; Nginx proxies in    stop_grace_period: 40s              # the SIGKILL grace period; must exceed the drain budget    healthcheck:      test: ["CMD", "curl", "-fsS", "http://localhost:8080/actuator/health/readiness"]      interval: 15s      timeout: 5s      retries: 5      start_period: 40s                # give the JVM time to boot    logging:      driver: json-file      options: { max-size: "20m", max-file: "5" }    deploy:      resources:        limits: { cpus: "2.0", memory: 1g }    networks: [ bee-net ]  mysql:    image: mysql:8.0    restart: unless-stopped    command: --character-set-server=utf8mb4 --collation-server=utf8mb4_unicode_ci    environment:      MYSQL_DATABASE: beeorder      MYSQL_ROOT_PASSWORD: ${MYSQL_ROOT_PASSWORD}      TZ: Asia/Shanghai    volumes:      - mysql-data:/var/lib/mysql      - ./backup:/backup                 # where the backup script writes    healthcheck:      test: ["CMD", "mysqladmin", "ping", "-h", "localhost", "-p${MYSQL_ROOT_PASSWORD}"]      interval: 10s      retries: 6    networks: [ bee-net ]  redis:    image: redis:7-alpine    restart: unless-stopped    command: ["redis-server", "--appendonly", "yes", "--requirepass", "${REDIS_PASSWORD}"]    volumes: [ redis-data:/data ]    healthcheck:      test: ["CMD", "redis-cli", "-a", "${REDIS_PASSWORD}", "ping"]      interval: 10s      retries: 6    networks: [ bee-net ]networks:  bee-net: { driver: bridge }volumes:  mysql-data:  redis-data:
Notes
  • restart: unless-stopped plus healthcheck gives a deterministic answer for both crashes and not-yet-ready dependencies
  • start_period: 40s is a grace period for the JVM's cold start, so it is not declared dead at boot
  • stop_grace_period: 40s must be larger than Spring's timeout-per-shutdown-phase, otherwise the drain is cut off by SIGKILL (Section 8)
  • logging caps log size and deploy.resources caps CPU/memory, so one service cannot sink the whole box
  • The mysql-data volume keeps data across container rebuilds, and ./backup is mounted out so mysqldump can write backups
82 / 180

A MySQL backup is one cron entry and one command:

83 / 180
bash
# /opt/beeorder/scripts/backup.shdocker compose exec -T mysql \  mysqldump -uroot -p"$MYSQL_ROOT_PASSWORD" --single-transaction beeorder \  | gzip > "/backup/beeorder-$(date +%F-%H%M).sql.gz"find /backup -name '*.sql.gz' -mtime +7 -delete   # keep 7 days
84 / 180
Table
DimensionDev ComposeProd Compose
Image sourceLocal build or DockerfilePulled from the registry at a fixed tag
Ports8080 exposed directlyBound to 127.0.0.1 only, Nginx fronts it
VolumesDisposableNamed volumes + scheduled backups
Resource limitsNonecpus / memory both capped
LoggingDefault stdoutjson-file rotation, 20m × 5
Health checksUsually omittedOn every service, with start_period
Shutdown graceThe default 10 secondsAn explicit stop_grace_period, aligned with the drain
ConfigPlaintext application-devenv_file + secrets, config externalized
85 / 180
Trap

when the container's time zone is not set, MySQL and the app each follow their own default and order times end up 8 hours off. In production, set TZ explicitly on the app, the database and Redis alike, and agree on one convention for storage and display (store UTC, or one shared local zone — pick one as a team and never mix).

86 / 180
Section
8. Graceful shutdown: pull traffic → drain → exit
87 / 180

Frames 4, 5 and 6 of the animation in Section 0 all belong here. Graceful shutdown is a three-leg relay, not a setting: the orchestrator pulls the instance from the pool, the app waits for in-flight requests to finish, and the signal genuinely reaches the JVM. Break any leg and the symptom is identical — "not one shutdown line in the log, and requests got severed".

88 / 180
yaml
# application-prod.ymlserver:  shutdown: graceful              # stop accepting new requests, let in-progress ones finishspring:  lifecycle:    timeout-per-shutdown-phase: 25s   # the drain budget; must be less than stop_grace_periodmanagement:  endpoint:    health:      probes:        enabled: true             # expose /actuator/health/liveness and /readiness      group:        readiness:          include: readinessState,db,redis   # dependencies go to readiness only: pull traffic, do not kill        liveness:          include: livenessState             # only the process's own deadlocks / OOM belong here  health:    livenessstate:      enabled: true    readinessstate:      enabled: true
89 / 180

The three config blocks map to the three legs, and none is optional:

90 / 180
Table
LegWho owns itSymptom when misconfiguredThe right move
Pull trafficThe orchestrator (K8s readiness, or Compose + Nginx upstream health)New requests keep arriving, so the drain never finishesFlip readiness red first, wait 2-5s, then start closing
DrainSpring server.shutdown=gracefulIn-flight requests die with Connection resetGive timeout-per-shutdown-phase room, and keep long transactions out
Signal deliveryThe exec form of ENTRYPOINT, or tiniNot one shutdown line in the log, hard kill at 10sUse the exec array form, or write exec java -jar … in shell form
91 / 180

Both grace periods must line up — this is the most common mismatch. Spring is configured for a 25-second drain while docker stop waits only 10 seconds before SIGKILL, so the last 15 seconds of requests are severed anyway.

92 / 180
bash
# three correct ways to waitdocker stop -t 40 beeorder-app            # a single container: raise the grace period explicitlydocker compose stop --timeout 40 app      # Compose: same, spell it out# K8s: terminationGracePeriodSeconds must be greater than timeout-per-shutdown-phase
93 / 180

Misconfigured and aligned side by side is the whole content of this comparison figure — on the left, the fifteen seconds that get severed; on the right, what the same numbers look like once they agree:

94 / 180
Diagram
Figure · Align the grace periods: 10s SIGKILL vs 40s drain
Figure · Align the grace periods: 10s SIGKILL vs 40s drain
95 / 180

The verification method is crude but it works: run docker compose stop --timeout 40 app in one terminal and docker compose logs -f app in another. It only counts if you see Commencing graceful shutdown and Waiting for requests to complete. This lab walks the whole signal chain and all three shutdown postures:

96 / 180
Kernel lab
TeaVMRollout and graceful shutdown: pull traffic → drain → exitidle
Step through ready / drain / kill: ready shows how the two probes divide the work, drain shows 80 in-flight requests being waited out, kill shows where the deducted-stock-but-no-order requests end up when the process is hard-killed
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
97 / 180
Analogy

a cinema closing for the night. Turning off the lights is "no longer selling tickets" (readiness goes red, new requests stop); the audience already inside must finish this scene before leaving (in-flight requests drain); and the announcement only works if it actually reaches the usher on duty (the signal has to reach PID 1). Ten minutes after closing, security clears the hall anyway (SIGKILL) — and the popcorn ends up on the floor.

98 / 180

The signal leg inside an image is where things most often break. dockerimg's signal argument demonstrates exactly how the shell form of ENTRYPOINT parks /bin/sh on PID 1 and swallows SIGTERM along the way:

99 / 180
Kernel lab
TeaVMDoes SIGTERM actually reach the JVMidle
Switch to signal: frames 2 and 3 contrast the exec and shell forms, frame 6 shows why both grace periods have to line up
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
100 / 180

Chain the three legs together and you get the full shutdown timeline below — from readiness turning red to the process actually exiting, every beat shows who waits for whom, and which beat leaves which symptom when it breaks:

101 / 180
Animation
Animation · The seven-step shutdown relay: from pulling traffic to exit
Animation · The seven-step shutdown relay: from pulling traffic to exit
102 / 180

Rather than reading this chain twice, walk it in the inner-kernel console — these commands play the probe division of labour through to hard-kill versus graceful side by side, and afterwards every line of the config looks familiar:

103 / 180
Console
104 / 180
Section
9. Actuator: the only place probes get configured
105 / 180

Who supplies the probes config from the last section? Actuator. During delivery it is the app's one and only operations interface, and it is also the item security audits love to name.

106 / 180
yaml
management:  server:    port: 9090                     # a separate management port: port 8080 exposes no actuator at all  endpoints:    web:      exposure:        include: health,info,metrics,prometheus   # an allowlist — never write *  endpoint:    health:      show-details: when-authorized  info:    env:      java:        enabled: true
107 / 180
Code
Codejava
// Tag your metrics with business meaning: checkout success rate beats raw QPS@Componentpublic class OrderMetrics {    private final MeterRegistry registry;    public OrderMetrics(MeterRegistry registry) {        this.registry = registry;    }    public void recordCreate(String outcome) {          // outcome: success / stock / validation        registry.counter("beeorder.order.create", "outcome", outcome).increment();    }}
Notes
  • With management.server.port=9090, Nginx only needs to proxy 8080 and Actuator never leaves the intranet
  • include: * is the most common high-severity misconfiguration: /actuator/env leaks the database password, /actuator/heapdump lets a stranger download your entire process memory, keys included
  • show-details: when-authorized keeps the outside world at UP/DOWN while logged-in operators still see which dependency is down
  • Populate /actuator/info from the build plugin's git sha — "which version is production on?" should be one curl away

Analogy: do not hang the OPEN sign while the kitchen is still prepping. Readiness is that sign: it asks "is the food prepped, can we seat people", so a dead database must turn it red (stop seating guests). Liveness asks "does the building still exist" — unless the roof has come down, you should not say the shop died. Wire the database into liveness and a 3-second failover makes you demolish and rebuild all six branches at once: the sign was red, but the shops were fine.

108 / 180
Kernel lab
TeaVMActuator endpoints and health aggregationidle
Start with health to see how the liveness and readiness groups aggregate; then risk to see exactly what /env and /heapdump leak when exposure is *; finish on metrics to see how they get scraped
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
109 / 180

Probes, grace periods, image tags — delivery terminology all sounds alike, yet the difference is entirely "which question does it answer", and one miswire is one incident. Time for a terminology match: click a term on the left, then the question it actually answers on the right:

110 / 180
Match
MatchDelivery terminology: every term answers one questionMatched 0/6 · Missed 0
All six terms live in this article's configuration; each one guards a specific boundary
Pick a card on the left first
111 / 180
Decision
Decisionshould `/actuator/health` count the database and Redis?
112 / 180
Section
10. Nginx reverse proxy configuration
113 / 180

Port 8080 inside the container is never exposed to the public. Nginx handles TLS, compression, real-client-IP forwarding and the upload size cap — none of which the application should do for itself.

114 / 180
Code
Codenginx
server {    listen 80;    server_name api.beeorder.example.com;    # send all HTTP to HTTPS    return 301 https://$host$request_uri;}upstream beeorder {    server app:8080 max_fails=3 fail_timeout=10s;   # one backend on a single host; add lines to scale    keepalive 32;}server {    listen 443 ssl;    server_name api.beeorder.example.com;    ssl_certificate     /etc/nginx/certs/fullchain.pem;    ssl_certificate_key /etc/nginx/certs/privkey.pem;    client_max_body_size 20m;        # upload cap for product images    gzip on;    gzip_types application/json text/plain text/css application/javascript;    gzip_min_length 1024;    location / {        proxy_pass http://beeorder;  # "app" is the Compose service name, resolved by internal DNS        proxy_http_version 1.1;        # forward the real client so app logs stop showing only Nginx's internal IP        proxy_set_header Host              $host;        proxy_set_header X-Real-IP         $remote_addr;        proxy_set_header X-Forwarded-For   $proxy_add_x_forwarded_for;        proxy_set_header X-Forwarded-Proto $scheme;        proxy_connect_timeout 5s;        proxy_read_timeout    30s;        # the backend may be draining right now: retry elsewhere instead of throwing 502        proxy_next_upstream error timeout http_502 http_503;    }    # health probing stays internal, never public    location = /healthz {        proxy_pass http://beeorder/actuator/health/readiness;        access_log off;    }}
Notes
  • proxy_pass http://app:8080 connects by service name via Compose internal DNS — no hardcoded IP
  • The X-Real-IP / X-Forwarded-For / X-Forwarded-Proto trio tells the app the true origin and protocol
  • proxy_next_upstream is the option almost nobody sets, and it is the one that matters during a release: while the old instance drains, Nginx tries a peer instead of handing the user a 502
  • Production also mounts the certificate directory (/etc/nginx/certs) read-only into the Nginx container and renews with certbot or acme.sh
115 / 180

On the application side you must turn on server.forward-headers-strategy=framework, and the reason is plain: by default Spring Boot does not trust these proxy headers. Without it, redirect URLs the app generates come out as http:// instead of https://, request.getRemoteAddr() reports Nginx's internal address, and absolute-URL assembly goes wrong.

116 / 180
Code
Codeyaml
server:  forward-headers-strategy: framework   # make Spring read and trust X-Forwarded-* headers  tomcat:    remoteip:      remote-ip-header: X-Forwarded-For      protocol-header: X-Forwarded-Proto
Notes

Attention: only enable forward-headers-strategy when you are sure all traffic passes through a trusted proxy. If the app's port is exposed to the internet, a client can forge X-Forwarded-For and you inherit IP-spoofing risk — in production this must go hand in hand with Section 7's "bind 127.0.0.1 only".

117 / 180
Section
11. The go-live checklist: hand your memory to the list
118 / 180

Nine out of ten go-live incidents are "one thing that should have been checked got forgotten in the heat of the moment." A checklist's job is to let the list replace your memory.

119 / 180
Animation
Animation · Seven steps to go live
Animation · Seven steps to go live
120 / 180
Table
#CheckPass criterion
1Config externalizedNo environment-specific config baked into the jar or code
2Secret managementDB/Redis/payment secrets come from secrets, never from Git
3JVM flagsMaxRAMPercentage and GC set; no default heap
4Log persistence & rotationContainer log driver capped; app logs roll daily
5Time zoneApp, MySQL and Redis all in the same zone
6Health check/actuator/health/readiness reachable and UP
7Connection poolMax size derived from DB limits and verified under load
8Slow-query switchSlow SQL logging on (1s threshold), traceable in production
9Backup strategyDaily full backup + binlog, with a restore drill
10Rollback planThe last good image tag is known and the rollback was rehearsed
11Monitoring & alertsThreshold alerts on QPS/error rate/P95/pool/disk
12Actuator port isolation/actuator/env, /heapdump and friends never exposed
13DB migrationDDL applied and re-runnable (idempotent scripts)
14Dependencies readyThe app starts only after MySQL and Redis are healthy
15Shutdown gracestop_grace_period > timeout-per-shutdown-phase, with an exec-form entrypoint
16Smoke testOrder → pay → query walked through by hand, once
17Observation windowOn watch for 30 minutes after go-live, ready to roll back
121 / 180
Key point

a checklist is not "the longer the better". What deserves a line is anything you would definitely forget under pressure and whose failure is very expensive — items 15 (grace-period alignment) and 12 (Actuator exposure) are exactly that kind.

122 / 180
Section
12. Canary releases and rollback
123 / 180
Table
StrategyHowProConBest for
Blue-greenTwo environments, switch trafficInstant rollbackDouble the resourcesSingle host / critical systems
RollingReplace instances one by oneResource efficientOld and new coexist brieflyMulti-instance clusters
CanarySend 5% of traffic firstSmallest riskNeeds traffic controlGateways / service meshes
By user cohortWhitelisted accounts get the new build firstReal business validated by nameBiased sampleWhen internal accounts exist
124 / 180
Analogy

open one branch restaurant first. Print the new dish on the menu of all 20 stores and one bad ingredient is a company-wide incident. The clever move is to open one small branch near the university, validate taste, table turnover and plating speed on 5% of your traffic, then spread to the other stores — and if it goes wrong, close that one branch. A canary is not really "in batches"; it is "making the blast radius boundable."

125 / 180

On a single-host Compose setup, the most practical choice is image-tag switching. Record the previous tag; when something breaks, switch back.

126 / 180
Code
Codebash
# Deploy a new version: write the tag into .env, then bring it upecho "IMAGE_TAG=v1.2.0" >> .envdocker compose pull appdocker compose up -d appdocker compose logs -f --tail=100 app     # watch the logs for a while# Something broke — roll back to the previous version (say v1.1.9): only the tag changessed -i 's/^IMAGE_TAG=.*/IMAGE_TAG=v1.1.9/' .envdocker compose pull appdocker compose up -d appcurl -fsS http://127.0.0.1:8080/actuator/health/readiness   # confirm recovery
Notes
  • Rollback is fundamentally "switch the image tag", so every release must keep the previous good tag, and the rollback path must have been rehearsed
  • Changing one line in .env is far less error-prone than editing yml
  • Verify the health endpoint right after rolling back; do not mistake "the command succeeded" for "the service recovered"
  • Schema and code move together: if the new DDL dropped a column, rolling back the image does not roll the database back — which is why #44's migrations must be backward compatible (add before you remove, never change semantics in place, split across two releases)
127 / 180

During a canary, old and new versions run side by side, and the failure you actually get is not a 500 but "the same request landed on two different versions" — so the precondition for canary-ing is that the idempotency key, the state machine and the schema can all hold both versions at once. This lab shows the 5% instance coexisting and being switched back:

128 / 180
Kernel lab
TeaVMCanary: the new instance holding 5% of trafficidle
Switch to gray: watch a latent slow query on the new build keep the error rate inside 5% of traffic, and see why rolling back is only setting the weight back to zero
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
129 / 180
Trap

if healthcheck.interval is too short (say 2s) and start_period is unset, the JVM is declared unhealthy before it finishes booting and the container enters a restart loop. Give start_period enough room (40s and up) before worrying about check frequency.

130 / 180
Section
13. Load testing primer: replace feelings with numbers
131 / 180

wrk is right for quickly getting a "good enough number"; JMeter is right for complex scenarios (parameterization, ramp-up). Which one matters less than the rule that you define the target before you test.

132 / 180

The acceptance targets for BeeOrder's checkout endpoint:

133 / 180
Table
MetricTargetReading
QPS≥ 500Twice the peak order volume, as headroom
P95 latency< 200ms95% of requests return within this window
P99 latency< 500msThe tail must stay controlled
Error rate< 0.1%Any timeout or 5xx fails the run
Stock consistencyNo oversellReconcile after the run; the books must balance
134 / 180

A typical checkout load test (wrk):

135 / 180
bash
# 12 threads, 400 connections, 60 seconds, POST to the checkout endpointwrk -t12 -c400 -d60s --latency \  -s post-order.lua \  http://127.0.0.1:8080/api/orders
136 / 180

Bottleneck hunting follows one fixed path, from the outside in:

137 / 180
  • Application: is the thread pool saturated (Tomcat max-threads), is there lock contention
  • JVM: is GC frequent (jstat -gcutil), is the heap large enough
  • Database: are new slow queries in the log, does EXPLAIN use the index
  • Connection pool: is HikariCP's active connection count pinned at maximumPoolSize
138 / 180
Note

when P95 jumps during a load test, it is usually not that the app is slow, but that it is stuck on the database or the pool. Check the pool's activeConnections first — it is usually the earliest tell.

139 / 180

The "slow request and timeout" path is the one most often misattributed to the application while it is really being dragged down by something downstream. This lab lines all five outcomes (success, validation failure, business error, database failure, slow request) up on the same pipeline; watch where the queue forms and which layer times out first:

140 / 180
Kernel lab
TeaVMOne checkout request across the whole site: where each of five paths endsidle
Run happy first to set a baseline, then slow to see which workstation the request sticks at and who times out first; db is where that Connection is not available row in Section 15 comes from
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
141 / 180
Section
14. Sandbox: how much traffic should this release take?
142 / 180

"How much do I switch over" and "how does the old process leave" are the only two dials in a rollout, and they are the scene of every release incident. This sandbox puts both dials on one panel, and each cell reports the same six numbers for one load test: new-version error rate, P99, in-flight requests severed, corrupted rows created, and time to roll back.

143 / 180
Sandbox
SandboxRollout dial sandbox: how much traffic × how the old process leaves
Result
The new build takes 25 QPS (5% of 500 total)
The latent slow query surfaces: new-version error rate 0.4%, old instances 0%
P99: 1.2s on the new build vs 180ms on the old — only 5% of users are affected
readiness flips red and the orchestrator hands all 25 QPS back to the old instances
The old instances drain 80 in-flight requests, 0 severed
0 corrupted rows · rollback in 4s (weight back to 0)
This is the entire point of a canary: the fault stays inside 5% of traffic, so you have time to see it, judge it and back out. The old instances drained cleanly, not one request cut.
144 / 180
Note

the numbers here are simulated, the conclusions are not. Blast radius is set by the traffic percentage; corrupted data is set by the exit method. They are independent of each other and both must be configured. Canary is not "something only big companies do" — a single-host Compose setup can do 5% with a couple of Nginx upstream weight lines.

145 / 180
Section
15. Common errors, searchable by exact wording
146 / 180

Every "error text" column below can be copied and searched verbatim — do not paraphrase it, do not abbreviate it. The annoying part of the delivery stage is that many failures never surface as a Java exception: the symptom is a log that stops mid-line, or a port that refuses you.

147 / 180
Table
Error text (fragment)SymptomReal cause30-second self-rescue
java.lang.ClassNotFoundException: com.beeorder.BeeOrderApplicationRuns locally after repackage, then cannot find the main class on another machineA fat jar launched as java -cp app.jar <MainClass>: the main class lives in BOOT-INF/classes, not at the jar root, so the system classpath cannot see itJust use java -jar app.jar (it goes through JarLauncher); after a layertools split, launch with java org.springframework.boot.loader.launch.JarLauncher
java.lang.NoClassDefFoundError: org/springframework/... (dies seconds after starting)The class itself was found, but something it depends on is not on the classpathThe fat jar's dependencies sit in BOOT-INF/lib and are read by LaunchedClassLoader; unzip-and-rezip by hand, or a CI cache holding a stale jar, destroys that structureNever hand-edit the package: mvn -B clean package again, then confirm Main-Class really is JarLauncher with unzip -p app.jar META-INF/MANIFEST.MF
docker: 'java' is not a command (ENTRYPOINT ["java","-jar","app.jar"] will not start)The container exits immediately and the log is a single lineThe base image has no Java at all (distroless/base, or a bare alpine), or ENTRYPOINT ["java","-jar"] was written while the binary is not on PATHCheck with docker run --rm --entrypoint sh <image> -c 'command -v java'; switch the base to eclipse-temurin:21-jre-jammy, or stay on distroless's -java variant and give the absolute path /opt/java/openjdk/bin/java
java.lang.OutOfMemoryError: MetaspaceThe process kills itself after a while, or after some hot reloads, with a full Java stack in the logClass-loader leakage: repeated reloads, dynamic class generation from Groovy/script engines, or a -XX:MaxMetaspaceSize below real need (common once a container limit was lowered and the flag was not)Cap it explicitly with -XX:MaxMetaspaceSize=256m and add -XX:+HeapDumpOnOutOfMemoryError; check jcmd <pid> GC.class_stats, and confirm whether each hot reload grows the class set
java.lang.OutOfMemoryError: Java heap spaceReported by the JVM itself, and you can capture an hprofThe heap really is full: an unbounded collection, a query returning a million rows, a cache with no ceilingAdd -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/tmp and read the dominator tree in MAT; this failure never leaves an exit 137, so first work out who killed the process
The container log ends with Killed and Java threw no exception at allThe process was executed from outside: no scene left behind, no stackTotal usage (heap + metaspace + thread stacks + direct memory) crossed the cgroup limit and the kernel killed it; writing -Xmx equal to the container limit is the usual triggerReplace -Xmx with -XX:MaxRAMPercentage=75.0; confirm with docker inspect --format '{{.State.ExitCode}} {{.State.OOMKilled}}'; inspect off-heap with NMT plus jcmd <pid> VM.native_memory
OOMKilled plus Exit Code 137 (kubectl describe pod)Repeated restarts, business log cut off mid-line137 = 128 + 9, i.e. SIGKILL. The cgroup killed it for total usage — a different question from whether the heap was large enoughDo the whole-budget sum first: heap + ~200 Tomcat worker threads (~200MB of stacks) + CodeCache + Netty direct memory; only then discuss the percentage, and do not rush to raise the limit just to hide it
Web server failed to start. Port 8080 was already in use.Local startup exits immediatelyThe previous instance never exited, the IDE is still running, or something else owns 8080. Note that the container flavour of this message differs: Bind for 0.0.0.0:8080 failed: port is already allocatedWindows: `netstat -ano \findstr :8080 then taskkill /PID <pid> /F; Linux/macOS: lsof -i:8080; in development just use -Dserver.port=8081`, in production change the host-side port in Compose
Healthcheck failed: Read timed out, or the service sits unhealthy in ComposeThe deploy script hangs or the pipeline goes red, yet a manual curl works fineJVM cold start plus slow init (retrying an unreachable dependency) exceeds start_period; or timeout: 5s is below the real duration of one indicator; or the curl binary the healthcheck needs is not in the image at allRaise start_period to 40s and above, timeout to 5-10s; point the check at readiness rather than the aggregate /health; and run the exact test command yourself via --entrypoint sh inside the container
Client sees Connection refusedNot one request got in; fails instantlyNobody is listening on that port: the container never started, EXPOSE was written but not -p, or the app already exited on a startup failure. This is nobody answering at the network layer, unrelated to your business logicdocker ps and look for 0.0.0.0:8080->8080/tcp in the PORTS column; curl -v http://127.0.0.1:8080/actuator/health/readiness and read it stage by stage; docker logs to prove the process is alive
Client sees Connection reset by peer, or intermittent 502The request got in and was then cut off; retrying usually succeedsThe peer exited mid-request: an old instance was SIGKILLed before its drain finished, Nginx still holds a removed instance in upstream, or the backend closed the keep-alive connection firstIf it only happens during a release, add server.shutdown=graceful and stop_grace_period; add proxy_next_upstream error timeout http_502 to Nginx; sleep 2-5s in preStop so traffic propagation stops before you begin closing
Connection is not available, request timed out after 30000msEverything turns 500 as soon as load arrives, log full of thisThe pool is at its default 10 connections while transactions wait on external calls (long transactions), or a slow query is holding every connectionFirst check for "a remote call inside a transaction" (#45 Section 3); then derive maximum-pool-size from the database's max_connections divided by instance count — do not just jump to 200
UnsupportedClassVersionError: ... has been compiled by a more recent version of the Java Runtime (class file version 65.0)Explodes on the server's very first lineBuilt on JDK 21 locally (class 65), the server runs JRE 17 (class 61) — a textbook case of a build precondition that was never written downPin FROM eclipse-temurin:21-jre in the image; pin <java.version>21</java.version> and maven.compiler.release in pom.xml so CI owns the version, not the machine
148 / 180
Tip

three rows of this table (Killed, OOMKilled, Connection reset) never appear as a Java exception — the symptom is the log simply stopping. The first command for container triage is docker inspect <id> --format '{{.State.ExitCode}} {{.State.OOMKilled}} {{.State.Status}}', which tells you who killed it in one glance.

149 / 180

The ClassNotFoundException in row 1 deserves a scene of its own — what makes it confusing is that the build was green and the jar is right there in your hand, yet the message reads as if the class was never packaged. Do not look at the answer yet; click the frame you believe is the culprit:

150 / 180
Triage
Error triageClassNotFoundException: com.beeorder.BeeOrderApplication
The fat jar is right there, yet the main class cannot be found

CI is green and the image is pushed. An operator copies the build artifact to the server and launches it with a hand-written script: java -cp beeorder.jar com.beeorder.BeeOrderApplication — not one line of your code runs, and the JVM throws ClassNotFoundException.

Error: Could not find or load main class com.beeorder.BeeOrderApplication
Caused by: java.lang.ClassNotFoundException: com.beeorder.BeeOrderApplication
at java.base/jdk.internal.loader.BuiltinClassLoader.loadClass(BuiltinClassLoader.java:641)
at java.base/jdk.internal.loader.ClassLoaders$AppClassLoader.loadClass(ClassLoaders.java:188)
at java.base/java.lang.ClassLoader.loadClass(ClassLoader.java:525)
at com.acme.ops.LaunchScript.main(LaunchScript.java:9)
Click the frame you blame — guessing is allowed
No pressure: guess the exception first, then which line actually made the call.
151 / 180
Section
16. Check yourself
152 / 180

One warm-up first, testing the decision card in Section 9 and the probe configuration in Section 8:

153 / 180
Quiz
Check yourselfAfter a 3-second MySQL primary-replica failover, K8s restarted all 6 Pods and the rolling release stalled on health checks. The config shows the liveness probe pointed at the default /actuator/health. Which change fixes the root cause?
Pick one — you get feedback right away
154 / 180

Then a combined question tying Sections 4, 6 and 15 together:

155 / 180
Quiz
Check yourselfBeeOrder was packaged as app.war. A colleague boots it with java -jar app.war and it works; you copy the same file into an external Tomcat's webapps directory and get 404. Which statement about these two ways of running it is correct?
Pick one — you get feedback right away
156 / 180
Section
17. Practice in three levels
157 / 180
Section
Level 1 · Follow the steps: run one shutdown that actually drains
158 / 180

Goal: confirm with your own eyes that graceful shutdown works — not that it is "configured", but that it is visible in the log. You need one machine with Docker and a BeeOrder image that starts.

159 / 180

Step 1, bring the stack up with Section 7's Compose file and confirm it is healthy:

160 / 180
bash
docker compose up -ddocker inspect --format '{{.State.Health.Status}}' "$(docker compose ps -q app)"# expected: healthy   (if starting, wait out the 40s start_period and look again)
161 / 180

Step 2, generate some in-flight traffic, then stop:

162 / 180
bash
# terminal A: keep hitting the service, recording status and durationwhile true; do curl -s -o /dev/null -w '%{http_code} %{time_total}s\n' \  http://127.0.0.1:8080/api/orders/1; sleep 0.05; done# terminal B: give it an explicit 40-second grace perioddocker compose stop --timeout 40 app
163 / 180

Step 3, check that these three lines appear in order (if any is missing, go back to Section 8):

164 / 180
text
Commencing graceful shutdown. Waiting for active requests to completegraceful shutdown complete... Tomcat shutdown ...
165 / 180
Table
What you watchExpectedIf it is not, first check
Terminal A's status codes200 all the way, then Connection refused only at the very endAny 5xx/reset → the ENTRYPOINT is not in exec form
Commencing graceful shutdown in the logPresentAbsent → server.shutdown=graceful never applied, or the signal was swallowed by a shell
Time to exitAs long as the longest in-flight request, and under 40sExactly 10s → the grace period never reached the daemon (--timeout omitted)
166 / 180
Section
Level 2 · Variation: act the two probes out
167 / 180

Goal: use one deliberate database outage to watch readiness go red while liveness stays green.

168 / 180
bash
# 1. put the db indicator in the readiness group (Section 8's config), then restart the appdocker compose restart app# 2. record what both probes report right nowcurl -s http://127.0.0.1:9090/actuator/health/liveness  | python -m json.toolcurl -s http://127.0.0.1:9090/actuator/health/readiness | python -m json.tool# 3. take the database down on purpose and look again: liveness should stay UP, readiness go DOWNdocker compose stop mysqlsleep 20curl -s http://127.0.0.1:9090/actuator/health/readiness   # expect "status":"DOWN"curl -s http://127.0.0.1:9090/actuator/health/liveness    # expect "status":"UP"# 4. bring the database back and confirm readiness recovers on its own — that is "pull traffic, do not kill"docker compose start mysql
169 / 180

If liveness also went red in step 3, you pointed it at the aggregate /actuator/health. Go back to Section 9's group block.

170 / 180
Section
Level 3 · Open-ended: run a 30-minute release drill for BeeOrder
171 / 180

Write a plan no longer than one page that answers these five questions, each with a concrete "how I will verify it":

172 / 180
  1. How long does rollback take, and who presses the button? (Hint: switch the tag only, target under 60s.)
  2. During the drill you deliberately inject "a slow query on the new build at 5% traffic" — how do you spot it within five minutes? (Hint: the first two cells of Section 14's sandbox, and a version tag on every metric.)
  3. During the canary, one payment callback hits the old version and then the new one. Is idempotency still intact, and on what basis? (Hint: #45 Section 5's unique-index backstop.)
  4. Is the DDL forward-compatible? When you roll the image back, does the database come with it? (Hint: checklist item 13 — add before you remove, split across two releases.)
  5. Which three numbers are your only evidence for "widen the canary"? (Suggested: error rate, P99, pool active connections — the same three as Section 12's table.)
173 / 180

The acceptance criterion is a single sentence: the drill must actually press the rollback button, not "I know how". A rollback plan that has never been executed is not a rollback plan.

174 / 180
Section
18. Self-check: the six hard conclusions of this article
175 / 180
  • Delivery = writing the hidden preconditions down: build premises into the Dockerfile, runtime premises into Compose, process premises into branch protection and needs
  • Layer tests by risk, not by count: units for branches, slices for contracts, integration for SQL and concurrency; oversell can only be caught concurrently, and coverage never substitutes for assertion strength
  • What you ship is an image tag, not a jar and not source: no CI-verified sha, no deployment; rollback is switching the tag
  • Graceful shutdown is a three-leg relay: pull traffic → drain → exit, both grace periods aligned, and the ENTRYPOINT in exec form
  • Liveness carries no dependencies; readiness does: reversed, one database blip becomes a site-wide restart storm
  • A rollout has exactly two dials: the traffic percentage sets your blast radius, the exit method sets how much corrupted data you create — configure both
176 / 180

Looking back, BeeOrder's go-live is not the credit of any single article:

177 / 180
  • In #1, installing the JDK felt like installing software; it was really about the "bytecode + virtual machine" abstraction
  • Learning IoC/AOP in #6-#14 is what lets #45's checkout service gain transaction power from one annotation
  • The configuration and REST contract from #21-#25 are the real groundwork for #44's API contract
  • The connection pool in #28, transaction internals in #31 and caching in #32 all resurface in Section 13's load test and Section 11's checklist
  • The testing and deployment of #33-#37, plus this article's CI/CD, upgrade "it runs" into "I dare ship it"
178 / 180

One last question from #45 closes the series: checkout and stock deduction must share one transaction (both succeed or both roll back), writing the stock ledger shares it too (it is part of the books), but sending notifications and adding points must commit independently — a failed notification must not roll back the order. That boundary judgement runs through every trade-off from #31's transaction propagation, through #45's idempotency design, to this article's canary and rollback.

179 / 180
Decision
Decisionshould a small team starting a monolith like BeeOrder adopt Kubernetes from day one?
180 / 180
Summary

this delivery article closes the loop on the entire series. Remember three sentences: layered tests so that every class of risk has a keeper; a CI pipeline so that gatekeeping does not rely on human discipline; production orchestration, probes and a checklist so that "it runs" becomes "rollback-able, observable and trustworthy." From java -version in #1 to a single docker compose up -d here, what you ship is no longer a piece of code but a system someone else can take over, a machine can verify, and a bad day can roll back — that is what "going live" really means for a backend engineer.