Async Tasks, Thread Pools and Scheduling

bee2026-10-0862 min read0 views
Why is @Async bad by default? Custom executors, rejection policies, context propagation, graceful shutdown — then @Scheduled's three modes and duplicate execution in a cluster.
1 / 168
Section
0. The 30-second version
2 / 168

This article solves one problem: some work should not make the user wait. Writing the order to the database is the main flow — the user must wait for it. Sending the SMS, the email, recalculating points — none of that affects "the order succeeded", so move it to the background. Spring gives you two annotations: @Async (someone else runs it, I do not wait) and @Scheduled (run it when the clock says so).

3 / 168

But neither annotation creates a thread. They only hand the task over. What decides whether you survive is the thing that catches it — the thread pool. That is why half of this article is about pools: it is not a detour, it is the point.

4 / 168

Six terms, one line each (used throughout):

5 / 168
  • Thread: one line of execution. Your request normally runs on a Tomcat worker; going async means starting another line to do the work
  • Thread pool: a set of threads built in advance and reused, like a rostered crew; tasks are jobs, whoever is free picks them up
  • Core threads (corePoolSize): the permanent members of the pool — they stay even when idle
  • Work queue (queueCapacity): when every permanent member is busy, new tasks wait here for a free hand
  • Max threads (maxPoolSize): only once the queue is full does the pool hire temporary extra hands, up to this number
  • Rejection policy (RejectedExecutionHandler): what to do when both crew and queue are full — throw, make the submitter run it, drop it, or drop the oldest
6 / 168
类比|Analogy

pin the whole article on a food-delivery dispatch station. Core threads are the station's staffed riders (a fixed few, always on call). The work queue is the row of pickup tickets hanging on the wall during an order surge (jobs wait there until a rider returns). Max threads are the crowd-sourced riders you call in at peak — they only appear once the wall is full, and go home when it quiets down. The rejection policy is the station's four possible rules when it truly cannot take more: tell the customer "order refused" (throw), have the shop clerk ride it out themselves (the caller runs it), quietly tear the ticket up (discard), or cancel the oldest ticket and keep the newest (discard-oldest). And scheduled tasks? They are orders that open automatically at a fixed time every day: if the station has only one rider on shift and he takes a slow delivery, every later timed order waits with him — that is the entire truth behind Section 9's trap.

7 / 168
Diagram
Figure · Default executor versus a custom pool
Figure · Default executor versus a custom pool
8 / 168

That comparison is the first life-or-death line of this article: if @Async gets no configured pool, you are in the left column (a fresh new Thread() per task, no bound, no reuse) — Section 3 shows the source evidence. The animation below is how the right column actually behaves internally, and note the part almost everyone remembers backwards — when core threads are busy, tasks are enqueued first, not answered with new threads:

9 / 168
Animation
Animation · The four stages of submitting a task (queue first, then grow)
Animation · The four stages of submitting a task (queue first, then grow)
10 / 168

After this article you should be able to answer three questions:

11 / 168
  • A @Async method throws an exception — why does neither your log nor the HTTP response show it, and what must you add to catch it?
  • Calling your own @Async method via this.sendSms() inside the same class — why does it run synchronously?
  • You deploy three replicas — why did the daily reconciliation job run three times, and how do you make it run once?
12 / 168
Section
1. Start from one timed-out request
13 / 168

Production reported that "the order endpoint sometimes takes 8 seconds to return". Investigation found that right after writing to the database, the main flow also called two external services synchronously — sending an SMS and sending an email. Each takes 1–3 seconds, and neither has anything to do with the outcome: the user only needs to know the order succeeded; an SMS two seconds late is fine.

14 / 168
Table
ApproachMain-flow timeUXImpact of SMS failure
All synchronous1.5s + 2s + 3s ≈ 6.5sUsers want to close the pageAn SMS glitch fails the whole order
SMS / email async≈ 1.5sInstantFailure is just logged; the order stands
15 / 168

This is exactly what async tasks solve: move "unrelated to the main result and slow" work onto a background thread and return immediately. Spring offers two annotations: @Async (run asynchronously) and @Scheduled (run on a schedule). They look trivial and hide many traps; this article covers both from usage through thread pools, context propagation and duplicate execution in a cluster.

16 / 168
Diagram
Figure 1 · Two execution models
Figure 1 · Two execution models
17 / 168
Section
2. @Async in two lines
18 / 168

First add @EnableAsync to a config class, then mark the method @Async:

19 / 168
Code
Codejava
@Configuration@EnableAsync                       // turn on async supportpublic class AsyncConfig { }@Servicepublic class OrderService {    private final SmsClient smsClient;    public OrderService(SmsClient smsClient) {        this.smsClient = smsClient;    }    public Long createOrder(OrderCmd cmd) {        Long orderId = doCreate(cmd);          // main flow: the DB write must be synchronous        sendNotifyAsync(orderId, cmd.getPhone());  // notification: async, does not block        return orderId;                        // return right away    }    @Async                                 // this runs on another thread    public void sendNotifyAsync(Long orderId, String phone) {        smsClient.send(phone, "order placed: " + orderId);    }}
Notes
  • @EnableAsync is the master switch; without it @Async does nothing.
  • Spring generates a proxy for the annotated method and hands the body to a thread pool.
20 / 168

The return type decides what the caller can get — the easiest part to misuse:

21 / 168
Table
Return typeCaller can...Best for
voidIgnore the result; exceptions never come backFire-and-forget (SMS, logging)
Future<T>get() and blockLegacy, not recommended
CompletableFuture<T>Compose, call back, time outFirst choice when orchestrating results
22 / 168
Animation
Animation · Life of an async task
Animation · Life of an async task
23 / 168

Two annotations switch async on, but the part that actually ships is a block of configuration. Do not copy somebody else's yml — tick the options you need, generate the file, and read why each line exists and what breaks without it:

24 / 168
Generator
GeneratorConfigure async and scheduling in one passapplication.yml2 / 4
Tick only the async pool first to see the minimum usable config, then add the server and logging groups and map each line back to the shutdown chain in Section 7
Output
server:
  port: 8080
  servlet:
    encoding: { charset: UTF-8, enabled: true, force: true }
  compression: { enabled: true, min-response-size: 2048 }

spring:
  application:
    name: demo-service
  task:
    execution:
      pool: { core-size: 8, max-size: 32, queue-capacity: 200 }
  threads:
    virtual:
      enabled: false                     # JDK 21 打开后 @Async 走虚拟线程
Why each choice matters
serverserver.port loses to --server.port=8081 on the command line and to the SERVER_PORT env var.
taskspring.threads.virtual.enabled=true is what puts @Async on virtual threads (Boot 3.2+, JDK 21).
25 / 168
Section
3. Why you must define your own pool
26 / 168

@Async defaults to SimpleAsyncTaskExecutor, whose name is lovely and whose source is alarming: it calls new Thread() for every task, never reuses, and has no upper bound.

27 / 168
java
// Simplified: the core behavior of SimpleAsyncTaskExecutorprotected void doExecute(Runnable task) {    Thread thread = createThread(task, this.threadNamePrefix);   // new thread every time    thread.start();}
28 / 168

The consequences:

29 / 168
  • No concurrency limit: ten thousand requests spawn ten thousand threads and blow up CPU and memory, possibly OOM.
  • No reuse: every task is born and dies, wasting creation/destruction cost.
  • No observability: no queue, no rejection policy, no metrics.
30 / 168

So production must define a custom pool. Use ThreadPoolTaskExecutor:

31 / 168
Code
Codejava
@Configuration@EnableAsyncpublic class AsyncConfig implements AsyncConfigurer {    @Bean("notifyExecutor")    public ThreadPoolTaskExecutor notifyExecutor() {        ThreadPoolTaskExecutor executor = new ThreadPoolTaskExecutor();        executor.setCorePoolSize(8);                 // core threads: always alive        executor.setMaxPoolSize(16);                 // max threads: temporary peak capacity        executor.setQueueCapacity(200);              // queue: buffer while cores are busy        executor.setKeepAliveSeconds(60);            // idle above-core threads are reclaimed        executor.setThreadNamePrefix("notify-");     // business prefix for diagnosis        executor.setRejectedExecutionHandler(new ThreadPoolExecutor.CallerRunsPolicy());        executor.setWaitForTasksToCompleteOnShutdown(true);   // graceful, see Section 7        executor.setAwaitTerminationSeconds(30);        executor.initialize();        return executor;    }    /** Executor used when @Async names no pool */    @Override    public Executor getAsyncExecutor() {        return notifyExecutor();    }}
Notes
  • @Async("notifyExecutor") picks a specific pool; give different businesses their own pools so they don't interfere.
  • The thread-name prefix is a lifesaver: seeing notify-3 in logs immediately tells you it's an async notification thread.
32 / 168
Section
4. Pool parameters and rejection policies
33 / 168

The pool's working order is: core threads → queue → max threads → rejection policy. Many people wrongly assume "more tasks means new threads up to max":

34 / 168
Table
StageTriggerBehavior
1threads < coreCreate a core thread
2threads = core and queue not fullEnqueue the task
3queue full and threads < maxCreate a temporary thread
4queue full and threads = maxApply the rejection policy
35 / 168

Those four rows are literally the script of the animation below; watch frames ②→③ in particular: what separates them is "the queue is full", not "the core threads are busy" — once that lands, every later discussion about capacity stops being confusing.

36 / 168
Animation
Animation · The four stages of submitting a task (queue first, then grow)
Animation · The four stages of submitting a task (queue first, then grow)
37 / 168
Trap

an unbounded queue makes maxPoolSize permanently useless. Set queue capacity to Integer.MAX_VALUE (the default for LinkedBlockingQueue) and tasks pile up forever, never reaching the "queue full → grow to max" step; max threads is a phantom and memory eventually blows up. Always give the queue a finite capacity.

38 / 168

The four rejection policies:

39 / 168
Table
PolicyBehaviorBest for
AbortPolicy (default)Throws RejectedExecutionExceptionFail fast and let the caller see it
CallerRunsPolicyThe submitting thread runs the task itselfDegrade to synchronous; don't lose tasks
DiscardPolicySilently dropsTasks you can afford to lose
DiscardOldestPolicyDrops the oldest queued task, then retriesOnly the newest data matters
40 / 168

The table above writes those four stages as four lines of prose, but "which one happens first" is exactly the thing prose cannot teach. Every cell below is clickable, and clicking it tells you what that stage does:

41 / 168
Diagram
FlowThe four stages of submitting a task (click through them)1 / 5
Go from ① to ⑤ and watch stage ②: the gate is "the queue is full", not "the core threads are busy"
→
→
→
→
① Threads < core
Every task spawns a new core thread until corePoolSize is reached. No surprises here — and this is the stage people wrongly assume keeps going.
All clearYou only have two levers: the queue decides how long work waits, max decides how many run at once.
42 / 168

To feel how those three numbers fight each other, drag the queue capacity. The same @Async method has a completely different fate between 0 and 20000:

43 / 168
Tuner
TunerDrag the queue capacity, watch the behavior change
spring.task.execution.pool.queue-capacity
100tasksNow 0 – 20000
Small queue: absorb a short burst, then add threads
  • A brief spike is buffered without spawning extra threads
  • Once the queue fills, max finally gets used — that is its only real value
  • Queueing delay has a bound you can predict and alert on
  • The capacity most production services actually want
Shrink the queue if you want work handled sooner, cap it if you want the process to survive — you have to manage both ends.
44 / 168
Section
5. Every way @Async can fail
45 / 168

Like @Transactional, @Async depends on a proxy; if the call bypasses the proxy, the annotation vanishes:

46 / 168
Table
FailureCauseFix
Self-invocationthis.method() skips the proxyMove it to another bean, or inject self
private / final methodCannot be proxied/overriddenMake it public, non-final
Primitive return typeAsync can only wrap void/FutureUse CompletableFuture<T> or void
void method swallows exceptionsNo one receives the exceptionConfigure an AsyncUncaughtExceptionHandler
Not a Spring beanA newed object has no proxyLet the container manage it
Missing @EnableAsyncMaster switch offAdd the annotation
47 / 168

Exceptions from void methods disappear silently by default; catch them explicitly:

48 / 168
Code
Codejava
@Configuration@EnableAsyncpublic class AsyncConfig implements AsyncConfigurer {    @Override    public AsyncUncaughtExceptionHandler getAsyncUncaughtExceptionHandler() {        return (ex, method, params) ->            log.error("async task failed method={} params={}", method.getName(), params, ex);    }}
Notes
  • With this, exceptions from void async methods at least reach the logs instead of evaporating.
49 / 168

That table has six rows and nobody memorises six rows. An annotation silently not working is something you only remember after it has bitten you once. So play a round: click a symptom on the left, then the cause you believe fits — miss one and it explains itself right there.

50 / 168
Match
Match@Async failures: pair the symptom with the causeMatched 0/6 · Missed 0
Every symptom on the left comes from a real ticket; the right column holds the causes
Pick a card on the left first
51 / 168
Section
6. Context propagation: ThreadLocal does not cross threads
52 / 168

Recall Section 3's core fact: async tasks run on another thread. Yet many Spring / SLF4J / transaction contexts live in ThreadLocal, which does not travel automatically:

53 / 168
  • Lost MDC traceId: the main thread's logs carry a traceId while the async thread's don't, breaking the trace. Fix with a TaskDecorator that copies the MDC when submitting:
54 / 168
Code
Codejava
public class MdcTaskDecorator implements TaskDecorator {    @Override    public Runnable decorate(Runnable task) {        Map<String, String> context = MDC.getCopyOfContextMap();   // capture the submitting thread's MDC        return () -> {            if (context != null) MDC.setContextMap(context);       // restore on the worker thread            try {                task.run();            } finally {                MDC.clear();                                        // avoid polluting reused threads            }        };    }}// config: executor.setTaskDecorator(new MdcTaskDecorator());
Notes
  • Lost SecurityContext: the async method cannot read the current user. Wrap the pool in a DelegatingSecurityContextExecutor, which carries the security context automatically.
  • Transaction context cannot cross threads (the key point): a transaction's connection is bound to the ThreadLocal of the thread that began it. An async method runs on another thread and cannot get the caller's transaction; it either has none or opens its own independent one via its own @Transactional.

Key point: remember three boundaries — carry MDC with a TaskDecorator, carry SecurityContext with DelegatingSecurityContextExecutor, and never try to carry a transaction. Especially the last: don't expect "if the main flow rolls back, the async task rolls back too" — they are two independent transactions by design.

55 / 168

Those are the conclusions, but "why it breaks" is not something conclusions teach. The lab below really runs it in your browser: the same handful of lines, read once on exec-3 and once on task-1.

56 / 168
Kernel lab
TeaVMWhy context cannot cross the threadidle
Start on "Child reads null" to see which link snaps, then switch to "Ship with the task" for the fix, and finish on "Request in async" to watch the framework throw on purpose
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
57 / 168

Now spread the same story out as a debug session. The left pane is the code you are stepping through; the right pane refreshes the variables and the call stack. Press Step five times and watch the exact moment the thread name changes from exec-3 to task-1:

58 / 168
Stepper
StepperStepping through it: where user 42 gets lost1 / 6
Press Step five times and watch step 4 change the thread
Code under debug
1UserContext.set(userId); // we are on http-nio-8080-exec-3
2notifier.sendAsync(order); // this hop goes through the proxy
3executor.submit(wrap(order)); // handed to the pool; caller returns at once
4// ---- camera cuts to task-1 ----
5Long uid = UserContext.get(); // same ThreadLocal, different map
6log.info("send to {}", uid); // uid = null
Variables now
threadhttp-nio-8080-exec-3
UserContext42
singletonObjects184 beans
Call stack
1OrderController.pay
2- request thread
1An interceptor put the current user into this thread's own ThreadLocalMap. Note the shape of the storage: one map per thread, not one global table.
59 / 168

Those "through the proxy" and "handed to the pool" moves from steps ② and ③ deserve their own look — the lab below prints every move AsyncExecutionInterceptor makes:

60 / 168
Kernel lab
TeaVMHow a task actually leaves the calleridle
Pick Submit to see the full path from proxy to executor; Queue then explains the exact band you dragged to in the tuner above
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
61 / 168

The correct way to move context, drawn as one picture — all four moves matter, and the fourth (giving it back) is the one people skip:

62 / 168
Diagram
Figure · How context crosses the thread boundary
Figure · How context crosses the thread boundary
63 / 168
Section
7. Graceful shutdown: don't lose queued tasks
64 / 168

When the app receives a shutdown signal, killing the process outright loses every async task still in the queue. Two settings together form the full chain:

65 / 168
yaml
server:  shutdown: graceful        # 1. Web layer: stop accepting new requests, finish in-flight onesspring:  lifecycle:    timeout-per-shutdown-phase: 30s   # 2. time budget for in-flight requests
66 / 168
Code
Codejava
// 3. Pool layer: drain the queue before closingexecutor.setWaitForTasksToCompleteOnShutdown(true);   // process the queue on shutdownexecutor.setAwaitTerminationSeconds(30);              // wait up to 30s, then force
Notes
  • The first setting covers HTTP requests, the third covers async tasks — both are required.
  • Keep the wait shorter than your orchestrator's terminationGracePeriodSeconds (K8s defaults to 30s), or the container is killed and the wait is pointless.
67 / 168
Section
8. @Scheduled: three scheduling modes
68 / 168

Enable with @EnableScheduling, then annotate methods with @Scheduled. It has three triggers with very different semantics:

69 / 168
Table
ModeSemanticsWhen one run exceeds the interval
fixedRateFixed rate from the start of the previous runNo overlap, but it catches up immediately; tasks pile up
fixedDelayFixed time after the previous run endsNever piles up; drifts naturally
cronFires at specific points in timeSame cron semantics; watch the multi-instance issue
70 / 168

Timeline, with a 3-second task and a 2-second interval:

71 / 168
Code
Code
fixedRate:  start0  ——3s——  start3 (0+2 already due, run now) ——6s——fixedDelay: start0  ——3s——  wait 2s  start5  ——8s——  wait 2s  start10
Notes
  • fixedRate's rate is measured from the start, so long tasks cause a backlog (no overlap, but back-to-back runs).
  • fixedDelay measures from the end, so it never backs up — usually the intuitive choice.
72 / 168

A cron expression has six fields: second minute hour day-of-month month day-of-week (Java's cron adds the "second" field that Linux's lacks):

73 / 168
Table
ExpressionMeaning
0 0 3 ?Every day at 03:00
0 /5 ?Every 5 minutes
0 0 9-18 MON-FRIHourly on the hour, 9–18, weekdays
0 30 2 1 * ?Monthly on the 1st at 02:30
0 0/30 * ?Every 30 minutes
74 / 168

initialDelay postpones the first run, handy for "wait until the app is fully ready":

75 / 168
java
@Scheduled(initialDelay = 60_000, fixedDelay = 300_000)public void cleanExpiredTokens() {    // first run 60s after startup, then every 5 minutes (measured from the end)}
76 / 168
Section
9. The scheduler's single thread
77 / 168

@Scheduled uses a single-threaded TaskScheduler by default. That means all scheduled tasks share one thread, and whoever runs longest blocks the rest.

78 / 168
Trap

a reporting task that runs every 5 minutes but takes 4 minutes makes every other task "late", because they all queue on the same thread.

79 / 168

Fix it with a custom multi-threaded TaskScheduler:

80 / 168
Code
Codejava
@Configuration@EnableSchedulingpublic class ScheduleConfig {    @Bean    public TaskScheduler taskScheduler() {        ThreadPoolTaskScheduler scheduler = new ThreadPoolTaskScheduler();        scheduler.setPoolSize(4);                        // 4 threads, tasks don't block each other        scheduler.setThreadNamePrefix("schedule-");        scheduler.setWaitForTasksToCompleteOnShutdown(true);        scheduler.setAwaitTerminationSeconds(30);        return scheduler;    }}
Notes
  • Size the pool at least as large as the number of tasks you expect to run in parallel, or you're back to blocking.
81 / 168

The animation below plays this trap out frame by frame — pay attention to step ③: task B has reached its trigger time and nobody runs it. It did not fail; it is waiting for the one and only thread.

82 / 168
Animation
Animation · Scheduled tasks share one thread by default
Animation · Scheduled tasks share one thread by default
83 / 168

Dragging this number around makes what "1" costs obvious:

84 / 168
Tuner
TunerHow many threads does the scheduler get
spring.task.scheduling.pool.size
1threadsNow 1 – 16
Everything serial: one slow job makes all of them late
  • The default is exactly 1 — every @Scheduled shares scheduling-1
  • One slow HTTP call inside a job pushes back every later job
  • It looks like "the job sometimes does not run", but it ran late, and the thread name in the log is always the same
  • Your cron expression is fine; the problem is there is only one thread
Jobs held up88%
Thread utilisation100%
Count the long-running jobs first, then choose the thread count — never the other way round.
85 / 168
Section
10. Scheduled tasks in a cluster: the duplicate-run nightmare
86 / 168

Everything above assumes one instance. Deploy several replicas and every replica fires the same scheduled task: a 5-minute report runs three times, duplicate bills go out. Three solutions:

87 / 168
Table
SolutionMechanismPros/cons
DB optimistic lockUpdate with a version; only one instance succeedsSimple, but DB pressure and coarse grain
Redis distributed lockSET key value NX EX ttl to claim the right to runLight and fast; must handle expiry/renewal
Scheduler center (XXL-JOB)A central service dispatches to a chosen executorFull-featured and observable, extra ops cost
88 / 168

A lightweight Redis SETNX de-dup implementation:

89 / 168
Code
Codejava
@Componentpublic class ReportJob {    private final StringRedisTemplate redis;    public ReportJob(StringRedisTemplate redis) {        this.redis = redis;    }    @Scheduled(cron = "0 0/5 * * * ?")    public void generateReport() {        // SETNX with an expiry: only one instance wins the lock        String lockKey = "lock:report:" + ZonedDateTime.now().format(DateTimeFormatter.ofPattern("yyyyMMddHHmm"));        Boolean got = redis.opsForValue()                .setIfAbsent(lockKey, "1", Duration.ofMinutes(4));   // TTL slightly less than the interval        if (!Boolean.TRUE.equals(got)) {            log.debug("report task already running elsewhere, skipping");            return;                         // didn't win the lock, do nothing        }        try {            doGenerateReport();             // do the real work        } finally {            // The TTL releases it automatically; don't delete here, or you might remove a newer lock        }    }}
Notes
  • The lock key includes the time granularity (to the minute) so a given point fires once, while the next point gets a fresh key.
  • The TTL must be slightly less than the interval, or a leftover lock blocks the next round.
  • If a run can exceed the TTL, you must add renewal (a watchdog), or early expiry causes concurrent runs.
90 / 168
Kernel lab
TeaVMThreads versus transaction boundaries: a controlled experimentidle
Think: does an @Async method see the caller's transaction? Use propagation to see the boundary
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
91 / 168
Section
11. Put the pool in your hands: four labs you can press
92 / 168

Sections 1-10 were reading material; this section is buttons. First the analogy that should stay with you throughout:

93 / 168
类比|Analogy

why do exceptions from @Async always seem to vanish? Picture a parcel you mailed that got lost in transit — nobody phones you to say so. The async worker thread is the recipient: after the main flow hands the parcel over it walks away immediately, leaving no phone number (the return type is void) and no delivery receipt (no Future). If the parcel burns halfway (an exception is thrown), apart from the courier's internal record (your log), nobody tells you. So Section 5's "you must configure an AsyncUncaughtExceptionHandler" is not a style preference — it is the only way to install that missing phone call.

94 / 168
Section
11.1 Submit, pile up, reject: one curve through all four stages
95 / 168

Section 4's table says core → queue → max → reject, but "enqueue before growing" is counter-intuitive enough that almost nobody gets it right on the first try. Press it out with a lab:

96 / 168
Kernel lab
TeaVMThe four stages of submitting a taskidle
Cycle submit / queue / reject: submit shows resident threads being created one by one, queue shows that busy cores send work into the queue instead of spawning threads, reject shows which policy fires once crew and queue are both full
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
97 / 168

Once you have seen that curve, Section 3's conclusion gets concrete: the default per-task executor has no stage ② at all — it never queues, it just opens a thread per arrival, so every parameter described in Section 4 is decoration on it. The 200/200 position of the Section 12 sandbox — "threads pile up, downstream collapses, shutdown cannot drain" — is exactly that default behaviour scaled up.

98 / 168
Section
11.2 "Where did the exception go?" and "who blocked my scheduled job?"
99 / 168

These two dominate async incident reports, and both are silent:

100 / 168
Kernel lab
TeaVMWhat actually happens when an async method throwsidle
Pick exc: compare where the exception lands for a void method versus a CompletableFuture-returning one — this is the parcel analogy above plus Section 5's handler
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
101 / 168
Kernel lab
TeaVMOne slow job delays every scheduled jobidle
Pick sched: the scheduler has a single thread, so while the report runs for 4 minutes every other task queues on that same line — the animated version of Section 9's trap
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
102 / 168
Section
11.3 Events for decoupling, transactions inside async work
103 / 168

A very common use of @Async is pairing it with event listeners: the main flow publishes, side effects move to listeners. Mind the default semantics of publishing — many beginners assume "publish and forget", yet by default listeners run one after another on the same thread:

104 / 168
Kernel lab
TeaVMIs publishing an event synchronous or asynchronousidle
Start with sync to see publishEvent block the publisher, then the @Async listener to see who moves to another thread, and finally AFTER_COMMIT timing to confirm 'notify only once the data is really committed'
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
105 / 168

And as soon as an async method writes to the database, you hit Section 6's boundary head-on: a transaction cannot be carried across. Act it out with propagation types:

106 / 168
Kernel lab
TeaVMDoes @Transactional inside an async thread join or start a transactionidle
Contrast REQUIRES_NEW with NOT_SUPPORTED: an async method can never inherit the caller's transaction — it either opens its own or runs without one
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
107 / 168
Section
11.4 The borrow-and-return scene: why this model rescues the main flow
108 / 168

Finally, back to the resource itself. Async protects the order endpoint's latency by leaning on bounded resources plus queuing; the same model has a more famous name on database connections — the connection pool. Watch it borrow and return, and "the queue" stops being abstract:

109 / 168
Kernel lab
TeaVMBorrow, queue and leak in a connection poolidle
Walk idle hit, create when empty, queue at the limit, wait timeout: the first three map onto core / grow / queue in a thread pool, and the fourth tastes like forgetting awaitTerminationSeconds
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
110 / 168

That is enough clicking on buttons — time to type. The console below talks to the same in-browser container, and every answer is computed by the Java kernel rather than read from a script. Start with beans, then run the lab commands one by one:

111 / 168
Console
112 / 168
Note

run lab threadlocal lost and lab threadlocal fix back to back — the first timeline lands on null at step ③, the second reports 'replayed'. Same @Async, the only difference is whether a TaskDecorator is hung on the executor.

113 / 168
Section
12. Sandbox: how the three pool numbers interact
114 / 168

corePoolSize / maxPoolSize / queueCapacity bite each other, and a table gives you no feel for it. The sandbox exposes two switches (the third value is fixed at a queue of 500); every position shows throughput, backlog, rejections and shutdown loss together:

115 / 168
Sandbox
SandboxThread pool tuning sandbox
Result
activeCount: 16 queuedTasks: 42 completedTaskCount: 1.2k/min
P99 order latency: 1.6s
# SMS cost is isolated on notify-* threads, invisible to the main flow
Balanced: the main flow answers instantly, the queue absorbs the peak, overflow fails fast and upstream notices.
116 / 168
Note

the numbers are illustrative, but three conclusions are real — more threads is not more throughput (past what downstream can take, throughput falls and you knock over dependencies); queue length decides how much latency you tolerate, not whether tasks survive; and picking a rejection policy is choosing who pays when things overflow: the caller, the business, or silence.

117 / 168

Then make the division of labour explicit: annotations only hand work over; the pool decides how it finishes — the left column below is the code you write, the right column is where every parameter, trap and tuning knob actually lands:

118 / 168
Diagram
Figure · Two execution models
Figure · Two execution models
119 / 168
类比|Analogy

the default executor is hailing a cab off the street — find a fresh car per passenger and scrap it after the ride, so a surge means scrambling for cars everywhere (unbounded threads). A custom pool is a rostered taxi fleet — a few cars permanently on duty (core), a waiting lane (queue), extra cars called in at peak (max), and refuse-or-reassign rules when even that is full (rejection policy). Every switch you flip in Section 12 only edits this fleet's roster.

120 / 168
Section
13. Check yourself
121 / 168

A warm-up on the order of stages that people remember backwards, straight from Section 4:

122 / 168
Quiz
Check yourselfA ThreadPoolTaskExecutor is configured with corePoolSize=8, maxPoolSize=16, queueCapacity=500. All eight core threads are busy. Where does the ninth submitted task go?
Pick one — you get feedback right away
123 / 168

Now a combined question threading Sections 5, 6, 9 and 10:

124 / 168
Quiz
Check yourselfThe order main flow, running inside an @Transactional method, calls an @Async void sendNotify(); that async method in turn contains logic annotated @Transactional. Which statement is correct?
Pick one — you get feedback right away
125 / 168
Section
14. Common errors, searchable by exact wording
126 / 168

Copy every snippet below straight into a search box — do not paraphrase or shorten it. Beginners stall in three places: tasks rejected, exceptions nowhere to be found, schedules drifting.

127 / 168
Table
Error text (fragment)What really happened30-second fixRead more in
org.springframework.task.support.TaskRejectedException: Executor java.util.concurrent.ThreadPoolExecutor@...[Running, pool size = 16, active threads = 16, queued tasks = 500, completed tasks = 8123] did not accept task: ...Threads reached max and the queue is full, so the default AbortPolicy refused the task. Note it is Spring's wrapper around RejectedExecutionExceptionRead the three numbers inside the brackets: queued tasks pinned means the queue is too small, active threads pinned means concurrency is too small; then decide between growing the pool and switching to CallerRunsPolicySection 4 · Section 12 sandbox position `8/16\AbortPolicy`
java.util.concurrent.RejectedExecutionException: Task ... rejected from java.util.concurrent.ThreadPoolExecutor[...]Same root cause, but you called executor.submit() by hand or used a raw JDK pool outside the containerLet ThreadPoolTaskExecutor manage it and give the pool a threadNamePrefix, otherwise logs cannot tell which pool refusedSection 3
An @Async method threw, yet the console shows no error at all and the endpoint still returns 200The method returns void, so the exception has no channel back; by default it is swallowed or reduced to one error lineImplement AsyncConfigurer#getAsyncUncaughtExceptionHandler (code in Section 5), or return CompletableFuture<T> and handle it with exceptionally on the caller sideSection 5 · Section 11 exc lab
You added @Async, but the SMS log line still prints on http-nio-8080-exec-1 (no thread switch)One of three cases: ① same-class this.method() self-invocation bypassing the proxy; ② the method is private/final; ③ @EnableAsync is missingTrust the thread name first: async must land on notify-*; then verify the injected bean with AopUtils.isAopProxy(bean); finally check the master switchSection 5's failure table
Scheduled jobs are "occasionally late" and the log shows the previous run had not finishedThe default scheduler has one thread, so a long job pushes every later job back (they queue, they are not lost)Register a multi-threaded ThreadPoolTaskScheduler (code in Section 9) with poolSize ≥ the number of jobs you expect to overlapSection 9 · Section 11 sched lab
The same @Scheduled(cron=...) job ran on every replica in a multi-instance deployment and bills went out twice@Scheduled is an in-process scheduler; replicas coordinate with each other not at allShort term: Redis SETNX with a time-granular key (code in Section 10); properly: ShedLock (@SchedulerLock) or a scheduler centre such as XXL-JOBSection 10
ShedLock reports Cannot acquire lock yet the job never succeeds oncelockAtMostFor is shorter than the job really takes, or lockAtLeastFor makes the current round skip outright, or replica clocks disagreeSet lockAtMostFor to "worst-case duration + margin", enable NTP, and on first rollout confirm via logs which instance won the lockSection 10's comparison table
After a restart, a batch of queued async tasks vanished with no traceShutdown did not wait: waitForTasksToCompleteOnShutdown was off, or your wait exceeded the orchestrator's terminationGracePeriodSeconds and the container was killedConfigure the trio together: server.shutdown: graceful + timeout-per-shutdown-phase + pool-level setAwaitTerminationSeconds, keeping the pool wait below the platform grace periodSection 7
Async task logs carry no traceId and the trace breaks in halfThe MDC lives in a ThreadLocal and does not travel across threads automaticallySet setTaskDecorator(new MdcTaskDecorator()) on the pool (code in Section 6) and remember MDC.clear() afterwards to avoid polluting reused threadsSection 6
java.lang.OutOfMemoryError: unable to create new native thread, with SimpleAsyncTaskExecutor in the stackProduction kept the default executor: one new thread per task, no reuse, no upper bound, so a traffic burst stacks threads until the process diesSwitch to ThreadPoolTaskExecutor with explicit core/max/queue — the entire content of Section 3Section 3 · Section 12 sandbox
128 / 168
Tip

if you print one row of this table, make it the first — the bracket after did not accept task already contains pool size / active threads / queued tasks as live numbers. That is your thread-pool dashboard: one glance tells you whether to grow threads or grow the queue.

129 / 168

Section 6 promised that touching the request from async code always blows up, so let a real stack trace close the argument. This is the highest-frequency variant in production — read it, then pick the frame you think is guilty before looking at the answer:

130 / 168
Triage
Error triageIllegalStateException: No thread-bound request found
Reading the request from an async method

The order is created, an SMS is sent asynchronously, and an exception appears in the log — while the checkout endpoint returned 200 the whole time.

java.lang.IllegalStateException: No thread-bound request found: Are you referring to request attributes outside of an actual web request, or processing one outside the normally exposed via ContextLoaderListener?
at org.springframework.web.context.request.RequestContextHolder.currentRequestAttributes(RequestContextHolder.java:131)
at com.bee.order.notify.SmsSender.fillLocale(SmsSender.java:58)
at com.bee.order.notify.SmsSender.sendAsync(SmsSender.java:41)
at java.base/java.lang.Thread.run(Thread.java:840)
Caused by: no RequestAttributes bound to thread task-3
Click the frame you blame — guessing is allowed
No pressure: guess the exception first, then which line actually made the call.
131 / 168
Section
15. Practice in three levels
132 / 168
Section
Level 1 · Follow along
133 / 168

Goal: with a small runnable Spring project, print the full chain "main flow returns instantly → the work happens on an async thread → the exception is visible only on that thread", and watch self-invocation degrade async into sync.

134 / 168

Step one, pom.xml (Java 17; swapping in Boot's spring-boot-starter works identically):

135 / 168
xml
<dependencies>    <dependency>        <groupId>org.springframework</groupId>        <artifactId>spring-context</artifactId>        <version>6.1.8</version>    </dependency>    <dependency>        <groupId>org.springframework</groupId>        <artifactId>spring-tx</artifactId>          <!-- TaskExecutor lives here -->        <version>6.1.8</version>    </dependency>    <dependency>        <groupId>ch.qos.logback</groupId>        <artifactId>logback-classic</artifactId>        <version>1.5.6</version>    </dependency></dependencies>
136 / 168

Step two, the pool plus the exception safety net (src/main/java/com/example/async/AsyncConfig.java) — one class supplies both hooks of AsyncConfigurer:

137 / 168
java
package com.example.async;import org.slf4j.Logger;import org.slf4j.LoggerFactory;import org.springframework.aop.interceptor.AsyncUncaughtExceptionHandler;import org.springframework.context.annotation.Bean;import org.springframework.context.annotation.Configuration;import org.springframework.scheduling.annotation.AsyncConfigurer;import org.springframework.scheduling.annotation.EnableAsync;import org.springframework.scheduling.concurrent.ThreadPoolTaskExecutor;import java.util.Arrays;import java.util.concurrent.Executor;@Configuration@EnableAsync                                   // master switch: without it @Async does nothingpublic class AsyncConfig implements AsyncConfigurer {    private static final Logger log = LoggerFactory.getLogger(AsyncConfig.class);    @Bean("notifyExecutor")    public ThreadPoolTaskExecutor notifyExecutor() {        ThreadPoolTaskExecutor ex = new ThreadPoolTaskExecutor();        ex.setCorePoolSize(2);        ex.setMaxPoolSize(4);        ex.setQueueCapacity(10);        ex.setThreadNamePrefix("notify-");      // key: a business prefix makes logs readable        ex.setWaitForTasksToCompleteOnShutdown(true);        ex.setAwaitTerminationSeconds(10);        ex.initialize();                        // plain container: call it yourself; Boot does it for you        return ex;    }    /** Executor used when @Async names no pool */    @Override    public Executor getAsyncExecutor() {        return notifyExecutor();    }    /** The only place a void async exception can be caught */    @Override    public AsyncUncaughtExceptionHandler getAsyncUncaughtExceptionHandler() {        return (ex, method, params) ->                log.error("async task failed method={} params={}",                        method.getName(), Arrays.toString(params), ex);    }}
138 / 168

Step three, the async service (NotifyService.java). It is called from another bean, so those calls pass through the proxy:

139 / 168
java
package com.example.async;import org.springframework.scheduling.annotation.Async;import org.springframework.stereotype.Service;@Servicepublic class NotifyService {    @Async("notifyExecutor")    public void sendNotify(Long orderId) {        System.out.println("[notify] thread=" + Thread.currentThread().getName()                + " sending SMS orderId=" + orderId);    }    @Async("notifyExecutor")    public void failingAsync() {        throw new IllegalStateException("the SMS gateway exploded");   // an exception nobody receives    }}
140 / 168

Step four, the main flow (OrderService.java), containing both the good case and the counter-example: cross-bean call vs same-class self-invocation.

141 / 168
java
package com.example.async;import org.springframework.scheduling.annotation.Async;import org.springframework.stereotype.Service;@Servicepublic class OrderService {    private final NotifyService notifyService;    public OrderService(NotifyService notifyService) {        this.notifyService = notifyService;      // constructor injection: complete at birth    }    public void createOrder(Long orderId) {        System.out.println("[main] thread=" + Thread.currentThread().getName()                + " writing order orderId=" + orderId);        notifyService.sendNotify(orderId);       // cross-bean -> through the proxy -> truly async    }    public void selfInvokeDemo() {        System.out.println("[main] self-invocation starts thread=" + Thread.currentThread().getName());        sendNotifyInline(2002L);                 // this.xxx() bypasses the proxy -> synchronous        System.out.println("[main] self-invocation ends thread=" + Thread.currentThread().getName());    }    @Async("notifyExecutor")    public void sendNotifyInline(Long orderId) {        System.out.println("[inline] thread=" + Thread.currentThread().getName()                + " sending SMS orderId=" + orderId);    }    public void boom() {        notifyService.failingAsync();            // handed to an async thread; it will not come back here    }}
142 / 168

Step five, the launcher (Main.java):

143 / 168
java
package com.example.async;import org.springframework.context.annotation.AnnotationConfigApplicationContext;public class Main {    public static void main(String[] args) throws InterruptedException {        try (AnnotationConfigApplicationContext ctx =                     new AnnotationConfigApplicationContext(AsyncConfig.class,                             OrderService.class, NotifyService.class)) {            OrderService svc = ctx.getBean(OrderService.class);            long t0 = System.currentTimeMillis();            svc.createOrder(1001L);                 // external call: through the proxy            System.out.println("createOrder returned in "                    + (System.currentTimeMillis() - t0) + "ms");            svc.selfInvokeDemo();                   // internal call: bypasses the proxy            svc.boom();                             // void async method throws            Thread.sleep(2000);                     // give the async threads a moment        }        System.out.println("---- container closed ----");    }}
144 / 168

Run main. Expected output — the thread names are the only evidence, match them line by line (the [notify] line may float slightly depending on scheduling):

145 / 168
text
[main] thread=main writing order orderId=1001createOrder returned in 2ms[main] self-invocation starts thread=main[inline] thread=main sending SMS orderId=2002        <- self-invocation degraded into sync![main] self-invocation ends thread=main[notify] thread=notify-1 sending SMS orderId=1001     <- cross-bean call is the real asyncERROR c.example.async.AsyncConfig - async task failed method=failingAsync params=[]java.lang.IllegalStateException: the SMS gateway exploded	at com.example.async.NotifyService.failingAsync(NotifyService.java:19)---- container closed ----
146 / 168

Acceptance checklist: ① point at the thread-name difference between the [inline] and [notify] lines and say which call really went async; ② remove @EnableAsync and rerun — the container then creates no async proxy for these beans at all, so you will observe all three notification lines printing on main in strict sequence, while IllegalStateException no longer appears in AsyncConfig's error log but propagates up the call stack into main and kills the program (proof that "whether anyone catches the exception" depends on the proxy existing, not on how many handlers you wrote); ③ explain why "createOrder returned in 2ms" and the SMS still went out.

147 / 168
Section
Level 2 · Variants
148 / 168

Change exactly one thing per run and the conclusion flips:

149 / 168
  1. Set queueCapacity to Integer.MAX_VALUE and submit 5000 tasks in a loop. You will observe: active threads stays frozen at corePoolSize=2, maxPoolSize=4 never engages, everything piles up in the queue and memory climbs — Section 4's "an unbounded queue makes max useless", reproduced.
  2. Restore queueCapacity=2 with corePoolSize=2, maxPoolSize=4, then submit 100 tasks concurrently. You will observe: TaskRejectedException ... did not accept task in the log with queued tasks = 2 and active threads = 4 both pinned — read those three numbers against the 8/16|AbortPolicy position of the Section 12 sandbox.
  3. Swap the rejection handler for new ThreadPoolExecutor.CallerRunsPolicy(), changing nothing else. You will observe: rejections disappear but the main thread starts running tasks itself, and createOrder returned in jumps from 2ms to hundreds — "lossless" billed to the submitter, i.e. the backpressure the sandbox describes.
  4. Add a @Scheduled(fixedRate = 1000) method that sleeps 4 seconds, plus a lightweight job printing every 2 seconds. You will observe: the lightweight job now prints only every ~4 seconds — the default scheduler owns a single thread. Register the ThreadPoolTaskScheduler(poolSize=4) from Section 9 and the two lines separate again.
  5. Annotate failingAsync() with @Transactional, insert a row and then throw, while the main flow also runs in a transaction. You will observe: the two rollbacks ignore each other and the async write follows its own transaction's rules — empirical proof of Section 6's boundary; watch it alongside the REQUIRES_NEW case in the txprop lab.
150 / 168

Tip: after variants 2 and 3, return to the Section 12 sandbox, flip both switches to the matching positions, and confirm the readings agree.

151 / 168
Section
Level 3 · Build one
152 / 168

Write an "async observability" component so any future project's pool state is visible at a glance and safe to shut down.

153 / 168

Requirements:

154 / 168
  • A PoolMetricsReporter printing, every 5 seconds, for each ThreadPoolTaskExecutor: poolSize / activeCount / queueSize / queueRemainingCapacity / completedTaskCount / rejectCount
  • Do not guess rejectCount: wrap a custom RejectedExecutionHandler that increments a counter and then delegates to the original policy (keep Abort / CallerRuns / DiscardOldest selectable)
  • Allow per-pool parameters from configuration, and validate at startup: warn loudly if queueCapacity == Integer.MAX_VALUE or maxPoolSize <= corePoolSize (turn the Section 4 and Section 12 traps into a boot self-check)
  • Install a TaskDecorator carrying the MDC, log the submission timestamp so queue wait is computable as "started at − submitted at"
  • On graceful shutdown print "tasks remaining / seconds actually waited", and alert explicitly if awaitTerminationSeconds is exceeded
155 / 168

Acceptance checklist: ① under load, queueSize tracks traffic and rejectCount increments at the exact moment of a rejection; ② deliberately set the queue to Integer.MAX_VALUE and the startup WARN must appear; ③ shutdown logs show "N remaining, waited Xs"; ④ the async log's traceId matches the HTTP request that triggered it (MDC carried successfully); ⑤ not one business method changes.

156 / 168
Section
16. Self-check
157 / 168
Self-check

from memory, name the four stages a submitted task passes through, and explain why "queue first, then grow" makes maxPoolSize useless behind a large queue.

158 / 168
Self-check

after a void @Async method throws, which objects does that exception travel through and where does it finally land? How many legitimate ways do you have to catch it?

159 / 168
Self-check

what happens when a class calls its own @Async method? Is that the same root cause as @Transactional self-invocation failing?

160 / 168
Self-check

among MDC, SecurityContext and the transaction context — which cross threads, with what mechanism, and which cannot? What is the underlying reason?

161 / 168
Self-check

fixedRate versus fixedDelay differ in "measured from which instant". When a run exceeds the interval, how does each behave?

162 / 168
Self-check

why does @Scheduled fire once per replica in a cluster? Why must the Redis lock key carry a time granularity, and why should its TTL sit slightly below the interval?

163 / 168

Mantra: **annotations hand work over; the pool decides its fate — queue first, hire next, refuse at the cap; a void exception is a lost parcel nobody reports; threads carry data but never a transaction; in a cluster, grab the lock before doing the work.**

164 / 168
Section
17. Two traps worth memorizing
165 / 168
Trap

an exception inside a @Scheduled task does not stop future runs — that's good; but it also doesn't stop the task, so always try-catch and log inside the task, or you'll see "the task seems not to run" when in fact it silently fails every time.

166 / 168
Trap

fixedRate means "a rate from the start instant", so a long task exceeding the interval runs back-to-back and accumulates backlog (note: no overlap, since the single-threaded scheduler waits for the previous run). If you don't want a backlog, switch to fixedDelay, which measures from the end of the previous run.

167 / 168
Decision
Decisiona small project's daily reconciliation job — should it adopt a scheduler center (XXL-JOB)?
168 / 168
Summary

four takeaways — @Async / @Scheduled only submit; the thread pool (or scheduler) decides the behavior; @Async needs a custom pool, must handle void exceptions, and fails on self-invocation; contexts (MDC / SecurityContext / transactions) do not cross threads automatically; in a cluster, scheduled tasks run repeatedly and need a distributed lock or scheduler center. Nail these four and async plus scheduling stop being a "sometimes-works mystery".