Redis and Spring Cache: The Caching Abstraction in Practice

bee2026-10-0856 min read0 views
From penetration, breakdown and avalanche to a complete Spring Cache + Redis setup: key design, serialization, expiry and the consistency trade-off.
1 / 154
Section
0. The 30-second version
2 / 154

Picture the little convenience store behind your school campus. Buy a bottle of water and it is either on the shelf — three seconds — or the clerk has to fetch it from the back-room warehouse, which takes two minutes; if the warehouse has none either, they phone the supplier, and that is half an hour away. Asking a database for a row is that errand: slow, and during a sale tens of thousands of people send the same errand at once, which is how a warehouse gets carried off. A cache does something very plain: put the frequently asked items on the shelf so most requests never leave the shop. The price is that once the shelf and the warehouse disagree (the data changed, the shelf did not), you now have to manage two copies of the truth. That is what this article is about: how to stock the shelf (read/write paths and annotations) and what to do when stocking goes wrong (penetration, breakdown, avalanche, consistency).

3 / 154

Five words, one line each (used throughout):

4 / 154
  • Cache: a copy that is easy to grab but may be stale, kept somewhere far faster than the database (usually in memory)
  • key: the address label of that copy, e.g. app:user::42. Every lookup depends on it — get the label wrong and you hand someone else's data to the wrong person
  • TTL: the expiry date on that label (time to live). When it lapses the copy is voided and the next read must re-check reality; no TTL means permanently stale
  • Refill (go back to source): the trip to the warehouse when the shelf is empty. The whole craft of caching is "refill as rarely as possible, but always from the right place"
  • Serialization: packing a Java object into bytes or JSON so it fits into Redis, and unpacking it again. Pack it badly and redis-cli shows you nothing but garbage
5 / 154
类比

the whole caching system is that shop. Redis is the shelf (within arm's reach, but small and perishable), MySQL is the back-room warehouse (has everything, slow to visit), other services or master/replica sync are the supplier (furthest away and most likely to let you down). "Penetration" is customers repeatedly asking for a SKU the shop has never sold, forcing a warehouse trip every single time; "breakdown" is the one slot for the best-seller being empty right now, with thirty customers crowding the back door at once; "avalanche" is the entire shelf being cleared simultaneously while everyone runs into the warehouse aisle together. The three failures differ by one sentence only: the thing asked for does not exist / one hot item is gone / a whole batch vanished at the same moment.

6 / 154
Animation
Animation · Five endings of one request
Animation · Five endings of one request
7 / 154

That animation draws the happy path and four failure scenes on one route: the first two steps are identical for every request, and the fork happens at step two — "is the key actually in the cache?". To locate a production incident fast, first recognise which branch you are on; the table in Section 2 hands out the matching cure.

8 / 154

After this article you should be able to answer three questions:

9 / 154
  • I added a cache, so why is the endpoint still slow? (Is it genuinely missing, or hitting yet still querying the database?)
  • Operations changes a price in the admin console — when will users see it, and how do I force that delay into a range I choose?
  • After rewriting @Cacheable as a this.getXxx() self-invocation it silently stops working. What single log setting proves it on the spot?
10 / 154
Section
1. Why cache at all: what a database query really costs
11 / 154

Caching works because of one plain fact: memory access and disk access differ by three to five orders of magnitude. Put the common operations side by side and the gap is obvious:

12 / 154
Table
OperationTypical latencySense of scale
CPU register access~0.3 ns1x
Read L1 / L2 cache~1–10 nstens of times
A single Redis GET (same DC)~0.1–0.5 msmillions of times
A MySQL primary-key lookup (index hit)~1–5 mstens of millions
A complex MySQL JOIN / aggregation~50–500 mseven slower
13 / 154

A MySQL query walks the whole chain: parse the SQL, build an execution plan, descend the B+ tree, possibly do a table lookup, then ship the result over the network. Redis is a single hash lookup in memory. The difference is not "a bit faster" — it gets amplified by concurrency: at 1000 QPS, hitting MySQL means the database serves a thousand queries a second, while hitting Redis may never touch the database at all.

14 / 154
Note

Redis is fast not because of black magic, but because it keeps everything in memory, runs single-threaded to avoid lock contention, and uses IO multiplexing to handle concurrency. Still, a large share of its latency is network round-trip — 0.1 ms in the same datacenter is excellent, and across datacenters it can degrade to several milliseconds. The real value of a cache is that it holds read pressure outside the database door.

15 / 154
Diagram
Figure 1 · Read and write paths
Figure 1 · Read and write paths
16 / 154

What does "slow" actually look like inside a real request? The lab below carries one request from the entry point all the way to the database. Switch to the "Slow request" position and watch which layer the latency piles into — see that clearly and you know which segment the cache is supposed to shield, instead of slapping @Cacheable onto every endpoint you meet:

17 / 154
Kernel lab
TeaVMOne request through every layer: where the time goesidle
Walk the happy path first to feel the normal rhythm, then this position to see a slow request holding up everyone upstream; finish with 'Database down', which is what the database faces alone once the cache stops covering it
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
18 / 154

But there is no free lunch — once you introduce a cache, the data exists in two places, and three classic failures come with it.

19 / 154
Section
2. Three classic problems and their fixes
20 / 154
Table
ProblemSymptomRoot causeCommon fixesCost
PenetrationRequests hammer a key that does not existNeither cache nor database has it, so every request hits the DBNull caching, Bloom filter, input validationNulls take memory / Bloom is probabilistic
BreakdownA hot key expires at the worst momentHigh concurrency misses at once and stampedes the DBMutual-exclusion rebuild, logical expiryLocking adds wait / logical expiry is complex
AvalancheA large batch of keys expires togetherUniform TTLs fire at once, or Redis goes downRandomized TTL, multi-level cache, rate limitingRandomization range must be tuned
21 / 154

Three rows are quick to read, but during an incident all you have is one sentence of symptoms. So here the two columns are laid out side by side — failure fingerprint on the left, matching first medicine on the right — worth pinning above your desk rather than memorising three times:

22 / 154
Diagram
Figure · Fingerprints and the first medicine
Figure · Fingerprints and the first medicine
23 / 154

The breakdown scene: at 8 p.m. sharp, the sale starts. A hit product's cache key has a 3600-second TTL, and it was written exactly one hour ago — so on the hour, that key expires. Tens of thousands of requests discover an empty cache in the same millisecond and all stampede toward the database for the same row, which saturates the database with a single-row query, exhausts the connection pool, and takes the product page down. Note: the database was not broken; it was sold out by its own caching strategy.

24 / 154

Here is the code behind two of the fixes. First, penetration and breakdown:

25 / 154
java
// Penetration fix 1: null caching — cache a short-lived "NULL" marker to block repeatspublic User findById(Long id) {    String key = "user:" + id;    String cached = redis.opsForValue().get(key);    if (cached != null) {        return "NULL".equals(cached) ? null : JSON.parseObject(cached, User.class);    }    User user = userMapper.selectById(id);    if (user == null) {        // short TTL — keep it brief or it becomes "cache pollution"        redis.opsForValue().set(key, "NULL", Duration.ofMinutes(2));        return null;    }    redis.opsForValue().set(key, JSON.toJSONString(user), Duration.ofMinutes(30));    return user;}
26 / 154
java
// Breakdown fix: mutual-exclusion rebuild — only one thread queries the DB, others waitpublic User findByIdWithLock(Long id) {    String key = "user:" + id;    User user = getFromCache(key);    if (user != null) return user;    String lockKey = "lock:user:" + id;    // SETNX: the winner rebuilds, the losers retry the cache shortly after    Boolean locked = redis.opsForValue().setIfAbsent(lockKey, "1", Duration.ofSeconds(10));    try {        if (Boolean.TRUE.equals(locked)) {            user = userMapper.selectById(id);                     // only the lock holder hits the DB            redis.opsForValue().set(key, JSON.toJSONString(user), Duration.ofMinutes(30));        } else {            Thread.sleep(50);                                     // wait a beat, then re-read the cache            return getFromCache(key);        }    } finally {        if (Boolean.TRUE.equals(locked)) redis.delete(lockKey);    }    return user;}
27 / 154

An avalanche usually needs no code — strategy is enough to blunt it:

28 / 154
  • Randomized TTL: write with TTL + random(0, 300) seconds to spread expiries instead of a synchronised burst
  • Multi-level cache: local Caffeine plus Redis, so a Redis hiccup still leaves the local layer shielding the DB
  • Rate limiting and fallback: add a circuit breaker (e.g. Resilience4j) in front of the database; better to reject a few requests than to avalanche it
29 / 154
Trap

null caching must use a short TTL. Cache a non-existent id for 30 minutes and an attacker can spray id=-1, -2, -3... to flood Redis with meaningless markers — the moment "null caching" becomes "cache pollution".

30 / 154

The three fixes above are scattered through code, while a real ticket hands you a single sentence of symptoms. So instead of another table, play a round: the left column quotes what tickets actually say, the right column is the diagnosis plus the first medicine — a wrong pick explains itself on the spot, which beats memorising rows.

31 / 154
Match
MatchCache incidents: one symptom, one causeMatched 0/7 · Missed 0
All seven are phrased the way tickets read. The right column carries both the diagnosis and the first medicine. Do not use positions — both columns are shuffled
Pick a card on the left first
32 / 154
类比

the shop tells these three apart just as easily. Penetration is a customer repeatedly asking "do you sell size-25 shoes?" — the shop never has sold shoes, yet each time the clerk walks to the warehouse, comes back and says no; the cure is a poster on the door listing what we do not carry (a Bloom filter), or a sticky note saying "asked ten minutes ago: no" (null caching). Breakdown is the single crate of the best-selling water being emptied at exactly this moment, with thirty customers at the back door — only one clerk may go to the warehouse, everyone else waits (a mutex). Avalanche is every price tag on the shelf expiring at the same instant, so the whole crowd surges into the warehouse aisle — either stagger the expiry dates you write (jittered TTL) or control the flow at the warehouse door first (circuit breaking and fallback).

33 / 154
Section
3. Redis in two minutes: five structures and Boot integration
34 / 154

Meet Redis's "five blades"; together they cover nearly every caching scenario:

35 / 154
Table
StructureIn a sentenceTypical use
StringThe basic key-valueObject JSON, counters (INCR), distributed locks
HashField-value mapPart of an object, shopping carts
ListOrdered, push/pop both endsMessage queues, latest N items
SetUnordered, deduplicatingLike dedup, mutual friends (intersection)
ZSetOrdered set with scoresLeaderboards, delay queues (scored by timestamp)
36 / 154

Integration is one starter and a few lines of config:

37 / 154
xml
<dependency>    <groupId>org.springframework.boot</groupId>    <artifactId>spring-boot-starter-data-redis</artifactId></dependency>
38 / 154

That one line is the minimum. A real project also needs the datasource, the serializers and the monitoring endpoint next to it — tick the boxes and watch what the dependency tree grows, paying attention to the scope difference between mysql and h2 (the second one should never live outside tests):

39 / 154
Generator
GeneratorTicking the dependency set for a caching projectpom.xml3 / 8
Start with Redis and see that Lettuce plus a pool arrive with it; then add JDBC / MySQL and match them against the config keys in Section 3; that Actuator line is the precondition for the 'watch Redis through the health endpoint' lab in Section 10
Output
<?xml version="1.0" encoding="UTF-8"?>
<project xmlns="http://maven.apache.org/POM/4.0.0"
         xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
         xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 https://maven.apache.org/xsd/maven-4.0.0.xsd">
    <modelVersion>4.0.0</modelVersion>

    <parent>
        <groupId>org.springframework.boot</groupId>
        <artifactId>spring-boot-starter-parent</artifactId>
        <version>3.3.4</version> <!-- 版本由 BOM 统管,子依赖不写 version -->
        <relativePath/>
    </parent>

    <groupId>com.example</groupId>
    <artifactId>demo-service</artifactId>
    <version>0.0.1-SNAPSHOT</version>

    <properties>
        <java.version>17</java.version>
        <project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
    </properties>

    <dependencies>
        <dependency>
            <groupId>org.springframework.boot</groupId>
            <artifactId>spring-boot-starter-web</artifactId>
        </dependency>
        <dependency>
            <groupId>org.springframework.boot</groupId>
            <artifactId>spring-boot-starter-data-redis</artifactId>
        </dependency>
        <dependency>
            <groupId>org.springframework.boot</groupId>
            <artifactId>spring-boot-starter-actuator</artifactId>
        </dependency>
    </dependencies>

    <build>
        <plugins>
            <plugin>
                <groupId>org.springframework.boot</groupId>
                <artifactId>spring-boot-maven-plugin</artifactId>
            </plugin>
        </plugins>
    </build>
</project>
Why each choice matters
parentInheriting 3.3.4 starter-parent means no spring-boot-starter-* needs a version; the moment someone adds an explicit version to one starter, that one wins — the most common source of dependency drift.
WebAnything that serves HTTP needs it: DispatcherServlet, embedded Tomcat and JSON mapping come inside this starter.
Data RedisBrings Lettuce and RedisTemplate; switching to Jedis means an exclusion plus a new dependency.
Actuatorhealth/metrics/info endpoints; expose a whitelist, never *.
40 / 154
yaml
spring:  data:    redis:      host: localhost      port: 6379      password: ${REDIS_PASSWORD:}      database: 0      timeout: 2s      lettuce:        pool:          max-active: 16      # pool cap — keep it small, Redis itself is single-threaded          max-idle: 8          min-idle: 2
41 / 154

Writing the whole file is harder than it looks: address, timeout, pool and serialization land in different sections. **Tick Redis + datasource + logging + profile and see where spring.data.redis. and logging.level each end up* — the pool cap in Section 3 and the "mass invalidation" row of the error table both start on these lines.

42 / 154
Generator
GeneratorRedis plus the datasource, configured in one goapplication.yml2 / 5
Redis alone gives the minimum working file; adding DataSource shows two pools living side by side, which is exactly what the queuing lab in Section 10 demonstrates; Profile shows how far apart the dev and prod addresses and timeouts should be
Output
server:
  port: 8080

spring:
  application:
    name: demo-service
  datasource:
    url: jdbc:mysql://127.0.0.1:3306/bee_order?useSSL=false&serverTimezone=Asia/Shanghai&allowPublicKeyRetrieval=true
    username: ${DB_USER:root}          # ${} 占位符:环境变量优先,冒号后是默认值
    password: ${DB_PASS:}
    hikari:
      maximum-pool-size: 20
      minimum-idle: 5
      connection-timeout: 30000
      max-lifetime: 1740000            # 必须小于 MySQL 的 wait_timeout
      pool-name: beeHikari
  data:
    redis:
      host: ${REDIS_HOST:127.0.0.1}
      port: 6379
      timeout: 3000ms
      lettuce:
        pool: { max-active: 16, max-idle: 8, min-idle: 2, max-wait: 2000ms }
Why each choice matters
datasourcePool settings only apply here; constructing HikariDataSource in code ignores every one of them.
redisBoot 3 does not pool by default; lettuce.pool.* requires commons-pool2 on the classpath.
43 / 154

Spring Boot auto-wires two templates, and their difference is the first trap a beginner hits:

44 / 154
Table
Templatekey / value serializationBest for
RedisTemplate<K, V>JDK serialization by default (binary)Storing Java objects; configure the serializer yourself
StringRedisTemplateEverything as String (UTF-8)Text keys/values that must match redis-cli
45 / 154
Tip

RedisTemplate<Object, Object> uses JdkSerializationRedisSerializer by default, which turns your object into a blob of binary — you will see it with your own eyes in the next section.

46 / 154
Section
4. The serialization scene: why redis-cli shows garbage
47 / 154

Many people start with this:

48 / 154
java
@Autowiredprivate RedisTemplate<String, Object> redisTemplate;public void save(User user) {    redisTemplate.opsForValue().set("user:1", user);   // store a Java object}
49 / 154

Then they open redis-cli and see:

50 / 154
text
127.0.0.1:6379> GET "user:1""\xac\xed\x00\x05sr\x00\x04com..User\x..."
51 / 154

That leading \xac\xed is exactly the Java serialization magic number (0xACED). JDK serialization writes the fully qualified class name and field metadata alongside the data, so problems pile up: unreadable content, bloated size, and the moment a class name or field changes, old data fails to deserialize — plus it is a notorious source of deserialization vulnerabilities.

52 / 154

The fix is JSON serialization:

53 / 154
java
@Configurationpublic class RedisConfig {    @Bean    public RedisTemplate<String, Object> redisTemplate(RedisConnectionFactory factory) {        RedisTemplate<String, Object> template = new RedisTemplate<>();        template.setConnectionFactory(factory);        // generic JSON serializer: writes @class info so types survive a round-trip        GenericJackson2JsonRedisSerializer jsonSerializer =                new GenericJackson2JsonRedisSerializer();        StringRedisSerializer keySerializer = new StringRedisSerializer();        template.setKeySerializer(keySerializer);        // keys stay plain strings, readable in redis-cli        template.setHashKeySerializer(keySerializer);        template.setValueSerializer(jsonSerializer);     // values as JSON        template.setHashValueSerializer(jsonSerializer);        template.afterPropertiesSet();        return template;    }}
54 / 154

Now the same key in redis-cli becomes:

55 / 154
Code
Codetext
127.0.0.1:6379> GET "user:1""{\"@class\":\"com.example.User\",\"id\":1,\"username\":\"alice\"}"
Notes

Key point: GenericJackson2JsonRedisSerializer writes an extra @class field, which lets deserialization restore the original type; the downside is that this JSON is tightly coupled to your package names, so renaming or moving a class breaks old data too. If you want "pure JSON, no type info", use Jackson2JsonRedisSerializer<T> with an explicit target type — more readable, but the caller must know the type.

56 / 154
Section
5. The Spring Cache abstraction: annotations are the daily driver
57 / 154

What makes caching pleasant in practice is Spring Cache, an annotation-based abstraction that removes the "check cache -> query on miss -> refill" boilerplate entirely. It is enabled with one annotation:

58 / 154
java
@Configuration@EnableCaching        // switches on the cache abstraction (proxy-based, same origin as @Transactional)public class CacheConfig {}
59 / 154

Four core annotations, each with its own trigger timing:

60 / 154
Table
AnnotationWhen it firesTypical use
@CacheableChecks the cache before the call; a hit skips the methodQuery methods
@CachePutAlways runs the method, then writes the result to cacheRefresh cache after an update
@CacheEvictDeletes the cache after the method runsDeletes, or delete-on-update
@CachingCombines several cache operationsDelete multiple keys in one update
61 / 154

The most common shape:

62 / 154
java
@Servicepublic class UserService {    // key built with SpEL: user::42    @Cacheable(cacheNames = "user", key = "#id")    public User findById(Long id) {        return userMapper.selectById(id);      // this line runs only on a miss    }    // refresh after update: the return value is written to user::#user.id    @CachePut(cacheNames = "user", key = "#user.id")    public User update(User user) {        userMapper.updateById(user);        return user;    }    // delete: remove the specific key    @CacheEvict(cacheNames = "user", key = "#id")    public void delete(Long id) {        userMapper.deleteById(id);    }}
63 / 154

The SpEL key syntax is where people stumble; the common expressions:

64 / 154
Table
SpEL expressionMeaning
#idThe value of the parameter named id
#user.idA property of the parameter object
#p0 / #a0The first parameter (by index; p or a both work)
#root.methodNameThe current method name
#root.targetClassThe target class
#result.idThe method return value (only for @CachePut / @CacheEvict)
#id + ':' + #typeA concatenated composite key
65 / 154
Trap

when key is omitted, Spring uses SimpleKeyGenerator to build a key from the arguments. A no-argument method yields a constant SimpleKey.EMPTY, so several no-arg methods sharing one cacheName will overwrite each other. Always write the key explicitly for no-arg methods.

66 / 154

All of those rules come from one execution path. Spread it out as a single-step run: the seven lines on the left, live variables and call stack on the right. Press step repeatedly and watch box ⑤ — on a hit your method body does not run at all — because that single fact is the origin of the first trap in Section 9 and of the data-mixing cell in the sandbox:

67 / 154
Stepper
StepperStep through @Cacheable: does your code run on a hit?1 / 7
Walk ①→⑦. Who builds the key in box ③, why box ⑤ returns immediately, and why box ⑦ must carry a TTL
Code under debug
1userService.findById(42L); // ① the call lands on the cache proxy, not on your Service
2CacheInterceptor.execute() // ② annotation metadata becomes one CacheOperation
3key = evaluator.key("#id") -> app:user::42 // ③ the SpEL expression is evaluated here
4cache = cacheResolver.resolve("user") // ④ cacheNames is bound to a concrete RedisCache
5value = cache.get(key) // ⑤ a hit returns right here — the body never runs
6result = method.invoke(target) // ⑥ only on a miss does the database get queried
7cache.put(key, result); return result // ⑦ refill (with a TTL) and hand the value back
Variables now
userServiceUserService$$SpringCGLIB$$0
caller threadhttp-nio-8080-exec-5
Call stack
1UserController.detail
2proxy.findById
1The same opening as @Transactional: what was injected is a shell, not the object you newed. If the class name carries no SpringCGLIB, none of the other six steps will ever happen — that is the entire mechanism behind the first trap in Section 9.
68 / 154

The annotations are pleasant, but they have one hard boundary: a cache annotation can only cache "the return value of one method". The moment you need an operation rather than a result — reading many keys at once (MGET/Pipeline), renewing one specific key, grabbing a lock to rebuild exclusively, patching just two fields of an object, or running a Lua script atomically — you go back to RedisTemplate. The picture below sets both roads side by side; note especially the fourth line on the left, because annotations ride on a proxy, self-invocation kills them silently; hand-written code has no proxy and therefore no such trap:

69 / 154
Diagram
Figure · @Cacheable everywhere versus hand-driven RedisTemplate
Figure · @Cacheable everywhere versus hand-driven RedisTemplate
70 / 154
Tip

one sentence decides it — one method = one cached object → annotate; you need operations, not just results → write the code. The two styles coexist happily: main lookups via @Cacheable, exclusive rebuild for hot keys and distributed locks via RedisTemplate. Do not jam a hand-written lock inside an annotation just for stylistic unity.

71 / 154
Section
6. Cache configuration: TTL, key prefix and expiry strategy
72 / 154

The default RedisCacheManager never expires — a frequent source of incidents. To set a TTL, a shared key prefix, or per-namespace policies, customize it:

73 / 154
Code
Codejava
@Configuration@EnableCachingpublic class CacheConfig {    @Bean    public RedisCacheManager cacheManager(RedisConnectionFactory factory) {        RedisCacheConfiguration base = RedisCacheConfiguration.defaultCacheConfig()                .entryTtl(Duration.ofMinutes(10))                 // default TTL: 10 minutes                .prefixCacheNameWith("app:")                      // key prefix: app:user::42                .serializeKeysWith(RedisSerializationContext.SerializationPair                        .fromSerializer(new StringRedisSerializer()))                .serializeValuesWith(RedisSerializationContext.SerializationPair                        .fromSerializer(new GenericJackson2JsonRedisSerializer()))                .disableCachingNullValues();                      // do not cache null; handle penetration yourself        Map<String, RedisCacheConfiguration> perCache = Map.of(                "user", base.entryTtl(Duration.ofMinutes(30)),    // user info: 30 minutes                "config", base.entryTtl(Duration.ofHours(6))       // config: 6 hours        );        return RedisCacheManager.builder(factory)                .cacheDefaults(base)                .withInitialCacheConfigurations(perCache)                .build();    }}
Notes
  • entryTtl: mandatory. A cache with no TTL is a memory leak plus permanent inconsistency
  • prefixCacheNameWith: a uniform prefix that makes bulk cleanup by business easy and avoids key collisions when several apps share one Redis
  • disableCachingNullValues: the opposite of null caching — here nulls are not cached and penetration defence falls to a Bloom filter or validation
  • Different TTLs per business, instead of a "one size fits all" that ties hot and cold data together
74 / 154

How long a TTL should be is the one genuine knob in caching, and its two ends each hold a different failure hostage: the left end bites memory and freshness, the right end bites the database:

75 / 154
Tuner
TunerDefault TTL: stale data on one side, the database on the other
spring.cache.redis.time-to-live
1800secondsNow 0 – 172800
The sweet spot for most read-heavy data
  • Thirty minutes to two hours is the usual band, balancing freshness against hit rate
  • Add jitter: TTL + random(0,300), so nothing aligns with the hour
  • @CacheEvict still has to be there — the TTL is a backstop, not the main defence
  • Watch hit rate and refill QPS; the number itself means nothing
Refill pressure18%
Staleness window35%
Ask first 'how stale may this data get before something breaks' — that answer is the ceiling; then push the risk down with jitter and CacheEvict.
76 / 154

Attention: `@Cacheable` and `RedisCacheManager` relate as abstraction and implementation — you write annotations, and the `CacheManager` decides which `Cache` implementation, TTL and serialization to use. Swap the implementation (say Redis for Caffeine) and the business annotations **do not change at all**. To pick a cache implementation per method dynamically, implement a `CacheResolver` that returns a different `Cache` based on method metadata.

77 / 154
Section
7. Cache and data consistency: why Cache-Aside wins
78 / 154

First, draw the two paths clearly (see Figure 1 at the top). The core conclusion in one line: reads go "cache first", writes go "update the database, then delete the cache" — that is the Cache-Aside pattern.

79 / 154
Animation
Animation · Cache-Aside sequence
Animation · Cache-Aside sequence
80 / 154

Why "delete the cache" rather than "update the cache"? One concurrent scenario explains it:

81 / 154
Table
MomentThread A (write)Thread B (read)
T1Updates database = 20
T2Reads the old value 10 (A's write is not visible yet)
T3Updates cache = 20
T4Writes cache = 10 (overwriting 20)
Resultdatabase 20, cache 10 -> inconsistent
82 / 154

"Update the cache" copies the database's write concurrency straight onto the cache, and the ordering of two writers cannot be guaranteed, so stale data appears. "Delete the cache" reduces the problem to a single atomic action, at the cost of one refill on the next read — one refill buys away the inconsistency risk.

83 / 154

That T1–T4 table is four lines of text; the animation plays it in order. Watch frame ④ — the overwrite happens with no exception and no log line, which is exactly why this class of bug is the hardest to catch:

84 / 154
Animation
Animation · Update versus delete: one concurrent overwrite
Animation · Update versus delete: one concurrent overwrite
85 / 154

Then, is "delete the cache first, then update the database" fine? It has its own race (a reader refills the old value after the delete but before the update). So the more robust order is Cache-Aside's "update the DB first, then delete the cache", backed by delayed double delete:

86 / 154
Code
Codejava
public void updateUser(User user) {    userMapper.updateById(user);                  // 1. persist first    redis.delete("app:user::" + user.getId());    // 2. delete the cache immediately    // 3. delete again after a delay, to overwrite a stale value refilled between delete and commit    delayedExecutor.schedule(            () -> redis.delete("app:user::" + user.getId()),            500, TimeUnit.MILLISECONDS);}
Notes

Note: cache consistency cannot be strongly consistent — only eventually consistent — unless you also add a distributed lock on the read path, which is usually not worth the cost. The engineering trade-off is to accept "a reader may see a stale value within a window" and shrink that window to milliseconds with TTL and delayed double delete, rather than chase theoretical perfection.

87 / 154

Delete-versus-update is only one cell inside five read/write strategies. Spring Cache defaults to the first (aside); the other four each have their own fit and their own way of failing. Click through them, and each box answers the one question that matters: who is responsible for filling the cache?

88 / 154
Diagram
FlowFive read/write strategies: who puts data into the cache1 / 5
Go ①→⑤. The deciding question is only whether the application, the cache component, or the database owns the load and the write
→
→
→
→
① Cache-aside (Spring Cache's default)
The **application** owns both paths: on a miss it queries the database and refills; on a write it persists first and deletes afterwards. Simplest, and the application survives Redis being down entirely. The price: nothing ever prefetches for you, so the hit rate is exactly as disciplined as your code.
All clearOne question decides the pick: who should own the load and the write — the application, the cache component, or the database?
89 / 154
Section
8. A first taste of distributed locks: SETNX + unique value + Lua delete
90 / 154

The mutual-exclusion cache rebuild above is a distributed lock in embryo. A usable lock has three elements: atomic acquisition, expiry as a backstop, and the ability to delete only your own lock:

91 / 154
Code
Codejava
public class RedisLock {    private final StringRedisTemplate redis;    private static final String UNLOCK_LUA =            "if redis.call('get', KEYS[1]) == ARGV[1] then " +            "  return redis.call('del', KEYS[1]) " +            "else return 0 end";    public boolean tryLock(String key, String token, Duration expire) {        // SETNX plus expiry in one write (never split setnx and expire: a crash in between deadlocks)        Boolean ok = redis.opsForValue().setIfAbsent(key, token, expire);        return Boolean.TRUE.equals(ok);    }    public boolean unlock(String key, String token) {        // Lua keeps "compare + delete" atomic: delete only when the token matches        Long r = redis.execute(                new DefaultRedisScript<>(UNLOCK_LUA, Long.class),                Collections.singletonList(key), token);        return r != null && r > 0;    }}
Notes
  • Atomic acquisition: setIfAbsent(key, token, expire) maps to SET key value NX PX ms, doing "set if absent plus expiry" in one step
  • Expiry as a backstop: if the lock holder crashes, the lock expires and releases itself — no permanent deadlock
  • Delete by unique value: token is a unique string generated for this acquisition; unlock compares then deletes, so that "after A's lock times out and B acquires it, A does not delete B's lock"

Trap: a hand-rolled distributed lock will almost always miss an edge case — renewal (when the work outlives the lock's TTL), reentrancy, and lock loss on a primary/replica failover. In production use Redisson's RLock, which ships a watchdog that renews automatically plus reentrancy and RedLock semantics, filling all these holes. The value of the hand-rolled lock here is to understand the mechanics, not to ship it.

92 / 154
Section
9. Three traps you must know + prove why the annotation dies
93 / 154
Trap

@Cacheable silently dies on self-invocation within the same class. Exactly like @Transactional — the annotation relies on an AOP proxy, and this.method() goes through the raw object, bypassing the proxy, so no cache lookup or refill ever happens. Fix it by moving the method into another bean, or injecting your own proxy.

94 / 154
Trap

cached nulls need a short TTL too. Otherwise a single "not found" calcifies into "never found" for tens of minutes — the data is later restored, but the cache still says it is missing.

95 / 154
Trap

big keys and hot keys need dedicated treatment. A multi-megabyte value slows the single-threaded Redis and blocks other commands; a key hammered at tens of thousands of QPS can saturate a single shard. Remedies: split big keys (shard into a Hash), shield hot keys behind a local cache, and spread the load across key replicas.

96 / 154

Cache, transaction and log failures all stem from the same proxy mechanism; the demo below lets you see exactly what self-invocation bypasses:

97 / 154
Kernel lab
TeaVMWhy the cache annotation diesidle
Switch to 'self-invocation' and watch the cache, transaction and logs all die together once the proxy is bypassed
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
98 / 154
Section
10. Hands on: press through all three failures and one silent death
99 / 154

The three traps in Section 9 are prose, but each has a shape you can actually see. The five labs below answer, in order: how do the three failures differ in timing? What does it look like when the cache stops shielding the database? How do you learn that Redis is alive before your users do? Which of the three rate limiters survives an avalanche? And what is the relationship between those MyBatis cache layers and Spring Cache?

100 / 154

The first lab is the star. Cycle through hit → miss → key → pen → bust while replaying the matching game in Section 2:

101 / 154
Kernel lab
TeaVMHit, miss, penetration, breakdown, avalanche liveidle
Start with 'Hit' to see the database never being touched; then 'Key generation' to understand the root of cross-user data mixing; the last two positions separate one hot key expiring from a whole batch expiring
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
102 / 154

The second turns the sentence "the cache holds read pressure outside the database" into a curve you can watch. Compare "Idle hit" with "Queue at the limit" and feel how the waiting queue grows as the hit rate falls from 95% to zero:

103 / 154
Kernel lab
TeaVMWhen the cache stops holding, the pool starts queuingidle
Use 'Idle hit' to see smooth borrow-and-return, then 'Queue at the limit' to simulate every request piling onto the database channel after a mass invalidation
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
104 / 154

The third answers a very practical anxiety: Redis dies — how do I know before my users do? The Actuator health endpoint aggregates the Redis connection state into a DOWN, which with monitoring becomes the earliest alarm you can get:

105 / 154
Kernel lab
TeaVMWatch Redis through the health endpointidle
Pick 'Health aggregation' to see how the redis row drags the overall status to DOWN, then 'Exposure' to confirm that endpoint is not naked on the public port
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
106 / 154

The fourth turns Section 2's "put a limiter in front of the database" into concrete algorithms. When an avalanche hits, a fixed window lets roughly double the traffic through at its boundary, a leaky bucket is always smooth but cannot absorb bursts, and a token bucket permits bursts provided you size it — switch to "Choosing and bursts" and the three shapes are drawn side by side:

107 / 154
Kernel lab
TeaVMThree rate limiters: which one survives the avalanche waveidle
Look at the shape of the fixed window and the leaky bucket first, then this position to compare all three under a burst — this is the implementation behind the avalanche cell of Section 2
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
108 / 154

The fifth settles a question that never stops appearing in comment sections: what is the relationship between MyBatis's first/second-level caches and Spring Cache? On the shared-cache position you see it surviving across sessions and, by default, invalidating a whole namespace — a different model from @CacheEvict deleting one precise key, and running both layers at once can hide stale reads:

109 / 154
Kernel lab
TeaVMThere is more than one cache in this stackidle
Start with the local cache to see reuse inside one session, then this position for sharing across sessions; read it against the third trap in Section 9 — beyond big keys and hot keys there is this invisible layer
Scenario
Click “Run demo” to execute the AOT-compiled Java kernel right in your browser, step by step.
110 / 154

Once you are bored of buttons, type the commands yourself. This console talks to the same in-browser Java kernel and every reply is computed there — start with ping to confirm the container is up, then run the three failures and one connection timeout:

111 / 154
Console
112 / 154
Trap

run lab pool queue immediately after lab cache bust — only then does the downstream of an avalanche become visible. The first timeline shows hundreds of thousands of requests missing at once; the second shows the queue length the database actually faces.

113 / 154
Section
11. Sandbox: which caching strategy does this data deserve
114 / 154

A wrong strategy shows up late, and by then the database is already taking the hits. This sandbox turns "how stale is acceptable" into one switch; each position shows the TTL, the concurrency protection and the metric to watch, side by side:

115 / 154
Sandbox
SandboxCache strategy picker: set the tolerance, then pick the fix
Result
TTL = 1800s + random(0,300) <- jitter prevents synchronised expiry
rebuild on miss: SETNX mutex, exactly one thread refills
watch: per-key QPS, lock-wait duration
# test: high read/write ratio concentrated on one row -> protect that key, do not add machines
The classic sale best-seller. Breakdown protection is the core; logical expiry (store an expireAt column, serve the old value past expiry and refresh asynchronously) is the more thorough alternative.
116 / 154
Note

the four positions map to four real demands, and the order of reasoning never changes — first measure "how long may this data be stale", then decide which failure to defend against. Starting instead from "should I use a Bloom filter?" usually means you have not yet identified which attack you are under.

117 / 154
Section
12. Checkpoints
118 / 154

A warm-up on the pair of concepts most often stated backwards in interviews:

119 / 154
Quiz
Check yourselfAt the top of a sale, one product-detail endpoint pushes the database to 100% CPU, yet the overall Redis hit rate is still 97%. Most likely cause and matching fix?
Pick one — you get feedback right away
120 / 154

Now a combined question tying the annotation mechanics of Section 5 to the trap in Section 9:

121 / 154
Quiz
Check yourselfIn `UserService`, `detail()` carries `@Cacheable(cacheNames="user", key="#id")`, and another method `batch()` calls it as `this.detail(id)`. The project does declare `@EnableCaching`. Which statement is correct?
Pick one — you get feedback right away
122 / 154
Section
13. Error quick-reference
123 / 154

Beginners hit four walls here: cannot connect, cannot read (garbage), reads the wrong thing (mixed data), and the nastiest of all — nothing happens at all. Every fragment below can be pasted verbatim into a search engine.

124 / 154
Table
Error text (fragment)Real cause30-second fixWhere to dig deeper
org.springframework.data.redis.RedisConnectionFailureException: Unable to connect to RedisAddress/port unreachable: Redis not running, spring.data.redis.host points elsewhere, or localhost inside a container resolving to the container itselfRun redis-cli -h <host> -p 6379 ping first; in Docker use the service name as host; then check whether password is an empty string versus unsetSection 3
io.lettuce.core.RedisConnectionException: Unable to connect to localhost/<unresolved>:6379Lettuce could not even establish a connection — usually DNS or network policy, not the password<unresolved> means the hostname never resolved: inspect the compose network and service nameSection 3
SerializationException: Could not write JSON: Java 8 date/time type java.time.LocalDateTime not supported by defaultJackson lacks the JSR-310 module, so LocalDateTime blows up on writeRegister JavaTimeModule and disable WRITE_DATES_AS_TIMESTAMPS; or store a string / legacy Date insteadSection 4
SerializationException: Cannot deserialize value of type ... unknown java class id com.example.UserOld entries were written by GenericJackson2JsonRedisSerializer, which embeds the fully qualified @class name — and the class was renamed or movedNever rename a class that already lives in the cache; if you must, purge first, or switch to a serializer with an explicit target typeSection 4
GET user:1 in redis-cli returns "\xac\xed\x00\x05sr\x00..."The default JdkSerializationRedisSerializer magic bytes — not an error, but effectively unoperatableUse StringRedisSerializer for keys and a JSON serializer for values; existing data must be discarded and rebuiltSection 4
@Cacheable added, nothing happens, zero errors① missing @EnableCaching; ② self-invocation bypassing the proxy; ③ the bean was newed and never entered the containerTurn on logging.level.org.springframework.cache: TRACE; no cache-related line means the interceptor never ran — check ①②③ in orderSections 5 and 9
Different users or tenants see each other's dataSeveral methods share one cacheNames without an explicit key, collapsing to SimpleKey.EMPTY; or the key omits the discriminating dimensionAlways write a key for no-arg methods; compose it as #userId + ':' + #type; print the actual generated key while debugging locallySection 5
Operations changed a price, the storefront keeps showing the old one for agesOnly the database was updated, with no @CacheEvict, or the evict key expression differs from the one used on writeUpdate paths must come in pairs (@CachePut/@CacheEvict with identical key rules); shorten the TTL as a backstopSection 7
The database pool keeps queuing and Pending requests never dropsMass cache invalidation (avalanche) sending every request back to the source, dumping the full read load on the databaseThrottle first to save the database, then warm the critical keys gradually; afterwards add TTL jitterArticle 28
125 / 154
Tip

when searching these, use only the first English phrase after the colon (e.g. Unable to connect to Redis). Different Spring Boot versions append extra clauses to the tail, so full-sentence searches often return nothing.

126 / 154

The first row is the one beginners hit most and misjudge most: it looks like a wrong password or a dead Redis, when it is usually just a localhost inside a container pointing at the container itself. Do not read the answer — click the frame you think is guilty:

127 / 154
Triage
Error triageRedisConnectionFailureException: Unable to connect to Redis
The application cannot reach Redis: one word, localhost

It worked perfectly on the laptop. The moment it ships to a container every endpoint returns 500, this sentence repeats in the log, and changing the password, the version or restarting Redis accomplishes nothing.

org.springframework.data.redis.RedisConnectionFailureException: Unable to connect to Redis
at org.springframework.data.redis.connection.lettuce.LettuceConnectionFactory$ExceptionTranslatingConnectionProvider.translateException(LettuceConnectionFactory.java:1689)
at org.springframework.data.redis.connection.lettuce.LettuceConnectionFactory$ExceptionTranslatingConnectionProvider.getConnection(LettuceConnectionFactory.java:1595)
at org.springframework.data.redis.connection.lettuce.LettuceConnectionFactory.getConnection(LettuceConnectionFactory.java:622)
at org.springframework.data.redis.core.RedisTemplate.execute(RedisTemplate.java:224)
at com.bee.catalog.service.ProductService.getDetail(ProductService.java:52)
Caused by: io.lettuce.core.RedisConnectionException: Unable to connect to localhost/<unresolved>:6379
Caused by: java.net.ConnectException: Connection refused
Click the frame you blame — guessing is allowed
No pressure: guess the exception first, then which line actually made the call.
128 / 154
Section
14. Decision card
129 / 154
Decision
DecisionA product-detail endpoint is read-heavy (read:write ~ 100:1), but occasionally a hot product gets slammed by sale traffic on the hour. Which consistency scheme do you choose?
130 / 154
Section
15. Exercises
131 / 154
Section
Tier 1 · Follow along
132 / 154

Build the smallest possible "two reads, one query" scene. No Redis required — Spring's built-in ConcurrentMapCacheManager shows every mechanism of these annotations. Fully runnable code:

133 / 154
java
package com.example.cache;public record Product(Long id, String name, java.math.BigDecimal price) {}
134 / 154
java
package com.example.cache;import org.springframework.stereotype.Service;@Servicepublic class ProductService {    private int dbHits = 0;    // key written explicitly as #id, so no SimpleKey.EMPTY collision can occur    @org.springframework.cache.annotation.Cacheable(cacheNames = "product", key = "#id")    public Product find(Long id) {        System.out.println("[db] SELECT * FROM product WHERE id=" + (++dbHits));        return new Product(id, "Bee plush", new java.math.BigDecimal("39.00"));    }    @org.springframework.cache.annotation.CachePut(cacheNames = "product", key = "#id")    public Product rename(Long id, String newName) {        System.out.println("[db] UPDATE product SET name=? WHERE id=" + id);        return new Product(id, newName, new java.math.BigDecimal("39.00"));    }    @org.springframework.cache.annotation.CacheEvict(cacheNames = "product", key = "#id")    public void evict(Long id) {        System.out.println("[cache] evict product::" + id);    }    public int dbHits() { return dbHits; }}
135 / 154
java
package com.example.cache;import org.springframework.boot.SpringApplication;import org.springframework.boot.autoconfigure.SpringBootApplication;import org.springframework.cache.annotation.EnableCaching;import org.springframework.context.ConfigurableApplicationContext;@SpringBootApplication@EnableCaching          // <- delete this line and the four calls below become four [db] lines, with not one error loggedpublic class CacheDemo {    public static void main(String[] args) {        try (ConfigurableApplicationContext ctx = SpringApplication.run(CacheDemo.class, args)) {            ProductService svc = ctx.getBean(ProductService.class);            svc.find(1L);               // miss -> query and refill            svc.find(1L);               // hit -> the method does not run at all            svc.rename(1L, "Limited");  // @CachePut -> runs and overwrites the entry            svc.find(1L);               // hit -> returns the renamed object            System.out.println("db hits = " + svc.dbHits());        }    }}
136 / 154

Expected output (db hits must be 1):

137 / 154
text
[db] SELECT * FROM product WHERE id=1[db] UPDATE product SET name=? WHERE id=1db hits = 1
138 / 154

The absence of a third [db] SELECT is the proof that the second read was served from cache. Comment out @EnableCaching, rerun, and you get db hits = 2 with not a single error in the log — the live version of the silent failure listed in Section 13.

139 / 154
Section
Tier 2 · Variant
140 / 154

Goal: prove that self-invocation bypasses the cache proxy.

141 / 154

Hint: add exactly one method to ProductService: public java.util.List<Product> findTwice(Long id) whose body is return List.of(find(id), find(id)); (calling find directly is this.find(id)), then call svc.findTwice(9L) twice from main.

142 / 154

You should observe: the first call prints two [db] lines where you expected one, and the second call prints two more — four database queries in total. The cache is not broken; those four calls never reached the proxy. Now move findTwice into a separate SearchFacade bean that injects ProductService, and the [db] count collapses to one. Write both SQL-line counts in your notes — that is the evidence for the first trap in Section 9.

143 / 154
Section
Tier 3 · Build one
144 / 154

Replace Tier 1's map cache with real Redis, then defend against each of the three failures using the table in Section 2.

145 / 154

Acceptance checklist:

146 / 154
  • [ ] KEYS app:product* in redis-cli shows readable JSON with a prefix, not \xac\xed binary
  • [ ] TTL works: TTL app:product::1 returns a positive number, and different cacheNames carry different TTLs
  • [ ] Penetration defence works: after querying 20 non-existent ids, Redis holds null markers for them with TTL ≤ 2 minutes
  • [ ] Breakdown defence works: delete the hot key, read it from 50 threads at once, and the SELECT log line prints exactly once (the mutex held)
  • [ ] Avalanche mitigation works: sample ten batch-written keys and their TTLs differ (visible jitter)
  • [ ] Updates follow "write the database, then delete the cache", and you can explain why updating the cache is wrong (see the T1–T4 table in Section 7)
  • [ ] /actuator/health lists the redis component; stopping Redis flips it to DOWN without the whole application returning 500 because of the cache
147 / 154
Section
16. Self-check
148 / 154
自检

without looking anything up, match these three symptoms to their failure — "hit rate stays high but one row is hammered", "the overall hit rate falls off a cliff", "requests ask for ids that do not exist" — plus the first medicine for each. Miss one and replay the matching game in Section 2.

149 / 154
自检

what three preconditions does @Cacheable need to take effect? The answer must include @EnableCaching, the call going through the proxy, and the method being public. Missing one sends you back to Sections 5 and 9.

150 / 154
自检

why is the default serializer unreadable in redis-cli, and how many serializers must you configure to switch to JSON (key / hashKey / value / hashValue)? Count them in Section 4.

151 / 154
自检

can you explain Cache-Aside's "delete the cache, do not update it" using the T1–T4 timeline table in Section 7? Say it out loud to a colleague.

152 / 154
自检

should a per-user personalised endpoint be @Cacheable as a whole? If not, which mistake is it — the key dimension or the caching granularity? Check the last position of the sandbox in Section 11.

153 / 154
口诀

the shelf stands in front of the warehouse and the TTL is its best-before date; what does not exist penetrates it, what is hot and expires breaks it, what expires in unison avalanches it; annotations ride the proxy, so self-invocation mutes them; write the database first, delete the cache second, and settle for eventual consistency.

154 / 154
Summary

Compress this article into a few lines — the value of a cache is not a faster single request but holding read pressure outside the database; penetration uses null caching or a Bloom filter, breakdown uses mutual-exclusion rebuild or logical expiry, avalanche uses randomized TTL and multi-level caching; default JDK serialization is garbage in redis-cli, so switch to a JSON serializer; Spring Cache removes boilerplate with annotations, but write the key explicitly, always set a TTL, and remember self-invocation kills it; consistency follows "update the DB then delete the cache" plus delayed double delete, aiming for eventual rather than strong consistency; do not hand-roll distributed locks — use Redisson in production. Stand on these and your cache will not be a mine buried under the database.