Spring Data JPA in Full: Repository Abstraction and JPQL
Kill the misconception first: Spring Data JPA is not a machine that writes SQL for you — it is an object-state management system. What you really have to watch is never the shape of a statement but whether this object is currently under the eye of the persistence context. Once that clicks, nine tenths of JPA's "black magic" collapses into common sense: why the row changed although you wrote no UPDATE, why nothing happened right after save, and why you get no Session the moment the transaction ends.
Six words, one line each (used throughout):
- JPA: the specification (Jakarta Persistence) — annotations and interfaces only, no implementation, exactly like the JDBC spec
- Hibernate: Spring Boot's default implementation, the engine that really generates SQL and runs sessions and caches
- Spring Data JPA: the declarative abstraction on top of JPA: you declare interface methods, proxies supply the implementation
- Persistence context: the register of objects under watch during one watching period, whose lifetime normally coincides with a transaction;
EntityManageris its entry point - First-level cache: not a cache you bolt on — it is the persistence context itself. Inside one transaction, one id costs one statement and one instance
- Dirty checking: field values are snapshotted when the entity loads; at flush each field is compared and only the difference becomes an UPDATE
the persistence context is the shopping trolley in your hands, and the EntityManager is the shop assistant. Dropping goods into the trolley (persist) does not mean you bought them — the trolley is still yours and you can put anything back on the shelf (remove). Only at checkout (flush) does the order become real: the assistant compares what was added and what was swapped, scanning item by item, and that scan is exactly where the UPDATE comes from. So repository.save(entity) literally means "put it in the trolley", not "write it to the database"; money moves only when the transaction commits. And a trolley you already pushed out of the store (a detached entity) — nobody is tallying it any more; swap the instant noodles for a steak at home and the shop's books stay untouched. That is the truth behind "I set the field but the database did not change". Lazy loading, meanwhile, is judging a book by its cover before deciding to borrow it: when you actually want to read it, the librarian walks to the warehouse for the volume (one extra SELECT). One book is fine; a hundred books, each needing its own trip, is N+1. Hold on to "who is watching" and "when do we reconcile", and every trap later in this article gets a coordinate.

That ring is the map of this article: transient → managed → detached → removed. The four states are decided by two questions (is the object admitted by the context? does it have an identifier?) and the three transitions are persist / find, merge and remove. The honest probe is entityManager.contains(u) — never guess. Section 4 opens every cell.
After this article you should be able to answer three questions:
- I only wrote
u.setNickName("new name"), nosave, no UPDATE — why did the database really change? - When does
repository.save(entity)actually send SQL? Why does a non-null id cost an extra SELECT? - One list request logged 101 statements. Which cure fits that N+1, and why is
open-in-viewnot a cure?
Feel the most fragile combination in the whole article first — fetch strategy × moment of access. Flip the switches and the Hibernate SQL log on the right changes immediately:
Hibernate: select o.id, o.order_no, o.user_id from t_order o where o.status = ?Hibernate: select u.id, u.user_name from t_user u where u.id = ? -- 99 more of these-- total: 1 + 100 = 101 statements
read all six cells in one go and you will see that only the two bottom ones are both cheap and harmless (declared fetching). The first pair buys correctness with 101 statements or with an exception; the EAGER pair spreads the cost over every query in the project. Two disciplines: write LAZY explicitly on entities, declare the associations you need per method.
The biggest confusion when learning JPA is the clashing names: JPA, Hibernate, Spring Data JPA. They are not three competing products but three layers: specification / implementation / abstraction.

| Layer | Role | What it does | Analogy |
|---|---|---|---|
| JPA (Jakarta Persistence) | Specification / standard | Defines annotations and interfaces (@Entity, EntityManager), no implementation | The JDBC spec |
| Hibernate | Implementation | Realises the spec: actually generates SQL, manages sessions and caches | A MySQL driver |
| Spring Data JPA | Abstraction | Wraps JPA once more: declare an interface method and get a query | A secretary who writes methods for you |
In one line: JPA is the contract, Hibernate is the worker, Spring Data JPA is the secretary. One real call chain is worth memorising now, because every later section dissects one of its segments: a Controller line userService.activate(42L) → Spring Data's dynamic proxy (PartTree parsing or SimpleJpaRepository) → EntityManager.find (implemented by Hibernate's Session) → first-level-cache lookup (a hit returns the same instance, no SQL) → on a miss, borrow a connection from HikariCP and send the select (#28).
- You call an interface method, not an implementation you wrote — section 5 shows where the implementation comes from
- Whether the returned object is managed depends on whether the call ran inside a transaction — section 4 is about state
- Whether SQL is sent is decided by dirty checking and the flush moment, not by whether you wrote
save— section 10 covers flush
swapping the implementation (EclipseLink, say) leaves your business code untouched, because the upper layer depends only on the JPA spec — that is the payoff of layering. Derived queries, Pageable and Specification, however, belong to Spring Data JPA and are implementation-independent. This article uses Hibernate 6, the Spring Boot default; note the package moved from javax.persistence to jakarta.persistence.
Three lines of dependency buy the whole stack, no hand-written DAO implementation needed:
<dependency> <groupId>org.springframework.boot</groupId> <artifactId>spring-boot-starter-data-jpa</artifactId></dependency><dependency> <groupId>com.mysql</groupId> <artifactId>mysql-connector-j</artifactId> <scope>runtime</scope></dependency>The starter really brings four things: spring-data-jpa (the Repository abstraction), hibernate-core (the JPA implementation), spring-orm (the glue plus transaction managers) and HikariCP (the default pool, see #28). Auto-configuration enables @EnableJpaRepositories for you, which is why you never write it — but knowing it exists is what lets you debug "my repository was not picked up".
The star of configuration is ddl-auto — it decides how the framework treats your schema at startup, and it is the hot zone for production incidents:
spring: datasource: url: jdbc:mysql://localhost:3306/demo?useSSL=false&serverTimezone=Asia/Shanghai username: root password: secret jpa: hibernate: ddl-auto: validate # production uses validate or none, never update open-in-view: false # switch this off explicitly; section 12 explains why show-sql: true # print generated SQL (turn off in production) properties: hibernate: format_sql: true jdbc.batch_size: 50 # batching: only pays off together with flush/clear (section 10) order_inserts: trueddl-auto | Meaning | Where to use |
|---|---|---|
none | Do nothing (default) | Production (with Flyway / Liquibase) |
validate | Only check schema vs entities; fail startup on mismatch | Recommended for production |
update | Auto-add new tables/columns, but never drops columns or changes types | Local dev only |
create / create-drop | Rebuild tables each start / drop on shutdown | Tests / in-memory DB (Boot's default there) |
Take the most common situation: the entity gained a remark field and the table did not. update prints one more line at startup, alter table t_user add column remark varchar(255) — it only ever adds: new columns, indexes and tables are fine, dropping a column, changing a type or renaming something is not, and on a shared database whoever boots first owns the schema. validate fails startup outright (Schema-validation: missing column [remark] in table [t_user]), which stops the drift before traffic arrives; none does nothing, so the error slides to the first runtime statement that uses the column — clean boot, random 500s later; create-drop rebuilds on every start, so the schema is right and the data is gone.
ddl-auto: update is destructive in production. It will not drop columns, but it adds columns and indexes uninvited; renaming an entity leaves the old column forever; and when several people share a database, whoever starts first changes the schema. In production use validate or none and manage changes with Flyway or Liquibase, reviewing the migration scripts like code. When you do see Schema-validation, the fix is a new V8__add_user_remark.sql — not quietly switching to update.
The naming strategy trips people up too: CamelCaseToUnderscoresNamingStrategy (Boot's physical naming strategy) maps the property userName to the column user_name; change the strategy or hit a case-sensitive database and you get "table/column does not exist". One discipline covers it: columns snake_case, properties camelCase, let the naming strategy convert, and pin @Column(name = "...") explicitly when you inherit a legacy schema.
That listing is a conclusion, not an exercise. Generate one yourself instead: does datasource alone boot anything, which extra lines appear the moment JPA is ticked, where do those Hibernate: log lines come from, and why the prod profile must differ?
server:
port: 8080
spring:
application:
name: demo-service
datasource:
url: jdbc:mysql://127.0.0.1:3306/bee_order?useSSL=false&serverTimezone=Asia/Shanghai&allowPublicKeyRetrieval=true
username: ${DB_USER:root} # ${} 占位符:环境变量优先,冒号后是默认值
password: ${DB_PASS:}
hikari:
maximum-pool-size: 20
minimum-idle: 5
connection-timeout: 30000
max-lifetime: 1740000 # 必须小于 MySQL 的 wait_timeout
pool-name: beeHikari
jpa:
open-in-view: never
hibernate:
ddl-auto: validate # 生产用 validate,别用 update/create
show-sql: false
properties:
hibernate.format_sql: true
hibernate.jdbc.batch_size: 50
logging:
level:
root: INFO
com.example.demoservice: DEBUG
org.springframework.jdbc.core.JdbcTemplate: DEBUG # 打 SQL 与参数
file:
name: logs/app.log
logback:
rollingpolicy: { max-file-size: 50MB, max-history: 14 }
Point: after generating, run one reverse self-check — take the `ddl-auto` and `open-in-view` lines from the jpa block and **say out loud what happens if each is deleted**. If you cannot, those two sections are still unread.
Entities are the heart of JPA; the annotations on the fields define the mapping. Here is one close to real production shape:
package com.example.demo.entity;import jakarta.persistence.*;@Entity@Table(name = "t_user", indexes = @Index(name = "idx_user_status_created", columnList = "status, created_at"))public class User { @Id @GeneratedValue(strategy = GenerationType.IDENTITY) // auto-increment id: note the wrapper type Long private Long id; @Column(name = "user_name", nullable = false, length = 32, unique = true) private String userName; @Column(name = "nick_name", length = 32) private String nickName; @Enumerated(EnumType.STRING) // store the enum name, never ORDINAL @Column(nullable = false, length = 16) private UserStatus status; @ManyToOne(fetch = FetchType.LAZY) // many-to-one defaults to EAGER; say LAZY explicitly @JoinColumn(name = "dept_id") private Dept dept; @Version private Integer version; // optimistic-lock version, section 13 public User() { // Hibernate needs the no-arg constructor } @Override public boolean equals(Object o) { if (this == o) return true; if (!(o instanceof User other)) return false; return id != null && id.equals(other.id); // only the id takes part } @Override public int hashCode() { return getClass().hashCode(); // stable, independent of mutable fields } // getters / setters omitted}The annotation cheat sheet (the right column is all stuff you learn by falling into it):
| Annotation | Purpose | The bit everyone forgets |
|---|---|---|
@Entity | declare a persistent entity | needs a no-arg constructor; all-final fields leave Hibernate stuck |
@Table | table name / indexes / unique constraints | omit it and the naming strategy decides — the usual cause of "table does not exist" |
@Id | the primary key | missing it fails startup with No identifier specified for entity |
@GeneratedValue | key generation strategy | IDENTITY degrades batch inserts (see the Point below) |
@Column | name / length / nullability / uniqueness | its unique=true only creates a constraint when the DDL is generated; under validate it checks nothing |
@Enumerated | enum mapping | the default is ORDINAL — always write STRING |
@Lob | large text / binary | on MySQL also write columnDefinition = "TEXT", or @Lob byte[] may become MEDIUMBLOB |
@Transient | not mapped | a different thing from @JsonIgnore: one rules the database, the other the JSON |
@Version | optimistic lock | use Integer / Long; business code must never set it |
@EntityListeners | attach the auditing listener | without @EnableJpaAuditing on a configuration class the audit fields stay null (section 13) |
@GeneratedValuestrategies:IDENTITYuses the database auto-increment column (typical MySQL),SEQUENCEuses a sequence (Oracle / PostgreSQL — keepallocationSizeequal to the sequence increment),AUTOlets the framework choose,TABLEemulates one (essentially never used)equals/hashCodeon the id only: managed entities mutate as flushes happen, so an entity hashed on business fields disappears from aHashSetforever; treatingid == nullas "never equal" is the cheapest correct rule- In a bidirectional association (
User.orders↔Order.user) one side needs@JsonBackReference, or simply return DTOs — otherwise Jackson reportsInfinite recursionandtoString()throwsStackOverflowError
Quiz one — its damage never raises an error, which is why everybody falls for it:
Point: `@GeneratedValue(strategy = IDENTITY)` has a side effect — **an insert must immediately fetch the generated key, which blocks Hibernate's batch inserts** (each insert becomes its own round trip). For bulk imports either switch to `SEQUENCE` or fall back to JDBC batching; this is the precondition for section 10.
Back to the ring from the opening. In JPA an object only ever has one of four identities, and the identity decides the behaviour:
| State | How you enter it | em.contains(u) | Does changing a field send SQL? | Where you usually meet it |
|---|---|---|---|---|
| transient | new User() | false | No — it has nothing to do with the database | after a DTO conversion, a freshly built object |
| managed | persist / find / findById inside a transaction | true | Yes, dirty checking emits the UPDATE at flush | inside a @Transactional method |
| detached | the transaction ends / em.detach() / passed in from elsewhere | false | No, unless you merge it back | Controller return values, async threads, cached objects |
| removed | em.remove(u) while still managed | may still be true until close | the DELETE goes out at flush | deletion flows |
There is only one probe: entityManager.contains(u). The context flips the state, not the object — the object has been the same one in your hands the whole time; what changed is whether anybody is keeping the books.
package com.example.demo.repository;import com.example.demo.entity.User;import com.example.demo.entity.UserStatus;import jakarta.persistence.EntityManager;import org.springframework.stereotype.Repository;import org.springframework.transaction.annotation.Transactional;@Repositorypublic class UserStateProbe { private final EntityManager em; public UserStateProbe(EntityManager em) { this.em = em; } @Transactional // transaction boundary = persistence-context boundary public void probe() { User fresh = new User(); fresh.setUserName("alice"); System.out.println(em.contains(fresh)); // 1) false: transient em.persist(fresh); System.out.println(em.contains(fresh)); // 2) true: managed // note: no INSERT has necessarily gone out (IDENTITY is the exception — it needs the key now) User again = em.find(User.class, fresh.getId()); System.out.println(again == fresh); // 3) true: first-level cache, same instance, no SQL fresh.setStatus(UserStatus.ACTIVE); // 4) only a field in memory changed em.flush(); // 5) dirty checking diffs the snapshot; UPDATE now // 6) the method returns, the context closes: the object in your hand is now detached }}- One id, one instance per transaction: that is JPA's identity guarantee, and the reason "I changed it here and it is visible there" works
- The first-level cache is not a performance tool: it lives exactly as long as the transaction and saves nothing across requests. Cross-request reuse is the job of
@Cacheableand second-level caching (section 12) getReferenceById(id)(formerlygetOne) does not hit the database at all — it fabricates a proxy carrying only the id. Touch another field after the transaction closed and you getLazyInitializationException. Need real values? UsefindById- Stuffing a bulk job into one long transaction makes the context grow until memory hurts before the database does — flush and clear periodically
The third row of that table is the one that bites in production — detached. Its full shape is "endpoint returns 200, no UPDATE in the log, the row unchanged", with not one exception anywhere. Watch these six frames and you will know what to ask first when an update "occasionally does nothing" (the fix is merge in Section 10, the symptom is the last row of Section 14):

Run the four states yourself; it beats rereading the table ten times:
| Argument | What to look for | Corresponding text |
|---|---|---|
| Four entity states | how contains() flips across persist / context close / merge, plus the luggage-check-in analogy | the table above |
| First-level hit | two find calls in one transaction, one statement, the same instance; a new transaction queries again | the third bullet |
the counter-intuitive line in the lab is again == fresh being true. It means you did not get a copy of a row, you got the one and only in-memory stand-in for that row. With that understood, "why did it update although I never called save" and "why do two services see each other's edits" stop being things you memorise.
The ring diagram at the top gives the panorama; what actually has to stick is which action pushes the object into which box. Click through these five boxes — each states what contains() returns right then, and whether anyone is still bookkeeping your field writes:
a Spring Data repository is the order counter of a fast-food shop. You shout "one braised pork" (declare findByStatus) and the kitchen cooks it — you never stepped into the kitchen and you do not know who fried it. Staples printed on the menu (save / findById / deleteById) come from the central kitchen (SimpleJpaRepository); your new dishes are translated on the spot by the name parser (PartTree). And if you mumble a dish that does not exist — findByUname — you do not find out half way through cooking: the shop cannot even open that morning (startup fails). That is the friendliest thing JPA has compared with hand-written SQL: errors move forward.
Spring Data ships a chain of repository interfaces, each richer than the last:
| Interface | Extends | Capabilities |
|---|---|---|
Repository<T, ID> | — | A marker interface, no methods |
CrudRepository<T, ID> | Repository | save / findById / delete / count |
PagingAndSortingRepository<T, ID> | CrudRepository | Adds findAll(Pageable) / findAll(Sort) |
JpaRepository<T, ID> | Above + QueryByExampleExecutor | Adds flush / saveAndFlush / batch delete / getReferenceById |
JpaSpecificationExecutor<T> | separate interface (not on that chain) | Adds findAll(Specification, Pageable) for dynamic criteria |
Why is an interface enough, with no implementation class? At startup Spring Data creates a dynamic proxy for each repository interface (RepositoryFactory behind JdkDynamicAopProxy); the real worker is the built-in SimpleJpaRepository, and any method you add is translated by a name parser or by @Query.
package com.example.demo.repository;import com.example.demo.entity.User;import com.example.demo.entity.UserStatus;import org.springframework.data.jpa.repository.JpaRepository;import org.springframework.data.jpa.repository.JpaSpecificationExecutor;public interface UserRepository extends JpaRepository<User, Long>, // the full CRUD + paging + flush set JpaSpecificationExecutor<User> { // dynamic criteria, needed in section 8 List<User> findByStatus(UserStatus status);}- Extending
JpaRepositorygives you the JPA-specific methods (flush,saveAndFlush); if you need dynamic criteria, extendJpaSpecificationExecutoras well — it is not on the inheritance chain and is the easiest thing to forget - For custom queries, just declare more methods; the proxy weaves the implementation at runtime
SimpleJpaRepository.saveis simply apersistormergebranch — exactly what section 10's animation dissects- Scanning: by default only the package of
@SpringBootApplicationand below are picked up. Repositories living elsewhere need an explicit@EnableJpaRepositories(basePackages = "com.other.repo")
Note: want your own implementation (a Redis counter, say)? The convention is UserRepository plus UserRepositoryImpl — same name, Impl suffix, same package. The proxy routes any method it cannot derive to that fragment class. Both the class name and the package must match exactly, otherwise you get No property found or Could not create query rather than "implementation missing".
This is Spring Data JPA's most charming feature: the method name itself is the query, parsed by convention.
public interface UserRepository extends JpaRepository<User, Long> { // Equivalent to WHERE user_name = ? AND status = ? User findByUserNameAndStatus(String userName, UserStatus status); // Newest first, top 3 List<User> findTop3ByOrderByCreatedAtDesc(); // Fuzzy match on the username List<User> findByUserNameContaining(String keyword); // Status in a set, and username not null List<User> findByStatusInAndUserNameIsNotNull(Collection<UserStatus> statuses); // Path navigation works too: where d.name = ? List<User> findByDept_Name(String deptName); // If you only need existence, do not haul the rows back boolean existsByUserName(String userName); // Derived delete — still needs a transaction long deleteByStatus(UserStatus status);}The derived-keyword cheat sheet (these twelve cover daily work):
| Keyword | Condition | Example |
|---|---|---|
And / Or | AND / OR | findByAAndB |
After / Before | time comparisons | findByCreatedAtAfter |
Is / Equals | equality | findByUserName |
Between | range | findByAgeBetween |
LessThan / GreaterThan / NotNull | ordering and null checks | findByAgeLessThanEqual |
Like / Containing / StartingWith | fuzzy matching | findByUserNameContaining |
In / NotIn | inside / outside a collection | findByStatusIn |
True / False | booleans | findByDeletedTrue |
OrderBy | sorting (Asc / Desc) | findByStatusOrderByIdDesc |
Top / First | row limit | findTop3By... |
Exists / Count | existence / counting | existsByUserName, countByStatus |
Derived methods read left to right. The framework strips the prefix (find / read / query / get / count / exists / delete), then By, and parses the rest as property names separated by And / Or — the class doing that job is PartTree. Names must match entity properties (not column names): findByUname cannot find userName, so the application fails at startup with Failed to create query for method ... No property uname found for type User!. That is the friendliest property of derived queries — the error moves forward, and you cannot ship it.
Two nearby mistakes worth noting: findByDeletedFalse also requires deleted to be a real property; and passing null to findByStatusAndUserNameContaining does not "skip that criterion" — it produces like null and returns nothing. Optional conditions belong to Specification in section 8.
past three conditions a derived method is already hard to read, and Or introduces precedence ambiguity (AAndBOrC — is it (A∧B)∨C or A∧(B∨C)? you end up reading the generated SQL to be sure). As soon as the name becomes a tongue-twister, switch to @Query or Specification: readability beats saving effort.
When a method name cannot express the query, write it with @Query. JPQL is entity-oriented: it uses entity and property names, not table and column names:
public interface UserRepository extends JpaRepository<User, Long> { // JPQL: User is the entity, u.userName the property; named parameters use :status @Query("SELECT u FROM User u WHERE u.status = :status AND u.createdAt > :from") List<User> findActiveSince(@Param("status") UserStatus status, @Param("from") LocalDateTime from); // Native SQL: nativeQuery = true, and only here do you write real table/column names @Query(value = "SELECT * FROM t_user WHERE score > :min ORDER BY score DESC LIMIT 20", nativeQuery = true) List<User> findHighScore(@Param("min") int min); // Modifying query: @Modifying + @Transactional, neither optional @Modifying(clearAutomatically = true, flushAutomatically = true) @Transactional @Query("UPDATE User u SET u.status = :status WHERE u.id IN :ids") int updateStatusBatch(@Param("ids") List<Long> ids, @Param("status") UserStatus status);}- Named parameters
:namebeat positional ones: reordering arguments can no longer silently swap them - With
nativeQuery = trueyou write the real schema — less portable, but you get database-specific functions (JSON_EXTRACT, window functions). Paging through aPageablestill works, but theCOUNTstatement must be supplied withcountQuery @Modifyingmarks an update statement; it must run inside a transaction — without@Transactionalyou getTransactionRequiredException: Executing an update/delete queryclearAutomatically = trueempties the context after the update,flushAutomatically = trueflushes pending changes before it, so the context cannot keep an object the UPDATE has already invalidated
Projection is the badly underrated move: a list page needs three columns, yet the framework assembles whole entities — snapshot, version field, lazy proxies and all — which is slower and one serialization away from a lazy-loading accident.
// Closed interface projection: exactly these three columnspublic interface UserSummary { Long getId(); String getUserName(); String getDeptName(); // matches the alias d.name AS deptName in the JPQL}@Query("SELECT u.id AS id, u.userName AS userName, d.name AS deptName " + "FROM User u JOIN u.dept d WHERE u.status = :status")List<UserSummary> findSummaries(@Param("status") UserStatus status);// Or materialise a real DTO: SELECT new <fqcn>(...) — the constructor must match exactly@Query("SELECT new com.example.demo.dto.UserRow(u.id, u.userName, d.name) " + "FROM User u JOIN u.dept d WHERE u.status = :status")List<UserRow> findRows(@Param("status") UserStatus status);- In interface projection the property names act as aliases; the constructor expression needs the fully qualified class name — a wrong package gives
unable to locate Constructor - The real win is not fewer columns but no entity at all: no snapshot means no dirty checking, no proxy means no
LazyInitializationException, and the response type needs no@JsonIgnore
JPQL can also splice in values with SpEL. This one sorts by a supplied property while a whitelist keeps it safe:
@Query("SELECT u FROM User u WHERE u.status = :status ORDER BY u.#{#sortField} DESC")List<User> findByStatusSorted(@Param("status") UserStatus status, @Param("sortField") String sortField);// Before the call — property names, not column namesprivate static final Set<String> ALLOWED_SORT = Set.of("id", "userName", "createdAt");Warning: a SpEL-concatenated sort field carries the same risk as MyBatis ${}. sortField must be whitelisted server-side; never drop a request parameter straight into JPQL. The Sort inside a Pageable needs the same check — otherwise you get Query was given a Sort on property named xxx which does not exist, turning bad input into a 500, which is information leakage of its own.
JPA paging is "parameter injection": pass a Pageable as a method argument.
public interface UserRepository extends JpaRepository<User, Long> { Page<User> findByStatus(UserStatus status, Pageable pageable);}// Page 0 (Spring Data pages start at 0!), 10 rows, newest firstPageable pageable = PageRequest.of(0, 10, Sort.by(Sort.Direction.DESC, "createdAt"));Page<User> page = userRepository.findByStatus(UserStatus.ACTIVE, pageable);page.getTotalElements(); // total rowspage.getTotalPages(); // total pagespage.getContent(); // current page as List<User>Return a Page<User> as-is and the JSON shape is: content (the page array) plus totalElements, totalPages, number (zero-based), size, first, last. In production wrap it in your own DTO — returning entities exposes version, lazy proxies and the pageable metadata, and usually triggers LazyInitializationException on the way out.
The Page vs Slice difference deserves its own callout:
| Type | Queries the count? | Best for |
|---|---|---|
Page<T> | Runs an extra COUNT(*) | Tables that show "page N of M / total M rows" |
Slice<T> | No count, only whether a next page exists | Infinite scroll / "load more" |
Point: pages starting at **0** is the Spring Data rule, one off from the UI's "page 1"; convert at the API layer or you will always be one page off. Feeds that need no total should use `Slice` and skip a `COUNT` that is expensive on big tables; when a `@Query` has a complicated count, supply your own `countQuery`.
Once conditions multiply (any of them possibly empty), derived names stop working. The proper JPA answer is Specification — a wrapper over the Criteria API, with QueryDSL as the alternative dialect:
package com.example.demo.repository.spec;import com.example.demo.entity.User;import jakarta.persistence.criteria.Predicate;import org.springframework.data.jpa.domain.Specification;import java.util.ArrayList;import java.util.List;public final class UserSpecs { private UserSpecs() {} /** Any field may be null: present means added, absent means it never reaches the SQL */ public static Specification<User> query(String keyword, UserStatus status, LocalDateTime since) { return (root, cq, cb) -> { List<Predicate> ps = new ArrayList<>(); if (keyword != null && !keyword.isBlank()) { ps.add(cb.like(root.get("userName"), "%" + keyword + "%")); } if (status != null) { ps.add(cb.equal(root.get("status"), status)); } if (since != null) { ps.add(cb.greaterThanOrEqualTo(root.get("createdAt"), since)); } return cb.and(ps.toArray(new Predicate[0])); }; }}// The repository must also extend JpaSpecificationExecutor<User>Page<User> page = userRepository.findAll( UserSpecs.query(keyword, status, since), PageRequest.of(0, 20, Sort.by(Sort.Direction.DESC, "createdAt")));- Specifications compose (
spec1.and(spec2)), so "visibility scope" and "status filter" can live as small reusable, testable pieces; they are also type-safe, so a wrong property name fails at compile time - Cross-association predicates use
root.join("dept")— and that defaults to an inner join, silently dropping users without a department. Useroot.join("dept", JoinType.LEFT)to keep them; exactly the kind of detail that makes dynamic criteria harder to get wrong than derived names
决策|Decision: The admin "user list" must support six optional filters plus keyword, date range, sorting and paging, and any combination of them may be empty. Implemented with Spring Data JPA, which approach is the sanest?
- One findByKeywordAndStatusAndSince... derived method, passing null to mean "do not filter"
- One hand-written @Query using tricks like (:kw is null or u.userName like :kw) so conditions switch themselves off
- A Specification (or QueryDSL) that appends predicates only for parameters that are present, handed to findAll(spec, pageable), with the sort field whitelisted
- Concatenate JPQL strings straight from whatever the front end sends — maximum flexibility
Conclusion: C. Derived queries cannot express "optional" — a null parameter becomes a real = null condition and the result set goes empty, the classic null trap. The B trick degrades index choice on MySQL, collapses as soon as there are more columns, and reads badly; D is a SQL-injection incident waiting to happen. Specification was designed for exactly this: predicates added on demand, unit-testable, composable, while paging and counting stay with the framework. Only genuinely heavy statistics fall back to native SQL (see the card in section 17).
The first thing to know about associations is that the defaults are asymmetric:
| Association | Default fetch | Why that default is annoying |
|---|---|---|
@ManyToOne, @OneToOne | EAGER | "fetch a user along with the order" sounds harmless, but it applies to every query; one more association level and you have hidden JOINs everywhere |
@OneToMany, @ManyToMany | LAZY | Lazy by default protects you — until you touch it in a loop (N+1) or after the transaction (exception) |
So the house style is: make @ManyToOne explicitly LAZY, keep @OneToMany lazy and say so, and let the query — not the entity — decide how much to read.
@Entity@Table(name = "t_order")public class Order { @Id @GeneratedValue(strategy = GenerationType.IDENTITY) private Long id; @Column(name = "order_no", nullable = false, length = 32, unique = true) private String orderNo; @Enumerated(EnumType.STRING) private OrderStatus status; // many orders to one user: EAGER by default, LAZY by explicit choice @ManyToOne(fetch = FetchType.LAZY) @JoinColumn(name = "user_id") private User user; // one order, many items: LAZY anyway; mappedBy points at the other side @OneToMany(mappedBy = "order", cascade = CascadeType.ALL, orphanRemoval = true) private List<OrderItem> items = new ArrayList<>(); /** a bidirectional pair must be maintained on both sides */ public void addItem(OrderItem item) { items.add(item); item.setOrder(this); // forget this and the inserted user_id is null }}Now the classic N+1 problem. You want 100 orders with the user who placed each:
List<Order> orders = orderRepository.findByStatus(OrderStatus.PAID); // 1 statementfor (Order o : orders) { System.out.println(o.getUser().getUserName()); // each order tops up with 1 statement}// total: 1 + 100 = 101 statements. The database is fast; the 101 round trips are not类比|Analogy: N+1 is wanting to know each of your 100 students' homeroom teacher. The absurd-but-common way is to walk up to every student and ask "who is your homeroom teacher?" — a hundred errands, each bringing back one name. The right way is to ask the registry office for a single list (one query that brings the needed columns along). JOIN FETCH and @EntityGraph are that list; a projection is the list with only the three columns you need. Laziness itself is not the villain — it is reading the cover before deciding to borrow the book; what turns bad is asking for a hundred unboxings inside a loop.

Three cures, ordered by scope of change:
public interface OrderRepository extends JpaRepository<Order, Long> { // cure 1: @EntityGraph — this method fetches now, the entity stays lazy @EntityGraph(attributePaths = {"user"}) List<Order> findByStatus(OrderStatus status); // cure 2: JOIN FETCH — declared inside the JPQL for exactly this query @Query("SELECT o FROM Order o JOIN FETCH o.user WHERE o.status = :status") List<Order> findPaidWithUser(@Param("status") OrderStatus status); // cure 3: projection — no entity at all, only the columns you need @Query("SELECT o.id AS orderId, o.orderNo AS orderNo, u.userName AS userName " + "FROM Order o JOIN o.user u WHERE o.status = :status") List<OrderRow> findRows(@Param("status") OrderStatus status);}| Cure | Statements | Use it when | Price |
|---|---|---|---|
@EntityGraph | 1 | on a Spring Data repository method, declaratively | it can only fetch whole associations, not columns; watch row multiplication on collections |
JOIN FETCH | 1 | you are writing JPQL anyway | combining it with Pageable has its own traps |
| DTO projection | 1 | list pages, reports, API output | you cannot write back through it; updates go via entities or @Modifying |
open-in-view=true | still 101 | not a cure | moves the top-up queries into view rendering and pins a connection (#12) |
- The root cause is "lazy loading + access in a loop": every uninitialised association is a database hit. The only way to locate it is to count statements —
show-sql,logging.level.org.hibernate.SQL=debug, orhibernate.generate_statistics: true - Do not overuse FETCH JOIN on one-to-many: fetching
itemsanduserin the same graph multiplies rows (100 orders × 5 items = 500 objects), which is slower than N+1, not faster. A common compromise is oneJOIN FETCH o.userplus@BatchSize(size = 50)on the collection - Cascade and
orphanRemovaldecide who is saved and deleted with whom; they decide nothing about fetching. Keep the two questions apart
Run the scene in the kernel lab and watch the statement count collapse from 101 back to 1:
That "there is only one way to locate it: count the statements" line deserves its own experiment. On "how to spot it" you see what show-sql and generate_statistics each put in front of you; on "join fetch" and "batching" you watch 101 collapse to different numbers — and this is also where the nastiest trap appears, fetching two collections at once and meeting MultipleBagFetchException:
To really see how "101" grows one statement at a time, spread it out as a single-step run. The six lines on the left are ordinary code from your own project; the right panel updates the variables. Watch sqlCount climb from 1 to 101, and watch the object in step 2 that is not null but a proxy:
List<Order> list = repo.findByStatus(PAID); // one main query, 100 orders come backfor (Order o : list) { // every o is a managed entity String name = o.getUser().getUserName(); // this reads a proxy log.info("{} placed by {}", o.getOrderNo(), name); // 100 SELECTs live in this log} // loop over: statement count = 101// now rerun it after switching to JOIN FETCH o.user // count goes back to 1| sqlCount | 1 |
| list.size() | 100 |
| user in each element | only user_id came back |
OrderRepository.findByStatusHibernate queryQuiz two — this one happens in production logs every single day:
Now connect section 4's state machine with checkout. The truth about save: it persists an object without an id and merges one that has an id, and both only hand the object to the context — neither means SQL.

Those few famous lines inside SimpleJpaRepository are just this branch:
// org.springframework.data.jpa.repository.support.SimpleJpaRepository (simplified)@Transactionalpublic <S extends T> S save(S entity) { if (entityInformation.isNew(entity)) { // id null / 0 means "new" em.persist(entity); return entity; } return em.merge(entity); // id present: SELECT first, then diff, then UPDATE}isNewlooks at the identifier by default. If your entity uses the primitivelong(never null) it always takespersist, so saving the same row twice collides with the primary key — hence "entity ids are wrapper types, always". To decide newness by a business field, implementPersistable<T>and overrideisNew()instead of faking the id- The first step of
mergeis a SELECT by id (it must learn what the row currently holds and handle cascades). That is why "one UPDATE I wanted" costs an extra query: the price of carrying a detached entity as an update vehicle
The timing table — memorise it and you never have to guess write behaviour again:
| Your action | What happens right now | SQL? |
|---|---|---|
new User() | transient, nothing to do with the database | none |
repository.save(u) (no id) | persist: registered in the context as to-be-inserted | IDENTITY inserts immediately; others wait |
repository.save(u) (id present, detached) | merge: SELECT first, fields copied into a new managed instance | one SELECT |
u.setStatus(ACTIVE) (managed) | only memory and the snapshot diff change | none |
before the next query (FlushModeType.AUTO) | automatic flush, so reads observe your own writes | the difference becomes INSERT/UPDATE/DELETE |
em.flush() / saveAndFlush() | statements go out now (still inside the transaction) | yes, still rollback-able |
@Transactional method returns normally | the interceptor flushes and then commits | the change lands |
| exception triggers rollback | statements already flushed are undone as well | nothing survives |
That table is worth compressing into one picture: everything in the left column happens only inside your memory, while only the right column means the database received a statement. Nine tenths of beginner confusion ("why did nothing change", "why is there an extra SELECT") lands on one of these two boxes:

See these two parameters run and "save is not persistence" plus "why detached updates vanish" become things you watched happen:
This question is the interview classic and the production incident at the same time:
Batching is the other face of flushing: Hibernate does not batch automatically. The standard shape is em.flush(); em.clear(); every 50 items — flush lets the batch actually go out (with hibernate.jdbc.batch_size: 50), clear empties the context, otherwise a million entities sit in the heap and memory hurts before the database does.
not flushing inside the loop costs you exceptions far from the bug. Insert two rows with the same unique key and nothing complains — until the method returns and you get DataIntegrityViolationException: could not execute statement; constraint [uk_user_uname], with a stack pointing at the commit, where you can no longer tell which row was bad. Two habits for bulk work: flush/clear in batches, and use saveAndFlush in tests so the exception lands where the bug is.
That "with hibernate.jdbc.batch_size: 50" clause is the only number in this article you can literally drag, and the counter-intuitive part is its default: 0, meaning no batching at all. You assume the framework is collecting statements; it is firing one per row. Drag it and see when it actually buys anything:
- 50 is the number Hibernate's own documentation and Section 10 both land on
- Each full batch goes out, then flush + clear empties the context
- Prerequisite: the id strategy must not be IDENTITY — it forces an immediate insert per row and defeats batching entirely
- Imports, seed data and message persistence all sit comfortably in this band
IDENTITY and batching are natural enemies — to obtain the generated key, Hibernate must execute each insert immediately, so batch_size becomes decoration. When write volume matters, move to SEQUENCE (or a generator with a matching allocationSize); Section 3 and the Tier-2 experiments both let you verify this with your own eyes.
JPA decides when SQL goes out; the transaction decides whether the changes count at all; caching answers whether this data can be reused across requests.
save succeeded and the SQL went out, yet if the outer transaction rolls back, the database keeps nothing. Conversely, putting audit logs in a new transaction gives you the deliberate "the business failed but the log survived".
| Argument | What to watch | Relation to JPA |
|---|---|---|
| REQUIRED | inner and outer join one transaction; if the inner marks rollbackOnly, the outer's commit fails with UnexpectedRollbackException | two @Transactional methods calling each other — this is the picture |
| REQUIRES_NEW | the inner suspends the outer and takes its own connection; the outer's rollback cannot touch it | audit trails and operation logs; costs two connections at once |
| NESTED | works with savepoints; the inner rolls back only to the savepoint | "one bad row must not kill the batch" in imports |
| Rollback rules | by default only RuntimeException / Error roll back — checked exceptions do not | your save flushed, then the method throws IOException and returns normally: the data stayed, which may not be what you wanted |
@Transactional works through a proxy, so self-invocation inside one class bypasses it and the annotation becomes decoration (#14 has the mechanism). In JPA that is especially nasty: without a transaction the entity you receive is not managed, and your setters silently do nothing — exactly the detach scene from section 10.
Section 4's identity guarantee lives inside one transaction only. Reusing data across requests is the job of @Cacheable and Hibernate's second-level cache, which people constantly confuse — section 12 puts them in one table.
the JPA first-level cache and @Cacheable have nothing in common — the first lives and dies with the transaction and only guarantees identity; the second spans requests and genuinely sheds load. Confusing them produces the classic "I added a cache and it is still slow" plus "there is a detached entity in my cache". The connection-pool view (the pool lab) sits at the end of section 12, because its cause is open-in-view right there.
spring.jpa.open-in-view is one of the few defaults in Boot that bites you, and Boot even logs a WARN telling you to switch it off. What it does is simple: it stretches the persistence context from "the transaction" to "the whole HTTP request" — the EntityManager opens when the request arrives and closes after the view/serialisation is done.
spring: jpa: open-in-view: false # switch it off explicitly; Boot's default is true| Setting | Context lifetime | Reading a lazy field outside a transaction | Side effects |
|---|---|---|---|
open-in-view: true (default) | the whole request | "works" — Hibernate tops up during rendering | the request thread stays tied to a connection; N+1 is postponed into the view and runs serially; slow pages drain the pool |
open-in-view: false | the transaction method | immediate LazyInitializationException | problems surface in development and force you to read everything you need inside the transaction |
With it off, there are exactly three correct answers, all already introduced: read the needed fields inside the transaction, declare the fetch per method (@EntityGraph / JOIN FETCH), project into DTOs. The third has a bonus: pure data is trivially cacheable.
The context binds to the transaction and the transaction binds to a connection, so open-in-view, long transactions and slow queries all end up reporting themselves as connection-pool metrics. Switch to "Queue at the limit" and "Wait timeout" and match the active / idle / awaiting numbers in SQLTransientConnectionException against the table above:
Caching has to be split into clearly separate layers (four rows below — the query cache is only an appendage of the second level), otherwise "I added @Cacheable and it is still slow" stays unsolvable:
| Layer | Scope | Default | Stores | When it is the right tool |
|---|---|---|---|---|
| First-level cache (the context) | one transaction / one EntityManager | on, cannot be turned off | entity instances + snapshots | not an optimisation — an identity guarantee; saves nothing across requests |
| Hibernate second-level cache | SessionFactory level | off | field state of entities (serialised) | read-mostly dictionary tables; you must choose a RegionFactory, and each node still holds its own copy |
| Query cache | same | off | the list of matching ids | only meaningful next to the second-level cache |
@Cacheable (Spring) | your own cache names | needs @EnableCaching | method return values | the right answer for cross-request reuse; put Redis under it in production, see #32 |
@Service@RequiredArgsConstructorpublic class ProductService { private final ProductRepository repository; // cache the DTO, not the entity: out of the context an entity is detached and any lazy field explodes @Transactional(readOnly = true) @Cacheable(value = "product", key = "#id", unless = "#result == null") // needs @EnableCaching public ProductDetail detail(Long id) { Product p = repository.findById(id).orElseThrow(); return new ProductDetail(p.getId(), p.getName(), p.getPrice()); } @Transactional @CacheEvict(value = "product", key = "#id") public void changePrice(Long id, BigDecimal price) { repository.findById(id).orElseThrow().setPrice(price); // managed: flush does the rest }}- Annotation order matters: when
@Cacheablehits, the method body never runs — so no transaction starts either. Put the transaction annotation on an outer service, or be sure the cached method really is read-only - Caching entities is an accident factory: serialise one out and it is detached, so touching an uninitialised association throws
LazyInitializationException - Writes need
@CacheEvict, entries need a TTL; penetration, breakdown and avalanche are covered in #32
Tip: the mnemonic for the three layers is "the context guards identity, the second-level cache guards field state, @Cacheable guards return values". Only one question decides which layer to touch: across what boundary must this data be reused — one transaction (first level), several transactions in one process (second level), or requests across the cluster (@Cacheable + Redis)?
Enough labs — type the same things yourself. This console is wired to the same in-browser kernel and every reply is computed there: query walks Repository → JdbcTemplate once, beans and cond show how this bean got wired in the first place, and then the four themes of this article in order — the four states, the flush moment, N+1, and a pool being slowly drained:
run lab jpa flush and lab jpa detach back to back — the first calls only setters and the UPDATE really goes out; the second calls exactly the same setters and nothing happens. One line of code, two fates, and the only difference is "is someone still keeping the books", which is Section 4's contains().
Audit fields (created time / modifier) need no manual assignment: once auditing is on, the framework fills them during the entity lifecycle.
@EntityListeners(AuditingEntityListener.class) // on the entity@Entitypublic class User { @CreatedDate @Column(name = "created_at", updatable = false) private LocalDateTime createdAt; @LastModifiedDate @Column(name = "updated_at") private LocalDateTime updatedAt; @CreatedBy @Column(name = "created_by", updatable = false) private String createdBy;}@Configuration@EnableJpaAuditing(auditorAwareRef = "auditorProvider") // without this the fields stay nullpublic class JpaConfig { @Bean public AuditorAware<String> auditorProvider() { // real projects read the current user from SecurityContextHolder; empty when anonymous return () -> Optional.ofNullable(SecurityContextHolder.getContext().getAuthentication()) .map(Authentication::getName); }}Optimistic locking prevents concurrent overwrites. Add a @Version field and Hibernate carries the version condition into every update:
@Entity@Table(name = "t_account")public class Account { @Id @GeneratedValue(strategy = GenerationType.IDENTITY) private Long id; private BigDecimal balance; @Version private Integer version; // maintained by Hibernate; business code never touches it}类比|Analogy: optimistic locking is two people each photocopying the same contract. Both write down "version 3" when they copy it (they read version=3). A signs first, and the filing clerk bumps the copy to version 4. B arrives with a sheet that still says "version 3", the clerk sees the number has moved and refuses it on the spot (the UPDATE affected 0 rows, an exception is thrown) and asks B to re-copy the current version and redo the change. At no point did the clerk lock the cabinet (no row lock) — the check happens only at the end. That is why it suits "rarely two people on the same row". If everybody is fighting over the same row, switch to "one person at the filing cabinet" (pessimistic locking).
A concurrent-update failure looks exactly like this:
-- Users A and B both read id=1, version=3-- A commits first:UPDATE t_account SET balance = 100, version = 4 WHERE id = 1 AND version = 3; -- 1 row, success-- B commits next:UPDATE t_account SET balance = 200, version = 4 WHERE id = 1 AND version = 3; -- 0 rows!-- Hibernate: org.hibernate.StaleObjectStateException: Row was updated or deleted by another transaction-- Spring wraps it as: org.springframework.orm.ObjectOptimisticLockingFailureException-- The spec-level exception: jakarta.persistence.OptimisticLockExceptionWhere conflicts really are frequent, use a pessimistic lock — it takes the row lock, SELECT ... FOR UPDATE:
public interface AccountRepository extends JpaRepository<Account, Long> { @Lock(LockModeType.PESSIMISTIC_WRITE) @Query("SELECT a FROM Account a WHERE a.id = :id") Optional<Account> findByIdForUpdate(@Param("id") Long id);}| Aspect | Optimistic (@Version) | Pessimistic (PESSIMISTIC_WRITE) |
|---|---|---|
| Database cost | no lock; the version is just a condition on UPDATE | a row lock; other transactions wait inside the lock window |
| On conflict | exception, caller retries | blocking, slower transactions |
| Fits | read-heavy, rare conflicts (profiles, config, content) | balances, stock, voucher codes |
| Trap | retries need a bound and a back-off | the longer the lock window, the more pool pressure (#28) |
Quiz three — what you do after the conflict:
@Version misbehaves in three classic ways: ① business code calling setVersion(...) — the number is overwritten and you effectively have no lock; ② a bulk JPQL @Modifying UPDATE — it bypasses the entity lifecycle, so the version does not increment; write SET a.version = a.version + 1 yourself; ③ merging a detached entity with an old version — the message mentions unsaved-value mapping was incorrect, which looks like a configuration bug but is a version conflict.
| Message (fragment) | Symptom | Real root cause | One-line fix | Deep dive |
|---|---|---|---|---|
org.hibernate.LazyInitializationException: failed to lazily initialize a collection of role: demo.User.orders, could not initialize proxy - no Session | reading an association outside the transaction (Controller return, Jackson, async thread) | the context closed with the transaction; nobody tops the proxy up | read it inside the transaction, or @EntityGraph / JOIN FETCH / projection — do not paper over it with open-in-view | 9, 12 |
jakarta.persistence.EntityNotFoundException: Unable to find demo.User with id 42 | getReferenceById gave you an object that explodes when you read a field | that method only builds a proxy; the row is gone or never existed | use findById(...).orElseThrow(...) and turn "missing" into your own business exception | 4 |
org.hibernate.StaleObjectStateException / ObjectOptimisticLockingFailureException: Row was updated or deleted by another transaction (or unsaved-value mapping was incorrect): [demo.User#42] | concurrent updates fail now and then; a retry works | committed with an old version, or merged a detached entity whose version moved | retry outside the transaction (reload, recompute, commit); switch to pessimistic locking if conflicts are dense | 13 |
jakarta.persistence.NonUniqueResultException: query did not return a unique result: 2 | a method returning Optional<User> blows up | the predicate matched several rows: no unique index on user_name, or a JOIN multiplied rows | add the unique constraint, or return List / findTop1By... | 6 |
jakarta.persistence.TransactionRequiredException: Executing an update/delete query | the @Modifying method throws the moment it is called | missing @Transactional, or a self-invocation that bypassed the proxy | add @Transactional and make sure the call comes from another bean | 7, 11 |
InvalidDataAccessApiUsageException: No EntityManager with actual transaction available for current thread - cannot reliably call 'flush' | flush / saveAndFlush fail from the inside | no transaction anywhere on the call chain, hence no real context | put @Transactional on the service, do not call flush from the controller | 10 |
Failed to create query for method ... No property uname found for type User! | startup fails | derived-query property name does not exist on the entity (case, or you used the column name) | match the entity property; fall back to @Query when it gets clever | 6 |
Schema-validation: missing column [remark] in table [t_user] (usually wrapped by Unable to build Hibernate SessionFactory) | the release dies at startup | ddl-auto: validate detected schema drift | add a Flyway / Liquibase script; do not switch to update to make it quiet | 2 |
DataIntegrityViolationException: could not execute statement ... Duplicate entry 'alice' for key 'uk_user_uname' with the stack pointing at the method return | the loop stays quiet and everything explodes on commit | statements were deferred to flush, so the exception is far from its cause | flush/clear in batches; use saveAndFlush in tests to bring the exception back to the bug | 10 |
java.lang.StackOverflowError / Jackson Infinite recursion (class demo.Order) | the endpoint dies on return or floods the log | the bidirectional pair references itself; both toString and serialisation recurse | project to DTOs; if you must return entities, @JsonManagedReference + @JsonBackReference | 3 |
| no error at all, but the update did not happen | the endpoint returns 200 and the column keeps its old value | ① readOnly=true skipped dirty checking; ② you mutated a detached entity; ③ the context still holds the old object after a @Modifying | mutate a managed entity inside a transaction; @Modifying(clearAutomatically = true); keep readOnly on read methods only | 4, 10 |
the last row is the one to copy into your notes — JPA's most expensive failures are not exceptions but silent no-ops. The diagnostic move is always the same: turn on show-sql or org.hibernate.SQL=DEBUG, count the statements of one request, look at the parameters, and check whether an UPDATE appears at all. The statement count never lies.
The no Session in the first row is the message that frightens JPA newcomers most, because it reads like "the database connection disappeared" — and it has nothing to do with connections. Here is a real stack: do not read the analysis, just click the frame you think is guilty:
The controller hands the User returned by the service straight to Jackson. The endpoint 500s intermittently: local unit tests (with no transaction boundary of their own) pass, integration testing fails.
Goal: get User/Dept running and watch the statement count change with the fetch strategy. Dept has just a name field; User.dept is @ManyToOne(fetch = LAZY). UserRepository extends both JpaRepository<User, Long> and JpaSpecificationExecutor<User>, and declares four queries: the derived findByStatusAndUserNameContaining, the paged Page<User> findByStatus(UserStatus, Pageable), the declared-fetch findWithDept (@EntityGraph(attributePaths = "dept")) and the three-column projection findSummaries from section 7. Plus one test:
@SpringBootTestclass UserJpaTest { @Autowired UserRepository repository; @Autowired TestEntityManager entityManager; @Test void managedEntityNeedsNoSave() { User u = repository.save(newUser("alice")); // persist: an UPDATE has not necessarily gone out u.setNickName("Ali"); // memory only repository.flush(); // dirty checking emits the UPDATE here entityManager.clear(); // simulate the context closing assertThat(repository.findById(u.getId()).orElseThrow().getNickName()).isEqualTo("Ali"); }}Expected console output (with show-sql: true the shape must match):
Hibernate: insert into t_user (nick_name, status, user_name, dept_id, version) values (?, ?, ?, ?, default)Hibernate: select u.id,u.user_name,u.nick_name,u.status,u.dept_id,u.version from t_user u where u.id=?Hibernate: update t_user set nick_name=?, status=?, dept_id=?, user_name=?, version=? where id=? and version=?Hibernate: select d.id,d.name from t_dept d where d.id=? -- per-row top-up, N times-- with findWithDept it collapses to one:Hibernate: select u.id,..., d.id, d.name from t_user u inner join t_dept d on d.id=u.dept_id where u.status=?- after the
insertthere is no immediate update — it appears only atflush(). That is "save ≠ persistence" - the
updateends withand version=?and bumpsversion: optimistic locking in full - the projection statement never mentions
version, because no entity is materialised and dirty checking is out of the picture
- Set
open-in-view: falseand readuser.getDept().getName()in the Controller → you get the exactLazyInitializationExceptionfrom row one of section 14. The most valuable experiment here: the error is not the disaster, it moves the landmine from the end of the request back into the service. - Change
@ManyToOne(fetch = LAZY)to EAGER → 101 statements become 1, and now every statistic that only wanted an order number JOINs the user table too. Feel a local optimisation turning into a global tax. - Call
setVersion(3)by hand before saving → watch the version behave differently from what you expected (section 13). - Change
PageRequest.of(0, 10)toof(1, 10)→ the first page disappears; UI page 1 is index 0. - Call
setXxxinside a@Transactional(readOnly = true)method → clean commit, no UPDATE in the log, database unchanged. The bugs that never error are the hardest to find. - Set
ddl-auto: update, add a field to the entity and restart →alter table ... add columnappears. Switch back tovalidateand write the migration script instead.
Build GET /admin/orders?keyword=&status=&deptId=&from=&to=&page=1&size=20&sortField=createdAt&order=desc. Acceptance list:
- [ ]
Order.userexplicitlyLAZY;equals/hashCodeon id only;@Versionpresent and untouched by business code - [ ] Dynamic criteria via
Specification; when a parameter is absent its condition must not appear in the SQL (no:x is null or ...tricks) - [ ] The sort field passes a server-side whitelist (entity property names), illegal values fall back to
createdAt - [ ] Ship both a
Sliceversion (no count) and aPageversion (with count); paste the two SQL logs in the README and explain why feeds takeSlice - [ ] The list output is a projected DTO; serialising it produces neither per-id top-ups nor
LazyInitializationException - [ ]
PATCH /admin/orders/statusbulk update with@Modifying(clearAutomatically = true)+@Transactional; explain why that UPDATE does not bumpversionand how you handle it - [ ] Load-test 100 concurrent callers while watching
hikaricp.connections.pendingand the SQL log; decide whether the bottleneck is statement count, mapping overhead or the pool, in five lines
Finish all three tiers and your grasp of JPA moves from "I can write the annotations" to "I know who decided to send each statement, and which layer to debug when it goes wrong".
after repository.save(u), is that row in the database? Answer: not necessarily. save only persists or merges — it hands the object to the context. SQL goes out at flush, and flush happens by default before the next query and at commit.
no save, no UPDATE statement, you only changed a field on a managed entity — does the data change? Answer: yes. "Snapshot at load + dirty check before commit" is the whole mechanism.
two findById(1L) calls in one transaction — how many statements? Answer: one, and both calls return the same instance (== is true). A new transaction hits the database again.
name the three correct fixes for LazyInitializationException. Answer: read everything you need inside the transaction, declare the fetch per method (@EntityGraph / JOIN FETCH), or project to DTOs. open-in-view hides it; it does not fix it.
how does N+1 appear and how do you locate it? Answer: laziness plus access in a loop; locate it by counting statements (show-sql, org.hibernate.SQL=debug, or generate_statistics).
what should ddl-auto be in production, and how does a schema change get applied? Answer: validate (or none), and changes go through Flyway / Liquibase scripts that are reviewed like code.
save only puts it in the trolley; checkout (flush) is the real purchase. readOnly skips checkout, and detached trolleys record nothing. Laziness is the cover, not the book; touch it in a loop and you get N+1. @EntityGraph fixes one query, join fetch fixes one, a projection is cleanest. The context guards identity, the second-level cache guards fields, @Cacheable guards return values. Validate in production, script your migrations, and on a version conflict reread before you rewrite.
the skeleton of this article is three sentences. First: JPA manages object state, not SQL — transient / managed / detached / removed are decided by the persistence context, contains() is the only honest test, and the transaction boundary is where states flip. Second: SQL is decided by flush, not by how many lines you wrote — dirty checking explains "updated without an UPDATE", deferred flushing explains "the exception is nowhere near the bug", and readOnly explains "nothing happened and nothing complained". Third: laziness is a double-edged tool — it makes "judge by the cover" possible and turns loops into 101 statements; the cure is always to fetch what this call needs in one go: @EntityGraph, JOIN FETCH, projection. Add two engineering disciplines — never ddl-auto: update in production, switch open-in-view off explicitly — and JPA stops being black magic and becomes a tool you can predict, diagnose and tune.