Java Virtual Threads in Payment APIs
bankingOctober 8, 2026

Java Virtual Threads in Payment APIs

What Changes and What Doesn't

After years building and maintaining Java payment services for banks: payment initiation, routing to clearing systems, and the calls to fraud, sanctions and limits services that each payment waits on. Virtual threads come up in almost every architecture discussion now, usually with the question "should we just turn them on?" With Java 25 LTS and Spring Boot 4, my answer is mostly yes, but the platform team should know exactly which bottlenecks this removes and which ones it only moves somewhere else. 


What virtual threads actually change 


A typical Spring Boot payment API spends most of its time waiting. It waits on the database, on the core banking system, on the fraud scoring provider, on sanctions screening and on the payment scheme gateway. With classic platform threads, every waiting request holds an operating-system thread, so throughput is capped by the size of the Tomcat thread pool (200 by default) long before the CPU is busy. 


Virtual threads (final since JDK 21, JEP 444) are lightweight threads managed by the JVM. When one blocks on I/O, the JVM unmounts it from its carrier thread and runs another one in its place. Blocking code stays blocking and readable, but waiting no longer costs an OS thread. 


In Spring Boot 4, enabling them is one property: 


spring: 

  threads: 

    virtual: 

      enabled: true 

  


With that property set, Tomcat request handling, @Async methods and Spring's task executors run on virtual threads. The business code doesn't change. 


Where they help a payment service 


The biggest gains appear in the parts of the system that are I/O-bound and use a thread per request: 


Payment orchestration endpoints that call four or five downstream services in sequence. Each call that waits no longer occupies a scarce platform thread. 


Fan-out checks. Running sanctions, fraud and limits checks in parallel becomes simple, readable code: 


try (var executor = Executors.newVirtualThreadPerTaskExecutor()) { 

    Future<SanctionsResult> sanctions = executor.submit(() -> sanctionsClient.screen(payment)); 

    Future<FraudScore>      fraud     = executor.submit(() -> fraudClient.score(payment)); 

    Future<AccountLimits>   limits    = executor.submit(() -> limitsClient.fetch(payment.debtorAccount())); 

 

    return decisionEngine.decide(sanctions.get(), fraud.get(), limits.get()); 

} 

  


Traffic peaks such as salary days, end-of-month batches and retail events. These no longer cause request-thread exhaustion, which used to show up as rising latency across every endpoint, including the ones unrelated to the load. 


The less visible benefit is about the codebase. Many banks introduced reactive stacks such as WebFlux mainly to work around thread limits. For new I/O-heavy services, virtual threads deliver much of that scalability with plain imperative code that is easier to review, debug and staff for. 


What Java 25 fixed: synchronized no longer pins 


On Java 21, the main caveat was pinning. A virtual thread that blocked inside a synchronized block stayed attached to its carrier thread. With only as many carriers as CPU cores, a few pinned threads waiting on a slow downstream call could stall the whole service. This mattered in practice, because JDBC drivers, connection pools and older client libraries used synchronized widely. 


JEP 491, delivered in JDK 24 and therefore part of Java 25 LTS, removes this limitation. Virtual threads can now block inside synchronized methods and blocks, and in Object.wait(), without holding their carrier. For teams on Java 25, the advice to rewrite every synchronized as a ReentrantLock no longer applies. 


Where pinning still bites 


Pinning is now rare, but it hasn't disappeared. The cases that remain: 


Native code. A virtual thread that calls a native method through JNI, or through the Foreign Function & Memory API, stays pinned for the duration of the native call. In banking this is not an edge case: HSM client libraries and PKCS#11 providers used for PIN translation, MAC generation and payment signing are often native. A slow HSM round-trip pins a carrier. 


Class initialisation. Blocking inside a static initialiser, or waiting for another thread to finish initialising a class, still pins. 


Both cases are easy to find before production. JDK Flight Recorder records a jdk.VirtualThreadPinned event: 


java -XX:StartFlightRecording:settings=profile,filename=payments.jfr -jar payment-api.jar 

jfr print --events jdk.VirtualThreadPinned payments.jfr 

  


Run this during a load test that includes the signing and HSM paths. For confirmed native hot spots, the usual fix is to send those calls to a small, bounded pool of platform threads, which keeps carrier threads free for everything else. 


What doesn't change 


Virtual threads remove the thread bottleneck. They do not remove any other bottleneck, and in a payment service the others are what matter most. 


Your connection pool is now the limit. With 200 platform threads, a HikariCP pool of 40 connections rarely saw more than 200 requests waiting for it. With virtual threads, ten thousand concurrent requests can queue for those same 40 connections. Keep maximum-pool-size sized to what the database can actually handle, and set a short connection-timeout so overload fails fast rather than turning into long latency. 


Downstream services still have limits. A fraud scoring vendor with a contractual rate limit, or a core banking API sized for a fixed number of concurrent sessions, will not scale because your service can now send more requests. The thread pool used to limit concurrency without anyone noticing. That limit now has to be explicit: 


@Component 

public class FraudScoringClient { 

 

    private final Semaphore permits = new Semaphore(50); // agreed concurrency with the provider 

    private final RestClient restClient; 

 

    public FraudScoringClient(RestClient.Builder builder) { 

        this.restClient = builder.baseUrl("https://fraud.internal").build(); 

    } 

 

    public FraudScore score(PaymentRequest payment) { 

        try { 

            if (!permits.tryAcquire(200, TimeUnit.MILLISECONDS)) { 

                throw new DownstreamSaturatedException("fraud-scoring"); 

            } 

        } catch (InterruptedException e) { 

            Thread.currentThread().interrupt(); 

            throw new DownstreamSaturatedException("fraud-scoring"); 

        } 

        try { 

            return restClient.post().uri("/score").body(payment).retrieve().body(FraudScore.class); 

        } finally { 

            permits.release(); 

        } 

    } 

} 

  


The same principle applies to bulkheads. A Resilience4j semaphore bulkhead fits virtual threads naturally, while a thread-pool bulkhead brings back the platform threads you just removed. 


CPU-bound work gains nothing. Signature verification, ISO 20022 XML parsing and validation, and encryption keep the same throughput, because they need a CPU core, not a thread that can wait. 


ThreadLocal needs rethinking. Code that caches expensive objects per thread (formatters, parsers, buffers) multiplies that memory by the number of virtual threads. For request context such as the payment ID, correlation ID and tenant, Java 25 adds Scoped Values (JEP 506, now final). They are immutable and bounded to a defined scope, and they are cheap with virtual threads: 


private static final ScopedValue<PaymentContext> CONTEXT = ScopedValue.newInstance(); 

 

ScopedValue.where(CONTEXT, new PaymentContext(paymentId, tenantId)) 

           .run(() -> paymentProcessor.process(payment)); 

  


Existing MDC-based logging keeps working. Migrating request context to Scoped Values can follow gradually. 


What to leave alone 


Not every service should move: 


A stable WebFlux service that meets its SLAs. A rewrite to imperative code brings migration risk with little operational gain. Use virtual threads for new services and for the ones that are being re-engineered anyway. 


Kafka consumers. Their concurrency comes from partitions, not threads. Virtual threads don't increase throughput beyond the partition count. 


CPU-heavy batch jobs, such as end-of-day reconciliation and statement generation. Size these on cores, as before. 


Preview APIs in production. Structured concurrency (StructuredTaskScope) is still in preview in Java 25. It is worth trying in development, but most banks' change and risk processes will want it final before it reaches a payment path. 


A rollout plan that keeps risk low 


Upgrade to Java 25 and Spring Boot 4 with virtual threads disabled, and release that change on its own. 


Make concurrency limits explicit: connection pools, semaphores per downstream service, and timeouts on every client. 


Enable virtual threads in a performance environment. Run a load profile that includes the HSM and signing paths, with JFR recording. 


Review jdk.VirtualThreadPinned events and thread dumps (jcmd <pid> Thread.dump_to_file -format=json threads.json). 


Roll out one service at a time, starting with the most I/O-bound orchestration API, and compare p99 latency and error rates during a real peak. 


Virtual threads are one of the most useful changes in recent Java for payment services: simpler code, more headroom and less pressure to adopt reactive stacks. The gain is real when concurrency limits that the thread pool used to enforce implicitly are designed in explicitly from the start.