ByteScrollGet the app
☰ Topics
parallel-streams6 / 200‹›
JAVA / FUNCTIONAL3 minute read

Parallel streams — when do they actually help?

Hard

.parallel() splits the source and runs chunks on the common ForkJoinPool. It pays off for large, easily split data (arrays, ArrayList, ranges) with CPU-heavy, independent work per element. It hurts for small inputs, blocking I/O, LinkedList sources, ordering and shared state.

How it works

  1. Split. The stream asks the source's Spliterator to trySplit() again and again, building a tree of chunks.
  2. Run. Each chunk becomes a task on ForkJoinPool.commonPool(). That pool has availableProcessors() - 1 workers by default, and the calling thread helps too.
  3. Combine. Partial results are merged up the tree with the combiner (reduce, collect).
  4. The bill. Splitting, task scheduling, and merging all cost time. Parallel only wins when the per-element work across all elements is big enough to dwarf that overhead.

What helps:

  • Lots of elements, or few elements that are each expensive (pure CPU work).
  • Sources that split by index: arrays, ArrayList, IntStream.range.
  • Stateless operations and cheap, associative merges (sum, max, count).

What hurts:

  • Small collections. Overhead beats the gain.
  • Blocking I/O inside the pipeline. It ties up the shared pool for the whole JVM.
  • LinkedList, Stream.iterate, BufferedReader.lines(): they split badly.
  • Order-sensitive steps on ordered streams: limit, findFirst, sorted, forEachOrdered.
  • Boxing (Stream<Integer> instead of IntStream) and merges that copy a lot (toList of big chunks, grouping into maps).
IntStream.range(0,8_000_000)chunk 0..4Mchunk 4M..8M0..2M onworker 12M..4M onworker 24M..6M onworker 36M..8M oncaller thread++total
IntStream.range(0,8_000_000)chunk 0..4Mchunk 4M..8M0..2M onworker 12M..4M onworker 24M..6M onworker 36M..8M oncaller thread++total

Example

Example.javaJava
static boolean isPrime(int n) {
    if (n < 2) return false;
    for (int d = 2; (long) d * d <= n; d++) if (n % d == 0) return false;
    return true;
}

// Good fit: big range, CPU-only work, cheap merge
long primes = IntStream.range(0, 5_000_000).parallel().filter(Main::isPrime).count();

// Broken: many threads adding to a plain ArrayList
List<Integer> found = new ArrayList<>();
IntStream.range(0, 10_000).parallel().filter(Main::isPrime).forEach(found::add); // lost items or exceptions

// Correct: let the stream build the list
List<Integer> safe = IntStream.range(0, 10_000).parallel().filter(Main::isPrime).boxed().toList();

Measure with JMH or at least a warmed-up loop before claiming a speedup.

Edge cases

  • reduce(identity, op) must use a true identity. reduce(10, Integer::sum) adds 10 once per chunk in parallel, so the answer depends on how the data split.
  • findAny() and unordered() let a parallel stream skip the bookkeeping that keeps encounter order.
  • Running a parallel stream inside myPool.submit(...) makes it use myPool. This works in practice but is an implementation detail, not a documented guarantee.
  • The common pool's size can be set with -Djava.util.concurrent.ForkJoinPool.common.parallelism=N.

Common mistakes

  • Adding .parallel() to everything as a free speed boost.
  • Calling a remote service or database in a parallel stream. For blocking work, use an executor (virtual threads on Java 21+).
  • Mutating shared state from lambdas instead of using collect/reduce.

Likely follow-up

"Why is parallel().sorted().forEachOrdered() often no faster than sequential?" Sorting needs every element before it can emit any, and forEachOrdered forces results back into a single ordered line, so most of the parallel work gets serialized again.

Get every deep dive in the app

Coming soon to the App StoreComing soon to Google Play