You can use Collections.frequency:
numbers.stream().filter(i -> Collections.frequency(numbers, i) >1)
.collect(Collectors.toSet()).forEach(System.out::println);
Answer from Bao Dinh on Stack OverflowYou can use Collections.frequency:
numbers.stream().filter(i -> Collections.frequency(numbers, i) >1)
.collect(Collectors.toSet()).forEach(System.out::println);
Basic example. First-half builds the frequency-map, second-half reduces it to a filtered list. Probably not as efficient as Dave's answer, but more versatile (like if you want to detect exactly two etc.)
List<Integer> duplicates = IntStream.of( 1, 2, 3, 2, 1, 2, 3, 4, 2, 2, 2 )
.boxed()
.collect( Collectors.groupingBy( Function.identity(), Collectors.counting() ) )
.entrySet()
.stream()
.filter( p -> p.getValue() > 1 )
.map( Map.Entry::getKey )
.collect( Collectors.toList() );
simplify java stream to find duplicate properties - Stack Overflow
How to check if exists any duplicate in Java 8 Streams? - Stack Overflow
How to find duplicate values in a stream without loading in memory
HashSet containing duplicates!?
One solution is
var duplicate = users.stream()
.collect(Collectors.toMap(User::getName, u -> false, (x,y) -> true))
.entrySet().stream()
.filter(Map.Entry::getValue)
.map(Map.Entry::getKey)
.collect(Collectors.toSet());
This creates an intermediate Map<String,Boolean> to record which name is occurring more than once. You could use the keySet() of that map instead of collecting to a new Set:
var duplicate = users.stream()
.collect(Collectors.collectingAndThen(
Collectors.toMap(User::getName, u -> false, (x,y) -> true, HashMap::new),
m -> {
m.values().removeIf(dup -> !dup);
return m.keySet();
}));
A loop solution can be much simpler:
HashSet<String> seen = new HashSet<>(), duplicate = new HashSet<>();
for(User u: users)
if(!seen.add(u.getName())) duplicate.add(u.getName());
Group by the names, find entries with more than one value:
Map<String, List<User>> grouped = users.stream()
.collect(groupingBy(User::getName));
List<User> duplicated =
grouped.values().stream()
.filter(v -> v.size() > 1)
.flatMap(List::stream)
.collect(toList());
(You can do this in a single expression if you want. I only separated the steps to make it a little more clear what is happening).
Note that this does not preserve the order of the users from the original list.
Your code would need to iterate over all elements. If you want to make sure that there are no duplicates simple method like
public static <T> boolean areAllUnique(List<T> list){
Set<T> set = new HashSet<>();
for (T t: list){
if (!set.add(t))
return false;
}
return true;
}
would be more efficient since it can give you false immediately when first non-unique element would be found.
This method could also be rewritten using Stream#allMatch which also is short-circuit (returns false immediately for first element which doesn't fulfill provided condition)
(assuming non-parallel streams and thread-safe environment)
public static <T> boolean areAllUnique(List<T> list){
Set<T> set = new HashSet<>();
return list.stream().allMatch(t -> set.add(t));
}
which can be farther shortened as @Holger pointed out in comment
public static <T> boolean areAllUnique(List<T> list){
return list.stream().allMatch(new HashSet<>()::add);
}
Out of the three following one-liners
return list.size() == new HashSet<>(list).size();return list.size() == list.stream().distinct().count();return list.stream().allMatch(ConcurrentHashMap.newKeySet()::add);
the last one (#3) has the following benefits:
- Performance: it short-circuits, i.e. it stops on the first duplicate (while #1 and #2 always iterate till the end) โ as Pshemo commented.
- Versatility: it allows to handle not only collections (e.g. lists), but also streams (without explicitly collecting them).
Additional notes:
โข ConcurrentHashMap.newKeySet() is a way to create a concurrent hash set.
โข Previous version of this answer suggested the following code: return list.stream().sequential().allMatch(new HashSet<>()::add); โ but, as M. Justin commented, calling BaseStream#sequential() doesn't guarantee that the operations will be executed from the same thread and doesn't eliminate the need for synchronization.
โข Performance benefit from short-circuiting won't usually matter, because we typically use such code in assertions and other correctness checks, where items are expected to be unique; in contexts where duplicates are really a possibility, we typically not only check for their presence but also to find and handle them.
โข M. Justin's provides a clean (not using side-effects) short-circuiting solution for Java 22, but that's not a one-liner.
I am working with a large stream and need to find out whether any duplicate values exist in the stream. I can do this when the stream is small by loading everything in a `Seq` and calling distinct however this doesn't work with large streams since I can't hold everything in memory.
Here is an example:
I have a source like this:
1 | red | light | 10
2 | blue | dark | 20
1 | brown | light | 2
1 | red | light | 10
20 | grey | dark | 200
I want to find out (`true / false`) whether there are any identical items in the source. In the above stream `1 | red | light | 10` would be identical. This stream could be very big over 2M records. I can return `true` soon as the identical item is found (i.e. in the example above, we can avoid reading `20 | grey | dark | 200`). What is the best way to do this? I tried reading the entire source into a `List(String)` and ran distinct on it. This works ok, however, for large sources I start receiving a OOM error.
val restResult: Future[immutable.Seq[Color]] =
mySource(ctx)
.via(framing("\n"))
.map(_.utf8String)
.map(_.trim)
.map(s => ColorParser(s))
.collect {
case Right(color) => color
}
.runWith(Sink.seq)