Either is fine. Pick the one that seems more readable to you. If the calculation naturally decomposes, as this one does, then the multiple maps is probably more readable. Some calculations won't naturally decompose, in which case you're stuck at the former. In neither case should you be worrying that one is significantly more performant than the other; that's largely a non-consideration.
Answer from Brian Goetz on Stack OverflowJava 8 Stream multiple files flatmap to lines - Stack Overflow
sql - Parsing multiple line records using Java 8 Streams - Code Review Stack Exchange
java - multiple lines of code in stream.forEach - Stack Overflow
Java Streams - Map two string lines each to one object - Stack Overflow
Either is fine. Pick the one that seems more readable to you. If the calculation naturally decomposes, as this one does, then the multiple maps is probably more readable. Some calculations won't naturally decompose, in which case you're stuck at the former. In neither case should you be worrying that one is significantly more performant than the other; that's largely a non-consideration.
use filter before map.
data.stream().filter(Objects::nonNull).
.map(x -> x*x)
filter(Objects::nonNull) try this instead .
Simply put, you only need a Consumer<String>-implementing class that knows what lines to read as the table name, the columns and rows.
public class StatementGenerator implements Consumer<String> {
private static String START = " (";
private static String END = ") ";
private static String ROW_DELIMITER = "\n";
private String delimiter = "\t";
private String parameter = "%";
private ObjIntConsumer<List<String>> parameterSupplier =
(list, i) -> { list.add(parameter); };
private String table = null;
private List<String> columns = new ArrayList<>();
private List<String> rows = new ArrayList<>();
public void setDelimiter(String delimiter) {
this.delimiter = delimiter;
}
public void setParameter(String parameter) {
this.parameter = parameter;
}
public void setParameterSupplier(ObjIntConsumer<List<String>> parameterSupplier) {
this.parameterSupplier = parameterSupplier;
}
@Override
public void accept(String t) {
if (table == null) {
table = t;
return;
}
if (!t.contains(delimiter)) {
columns.add(t);
return;
}
rows.add(t);
}
public String getTableName() {
return table;
}
public List<String> getColumns() {
return columns;
}
public List<String> getRows() {
return rows;
}
public String getParameterizedStatement() {
StringBuilder result = new StringBuilder("INSERT INTO ");
result.append(constructSegment(getTableName(), getColumns())).append(
constructSegment("VALUES", IntStream.rangeClosed(1, getColumns().size())
.collect(ArrayList::new, parameterSupplier, List::addAll)));
return result.append(";").toString();
}
public List<String> getRawStatements() {
String placeholderStatement = getParameterizedStatement()
.replaceAll(Pattern.quote(parameter), "%s");
return getRows().stream().map(r -> String.format(placeholderStatement,
Pattern.compile(delimiter).splitAsStream(r)
.map(v -> "'" + v + "'").toArray()))
.collect(Collectors.toList());
}
public String getFullStatement() {
return getRawStatements().stream().collect(Collectors.joining(ROW_DELIMITER));
}
@Override
public String toString() {
return getParameterizedStatement();
}
private static String constructSegment(String prefix, List<String> list) {
return prefix + list.stream().collect(Collectors.joining(", ", START, END));
}
}
In my sample implementation above, I've made the delimiter configurable, but there's one other configuration I want to highlight - parameter.
SQL injection
The sample implementation provides a getParameterizedStatement() as the starting point, because by right you should let your database driver handle the escaping of quotes (and possibly other magic values). We certainly do not need Little Bobby Tables to hang around here.
The parameter list is driven by an ObjIntConsumer, because if I'm not mistaken, there are certain drivers allowing for an index-based parameter substitution. That will mean you can potentially override parameterSupplier with something like (list, i) -> { list.add("::" + i); }; (or whatever the placeholder format is).
With that said, assuming you can absolutely trust your file-based input (i.e. sanitized inputs), we can then proceed with using getRawStatements(), which performs a simple substitution on the parameter with our String.format()'s "%s" placeholder. This is more aligned with what you are asking for. getFullStatement() simply concatenates all the rows into a single String, if that is what you actually require.
Here's what I used in my main() code:
public class StatementGeneratorMain {
public static void main(String[] args) throws IOException, URISyntaxException {
StatementGenerator generator = new StatementGenerator();
generator.setDelimiter(",");
try (Stream<String> lines = Files.lines(Paths.get(
ClassLoader.getSystemResource("sqlRecords.txt").toURI()))) {
lines.forEach(generator);
}
System.out.println(generator.getTableName());
System.out.println(generator.getColumns());
System.out.println(generator.getParameterizedStatement());
generator.getRawStatements().forEach(System.out::println);
}
}
And the sample output:
STUDENTS
[ID, NAME]
INSERT INTO STUDENTS (ID, NAME) VALUES (%, %) ;
INSERT INTO STUDENTS (ID, NAME) VALUES ('1', 'Mike') ;
INSERT INTO STUDENTS (ID, NAME) VALUES ('2', 'Kimberly') ;
Wow, that's a long lambda. It deserves to be a method of its own.
There is a fundamental limit to how much can be done using functional programming, since the very act of reading from the BufferedReader alters its state. In Haskell, for example, all I/O is a bit of a hassle, using the IO monad to indicate that the function has side-effects of consuming input or producing output.
A further complication here is that the first few lines are to be treated differently. Moreover, the result from the first few lines is to be used in the rest of the response. Making a Stream consisting of all lines of the file is therefore not a great idea.
For those reasons, I recommend abandoning the idea of using streams for processing the header. Here's my suggestion:
import java.io.BufferedReader;
import java.io.InputStreamReader;
import java.io.IOException;
import java.util.ArrayList;
import java.util.stream.Collectors;
import java.util.stream.Stream;
public class StatementGeneratorMain {
private static BufferedReader getBufferedReader(String fileName) {
ClassLoader cl = StatementGeneratorMain.class.getClassLoader();
return new BufferedReader(new InputStreamReader(cl.getResourceAsStream(fileName)));
}
private static Stream<String> toSql(BufferedReader br) throws IOException {
String tableName = br.readLine();
ArrayList<String> columnNames = new ArrayList<>();
do {
br.mark(256); // max length of column name or first data column
String line = br.readLine();
if (line.indexOf("\t") >= 0) {
// End of header; this is a row of data
br.reset();
break;
}
columnNames.add(line);
} while (true);
String columns = columnNames.stream().collect(Collectors.joining(", "));
return br.lines()
.map(row -> row.replace("'", "''").replace("\t", "', '"))
.map(row -> String.format("INSERT INTO %s (%s) VALUES('%s');",
tableName, columns, row));
}
public static void main(String[] args) throws IOException {
try (BufferedReader br = getBufferedReader("STUDENTS.txt")) {
toSql(br).forEach(System.out::println);
}
}
}
Note that this approach is vulnerable to SQL injection. I've done some rudimentary escaping using row.replace("'", "''"), but that might not protect you against all special characters.
There are two completely different things you should ask here:
a) how do I place multiple lines of code in stream.forEach()?
b) what should I do to count the number of lines in a Stream?
The question b) is answered already by other posters; on the other hand, the general question a) has a quite different answer:
use a (possibly multi-line) lambda expression or pass a reference to multi-line method.
In this particular case, you'd either declare i a field or use a counter/wrapper object instead of i.
For example, if you want to have multiple lines in forEach() explicitly, you can use
class Counter { // wrapper class
private int count;
public int getCount() { return count; }
public void increaseCount() { count++; }
}
and then
Counter counter = new Counter();
lines.stream().forEach( e -> {
System.out.println(e);
counter.increaseCounter(); // or i++; if you decided i is worth being a field
} );
Another way to do it, this time hiding those multiple lines in a method:
class Counter { // wrapper class
private int count;
public int getCount() { return count; }
public void increaseCount( Object o ) {
System.out.println(o);
count++;
}
}
and then
Counter counter = new Counter();
lines.stream().forEach( counter::increaseCount );
or even
Counter counter = new Counter();
lines.stream().forEach( e -> counter.increaseCount(e) );
The second syntax comes in handy if you need a consumer having more than one parameter; the first syntax is still the shortest and simplest though.
The forEach method takes an instance of any class that implements Consumer. So here is an example of using a custom Consumer implementation that keeps up with the count. Later you can call getCount() on the Consumer implementation to get the count.
import java.util.ArrayList;
import java.util.List;
import java.util.function.Consumer;
public class ConsumerDemo {
public static void main(String[] args) {
List<String> lines = new ArrayList<String>();
lines.add("line 1");
lines.add("line 2");
MyConsumer countingConsumer = new MyConsumer();
lines.stream().forEach(countingConsumer);
System.out.println("Count: " + countingConsumer.getCount());
}
private static class MyConsumer implements Consumer<String> {
private int count;
@Override
public void accept(String t) {
System.out.println(t);
count++;
}
public int getCount() {
return count;
}
}
}
You can make your own Collector which temporarily stores the previous element/string. When the current element starts with a $, the name of the product is stored in prev. Now you can convert the price to a double and create the object.
private class ProductCollector {
private final List<Product> list = new ArrayList<>();
private String prev;
public void accept(String str) {
if (prev != null && str.startsWith("$")) {
double price = Double.parseDouble(str.substring(1));
list.add(new Product(prev, price));
}
prev = str;
}
public List<Product> finish() {
return list;
}
public static Collector<String, ?, List<Product>> collector() {
return Collector.of(ProductCollector::new, ProductCollector::accept, (a, b) -> a, ProductCollector::finish);
}
}
Since you need to rely on the sequence (line with price follows line with name), the stream cannot be processed in parallel. Here is how you can use your custom collector:
String[] lines = new String[]{
"Ice Cream", "$3.99",
"Chocolate", "$5.00",
"Nice Shoes", "$84.95"
};
List<Product> products = Stream.of(lines)
.sequential()
.collect(ProductCollector.collector());
Note that your prices are not integers which is why I used a double to represent them properly.
if you have the items in an array you can use vanilla java with an intStream and filter on even values and then in the next map you can use index and index+1. Maybe have look here
Generally collecting to anything other than standard API's gives you is pretty easy via a custom Collector. In your case collecting to 3 lists at a time (just a small example that compiles, since you can't share your code either):
private static <T> Collector<T, ?, List<List<T>>> to3Lists() {
class Acc {
List<T> left = new ArrayList<>();
List<T> middle = new ArrayList<>();
List<T> right = new ArrayList<>();
List<List<T>> list = Arrays.asList(left, middle, right);
void add(T elem) {
// obviously do whatever you want here
left.add(elem);
middle.add(elem);
right.add(elem);
}
Acc merge(Acc other) {
left.addAll(other.left);
middle.addAll(other.middle);
right.addAll(other.right);
return this;
}
public List<List<T>> finisher() {
return list;
}
}
return Collector.of(Acc::new, Acc::add, Acc::merge, Acc::finisher);
}
And using it via:
Stream.of(1, 2, 3)
.collect(to3Lists());
Obviously this custom collector does not do anything useful, but just an example of how you could work with it.
I have adapted the answer to this question to your case. The custom Spliterator will "split" the stream into multiple streams that collect by different properties:
@SafeVarargs
public static <T> long streamForked(Stream<T> source, Consumer<Stream<T>>... consumers)
{
return StreamSupport.stream(new ForkingSpliterator<>(source, consumers), false).count();
}
public static class ForkingSpliterator<T>
extends AbstractSpliterator<T>
{
private Spliterator<T> sourceSpliterator;
private List<BlockingQueue<T>> queues = new ArrayList<>();
private boolean sourceDone;
@SafeVarargs
private ForkingSpliterator(Stream<T> source, Consumer<Stream<T>>... consumers)
{
super(Long.MAX_VALUE, 0);
sourceSpliterator = source.spliterator();
for (Consumer<Stream<T>> fork : consumers)
{
LinkedBlockingQueue<T> queue = new LinkedBlockingQueue<>();
queues.add(queue);
new Thread(() -> fork.accept(StreamSupport.stream(new ForkedConsumer(queue), false))).start();
}
}
@Override
public boolean tryAdvance(Consumer<? super T> action)
{
sourceDone = !sourceSpliterator.tryAdvance(t -> queues.forEach(queue -> queue.offer(t)));
return !sourceDone;
}
private class ForkedConsumer
extends AbstractSpliterator<T>
{
private BlockingQueue<T> queue;
private ForkedConsumer(BlockingQueue<T> queue)
{
super(Long.MAX_VALUE, 0);
this.queue = queue;
}
@Override
public boolean tryAdvance(Consumer<? super T> action)
{
while (queue.peek() == null)
{
if (sourceDone)
{
// element is null, and there won't be no more, so "terminate" this sub stream
return false;
}
}
// push to consumer pipeline
action.accept(queue.poll());
return true;
}
}
}
You can use it as follows:
streamForked(Stream.of(new Row("content1", "client1", "location1", 1),
new Row("content2", "client1", "location1", 2),
new Row("content1", "client1", "location2", 3),
new Row("content2", "client2", "location2", 4),
new Row("content1", "client2", "location2", 5)),
rows -> System.out.println(rows.collect(Collectors.groupingBy(Row::getClient,
Collectors.groupingBy(Row::getContent,
Collectors.summingInt(Row::getConsumption))))),
rows -> System.out.println(rows.collect(Collectors.groupingBy(Row::getClient,
Collectors.groupingBy(Row::getLocation,
Collectors.summingInt(Row::getConsumption))))),
rows -> System.out.println(rows.collect(Collectors.groupingBy(Row::getContent,
Collectors.groupingBy(Row::getLocation,
Collectors.summingInt(Row::getConsumption))))));
// Output
// {client2={location2=9}, client1={location1=3, location2=3}}
// {client2={content2=4, content1=5}, client1={content2=2, content1=4}}
// {content2={location1=2, location2=4}, content1={location1=1, location2=8}}
Note that you can do pretty much anything you want with your the copies of the stream. As per your example, I used a stacked groupingBy collector to group the rows by two properties and then summed up the int property. So the result will be a Map<String, Map<String, Integer>>. But you could also use it for other scenarios:
rows -> System.out.println(rows.count())
rows -> rows.forEach(row -> System.out.println(row))
rows -> System.out.println(rows.anyMatch(row -> row.getConsumption() > 3))
Use Collectors.toMap() with identity() for the key and your calculation for the value:
public Map<String, String> extractPartitionsValues(Path marketFile) {
return partitions.stream()
.collect(toMap(identity(), p -> partitionValueFromFilePath(p, marketFile.toString())));
}
this should work:
Map<String, Object> collect =
partitions.stream().collect(Collectors.toMap(p -> p, p -> partitionValueFromFilePath(p, marketFile.toString())));
Unfortunately, Java 8 streams do not support such extraction of elements in-between two matches. In Java 9, you could use
Map<String,String> map;
try(Stream<String> stream = Files.lines(path)) {
map = stream
.dropWhile(s -> !s.equals("#DATA")).skip(1)
.takeWhile(s -> !s.equals("#DEND"))
.filter(Pattern.compile("^[^#].*:").asPredicate())
.map(item -> item.split(":", 2))
.collect(Collectors.toMap(parts->parts[0], parts->parts[1]));
}
// use the map
map.forEach((k,v)->System.out.println(k+" -> "+v));
dropWhile will drop all elements before the first matching element, skip(1) will skip the matching element, takeWhile effectively removes all elements after the first element matching the end criteria.
The next filter step using the pattern ^[^#].*: will skip all lines starting with # or not containing a :. The remaining steps are straight-forward. When specifying a limit of 2 to split, it will not search for subsequent :s after encountering the first :.
Under Java 8, extracting the part between the two matches can be implemented with a Scanner before the stream operation:
String part;
try(Scanner s = new Scanner(path)) {
part = s.findWithinHorizon("(?<=\\R#DATA\\R)(.|\\R)*(?=\\R#DEND\\R)", 0);
}
Map<String,String> map = Pattern.compile("\\R").splitAsStream(part)
.filter(Pattern.compile("^[^#].*:").asPredicate())
.map(item -> item.split(":", 2))
.collect(Collectors.toMap(parts->parts[0], parts->parts[1]));
// use the map
map.forEach((k,v)->System.out.println(k+" -> "+v));
If you see the lies between #DATA and #DEND contain ':' therefore I came up with the following solution -
File file = new File("Input.txt");
try {
Map<String,String> map = Files.lines(file.toPath())
.filter(list -> list.contains(":"))
.map(item -> item.split(":"))
.filter(arr -> arr.length > 1)
.collect(Collectors.toMap(parts->parts[0], parts->parts[1]));
System.out.println(map.values());
} catch (IOException e) {
e.printStackTrace();
}
The above code first filters only the lines containing colon ':', then split these lines based on the colon, after that, we filter only the list having length greater than 1 because if you see the input.txt file carefully, you can find that "#ENTRIES:" contains colon but do not contain any character after that as others do. Once we get the required data we create the HashMap.