Repository navigation
Conversation
|
I like it so far. It would be nice to provide an example pulling the data into pandas as .csv and pulling into pandas as .parquet and highlighting any efficiencies? Maybe amount of bits sent over the wire is less? Resultant data is smaller? Something to show the benefit of having parquet as an option. |
So, the parquet file is 1 order of magnitude smaller, no surprises there. However, the time it took for the server-side creation of the file + the download was longer than the csv! This is not a valid benchmark at all and your mileage may vary with different datasets, slices/filters, etc. I just wanted to mention this b/c when thinking about cloud infrastructure, if only the download is of concern then parquet is a huge win, if server computing time should be taken into account, then things are not so simple. |
|
ah, I see. Let's just highlight the download difference between the two responses. |
Closes #350
@MathewBiddle this is just a first draft to test if you are OK with us going in this direction for #350. Let me know what you think.
PS: Do you know any croissant expert? It would be nice to do something with the metadata we download beyond spec checking. Maybe attach the data and prepare a pytorch/tensorflow workflow? (Not really running anything, just get as close to running as possible to demonstrate it.)