Apache Parquet
SoftwareWikipedia
105
MENTIONS
20
EPISODES
16
PODCASTS
Search complete. 105 mentions across 20 episodes found for "Apache Parquet".
Sep 18, 2026
Decoding the Open Lakehouse - Sahil Walia
A
11:50Alexey GrigorevHOST
JSON events, I just dump them to S3 and then I run Presto on top of that.
A
11:55Alexey GrigorevHOST
Do I already have a lake house or there is more to lake house? Or let's say I take this JSON and convert it to Parquet.
A
12:03Alexey GrigorevHOST
Is it already a lake house or not yet? Like where does a lake house, what is actually a lake house and where does it start?
S
12:10Sahil WaliaGUEST
Okay, sounds good.
S
12:15Sahil WaliaGUEST
Yeah.
S
12:15Sahil WaliaGUEST
So when you ask that question, let's think through it.
S
12:18Sahil WaliaGUEST
What you are describing is a file storage, right? Parquet.
S
12:23Sahil WaliaGUEST
Although Parquet is very rich in metadata, it has...
From Financial Data to AI-Ready Decisions: Power BI, Microsoft Fabric, Semantic Models & AI Agents with Rishi Sapra [MVP]
R
45:32Rishi SapraGUEST
This is another compression technology that's columnular as well." Um, could we actually just build the entire VertiPaq engine into Park...
R
45:41Rishi SapraGUEST
against Parquet files? And then we don't need to actually do all the transformations.
R
45:45Rishi SapraGUEST
You don't need to go and convert your data, like with the tabular VertiPaq engine analysis services.
R
45:51Rishi SapraGUEST
Analysis services is its own layer, right? So you need to convert data from whatever spreadsheets, CRM systems, whatever that data's coming into, you need to convert, bring it, refresh your semantic model, and bring your data into there.
R
46:06Rishi SapraGUEST
You could do Direct Query, but you know, that's, that's horrible, right? [laughs] But if you're not...
R
46:11Rishi SapraGUEST
If you're doing import mode, you need to g- refresh your semantic model, which will take all the data, bring it in compression, and bring it into the semantic model.
R
46:18Rishi SapraGUEST
But actually, if your data coming into this layer is in delta format already, in, in Parquet, which is just Parquet files with a delta log, then you don't need to do that anymore.
R
46:28Rishi SapraGUEST
You don't need to go and transform your data because Power BI could read directly against those Parquet files, and actually it still has the same similar level of com- columnular compression, and therefore performance.
Ep. 30 - Telling Stories with Data
E
27:17Eugene MeidingerHOST
And it wasn't until I had to do some training that included some Fabric stuff, and until last year I actually had, you know, my first Fabric thing that a lot of this started to click.
E
27:26Eugene MeidingerHOST
But for anyone out there who's in a similar boat, I would say, like, learn about Parquet files and the whole delta lake kind of thing.
E
27:34Eugene MeidingerHOST
Understand what a lakehouse is, and have a peripheral awareness of some of these pieces.
E
27:40Eugene MeidingerHOST
But I think that'll help you the point that you decide to switch, maybe learn a little bit of Python.
From the One Billion Row Challenge to a Multi-Threaded Parquet Library
G
1:30Gunnar MorlingGUEST
Yes, exactly right.
G
1:31Gunnar MorlingGUEST
So, um, I mean, first of all, what is it? It's a library for reading and also writing Apache Parquet files, and I guess we should talk about what Parquet even is.
G
1:41Gunnar MorlingGUEST
But yes, um-
A
1:42Adam BienHOST
A- and, and why you got the idea to do such a thing?
September 11th 2026 - Google Cloud Update
S
10:18speaker_0HOST
This week brings several new sub-minor cluster image versions across Debian, Rocky Linux, and Ubuntu.
S
10:25speaker_0HOST
Key updates in these images include Apache Hudi 1.2.0 support in 3.0 images, Parquet footer caching enabled by default for the Lightning Engine using Velux, and Apache Iceberg 1.10 support in 2.2 images.
S
10:43speaker_0HOST
The Lakehouse catalog also supports autoloading for image versions 2.2 and later.
S
10:49speaker_0HOST
There is a breaking change to watch out for.
#562: DuckLake: The Lakehouse That's Just SQL and Parquet
M
0:00Michael KennedyHOST
How many files does your query read before it reads any data? On some data lakes, you go through JSON and metadata files first just to learn which Parquet files actually matter.
M
0:10Michael KennedyHOST
Duck Lake asks one SQL question instead.
M
0:13Michael KennedyHOST
The metadata lives in a real database.
M
0:15Michael KennedyHOST
The data stays in plain Parquet.
M
0:18Michael KennedyHOST
That's the entire format.
M
0:19Michael KennedyHOST
Pedro Holanda joined DuckDB in 2018 when it was still a research prototype at CWI.
39 MINS LATER
M
39:07Michael KennedyHOST
They've worked with databases, they've worked with APIs, and they've worked with storage like S3.
M
39:12Michael KennedyHOST
But the general idea, I guess maybe the Iceberg original idea was, well, what if we could just use all of S3 for scaling? Think how scalable that is.
AWS DuckDB Moves, Fixing "Lazy" AI Agents, and the 2026 AI Budget Reckoning | Agents of Dev Episode 37
B
2:14Brad ShimminHOST
If folks on the call are not familiar with them, they really should check out this company and spate of products.
B
2:22Brad ShimminHOST
Mitch and I have spoken about them before in regards to the work that I do for our Futurum Signal and how important it's been for me to use like DuckDB on top of Parquet files to basically stand up an RDF style basic graph that I can use to say, this is the truth.
B
2:43Brad ShimminHOST
This is what company owns which products.
B
2:46Brad ShimminHOST
Kind of important things.
B
3:06Brad ShimminHOST
Yeah, exactly.
B
3:07Brad ShimminHOST
So what they're going to do is basically embed DuckDB's vectorized query engine inside of their storage tier.
B
3:14Brad ShimminHOST
So basically you can get like on top of S3, like sub second, you know, really cheap query, you know, a query path across, you know, our beloved Parquet files and Apache Iceberg.
B
3:27Brad ShimminHOST
So it's a great acquisition.
AWS Bought DuckLabs and users want their own agents
D
13:38Dumky de WildeHOST
Um, and Lens is interesting 'cause basically there is a core extension in DuckDB, uh, for the Lens file format.
D
13:46Dumky de WildeHOST
There is obviously extensions for Parquet, for Iceberg as well.
D
13:49Dumky de WildeHOST
So DuckDB, and this is the mechanism that they're describing in this article, DuckDB is the great way actually to combine all of these into one place.
D
14:02Dumky de WildeHOST
And so you can query both your lake, b- your vector index, and the actual image files, for example, um, to make this work.
2025 FOSS4G NA | DuckDB + Rasters - Sina Kashuk
S
13:41Sina KashukGUEST
As you all may know, DocDB doesn't turn your browser to a database.
S
13:46Sina KashukGUEST
The challenge with it is how can I get my data to that, right? And that's where under the hood fuse is doing, is taking the data and converting it to Parquet in here and then passing it to the table.
S
13:57Sina KashukGUEST
So that's one way of doing it.
S
13:59Sina KashukGUEST
Of course, you may come back and say, well, I have terabyte of data.
5 MINS LATER
S
19:17Sina KashukGUEST
This is a perfect example of The bottleneck of this demo that you're seeing right now is the maximum speed that I can get from internet.
S
19:23Sina KashukGUEST
It's not the concurrency, and I can set the concurrency as many as I want.
S
19:28Sina KashukGUEST
This also allows you to have a smaller chunk when you're building your Parquet file, because that's another piece that you have to keep in mind if you take your parquet file and making it to be a smaller chunk, then you can actually have the data to come faster for the specific partition that you have.
S
19:45Sina KashukGUEST
You can align your partition over time, or you can align your partition for space, or combination of work.
FedGeoDay 2025 | Lightning Talk - Adam Timm
A
0:31Adam TimmGUEST
I'm the field CTO for the US public sector at Crunchy Data.
A
0:34Adam TimmGUEST
I'm going to talk very briefly about how you can bring Iceberg and Parquet into Postgres with a new product from Crunchy Data.
A
0:41Adam TimmGUEST
Alright, so I do want to spend a little bit of time just kind of posing some thought questions for everyone to kind of set the stage for what we feel this.
A
0:48Adam TimmGUEST
This the problem that this really solves for both small development teams as well as large enterprises.
10 more episodes mention Apache Parquet.
Create an account to see the whole feed, search across every transcript, and follow the entities you care about.