Apache Spark
SoftwareWikipedia
138
MENTIONS
20
EPISODES
18
PODCASTS
Search complete. 138 mentions across 20 episodes found for "Apache Spark".
Sep 18, 2026
Decoding the Open Lakehouse - Sahil Walia
A
1:19Alexey GrigorevHOST
So in order to create a project, in order to make sure a data science project runs, I needed to do a lot of transformation.
A
1:28Alexey GrigorevHOST
I needed to do a lot of work in Spark and Presto and similar things.
A
1:33Alexey GrigorevHOST
But I'm a bit out of touch with all these modern things.
A
1:37Alexey GrigorevHOST
So, Sahil, it's very nice to have you here so we can talk about all these things that modern data engineers need to use.
5 MINS LATER
S
6:56Sahil WaliaGUEST
There was no good practices around it.
S
6:59Sahil WaliaGUEST
So when we evolved, we now are in the current state, I would say, is interoperable open lake houses.
S
7:08Sahil WaliaGUEST
So if I talk about it, interoperable means that you want your data to be decoupled with the compute right so your storage decoupled from your compute and then what are you also trying to do is you're making it open when you're going through it and data engineering uh i'm pretty sure that's the case with ml i'm not ml guy data engineering is very much motivated by what's happening in open source a lot of things that get that happen in data engineering are inspired by open source let's say for example storage storage is Sparky, which is open source.
S
7:46Sahil WaliaGUEST
Let's say compute.
Databricks’ Ali Ghodsi Never Wanted to Be CEO. Now He’s Among the Best
A
1:59Ali GhodsiGUEST
I knew the identity of at least two people.
A
2:01Ali GhodsiGUEST
So, you know, it was I mean, 2015 was kind of a turbulent year for Databricks because we were having great success with our open source project, Apache Spark.
A
2:10Ali GhodsiGUEST
We were just not having commercial success.
A
2:12Ali GhodsiGUEST
Revenue was, you know, Gap revenue came in at like one and a half million that year.
8 MINS LATER
A
10:12Ali GhodsiGUEST
problem is because, you know, you most likely won't be able to actually unclog or remove this bottleneck.
A
10:17Ali GhodsiGUEST
So you should just turn all your attention to that.
A
10:19Ali GhodsiGUEST
So for me at that time, it was, you know, open source project, check, super successful, you know, tech, Super awesome.
A
10:28Ali GhodsiGUEST
No commercial success.
564: Data Literacy in the AI Era
T
2:45Tommy PugliaHOST
The support of connectors.
T
2:47Tommy PugliaHOST
There are Snowflake Databricks, BigQuery, Redshift, Dreamio, Spark, and generic ODBC LLDB connections.
T
2:57Tommy PugliaHOST
So Mike, when I first started with Power BI, most of our things were done via ODBC.
T
3:02Tommy PugliaHOST
And it was a pain in the gun because we had a virtual machine that allowed us to do so.
M
4:03Mike CarloHOST
It talks about the different drivers.
M
4:05Mike CarloHOST
It talks about connectors and the changing of drivers.
M
4:08Mike CarloHOST
So the idea is that, If you're connecting to Databricks, if you're doing Azure Databricks, if you're connecting to Google BigQuery, Hive, Impala, Snowflake, Spark, these are the recommended switches that you're able to apply.
M
4:21Mike CarloHOST
So it's not like I'm going back to SQL databases and I'm having to switch the driver on that.
#496 A lake house in Seattle
C
3:28Calvin Hendryx-ParkerHOST
You can easily go find sample data sets that are in that realm and that range and that size.
C
3:33Calvin Hendryx-ParkerHOST
The next thing they reach for is typically a commercial tool like Databricks, Snowflake, Dask, or some of these other things that are Spark.
C
3:41Calvin Hendryx-ParkerHOST
You're distributing the memory of that dataset across many machines or maybe even across one very large machine, but doing it in an attributed manner.
C
3:50Calvin Hendryx-ParkerHOST
most people probably don't need to go that far.
How Databricks engineers lose the final round (and the 4 places it happens)
J
1:29JR LasakHOST
where they check whether you really know the platform.
J
1:32JR LasakHOST
That belief is what sends you back to Spark Internals for a third weekend.
J
1:37JR LasakHOST
Now, if your last stage is still a coding or SQL round, go and study.
J
1:42JR LasakHOST
That one is an exam.
September 11th 2026 - Google Cloud Update
S
0:51speaker_0HOST
After that, we will head over to Kubernetes, discussing the release of GKE 1.37 and a container-eyed privilege escalation vulnerability.
S
1:02speaker_0HOST
Then we will cover image updates in Managed Service for Apache Spark, Cloud SQL's regional endpoints going GA, and big updates in networking, identity, and cloud storage.
S
1:14speaker_0HOST
We have a lot of ground to cover, so let's get right into it.
S
1:21speaker_0HOST
First up, let's talk about BigQuery.
8 MINS LATER
S
9:00speaker_0HOST
This includes CC Insights QA Scorecard, Content Warehouse Synonym Set, DevConnect Account Connector, Discovery Engine Engine, Discovery Engine Serving Config, GKE Hub Fleet, Model Armor Template, Network Security Offsize Policy, Rapid Migration Assessment Collector, Security Center Management Event Threat Detection Custom Module, Storage Insights Dataset Config, and Vector Search Collection.
S
9:34speaker_0HOST
Config Connector also added new fields like automated backup policy locations for Big Table Table, swap configs for container cluster and container node pools, instance flexibility policies for Dataproc cluster, and organization ref support for network security firewall endpoints 1.1.
S
9:57speaker_0HOST
Direct reconciliation support has also been added as an opt-in for Gemini Enterprise Agent Platform TensorBoard 1.1. Moving on to data engineering, let's talk about managed service for Apache Spark, which is the new name for Dataproc on Compute Engine and Serverless deployment.
S
10:18speaker_0HOST
This week brings several new sub-minor cluster image versions across Debian, Rocky Linux, and Ubuntu.
A Rust Framework to Simplify Distributed Systems
J
20:26Joe HellersteinGUEST
If you wanted to run, like, a big data analytics job, you can also use Hydro.
J
20:30Joe HellersteinGUEST
So it can do things like Spark or Flink would do.
J
20:32Joe HellersteinGUEST
It doesn't care about the performance shape of your events.
J
20:37Joe HellersteinGUEST
They can be many or few, they can be quick or in volume, but it allows you to program a distributed system in a way that, like, is quite general purpose.
J
25:42Joe HellersteinGUEST
Right.
J
25:42Joe HellersteinGUEST
This is where I think a lot of libraries trip actually, and why we've postponed a general purpose solution for all these years is that if you start to answer those questions like you're asking, what about durability? What about replay of dropped messages or things? All these kind of like specific problems typically lead people to specific solutions.
J
26:02Joe HellersteinGUEST
So if you look at Spark or Flink or Kafka or many of these frameworks, they only work in their narrow sweet spot.
J
26:08Joe HellersteinGUEST
They're distributed, but they only work at the stuff they're good at, and they're really bad at other stuff.
AWS Bought DuckLabs and users want their own agents
D
4:14Dumky de WildeHOST
Yeah
M
4:14Mehdi OuazzaHOST
... or Spark.
D
4:15Dumky de WildeHOST
Athena.
M
4:16Mehdi OuazzaHOST
Yeah.
5 MINS LATER
M
9:18Mehdi OuazzaHOST
There is so many, you know, vendors, uh, selling-
D
9:21Dumky de WildeHOST
Exactly
M
9:21Mehdi OuazzaHOST
... and for Spark.
M
9:24Mehdi OuazzaHOST
There is good and bad example.
P
Unknown podcast
Open Source Business Models: From Free Code to Paid Services
Sep 4 · 2 Mentions
S
3:48speaker_1UNKNOWN
That's where managed services come in.
S
3:51speaker_1UNKNOWN
Providers like Databricks built an entirely managed analytics platform on top of Apache Spark which is open source but complex to configure at scale.
S
4:01speaker_1UNKNOWN
Their pricing model usage-based subscription for compute and storage removes the burden from customers while generating predictable revenue streams.
S
4:11speaker_1UNKNOWN
One downside with SaaS on open source is that if a cloud provider starts offering a competing managed service, they could effectively commoditize what your company built.
6 MINS LATER
S
10:29speaker_1UNKNOWN
This misalignment led to Dockers founder Solomon Hikes stepping aside in twenty eighteen.
S
10:35speaker_1UNKNOWN
The company refocused its strategy around managed services and enterprise features eventually spinning out Mirantis for kubernetes specific tooling From that story we learned that Founders need a clear separation of Revenue generating functions from community driven Ones to keep investors happy without compromising open source ethos.
S
10:57speaker_1UNKNOWN
The decision of whether to go Fully Proprietary Or Keep the Core Open Also hinges on market positioning Some niches Like Security Require strong Compliance Controls that May Be hard to guarantee in a purely os model Take Palo Alto networks They Provide open source components but Their Flagship firewall platform Is entirely closed Their community Contributions Go Into open source libraries for developers but the core Product Remains a Lockin for enterprise Customers in this scenario the Business leverages a Dual approach it Offers An open Codebase to reduce cost Of Entry Then uses proprietary infrastructure to Generate recurring Revenue From Managed Services and Support Contracts conversely there Are Companies that stay Entirely oss Because Their Value Proposition IS Community Collaboration Itself Apache Spark Has Never Offered Paid Support Directly but Instead Relies on Vendor Partnerships Like Databricks or Cloudera those Vendors Monetize By Providing Managed Infrastructure and Specialized Consulting turning the community into a sales funnel rather Than a direct Revenue source The choice of model also affects how you build your distribution.
S
12:16speaker_1UNKNOWN
If your business relies on SaaS, you get immediate visibility of churn.
SE Radio 736: Sahil Walia on Apache Iceberg
S
9:12Sahil WaliaGUEST
You don't have to maintain multiple copies, right? And that's the way it's kind of headless.
S
9:17Sahil WaliaGUEST
And when you can bring your own compute, say Spark, or let's say you want to do Databricks, or you want to do, let's say Snowflake, you can connect any of them to that data and start running your, say, analytics workload on top of it.
R
9:32Robert BlumenHOST
Okay, I think we're ready now to talk about what is Apache Iceberg.
R
9:37Robert BlumenHOST
Where does it fit into this landscape we've been sketching out?
34 MINS LATER
S
43:48Sahil WaliaGUEST
Yeah, there are multiple engine options you have.
S
43:50Sahil WaliaGUEST
You have proprietary ones, the commercial ones, and you have the open ones.
S
43:53Sahil WaliaGUEST
You can use Spark, the Flink, the Trino, which is open source, but you can use Snowflake, you can use Databricks, you can use BigQuery, anything which is compatible to the Iceberg spec can use it.
R
44:07Robert BlumenHOST
You also mentioned there are multiple options for catalog.
10 more episodes mention Apache Spark.
Create an account to see the whole feed, search across every transcript, and follow the entities you care about.