Skip to main content
Apache Spark

Apache Spark

SoftwareWikipedia

Search complete. 138 mentions across 20 episodes found for "Apache Spark".

Sep 18, 2026

Alexey GrigorevHOST
1:19
So in order to create a project, in order to make sure a data science project runs, I needed to do a lot of transformation.
Alexey GrigorevHOST
1:28
I needed to do a lot of work in Spark and Presto and similar things.
Alexey GrigorevHOST
1:33
But I'm a bit out of touch with all these modern things.
Alexey GrigorevHOST
1:37
So, Sahil, it's very nice to have you here so we can talk about all these things that modern data engineers need to use.

5 MINS LATER

Sahil WaliaGUEST
6:56
There was no good practices around it.
Sahil WaliaGUEST
6:59
So when we evolved, we now are in the current state, I would say, is interoperable open lake houses.
Sahil WaliaGUEST
7:08
So if I talk about it, interoperable means that you want your data to be decoupled with the compute right so your storage decoupled from your compute and then what are you also trying to do is you're making it open when you're going through it and data engineering uh i'm pretty sure that's the case with ml i'm not ml guy data engineering is very much motivated by what's happening in open source a lot of things that get that happen in data engineering are inspired by open source let's say for example storage storage is Sparky, which is open source.
Sahil WaliaGUEST
7:46
Let's say compute.
Ali GhodsiGUEST
1:59
I knew the identity of at least two people.
Ali GhodsiGUEST
2:01
So, you know, it was I mean, 2015 was kind of a turbulent year for Databricks because we were having great success with our open source project, Apache Spark.
Ali GhodsiGUEST
2:10
We were just not having commercial success.
Ali GhodsiGUEST
2:12
Revenue was, you know, Gap revenue came in at like one and a half million that year.

8 MINS LATER

Ali GhodsiGUEST
10:12
problem is because, you know, you most likely won't be able to actually unclog or remove this bottleneck.
Ali GhodsiGUEST
10:17
So you should just turn all your attention to that.
Ali GhodsiGUEST
10:19
So for me at that time, it was, you know, open source project, check, super successful, you know, tech, Super awesome.
Ali GhodsiGUEST
10:28
No commercial success.
Tommy PugliaHOST
2:45
The support of connectors.
Tommy PugliaHOST
2:47
There are Snowflake Databricks, BigQuery, Redshift, Dreamio, Spark, and generic ODBC LLDB connections.
Tommy PugliaHOST
2:57
So Mike, when I first started with Power BI, most of our things were done via ODBC.
Tommy PugliaHOST
3:02
And it was a pain in the gun because we had a virtual machine that allowed us to do so.
Mike CarloHOST
4:03
It talks about the different drivers.
Mike CarloHOST
4:05
It talks about connectors and the changing of drivers.
Mike CarloHOST
4:08
So the idea is that, If you're connecting to Databricks, if you're doing Azure Databricks, if you're connecting to Google BigQuery, Hive, Impala, Snowflake, Spark, these are the recommended switches that you're able to apply.
Mike CarloHOST
4:21
So it's not like I'm going back to SQL databases and I'm having to switch the driver on that.
Calvin Hendryx-ParkerHOST
3:28
You can easily go find sample data sets that are in that realm and that range and that size.
Calvin Hendryx-ParkerHOST
3:33
The next thing they reach for is typically a commercial tool like Databricks, Snowflake, Dask, or some of these other things that are Spark.
Calvin Hendryx-ParkerHOST
3:41
You're distributing the memory of that dataset across many machines or maybe even across one very large machine, but doing it in an attributed manner.
Calvin Hendryx-ParkerHOST
3:50
most people probably don't need to go that far.
JR LasakHOST
1:29
where they check whether you really know the platform.
JR LasakHOST
1:32
That belief is what sends you back to Spark Internals for a third weekend.
JR LasakHOST
1:37
Now, if your last stage is still a coding or SQL round, go and study.
JR LasakHOST
1:42
That one is an exam.
speaker_0HOST
0:51
After that, we will head over to Kubernetes, discussing the release of GKE 1.37 and a container-eyed privilege escalation vulnerability.
speaker_0HOST
1:02
Then we will cover image updates in Managed Service for Apache Spark, Cloud SQL's regional endpoints going GA, and big updates in networking, identity, and cloud storage.
speaker_0HOST
1:14
We have a lot of ground to cover, so let's get right into it.
speaker_0HOST
1:21
First up, let's talk about BigQuery.

8 MINS LATER

speaker_0HOST
9:00
This includes CC Insights QA Scorecard, Content Warehouse Synonym Set, DevConnect Account Connector, Discovery Engine Engine, Discovery Engine Serving Config, GKE Hub Fleet, Model Armor Template, Network Security Offsize Policy, Rapid Migration Assessment Collector, Security Center Management Event Threat Detection Custom Module, Storage Insights Dataset Config, and Vector Search Collection.
speaker_0HOST
9:34
Config Connector also added new fields like automated backup policy locations for Big Table Table, swap configs for container cluster and container node pools, instance flexibility policies for Dataproc cluster, and organization ref support for network security firewall endpoints 1.1.
speaker_0HOST
9:57
Direct reconciliation support has also been added as an opt-in for Gemini Enterprise Agent Platform TensorBoard 1.1. Moving on to data engineering, let's talk about managed service for Apache Spark, which is the new name for Dataproc on Compute Engine and Serverless deployment.
speaker_0HOST
10:18
This week brings several new sub-minor cluster image versions across Debian, Rocky Linux, and Ubuntu.
Joe HellersteinGUEST
20:26
If you wanted to run, like, a big data analytics job, you can also use Hydro.
Joe HellersteinGUEST
20:30
So it can do things like Spark or Flink would do.
Joe HellersteinGUEST
20:32
It doesn't care about the performance shape of your events.
Joe HellersteinGUEST
20:37
They can be many or few, they can be quick or in volume, but it allows you to program a distributed system in a way that, like, is quite general purpose.
Joe HellersteinGUEST
25:42
Right.
Joe HellersteinGUEST
25:42
This is where I think a lot of libraries trip actually, and why we've postponed a general purpose solution for all these years is that if you start to answer those questions like you're asking, what about durability? What about replay of dropped messages or things? All these kind of like specific problems typically lead people to specific solutions.
Joe HellersteinGUEST
26:02
So if you look at Spark or Flink or Kafka or many of these frameworks, they only work in their narrow sweet spot.
Joe HellersteinGUEST
26:08
They're distributed, but they only work at the stuff they're good at, and they're really bad at other stuff.
Dumky de WildeHOST
4:14
Yeah
Mehdi OuazzaHOST
4:14
... or Spark.
Dumky de WildeHOST
4:15
Athena.
Mehdi OuazzaHOST
4:16
Yeah.

5 MINS LATER

Mehdi OuazzaHOST
9:18
There is so many, you know, vendors, uh, selling-
Dumky de WildeHOST
9:21
Exactly
Mehdi OuazzaHOST
9:21
... and for Spark.
Mehdi OuazzaHOST
9:24
There is good and bad example.

Unknown podcast

Open Source Business Models: From Free Code to Paid Services

Sep 4 · 2 Mentions

speaker_1UNKNOWN
3:48
That's where managed services come in.
speaker_1UNKNOWN
3:51
Providers like Databricks built an entirely managed analytics platform on top of Apache Spark which is open source but complex to configure at scale.
speaker_1UNKNOWN
4:01
Their pricing model usage-based subscription for compute and storage removes the burden from customers while generating predictable revenue streams.
speaker_1UNKNOWN
4:11
One downside with SaaS on open source is that if a cloud provider starts offering a competing managed service, they could effectively commoditize what your company built.

6 MINS LATER

speaker_1UNKNOWN
10:29
This misalignment led to Dockers founder Solomon Hikes stepping aside in twenty eighteen.
speaker_1UNKNOWN
10:35
The company refocused its strategy around managed services and enterprise features eventually spinning out Mirantis for kubernetes specific tooling From that story we learned that Founders need a clear separation of Revenue generating functions from community driven Ones to keep investors happy without compromising open source ethos.
speaker_1UNKNOWN
10:57
The decision of whether to go Fully Proprietary Or Keep the Core Open Also hinges on market positioning Some niches Like Security Require strong Compliance Controls that May Be hard to guarantee in a purely os model Take Palo Alto networks They Provide open source components but Their Flagship firewall platform Is entirely closed Their community Contributions Go Into open source libraries for developers but the core Product Remains a Lockin for enterprise Customers in this scenario the Business leverages a Dual approach it Offers An open Codebase to reduce cost Of Entry Then uses proprietary infrastructure to Generate recurring Revenue From Managed Services and Support Contracts conversely there Are Companies that stay Entirely oss Because Their Value Proposition IS Community Collaboration Itself Apache Spark Has Never Offered Paid Support Directly but Instead Relies on Vendor Partnerships Like Databricks or Cloudera those Vendors Monetize By Providing Managed Infrastructure and Specialized Consulting turning the community into a sales funnel rather Than a direct Revenue source The choice of model also affects how you build your distribution.
speaker_1UNKNOWN
12:16
If your business relies on SaaS, you get immediate visibility of churn.
Sahil WaliaGUEST
9:12
You don't have to maintain multiple copies, right? And that's the way it's kind of headless.
Sahil WaliaGUEST
9:17
And when you can bring your own compute, say Spark, or let's say you want to do Databricks, or you want to do, let's say Snowflake, you can connect any of them to that data and start running your, say, analytics workload on top of it.
Robert BlumenHOST
9:32
Okay, I think we're ready now to talk about what is Apache Iceberg.
Robert BlumenHOST
9:37
Where does it fit into this landscape we've been sketching out?

34 MINS LATER

Sahil WaliaGUEST
43:48
Yeah, there are multiple engine options you have.
Sahil WaliaGUEST
43:50
You have proprietary ones, the commercial ones, and you have the open ones.
Sahil WaliaGUEST
43:53
You can use Spark, the Flink, the Trino, which is open source, but you can use Snowflake, you can use Databricks, you can use BigQuery, anything which is compatible to the Iceberg spec can use it.
Robert BlumenHOST
44:07
You also mentioned there are multiple options for catalog.

10 more episodes mention Apache Spark.

Create an account to see the whole feed, search across every transcript, and follow the entities you care about.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.