Skip to main content
pandas

pandas

SoftwareWikipedia

Search complete. 29 mentions across 15 episodes found for "pandas".

Sep 23, 2026

Tim ScarfeHOST
3:11
And speaking from bitter experience, it's a nightmare to work with tabular data.
Tim ScarfeHOST
3:15
So anyone who's built a machine learning model with tabular data, you've-- I mean, obviously like Pandas and, you know, scikit-learn has got some stuff in there, but you have to deal with missing values.
Tim ScarfeHOST
3:23
You know, what do you do with the categorical features? How do you do some feature engineering? How do you do transformations that reflect the semantics of the problem? You, you, you need to really, really know what you're doing, and it feels kind of kludgy.
Tim ScarfeHOST
3:35
You always-- Even though you can make an ML model, you always feel kind of dirty afterwards because you feel like you've, um, done all of these shortcuts and kind of corrupted it in, in some way.
Avery SmithHOST
13:54
Because if you're expecting a hiring manager or a recruiter, to scroll through or whatever you know go through oh this is the business questions and the source and the queries and all these things before seeing the recommendations and the findings that's just not going to happen so get the buy-in from them by having the findings and recommendations up front in your portfolio project For this deadline, I postpone Python unless your target postings consistently require it.
Avery SmithHOST
14:19
If they do, substitute some second time, second project time with Kaggle's Python and Pandas lesson.
Avery SmithHOST
14:26
I actually think Kaggle's lessons aren't bad at all.
Avery SmithHOST
14:28
So good recommendation there.
Andrew ZiglerGUEST
9:43
Right.
Andrew ZiglerGUEST
9:43
You're not making your own pandas.
Andrew ZiglerGUEST
9:44
You're not making-
Dan GerlachHOST
9:45
Exactly
Calvin Hendryx-ParkerHOST
2:00
That's truly big data.
Calvin Hendryx-ParkerHOST
2:01
Most folks probably lie in the medium sized data, but Pandas definitely tops out.
Calvin Hendryx-ParkerHOST
2:06
I mean, he does some interesting benchmarks in here, gives a couple of good code examples, actually shows a really interesting post from Amazon Redshift team.
Calvin Hendryx-ParkerHOST
2:16
where they were looking at the composition of many of the tables that are out there in the Redshift environment.
Calvin Hendryx-ParkerHOST
3:00
A lot of blog posts have been produced, a lot of data science is based on pandas, but they're really based on data frames.
Calvin Hendryx-ParkerHOST
3:06
And there's more than one library out there to handle data frames and probably do it more efficient.
Calvin Hendryx-ParkerHOST
3:10
So if you actually looked at the chart here, they're basically saying, When you get up into like the 10 gigabyte range for data sizes, Pandas is probably still pretty good, but then there's a gap.
Calvin Hendryx-ParkerHOST
3:22
It falls off somewhere between 10 and 100 gigabytes of data.
Nigel GlendayGUEST
39:15
I started to be able to learn these patterns.
Nigel GlendayGUEST
39:16
Wes McKinney was the developer of a data manipulation package in Python called Pandas, right? Which is very well known as basically a way within Python to manipulate time series data.
Nigel GlendayGUEST
39:28
He built it because he was at a QR at a hedge fund and he saw people just crunching time series and spreadsheets and was like,
CJ GustafsonHOST
39:35
this is
speaker_1HOST
18:14
The agent autonomously unpacked the tabular structures.
speaker_1HOST
18:17
It wrote temporary Python scripts in the background using libraries like Pandas, executed them within the sandbox to analyze the headers and detect statistical anomalies, normalized the ledgers, and generated reconciled auto reports.
speaker_0HOST
18:29
All without human coding.
speaker_0HOST
18:31
And the crucial part here is the data sovereignty.
Christopher BaileyHOST
15:02
He has a, a deeper explanation of this and how this might work inside of SQL.
Christopher BaileyHOST
15:06
He has an article called Mastering DuckDB when you're used to Pandas or Polars.
Christopher BaileyHOST
15:10
So I'll include that link also.
Christopher BaileyHOST
15:13
So again, some differences here, and if you're bouncing around and moving between these two, which I think is really common.
Pedro HolandaGUEST
22:18
So you can do a bunch of tricks to avoid copying memory all over.
Pedro HolandaGUEST
22:23
So this is especially interesting for data science projects, because if you're using something like Pandas or NumPy, what is a NumPy array? It's literally a C array with some makeup on top.
Pedro HolandaGUEST
22:34
What is a DuckDB vector? It's literally an array with some makeup on top.
Pedro HolandaGUEST
22:38
So you can just change the makeup and then you can suddenly access the same data with constant cost, right? You don't have to transform your actual data.
Michael KennedyHOST
22:47
So when you do a query, you may be able to just return a piece of the in-memory chunk instead of going, OK, ours looks like this.
Michael KennedyHOST
22:57
But then we're going to copy a million floats over to this thing in this column and then send it back, right?
Pedro HolandaGUEST
23:01
so this is what was also like one of the things that uh was one of the realizations that database protocols right like so the way you transfer data from the server to the clients they're actually quite slow uh so this is one of the main frustrations we had seen with the data scientists is not only like oh it's clumsy to set it up and you'd like to start a server and create schemas and whatnot But it's also just to get your data from your NumPy or TensorFlow or Pandas or whatever you're running into the database system and back and forth was super slow.
Pedro HolandaGUEST
23:33
So you completely remove that boundary.
Quincy MorganGUEST
1:51
It's in GeoParquet format, and we only have two layers to start, buildings and highways, and we are doing a weekly update cadence.
Quincy MorganGUEST
2:01
So this allows you to quickly query a lot of OpenStreetMap data using tools like DuckDB, QGIS, you can use LawnBoard, Jupyter Notebook, shout out to Dev Seed, Pandas, even ArcGIS Analytics, to quickly do certain queries in the data.
Quincy MorganGUEST
2:25
So it just plugs OpenStreetMap data directly into all these tools.
Quincy MorganGUEST
2:29
There's a few other solutions out there, and we love them all.
speaker_2HOST
10:40
So moving beyond a single calculation, how do we wrangle all that real-world data without just drowning in it?
speaker_3HOST
10:46
To handle real scale, you turn to Pandas.
speaker_2HOST
10:48
The animal.
speaker_3HOST
10:49
The library, actually.
speaker_3HOST
10:50
Pandas introduces a data structure called a data frame.
speaker_3HOST
10:53
If NumPy is the forklift, Pandas is basically the central nervous system for your data.
speaker_3HOST
10:57
Okay, I like
speaker_2HOST
10:58
that.

5 more episodes mention pandas.

Create an account to see the whole feed, search across every transcript, and follow the entities you care about.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.