pandas
SoftwareWikipedia
29
MENTIONS
15
EPISODES
14
PODCASTS
Search complete. 29 mentions across 15 episodes found for "pandas".
Sep 23, 2026
How Deep Learning Finally Cracked Messy Tables - Frank Hutter
T
3:11Tim ScarfeHOST
And speaking from bitter experience, it's a nightmare to work with tabular data.
T
3:15Tim ScarfeHOST
So anyone who's built a machine learning model with tabular data, you've-- I mean, obviously like Pandas and, you know, scikit-learn has got some stuff in there, but you have to deal with missing values.
T
3:23Tim ScarfeHOST
You know, what do you do with the categorical features? How do you do some feature engineering? How do you do transformations that reflect the semantics of the problem? You, you, you need to really, really know what you're doing, and it feels kind of kludgy.
T
3:35Tim ScarfeHOST
You always-- Even though you can make an ML model, you always feel kind of dirty afterwards because you feel like you've, um, done all of these shortcuts and kind of corrupted it in, in some way.
229: i asked gpt-6 astra how to become a data analyst (it was wrong)
A
13:54Avery SmithHOST
Because if you're expecting a hiring manager or a recruiter, to scroll through or whatever you know go through oh this is the business questions and the source and the queries and all these things before seeing the recommendations and the findings that's just not going to happen so get the buy-in from them by having the findings and recommendations up front in your portfolio project For this deadline, I postpone Python unless your target postings consistently require it.
A
14:19Avery SmithHOST
If they do, substitute some second time, second project time with Kaggle's Python and Pandas lesson.
A
14:26Avery SmithHOST
I actually think Kaggle's lessons aren't bad at all.
A
14:28Avery SmithHOST
So good recommendation there.
Beads, Better Specs, and Less Rework
A
9:43Andrew ZiglerGUEST
Right.
A
9:43Andrew ZiglerGUEST
You're not making your own pandas.
A
9:44Andrew ZiglerGUEST
You're not making-
D
9:45Dan GerlachHOST
Exactly
#496 A lake house in Seattle
C
2:00Calvin Hendryx-ParkerHOST
That's truly big data.
C
2:01Calvin Hendryx-ParkerHOST
Most folks probably lie in the medium sized data, but Pandas definitely tops out.
C
2:06Calvin Hendryx-ParkerHOST
I mean, he does some interesting benchmarks in here, gives a couple of good code examples, actually shows a really interesting post from Amazon Redshift team.
C
2:16Calvin Hendryx-ParkerHOST
where they were looking at the composition of many of the tables that are out there in the Redshift environment.
C
3:00Calvin Hendryx-ParkerHOST
A lot of blog posts have been produced, a lot of data science is based on pandas, but they're really based on data frames.
C
3:06Calvin Hendryx-ParkerHOST
And there's more than one library out there to handle data frames and probably do it more efficient.
C
3:10Calvin Hendryx-ParkerHOST
So if you actually looked at the chart here, they're basically saying, When you get up into like the 10 gigabyte range for data sizes, Pandas is probably still pretty good, but then there's a gap.
C
3:22Calvin Hendryx-ParkerHOST
It falls off somewhere between 10 and 100 gigabytes of data.
Masterwork’s CFO Runs 12 AI Agents at Once | Nigel Glenday
N
39:15Nigel GlendayGUEST
I started to be able to learn these patterns.
N
39:16Nigel GlendayGUEST
Wes McKinney was the developer of a data manipulation package in Python called Pandas, right? Which is very well known as basically a way within Python to manipulate time series data.
N
39:28Nigel GlendayGUEST
He built it because he was at a QR at a hedge fund and he saw people just crunching time series and spreadsheets and was like,
C
39:35CJ GustafsonHOST
this is
The Sovereign Architect: OpenCode and the Autonomous Agent Frontier
S
18:14speaker_1HOST
The agent autonomously unpacked the tabular structures.
S
18:17speaker_1HOST
It wrote temporary Python scripts in the background using libraries like Pandas, executed them within the sandbox to analyze the headers and detect statistical anomalies, normalized the ledgers, and generated reconciled auto reports.
S
18:29speaker_0HOST
All without human coding.
S
18:31speaker_0HOST
And the crucial part here is the data sovereignty.
Django Developers Survey Results & Reproducible Python Builds
C
15:02Christopher BaileyHOST
He has a, a deeper explanation of this and how this might work inside of SQL.
C
15:06Christopher BaileyHOST
He has an article called Mastering DuckDB when you're used to Pandas or Polars.
C
15:10Christopher BaileyHOST
So I'll include that link also.
C
15:13Christopher BaileyHOST
So again, some differences here, and if you're bouncing around and moving between these two, which I think is really common.
#562: DuckLake: The Lakehouse That's Just SQL and Parquet
P
22:18Pedro HolandaGUEST
So you can do a bunch of tricks to avoid copying memory all over.
P
22:23Pedro HolandaGUEST
So this is especially interesting for data science projects, because if you're using something like Pandas or NumPy, what is a NumPy array? It's literally a C array with some makeup on top.
P
22:34Pedro HolandaGUEST
What is a DuckDB vector? It's literally an array with some makeup on top.
P
22:38Pedro HolandaGUEST
So you can just change the makeup and then you can suddenly access the same data with constant cost, right? You don't have to transform your actual data.
M
22:47Michael KennedyHOST
So when you do a query, you may be able to just return a piece of the in-memory chunk instead of going, OK, ours looks like this.
M
22:57Michael KennedyHOST
But then we're going to copy a million floats over to this thing in this column and then send it back, right?
P
23:01Pedro HolandaGUEST
so this is what was also like one of the things that uh was one of the realizations that database protocols right like so the way you transfer data from the server to the clients they're actually quite slow uh so this is one of the main frustrations we had seen with the data scientists is not only like oh it's clumsy to set it up and you'd like to start a server and create schemas and whatnot But it's also just to get your data from your NumPy or TensorFlow or Pandas or whatever you're running into the database system and back and forth was super slow.
P
23:33Pedro HolandaGUEST
So you completely remove that boundary.
FedGeoDay 2025 | Lightning Talk - Quincy Morgan
Q
1:51Quincy MorganGUEST
It's in GeoParquet format, and we only have two layers to start, buildings and highways, and we are doing a weekly update cadence.
Q
2:01Quincy MorganGUEST
So this allows you to quickly query a lot of OpenStreetMap data using tools like DuckDB, QGIS, you can use LawnBoard, Jupyter Notebook, shout out to Dev Seed, Pandas, even ArcGIS Analytics, to quickly do certain queries in the data.
Q
2:25Quincy MorganGUEST
So it just plugs OpenStreetMap data directly into all these tools.
Q
2:29Quincy MorganGUEST
There's a few other solutions out there, and we love them all.
Python for Civil and Structural Engineers
S
10:40speaker_2HOST
So moving beyond a single calculation, how do we wrangle all that real-world data without just drowning in it?
S
10:46speaker_3HOST
To handle real scale, you turn to Pandas.
S
10:48speaker_2HOST
The animal.
S
10:49speaker_3HOST
The library, actually.
S
10:50speaker_3HOST
Pandas introduces a data structure called a data frame.
S
10:53speaker_3HOST
If NumPy is the forklift, Pandas is basically the central nervous system for your data.
S
10:57speaker_3HOST
Okay, I like
S
10:58speaker_2HOST
that.
5 more episodes mention pandas.
Create an account to see the whole feed, search across every transcript, and follow the entities you care about.