Amazon · Databricks · Snowflake · eddie.codes
Pandas Should Go Extinct
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
◌ Single Source
This post covers material from a talk the reporter gave at Latency Conference.
Key facts
- Let’s analyze the NYC Taxi dataset (trips from 2009-present, stored in monthly Parquet files) to see if cash payments became less common during the pandemic (2019-2022)
- To arrive at the claim of “94.68% of tables contain less than 100GB” they sum the rows up to the 10^8 limit, giving them 94.68% of rows
- The 1 Billion Row Challenge was a challenge to write the fastest Java program which could compute the min, mean and max of a 1 billion row CSV containing weather station data
- The original challenge used a bare metal Hetzner AX161 server with 32 cores and 128GB of RAM running Debian 12
Summary
You read that correctly, Pandas should go extinct. Because Pandas’ inefficiencies force you to adopt distributed querying systems before your workloads justify the added complexity. To understand what the reporter is talking about they first must understand the typical adoption pathway for Pandas. The diagram below shows a rough guide of when you typically would consider adopting a given DataFrame library based on the data size you are working with.