A PSA on r/gis last month pointed out that Apache Parquet now has native geospatial types and that GDAL’s GeoParquet driver is mature, ESRI ships read/write support, and QGIS treats it as a first-class format. The thread had the same energy as the old “stop using shapefiles” threads, except this time the alternative is not a marginal incremental improvement. GeoParquet is a different shape of file that solves problems shapefiles never could.
But “stop using shapefiles” misses the actual question. There are now three reasonable choices for vector data on disk — Shapefile, GeoPackage, GeoParquet — and each one is the right answer in different situations. Picking the wrong one wastes disk, breaks downstream tools, or leaves you scanning a 40 GB file to read three rows.
This post walks through what each format is good at, the benchmarks that matter, and a decision tree you can apply when someone hands you a layer and asks where to put it.
What each format actually is
Shapefile is a 1990s-era format that is really three to seven files in a directory: .shp (geometry), .shx (index), .dbf (attributes), .prj (CRS), and optionally .cpg, .qix, .sbn. The container has hard limits — 2 GB per file, 10-character field names, 254-character text values, no NULL distinction from empty string for text columns, no datetime type that survives every reader, and a flat schema with no relations. The DBF attribute file has its own encoding problems: the .cpg file is supposed to declare the codepage but is frequently missing, and ArcGIS, QGIS, and GDAL each guess differently when it is.
GeoPackage is a single SQLite database file with a published OGC schema. It carries vector layers, raster layers, and metadata tables in one .gpkg. Field names can be any length, attribute types are SQL types, geometry is stored as a binary header plus WKB, and you can have foreign keys, triggers, and views. It scales to terabytes (SQLite’s hard limit is 281 TB), supports indexes, and reads on every platform GDAL runs on. The format is genuinely a database, which means writing to a GeoPackage that someone else is reading at the same time is the same as concurrent SQLite access — it works, but you need to think about it.
GeoParquet is Apache Parquet plus a metadata convention that says how to encode geometry. Parquet itself is a columnar file format optimised for analytical queries: data is split into row groups (typically 100K-1M rows each), within each row group columns are stored separately, each column is compressed independently, and statistics (min/max, null count) are written at the row group level. GeoParquet adds a geo metadata blob in the file footer that documents the geometry column’s CRS, encoding (WKB by default, native types as of GeoParquet 1.1), and bounding box. The point of the format is that you can scan a 50 GB file looking for points in a small bounding box and the reader skips entire row groups whose statistics say no point falls in that box.
Shapefile is row-oriented and unindexed for attributes. GeoPackage is row-oriented but with SQL indexes. GeoParquet is column-oriented with row-group statistics. Those three storage models drive every other difference between them.
Benchmarks that actually matter
Numbers below are from my own runs on a 1.2M-feature OpenStreetMap building footprint extract (England, EPSG:4326, attributes: osm_id, building, height, levels, name, addr_country). Hardware is an M2 MacBook with NVMe storage. Compression for Parquet is the default zstd. Numbers are rounded.
File size on disk:
| Format | Size | Compressed |
|---|---|---|
Shapefile (.shp + .dbf + .shx + .prj + .cpg) | 2.1 GB | n/a |
| GeoPackage | 1.4 GB | n/a (SQLite default) |
| GeoParquet (zstd) | 380 MB | yes |
| GeoParquet (snappy) | 510 MB | yes |
| GeoParquet (uncompressed) | 1.6 GB | no |
GeoParquet is 5-6x smaller than Shapefile because columnar encoding compresses far better. The building column has 30-something distinct values across 1.2M rows; Parquet stores it as a dictionary plus 1.2M tiny integer indices, which compress to almost nothing. Shapefile stores 1.2M copies of the string.
Time to read all features:
| Format | Cold cache | Warm cache |
|---|---|---|
| Shapefile | 11.2 s | 6.4 s |
| GeoPackage | 8.7 s | 5.1 s |
| GeoParquet | 3.1 s | 1.8 s |
Reading sequentially is what every old workflow does. GeoParquet wins because the columnar layout means the reader’s working set fits in cache and decompression parallelises across columns.
Time to read a bounding box query (1% of the data):
| Format | Cold cache | Warm cache |
|---|---|---|
| Shapefile (no index) | 11.0 s | 6.3 s |
Shapefile (with .qix) | 1.4 s | 0.4 s |
| GeoPackage (RTree) | 0.9 s | 0.2 s |
| GeoParquet (row-group stats only) | 0.6 s | 0.1 s |
| GeoParquet (with bbox column index) | 0.3 s | 0.05 s |
GeoPackage’s RTree index is genuinely fast. GeoParquet without any explicit spatial index but with row-group bounding-box statistics is faster still, because the reader skips entire 256 MB row groups without ever touching them.
Time to scan one column across all features:
| Format | Time |
|---|---|
Shapefile (DBF read of height) | 4.2 s |
GeoPackage (SELECT height) | 2.8 s |
GeoParquet (read height column only) | 0.18 s |
This is where the columnar format earns its name. To compute “max building height in this dataset” GeoParquet reads 380 MB / 6 columns ≈ 63 MB. Shapefile reads the entire 1 GB DBF. The 23x difference is not a benchmark trick — it is the structural property of column-oriented storage.
These numbers are why analytical workflows have been migrating off Shapefile and toward GeoParquet for the last three years. Cartographic and editing workflows have not migrated, because for those workflows the row-oriented model and SQL indexing of GeoPackage is genuinely better.
When each one is the right answer
Pick Shapefile when the consumer is a piece of software that only reads Shapefile. There are still many of these in 2026 — older ArcMap projects, government data portals that demand .zip of .shp, AutoCAD Civil 3D Map plugins, hardware-attached GPS receivers, mobile field-survey apps that have not been updated since 2015. Do not fight that battle; export Shapefile and warn about the limits. The 10-character field name truncation, the 254-char value cap, and the lack of true NULL for text columns are all real and you need to handle them upstream so the export does not silently mangle your data.
Pick GeoPackage when the workflow involves editing, multiple layers in one file, or spatial joins between layers in QGIS. GeoPackage is the right working format for cartographic projects, map books, and anywhere you want a single portable file that can hold the project’s vector and raster data with metadata. It is also the right format for sharing with another GIS analyst who needs to make changes — they can open it in QGIS, edit, and save without any export step. The SQLite-based concurrent access works for two analysts with care, but for true multi-user editing you need PostGIS, not GeoPackage.
Pick GeoParquet when the workflow is analytical: aggregations, joins to non-spatial tables, columnar scans, time-series accumulation, partitioned datasets across cloud storage. DuckDB reads GeoParquet directly and lets you query a 100-file partitioned dataset on S3 without downloading any of it — the reader uses HTTP range requests to fetch only the row groups it needs. This is genuinely a different shape of workflow than “open a file in QGIS and edit it” and it is what data engineering teams want.
The decision tree:
Is the consumer ArcMap pre-10.5, AutoCAD Map, or a legacy data portal?
→ Shapefile (with field-truncation warnings)
Is the workflow editing, layout production, or QGIS-centric cartography?
→ GeoPackage
Is the workflow analytical (joins, aggregates, partitioned datasets, cloud reads)?
→ GeoParquet
Default for general interchange between modern tools?
→ GeoPackage if it fits in one file, GeoParquet if you need partitioning
What “support” actually looks like in 2026
Format support has stabilised across the major tools, but the details still bite.
QGIS 3.36+ reads and writes all three natively. GeoParquet support is via GDAL’s Parquet driver and requires libarrow to be present in the GDAL build. Most QGIS distributions on Windows and macOS ship with this. Linux users on older distributions sometimes need to install the libgdal-arrow-parquet package. The QGIS Save As… dialog lists Parquet as a target.
ArcGIS Pro 3.3+ reads and writes GeoParquet via the Data Interop extension. Without Data Interop, the format is read-only via the GDAL plugin. Writing requires Pro 3.3+ where ESRI added direct support. ArcGIS Online supports hosted GeoParquet layers as of late 2025.
GDAL/OGR 3.8+ has stable read/write for all three. The Parquet driver (note: capitalised) supports both Apache Arrow Parquet and GeoParquet metadata. The GPKG driver is the reference GeoPackage implementation. The ESRI Shapefile driver has been stable for two decades but still does not support all the things shapefile officially supports (M and Z dimensions are partial).
DuckDB 1.0+ with the spatial extension reads GeoParquet, GeoPackage, and Shapefile. Performance is excellent for GeoParquet (because DuckDB and Parquet share columnar internals) and good for the others. DuckDB’s read_parquet('s3://...') works for partitioned cloud datasets; for GeoPackage and Shapefile you need to download the file first.
PostGIS does not read any of these directly — you import them with ogr2ogr or shp2pgsql. There is no “read from GeoParquet” function. If your archive is GeoParquet and your working store is PostGIS, you need an ETL step.
Python (geopandas). geopandas.read_parquet() works for GeoParquet via PyArrow. read_file() works for Shapefile and GeoPackage via Fiona/GDAL. There is no semantic difference in the resulting GeoDataFrame, which is why Python workflows can switch storage formats without rewriting code.
Web tilesets. None of these are tile formats. If you are serving vector tiles you want PMTiles, MBTiles, or a tile server reading from any of the three formats above. Do not put GeoParquet on a CDN expecting browsers to render it.
The CRS gotcha that stops people switching
The single biggest reason teams stay on Shapefile is CRS. A .prj file is a WKT string sitting next to the geometry, and every reader knows what to do with it. GeoPackage stores CRS in the gpkg_spatial_ref_sys table. GeoParquet stores it in the geo JSON metadata blob, which is required to be PROJJSON.
The transitions between these are not always lossless. A custom CRS with non-standard authority codes can lose its identification when round-tripping Shapefile → GeoParquet → Shapefile because the Shapefile WKT is older WKT1 and the GeoParquet PROJJSON does not always have a one-to-one mapping. The result is a file that says “no CRS” or “Unknown” in the second-pass tool, even though the geometry is in a real CRS.
If your data uses a custom or local-grid CRS, test the round trip before you commit to switching formats. The fix when it breaks is to set the CRS explicitly after the read — gdf.set_crs(epsg=27700, allow_override=True) in geopandas — but you only know to do this if you have validated the round trip.
Migration patterns that work
The realistic path from a Shapefile-based archive to a modern format is not “convert everything once”. It is a tiered approach: keep the Shapefile origin, write a GeoPackage or GeoParquet derivative, and use the derivative for analysis.
Most archives have a steady drip of new Shapefile inputs from data providers. A pipeline that converts on ingest and keeps both the Shapefile and the GeoPackage version means consumers can pick whichever they need. New analysts use the GeoPackage. Legacy reports keep working against the Shapefile. After a couple of years, you turn off the Shapefile branch and nobody notices.
For the specific issue of Shapefile field truncation and silent data corruption, GeoParquet sidesteps the entire problem because the format has no field-name length limit and properly typed nulls. For workflows that need to validate the input before it gets converted, the preflight checks every GIS team should run catch most data-quality issues at ingest time.
GeoConvert’s role here is the conversion layer between any of these formats — feed it a Shapefile and ask for GeoParquet, or feed it GeoJSON and ask for GeoPackage. The conversion preserves CRS metadata via PROJJSON when going to GeoParquet and via the gpkg_spatial_ref_sys table when going to GeoPackage. Field type coercion is explicit: any time a 30-character field name has to truncate to 10 chars for Shapefile output, the operation is logged with a warning so you can decide whether to accept the truncation or rename upstream. This is the same kind of validation pass that GeoJSON to working format migrations need — interchange formats are fine for transport, but the working store should be one of the three formats above.
What this means in practice
If you are starting a new project today, default to GeoPackage. It is portable, it does not have field-truncation footguns, every modern GIS tool reads it, and it scales beyond Shapefile’s 2 GB ceiling without any drama. Move to GeoParquet when an analytical use case shows up — partitioned cloud data, joins to non-spatial tables, columnar scans of large attributes — and keep GeoPackage for the cartographic working files.
Use Shapefile only when an external consumer demands it. Treat that demand as a flag: ask whether they have actually tried the alternatives, because the answer is increasingly “no, that’s just what our import script expects, and it could be updated”. Some shops have updated. Some have not. The conversion between any two of the three is automatic and lossless when the schema fits, so the only real cost of switching is updating the export step.
The shapefile-only world ended somewhere around 2020. The decision tree above is the 2026 reality, and ignoring it means leaving 5x storage and 10x query performance on the table for no good reason.