GeoJSON keeps showing up in places it shouldn’t. Someone drops a 2 GB .geojson into a processing pipeline, the script that used to run in 20 seconds runs for eight minutes, and the team spends a Thursday wondering whether the problem is their query planner, their RAM, or their patience. The problem is the format. GeoJSON is a perfectly good interchange format — human-readable, one-spec wonder, works everywhere a text editor opens. It’s a terrible working format as soon as your dataset has more than a few hundred thousand features. This post is about when to stop loading GeoJSON directly and what to convert to instead, based on what teams handling real-scale data actually reach for.
Why GeoJSON Breaks at Scale
Three characteristics of the spec cause most of the pain.
It’s one giant JSON document. A standard .geojson file is a single FeatureCollection object. To read a single feature, a parser has to parse the entire file into memory. A 1 GB file needs roughly 2-4 GB of RAM to deserialize, because every coordinate becomes a boxed floating-point number in whatever language you’re in. Parsing a 27 MB GeoJSON file into memory has been measured at several seconds and around 240 MB of RAM — two orders of magnitude more memory than the file on disk.
No index. Nothing in the format supports spatial filtering without reading everything. If you want all features inside a bounding box, you read every feature. Compare this to GeoPackage, which has an R-tree index built into SQLite, or GeoParquet, where row groups carry bounding-box metadata in their headers so DuckDB can skip entire chunks.
Text encoding of numbers. A double-precision coordinate takes 8 bytes in binary and averages 17-20 bytes as ASCII. A dataset with a few million points is two to three times larger as GeoJSON than as Parquet. You’re paying disk, paying bandwidth, and paying deserialization cost on every read.
These aren’t bugs in GeoJSON. They are the consequences of a format designed for web APIs that return small feature collections, being used as a data warehouse. The spec is doing exactly what it was designed to do. The mismatch is in the use case.
The Practical Threshold
A rough rule from teams handling real datasets: under 100 MB, GeoJSON is fine. Between 100 MB and 1 GB, you’ll feel the pain in parse times and editor crashes but the job still completes. Above 1 GB, you’re picking between conversion to a binary format or building a custom streaming parser. Above 10 GB, GeoJSON stops being viable on a single laptop and you have to either shard it or convert.
The conversion is nearly always cheaper than the engineering work to keep using GeoJSON. GeoConvert handles shapefile outputs with the field-truncation and CRS handling that trip up quick-and-dirty conversions, but the target format matters as much as the converter. The rest of this post is about picking the right target.
What Teams Actually Use
GeoParquet for columnar analytics
GeoParquet wraps Parquet’s columnar storage with WKB geometry columns and a metadata block describing the CRS. Three properties make it the default choice for analytical workloads:
- Column pruning. Reading only the columns you need avoids deserializing the rest. A dataset with 40 attribute columns where your query uses two reads around 5% of the file.
- Row group skipping. Row groups carry min/max bounding boxes in metadata. A bounding-box filter can skip every row group that doesn’t overlap. For spatial queries on partitioned data, this is effectively a built-in spatial index.
- Compression by default. Snappy or Zstd compression typically yields 3-6x size reduction on GIS attribute data, and coordinate data compresses well when stored as WKB rather than text.
Benchmarks from the cloud-native geospatial community put GeoParquet ingestion at ~11 seconds for a dataset that took 1 minute 40 seconds via GeoPackage and 1 minute 42 seconds via Shapefile on the same workload. The ratio holds up across most operations.
Where GeoParquet wins: read-heavy analytics, cloud storage, DuckDB or Athena or Spark workloads, anything you expect to query with SQL. Where it doesn’t: interactive editing. There’s no transactional mutation model. You rewrite the file to change it. If your workflow is “open the dataset, edit a polygon, save,” Parquet is the wrong choice.
GeoPackage for offline editing and mixed workloads
GeoPackage is SQLite with a spatial extension and a standardized schema. Because it’s SQLite, you get ACID transactions, you can edit individual features, and you can ship the file to a colleague who opens it in QGIS. It has an R-tree spatial index built in, so bounding-box queries are fast without extra setup.
GeoPackage is the format of choice when the workflow includes editing. QGIS uses it as the native save format for a reason — the index is there, the transactional model works, and you can store multiple feature classes and even raster layers in one file. For single-user or small-team workflows where the data isn’t massive but needs editing, GeoPackage beats Parquet and beats Shapefile by a wide margin.
Its weakness is analytical scan speed on very large files. Once you’re querying hundreds of millions of features, the row-by-row storage and per-row overhead show up. GeoPackage handles a million features easily. Ten million takes noticeable time. A hundred million and you want Parquet.
PostGIS for multi-user and server-side
When multiple people or services need concurrent access, you’re in database territory. PostGIS on PostgreSQL is the conventional choice and the path most serious GIS teams end up on. GiST indexes on geometry columns, rich spatial SQL, battle-tested concurrency, support for tiling via ST_TileEnvelope, mature tooling.
The cost is the operational overhead of running a database. For a small team with one or two analysts, PostGIS is overkill. For any application serving spatial data to users, it’s usually the right answer unless you specifically need the read-only, serverless model of object storage + Parquet.
DuckDB spatial as the middle ground
DuckDB with the spatial extension has become many teams’ default for “I need SQL on my GeoJSON but I don’t want to set up PostGIS.” It reads GeoJSON, GeoPackage, Shapefile, and GeoParquet directly, runs spatial SQL on them, and works as a single-binary install. For analytical workloads on a laptop or a single server, it’s often faster than PostGIS for pure scan-and-aggregate operations because it doesn’t have the transactional overhead.
The idiomatic pattern is: convert the GeoJSON to GeoParquet once, then query with DuckDB. You get columnar speed, spatial SQL, and no database to administer.
-- With DuckDB spatial
INSTALL spatial; LOAD spatial;
CREATE TABLE parcels AS
SELECT * FROM ST_Read('parcels.geojson');
COPY parcels TO 'parcels.parquet' (FORMAT 'parquet');
-- Subsequent queries
SELECT COUNT(*) FROM 'parcels.parquet'
WHERE ST_Intersects(geometry, ST_GeomFromText('POLYGON(...)'));
The second query is the one that was taking eight minutes on the raw GeoJSON. On Parquet with a bounding-box-aware scan it drops to seconds, because DuckDB reads row-group metadata and skips everything that doesn’t overlap.
Line-delimited GeoJSON as the escape hatch
If for whatever reason you must stay in JSON territory, there’s a middle ground: newline-delimited GeoJSON, also known as GeoJSONL, NDJSON, or JSON Text Sequences under RFC 8142. Each line is a single Feature object. Parsers can read one line at a time without loading the whole file, which solves the memory problem but keeps the text-encoding and no-index problems.
GeoJSONL is the right pick when you’re piping features through UNIX tools (grep, jq, awk), when a downstream system only accepts JSON, or when you’re streaming features through a message queue. It’s not a working format for analysis. It’s a transport format that happens to play well with line-oriented tooling.
Converting GeoJSON to GeoJSONL is mechanical — read the FeatureCollection, write each feature to a new line. The opposite direction is equally mechanical. Useful to know as an escape hatch, but don’t confuse it with solving the scale problem.
A Concrete Decision Tree
The simplification that covers 95% of cases:
- < 100 MB, single user, short-lived: keep GeoJSON. The overhead of conversion isn’t worth it.
- 100 MB to 10 GB, read-heavy, analytical: GeoParquet. Query with DuckDB.
- Any size, needs editing, single user or small team: GeoPackage.
- Multi-user, served over HTTP, application backend: PostGIS.
- Transport between systems, message queues, streaming: GeoJSONL.
- Interop with legacy GIS, ArcGIS, AutoCAD: Shapefile (with eyes open about the silent corruption risks and the preflight checks GIS teams should run).
Most teams end up with two or three of these. GeoPackage for editing, Parquet for analytics, Shapefile for the one downstream system that won’t accept anything else. GeoJSON is usually what arrives, not what stays.
Why GeoConvert Matters Here
GeoConvert’s job is the “GeoJSON arrives, Shapefile is what the downstream system accepts” moment. That conversion is where data silently corrupts — field names over 10 characters get truncated, attribute values over 254 characters get cut, multi-geometry features get split or rejected depending on the tool. Handling those correctly is table stakes for any conversion workflow, and doing it via ArcGIS for a one-off is an hour of GUI clicking. Running it through a validated converter with geometry repair and attribute-mapping rules is 30 seconds.
The broader point: the format you convert to should match your use case, not just the request. When a colleague asks for “a shapefile” but the actual workflow is analytical, sometimes the right answer is to convert to Parquet instead and hand them a DuckDB query. When the ask is “a GeoJSON of our parcels” but the file will be 1.5 GB, offering a GeoPackage or a link to a Parquet in object storage is a better answer. The conversion step is where those decisions get made.
What Changes When Data Lives in the Cloud
Cloud object storage changes the calculus for large datasets. GeoParquet on S3 or Azure Blob, queried via HTTP range requests, lets DuckDB or Athena read only the relevant row groups over the network. A 50 GB dataset can be queried with a 20 MB download if your query is selective enough. That model doesn’t work with GeoJSON — you’d need to download the entire file to read one feature.
The practical implication: teams building cloud-native pipelines converge on Parquet for storage and DuckDB/Athena/Sedona for query. The broader GIS scripting on-ramp eventually leads here for anyone whose datasets grow past laptop scale. The tooling has matured enough that it’s a week of migration work, not a quarter.
The Take
GeoJSON is fine. It’s the format your tiles come back in, the format the web API returns, the format every GIS tool can read. Use it for what it’s good at — moving data between systems. The moment the same file is being opened repeatedly for analysis, or being loaded into a process that runs more than once a day, convert it. The right target depends on your workload, but doing the conversion is almost always cheaper than tolerating GeoJSON’s scale ceiling for one more week.
Sources
- Performance Explorations of GeoParquet and DuckDB — Cloud-Native Geospatial Forum
- Introducing geoparquet-io — Cloud-Native Geospatial Forum
- Efficiently querying large geospatial datasets with DuckDB, GeoParquet and S3 (Madole)
- Geospatial Tools Compared: GeoPandas, PostGIS, DuckDB, Sedona, Wherobots (Matt Forrest)
- GeoJSONL: An optimized format for large geographic datasets (Interline)
- Tips for optimising large GeoJSON files (Open Innovations)
- Newline-delimited GeoJSON reference (ndgeojson)
- GeoParquet specification