A state environmental agency analyst posted on r/gis last week, describing the same problem every public-sector GIS team deals with: contractors, federal agencies, and utility companies all send shapefiles that are technically valid per the format spec, and yet every batch requires hours of cleanup. The analyst was scoping an open-source CLI tool to run preflight checks — field truncation, CRS mismatch, missing indexes, null geometry.
The tool they need already exists in pieces inside GDAL. The problem is that nobody publishes the checklist. Here is the full preflight — the specific failure modes, the ogrinfo and ogr2ogr commands that detect them, and the decisions each check forces.
What a shapefile actually is
A shapefile is not one file. It is a bundle of sibling files sharing a base name. If any of the required three are missing, the dataset is unusable; if any of the optional ones are missing, you lose capability silently.
| File | Required | Purpose |
|---|---|---|
.shp | yes | Geometry records. |
.shx | yes | Byte-offset index into .shp. |
.dbf | yes | Attribute table (dBase III/IV). |
.prj | no | WKT coordinate reference system. |
.cpg | no | DBF code page (UTF-8, ISO-8859-1). |
.sbn / .sbx | no | ESRI proprietary spatial index. |
.qix | no | Quadtree spatial index (GDAL, MapServer). |
A zip archive with county.shp, county.shx, county.dbf, and county.prj is the minimum viable delivery. A zip with only the .shp will open in some tools (QGIS warns but tries), and the .shx will be regenerated — but counties with null geometry records will shift indices, and your joins will fail silently three weeks later.
Check one: every .shp you ingest must have a matching .shx and .dbf of the same base name, with matching record counts. The .prj must exist and be non-empty.
The field-name truncation problem
The DBF format limits field names to 10 bytes. Every downstream tool truncates silently. GDAL, ArcGIS Pro, and QGIS all do it; only QGIS shows a visible warning on export, and none show it on import. If an input layer has fields BuildingHeightMeters and BuildingHectares, both collapse to BUILDINGHE. GDAL appends _1, _2 to duplicates:
BuildingHeightMeters → BUILDINGHE
BuildingHectares → BUILDINGH_1
ArcGIS handles the collision differently — a warning dialog, a fallback to the first field wins. By the time the data reaches analysis, nobody remembers which field is which.
Check two: before ingesting a shapefile from an external source, enumerate fields and assert length ≤ 10 and uniqueness on the 10-char prefix.
ogrinfo -ro -so layer.shp | grep "^[A-Z_]*:"
If two fields share their first 10 characters, rename them upstream.
The 254-byte limit on C (text) fields is a parallel concern. DBF C fields max out at 254 bytes per value, and with .cpg of UTF-8, that is bytes — not characters. A Cyrillic or Chinese value can truncate mid-codepoint, producing a .dbf that fails to decode in tools expecting clean UTF-8. The failure manifests as replacement characters or blank attributes. No warning is emitted.
The .prj integrity check
The .prj is a single line of ESRI-flavor WKT1. Four failure modes, each with different downstream consequences:
- Missing
.prj. GDAL reports CRS asUnknown. Most reprojection tools refuse to operate. You can reconstruct if you know the CRS:gdalsrsinfo -o wkt "EPSG:27700" > layer.prj. - Zero-byte
.prj. Identical to missing, but a directory listing looks correct. Catch this with a file-size check; do not trust the extension. - Valid WKT but no
AUTHORITY[...]token. The CRS parses but automated EPSG lookup fails. Downstream systems storing EPSG codes recordnull. - ESRI-only parameters — for example
Hotine_Oblique_Mercator_Azimuth_Natural_Origin. GDAL ≥ 3.0 handles these; older GDAL builds and non-ESRI tools fail. Always test with the target toolchain before ingestion.
Catch these in one command:
ogrinfo -ro -so layer.shp | grep -E "(Layer SRS WKT|Unknown|AUTHORITY)"
Absence of AUTHORITY["EPSG",...] is a yellow flag. Presence of Unknown is a red flag — reject and ask the supplier for a .prj.
GeoConvert’s own pipeline uses GDAL-powered CRS handling and refuses to write an output shapefile without an AUTHORITY-tagged CRS. This is not a polite behaviour — it is the only way to guarantee round-trip interoperability.
Geometry type must be uniform
The ESRI spec (§3) is explicit: “All the non-Null shapes in a shapefile are required to be of the same shape type.” Each record carries a shape-type header, but every non-null record must match the file’s declared type.
Shape type codes:
| Code | Type |
|---|---|
| 0 | Null |
| 1 | Point |
| 3 | PolyLine |
| 5 | Polygon |
| 8 | MultiPoint |
| 11, 13, 15, 18 | PointZ, PolyLineZ, PolygonZ, MultiPointZ |
| 21, 23, 25, 28 | PointM, PolyLineM, PolygonM, MultiPointM |
| 31 | MultiPatch |
MultiPolygon and MultiLineString have no dedicated codes. MultiPolygons live inside type 5 (Polygon) with multiple rings, and MultiLineStrings live inside type 3 (PolyLine) with multiple parts. You cannot distinguish single-part from multi-part geometry at the type level — the distinction appears in each record’s NumParts field.
The tricky case: an input layer stored as GeoJSON FeatureCollection with mixed geometry (some Points, some Polygons) cannot round-trip to shapefile. A conversion pipeline that silently writes only the first type — or splits the data into two files — is a bug. Flag this before ingestion:
ogrinfo -ro -al layer.shp | grep "Geometry:"
There should be exactly one geometry line. Anything else and the upstream supplier already lost data.
Null vs empty geometry
The spec distinguishes two concepts that downstream tools conflate:
- Null shape (type 0): a record in
.shpwith zero bytes of geometry payload. Attributes in.dbfstill exist and are valid. Use case: a feature whose location is unknown but whose attributes still matter. - Empty parts inside a multi-part geometry:
NumParts > 0with at least one part whereNumPoints = 0. This is always a data bug. ArcGIS Check Geometry reports it as “Empty part in multipart feature.” GDAL exposes it throughOGRGeometry::IsEmpty()on sub-geometries.
There is also the null-attribute problem. DBF has no true NULL bit. ArcGIS writes a sentinel value that its own tooling interprets as NULL, but any tool round-tripping through GDAL collapses it to 0 (for numerics) or empty string (for text). A Population column with unreliable data will become a column of zeros, distorting every aggregate statistic downstream.
Two checks, both with SQLite-dialect SQL:
# Count null geometries
ogrinfo -ro -q -dialect SQLITE -sql \
"SELECT COUNT(*) FROM layer WHERE geometry IS NULL" layer.shp
# Count invalid geometries
ogrinfo -ro -q -dialect SQLITE -sql \
"SELECT COUNT(*) FROM layer WHERE ST_IsValid(geometry) = 0" layer.shp
Both require GDAL compiled with SpatiaLite support. Most official GDAL binaries include it.
CRS plausibility — even when .prj looks fine
A valid .prj does not guarantee the coordinates are in that CRS. The most common failure: a file is reprojected to Web Mercator but the .prj is never updated, so EPSG:4326 claims to describe coordinates that are actually EPSG:3857. Every downstream reprojection pushes the data further off.
Sanity bounds by CRS family:
EPSG:4326(geographic WGS84): X ∈ [-180, 180], Y ∈ [-90, 90].EPSG:3857(Web Mercator): |X| and |Y| ≤ ~20,037,508.- UTM zones: X ∈ [166,021, 833,978], Y ∈ [0, 10,000,000].
- State Plane / national grids: supplier-specific, but always within a bounded rectangle that the CRS authority publishes.
Extract the extent and bounds-check:
ogrinfo -ro -so -al layer.shp | grep "Extent:"
A file that claims EPSG:4326 with Extent: (400000, 5600000) - (500000, 5700000) is lying. The numbers are UTM coordinates; the CRS tag is wrong. Reject the dataset and ask for a correct .prj.
Spatial indexes — not required, but performance-critical at scale
.shp and .shx alone let you read the file correctly. Indexes (.qix, .sbn/.sbx) affect speed, not correctness. GDAL writes .qix (quadtree) and reads .sbn/.sbx (ESRI proprietary) but does not write the ESRI form.
For files above ~100,000 features or ~100 MB, every spatial filter without an index triggers a full scan. On a contractor deliverable where you plan to do point-in-polygon queries, the absence of .qix means your web service will hit 10× slower than it needs to.
Create on demand:
# Add a quadtree index to an existing shapefile
ogrinfo mylayer.shp -sql "CREATE SPATIAL INDEX ON mylayer"
# Or build it at creation time
ogr2ogr -lco SPATIAL_INDEX=YES out.shp in.shp
Check: if feature count > 10,000 and the delivery has no .qix, add one before loading into your working environment. It is cheap up front and free forever after.
Field-type coercion risks
DBF field types have surprising limits. The types you should expect:
N(numeric): stored as ASCII text, width 1–19. Integer types cap at 9 digits (4-byte signed int); values beyond 2,147,483,647 wrap or clamp. Floats with width 13 and 4 decimals retain only ~8 significant digits — a risk if you store coordinates as attributes.D(date): 8-byteYYYYMMDDstring. No time component. Timestamps silently lose their time portion on round-trip.L(logical): 1 byte,T/F/?. GDAL converts to integer 0/1, losing the?unknown state.C(text): up to 254 bytes per value..cpgcontrols encoding.
Preflight:
ogrinfo -ro -so layer.shp | grep -E ":\s*(Integer|Real|String|Date)"
Review every field that might carry integers > 2 billion or floats needing > 8 digits. These often silently degrade on first ingest and are unrecoverable without the source.
Running -makevalid on invalid geometry
GDAL’s ogr2ogr has two flags for dealing with invalid geometry:
-validateskips invalid features.-makevalidrepairs them by re-noding all rings and extracting valid faces.
# Skip invalid features (destructive — records are silently dropped)
ogr2ogr -validate out.shp layer.shp
# Attempt repair (conservative — preserves record count, splits complex fixes)
ogr2ogr -makevalid out.shp layer.shp
-makevalid uses GEOS’s ST_MakeValid in LINEWORK mode by default. For self-intersecting polygons this splits the input into multiple valid polygons stored as a single multi-polygon. Record count stays the same but geometry complexity can increase significantly.
If ogr2ogr -validate drops more than 1% of your features, do not silently accept the output — go back to the supplier.
The full preflight script, in order
A concrete order-of-operations that catches every failure mode above:
- File-bundle check. Assert
.shp,.shx,.dbfall exist. Assert matching record counts (GDAL’sogrinfo -alwill fail loudly if they do not). .prjcheck. Assert file exists, non-empty, parseable WKT, containsAUTHORITY["EPSG",...]..cpgcheck. If anyCfield contains bytes >0x7Fand no.cpgexists, warn. ProposeUTF-8to the supplier.- Field schema check. Every field name ≤ 10 bytes, unique on 10-char prefix. Every
Cfield ≤ 254 width. Flag integers stored inNfields with width > 9. - Geometry type check. Exactly one geometry type per file. Enumerate type codes and assert uniformity.
- Geometry validity check. Count nulls and invalid geometries; reject if > 1% invalid.
- Extent check. Extract extent, cross-check against declared CRS family bounds.
- Spatial index check. If feature count > 10,000, confirm
.qixpresent or create it. - Smoke conversion. Round-trip through
ogr2ogrto a clean output; diff feature counts.
The whole thing fits in a 50-line bash script. Run it on every delivery before the data lands in production. Silent failures — field truncation, CRS mismatch, dropped nulls — are the ones that corrupt analysis three months later when nobody remembers who delivered the file.
Where GeoConvert fits
GeoConvert’s conversion pipeline runs most of these checks on input before it writes an output. Missing .prj is a hard error. Field names that would truncate on output are renamed with a documented mapping. Null geometries are preserved as null, not silently dropped. Invalid geometry triggers a prompt to repair, skip, or abort. This is a different philosophy from tools that prioritise throughput over correctness — and it is the reason GeoJSON to shapefile conversions silently corrupt data in most converters but not in a GDAL-backed pipeline that flags the problems instead of hiding them.
External data is not your data. Treat every delivery as untrusted input. Preflight every file. The 30 seconds it takes to run ogrinfo -al -so on a new shapefile will save you days of downstream debugging when a contractor’s UTM coordinates arrive in a file claiming EPSG:4326.
If you need to export your validated shapefile to DXF for use in Rhino or other CAD tools, see QGIS DXF export pitfalls after shapefile validation for the coordinate system and geometry issues that trip up the QGIS DXF exporter.