You convert a GeoJSON file to a shapefile and load it into ArcGIS. Your attribute table has changed. The field called municipality_name is now municipali. The field called building_type_code and the field called building_type_category are both now called building_t, except one of them got renamed to building_t0 by GDAL and you missed it. Your spatial join script that references building_type_code now silently returns nulls on every row.
This is not a GDAL bug. It is the shapefile format working exactly as designed since 1983.
Why the Limit Exists
The .dbf file is dBASE III format. When Esri adopted it as the attribute store for shapefiles in the early 1990s, dBASE III already defined field names as 10-byte, null-terminated ASCII strings — a fixed field in the file header. That limit was baked into the binary format specification.
GeoJSON, defined in RFC 7946 (2016), places no restriction on property names. Property names are JSON keys: any valid UTF-8 string. The format is structurally indifferent to whether a key is 4 characters or 40.
When you convert GeoJSON to shapefile, something has to resolve this mismatch. GDAL’s OGR library does it automatically, silently, and in a way that can corrupt your data’s schema without raising an error.
What GDAL Actually Does
GDAL 3.x (OGR’s shapefile driver) applies this algorithm when a GeoJSON property name exceeds 10 characters:
- Truncate to 10 characters.
municipality_name→municipali. - Check for collision. If the truncated name already exists (because another property also truncated to the same string), append a numeric suffix.
- Increment until unique.
building_type_code→building_t;building_type_category→building_t(collision) →building_t0. If there is a third collision:building_t1. And so on. - Write a warning to stderr. This is the only signal you get. If you are not capturing stderr — if you are running
ogr2ogrin a pipeline that discards stderr, or calling OGR from Python and not checking warnings — the truncation is invisible.
The resulting shapefile has different field names than your source GeoJSON. Any code, query, or join that references the original field names breaks silently.
Detecting Truncation Before It Happens
Run ogrinfo on your source GeoJSON before conversion and count field name lengths:
ogrinfo -al -so your_file.geojson | grep -E "^\s+\w+ \(.*\)"
The output lists each field with its type. Count characters in any field name longer than 10. Any field longer than 10 characters will be truncated.
For a programmatic check:
python3 -c "
import json, sys
with open('your_file.geojson') as f:
gj = json.load(f)
if gj['features']:
props = list(gj['features'][0]['properties'].keys())
long_names = [n for n in props if len(n) > 10]
if long_names:
print('Will be truncated:', long_names)
else:
print('All field names OK')
"
Run this before every GeoJSON-to-shapefile pipeline. If it prints names, you need to either rename fields in the source or use a field mapping at conversion time.
The ogr2ogr Field Mapping Approach
ogr2ogr accepts a configuration file (--config OGR_FIELDMAP_PATH) that maps source field names to output field names. Use this to control what the shapefile field names will be:
ogr2ogr \
-f "ESRI Shapefile" \
output.shp \
input.geojson \
-fieldmap "municipality_name=muni_name,building_type_code=bldg_code,building_type_category=bldg_cat"
The -fieldmap option syntax in GDAL 3.5+ accepts comma-separated source=dest pairs. For older GDAL versions, use a VRT (Virtual Dataset) file to define the field renames before conversion.
Alternatively, use Python’s osgeo.ogr directly and set output field names explicitly:
from osgeo import ogr, osr
src = ogr.Open('input.geojson')
src_lyr = src.GetLayer()
src_defn = src_lyr.GetLayerDefn()
drv = ogr.GetDriverByName('ESRI Shapefile')
dst = drv.CreateDataSource('output.shp')
dst_lyr = dst.CreateLayer('output', srs=src_lyr.GetSpatialRef())
# Rename long fields explicitly
field_map = {
'municipality_name': 'muni_name',
'building_type_code': 'bldg_code',
'building_type_category': 'bldg_cat',
}
for i in range(src_defn.GetFieldCount()):
fld = src_defn.GetFieldDefn(i)
name = fld.GetName()
new_name = field_map.get(name, name[:10]) # truncate only as fallback
new_fld = ogr.FieldDefn(new_name, fld.GetType())
new_fld.SetWidth(fld.GetWidth())
dst_lyr.CreateField(new_fld)
for feat in src_lyr:
new_feat = ogr.Feature(dst_lyr.GetLayerDefn())
new_feat.SetGeometry(feat.GetGeometryRef().Clone())
for i in range(src_defn.GetFieldCount()):
src_name = src_defn.GetFieldDefn(i).GetName()
dst_name = field_map.get(src_name, src_name[:10])
new_feat.SetField(dst_name, feat.GetField(i))
dst_lyr.CreateFeature(new_feat)
This gives you full control. The output shapefile has exactly the field names you specified.
Encoding Is a Separate Problem
Field name truncation is one issue. Field value encoding is another, and they can compound each other.
Shapefiles store attribute values in the .dbf using the encoding declared in the optional .cpg file. If your GeoJSON contains non-ASCII characters in property values (accented characters, CJK characters, Arabic text), the conversion needs to declare the encoding correctly. Without a .cpg file, ArcGIS assumes the platform default encoding (often Windows-1252 on Windows), and any UTF-8 character above U+007F becomes garbled.
Pass -lco ENCODING=UTF-8 to ogr2ogr to create the .cpg file with the correct encoding declaration.
A conversion that silently truncates field names while also silently mis-encoding field values is the scenario that produces the worst outcomes — your downstream analysis runs, produces numbers, and those numbers are wrong in a way that is not immediately obvious.
The Null Geometry Problem
One more silent failure mode: GeoJSON allows features with "geometry": null. Shapefile geometry is required — every record must have valid geometry. By default, GDAL drops null-geometry features silently. They disappear from the output without warning.
Check before conversion:
python3 -c "
import json
with open('input.geojson') as f:
gj = json.load(f)
null_count = sum(1 for f in gj['features'] if f.get('geometry') is None)
print(f'Null geometry features: {null_count}')
"
If the count is nonzero, decide whether to filter them explicitly (-where "geometry IS NOT NULL" in ogr2ogr) or handle them upstream.
Why This Matters for Repeatable Pipelines
The field truncation problem is worst in automated pipelines where the source GeoJSON changes over time. A field name that was 9 characters in the initial data might become 11 characters in a later version (a schema change upstream). The pipeline continues to run, the truncation happens, joins break, and the failure is in the query layer — not in the conversion step.
Treat field name validation as a schema contract: before any GeoJSON-to-shapefile conversion in a pipeline, assert that all field names are ≤ 10 characters. Fail fast at the assertion rather than silently at the join.
For format tradeoffs when shapefile is not actually required, see GeoParquet vs GeoPackage vs Shapefile: a decision tree — GeoPackage and GeoParquet both use full-length field names and avoid the 10-character problem entirely.
Using GeoConvert
GeoConvert handles GeoJSON-to-shapefile conversion and surfaces truncation warnings explicitly in the conversion result, including a field name mapping table that shows exactly what each source field was renamed to. This makes the schema change visible rather than buried in stderr output you have to know to look for.
For validating shapefiles after ingestion from any source, the shapefile validation preflight checklist covers CRS verification, DBF field constraints, geometry validity, and encoding — the full set of checks before loading external GIS data.
Summary
The .dbf 10-character field name limit is not going away — it is part of the binary format specification. When converting GeoJSON to shapefile:
- Check field names before conversion: any name longer than 10 characters will be truncated
- GDAL appends numeric suffixes on collision, but only warns to stderr — no error is raised
- Use explicit field mapping (
-fieldmapor Python OGR) to control output names - Add
-lco ENCODING=UTF-8to preserve non-ASCII character values - Check for null-geometry features before conversion — they are silently dropped
- In automated pipelines, assert field name length as a schema contract before conversion