You have a CSV of postal addresses — a membership list, a customer database, a survey response set. Someone wants to know where these people are concentrated. A heatmap is the right output: it shows density without exposing individual addresses, it scales from 100 to 100,000 records without changing the workflow, and it overlays cleanly on basemaps, zoning layers, or any other spatial data you have.
The full pipeline in QGIS: geocode addresses to coordinates using MMQGIS and Nominatim (OpenStreetMap’s free geocoding service), reproject the resulting point layer to a metric CRS, and run kernel density estimation in the QGIS Processing Toolbox. Three steps. The traps are in the rate limits, the CRS handling, and the distinction between two different “heatmap” tools inside QGIS that look alike and behave differently.
Preparing the CSV
Nominatim’s address parser is tolerant but structured. Get the CSV into this column layout before importing into MMQGIS:
| Column | Content | Notes |
|---|---|---|
| Name | Any identifier | Optional; preserved in output attributes |
| Street | Number + street name | “123 Main St” — not “Main St, 123” |
| City | City name | |
| State | State or province | Abbreviations work: “NY”, “ON”, “Bavaria” |
| Country | Country name or ISO code | “United States” or “US” |
The format sensitivity is higher than it appears. Abbreviations that differ by region (Ave vs Avenue vs Av), addresses where city is folded into the street field, and PO Box addresses with no geocodable street all reduce match rates. Clean the data first — a 15-minute review of the CSV saves hours of debugging failed geocodes.
Save the file as UTF-8. Excel’s default encoding on Windows is CP1252. Non-ASCII characters in city names (accented letters, umlauts, diacritics) cause silent geocoding failures if the file is not UTF-8. In Excel: File → Save As → CSV UTF-8 (Comma delimited).
Installing MMQGIS and geocoding
MMQGIS is a QGIS plugin maintained by Michael Minn. Install via Plugins → Manage and Install Plugins → search “MMQGIS” → Install Plugin. It adds a MMQGIS top-level menu.
Navigate to MMQGIS → Geocode → Geocode CSV with Web Service:
- CSV Input File: Your prepared CSV file
- Address / City / State / Country: Dropdown menus mapping each field to your CSV column names
- Web Service: Select
OpenStreetMap / Nominatim - Output Shapefile: Path for the geocoded point layer output
- Not Found Output List: Separate CSV for addresses that could not be matched
Click Apply. MMQGIS geocodes each row sequentially, pacing itself to respect Nominatim’s rate limits. A progress bar shows status. Do not close QGIS or the plugin dialog while it runs.
Output is a point shapefile containing all original CSV columns plus added Latitude and Longitude attributes. The Latitude and Longitude values are decimal degrees (WGS84). Load the output layer and spot-check a handful of known addresses before proceeding.
Nominatim rate limits — the constraint that catches everyone
Nominatim is the OSM Foundation’s public geocoding API. The usage policy is explicit: maximum 1 request per second for non-automated use. For bulk or recurring scripts, the limit is stricter.
What this means for batch geocoding:
| Address count | Minimum runtime |
|---|---|
| 1,000 | ~17 minutes |
| 3,600 | ~1 hour |
| 10,000 | ~2.75 hours |
| 50,000 | Not suitable for public Nominatim |
MMQGIS enforces the pacing automatically. You cannot speed it up by changing a setting. Plan for the runtime and do not interrupt a job partway through — the “not found” CSV only captures addresses that completed processing (as failures), so a partial run leaves you with no clean record of what failed versus what was not attempted.
For lists over 10,000 addresses, two practical options:
Host your own Nominatim instance. A Docker-based Nominatim instance handles 10+ requests per second with no rate limit. The initial database import takes 24–48 hours for a full planet extract or 1–2 hours for a single country. If you geocode regularly, this is the right infrastructure investment. The official Docker image at mediagis/nominatim is the standard starting point.
Use a commercial geocoding API. MMQGIS supports Google Maps and other web services via the web service dropdown. Google’s Maps Geocoding API handles 50 requests per second and provides a confidence score per result. The free tier covers 300 USD/month of usage (roughly 40,000 geocodes). For one-off large jobs, this is often cost-effective.
Handling unmatched addresses
Every geocoding run produces failures. Typical failure rates:
- Urban US or European lists, clean format: 3–8%
- Rural or mixed-region lists: 10–20%
- International lists with variable formatting: 15–30%
MMQGIS writes all failures to the “Not Found Output List” CSV. Open it and categorise the failures before deciding how to handle them:
Correctable problems: Typos in street names, missing house numbers that OSM requires for disambiguation, wrong zip codes used as city names. Fix these in your source CSV and re-run only the failed rows.
Structural non-geocodables: PO Box addresses, building names without street addresses, military APO/FPO designators. These cannot be geocoded to a point by any street-based geocoder. Either remove them from the analysis or treat them as a separate category.
Coverage gaps: Nominatim’s coverage mirrors OSM’s coverage. OSM is dense in Western Europe, North America, Japan, and Australia. Rural areas of South America, sub-Saharan Africa, and parts of Southeast Asia have sparse or outdated data. High failure rates in specific regions may reflect data quality in OSM rather than problems with your addresses.
The spatial distribution of your failures matters analytically. If unmatched addresses cluster geographically — all rural, all in one neighbourhood, all in one demographic group — their absence introduces bias into the heatmap. Note this in any analysis you deliver.
Heatmap renderer vs the Processing algorithm
QGIS has two things named “heatmap” and they are not the same tool.
Heatmap renderer (in layer symbology): Right-click the point layer → Properties → Symbology → change the dropdown from “Single Symbol” to “Heatmap”. This renders the layer as a kernel density visualization in the map canvas. It is a display mode. It updates in real time as you pan and zoom. You cannot export it as a raster file, you cannot generate a legend with density values, and you cannot use it in processing pipelines downstream. Use it for visual exploration.
Heatmap (Kernel Density Estimation) in the Processing Toolbox: Processing → Toolbox → Interpolation → Heatmap (Kernel Density Estimation). This is a geoprocessing algorithm that writes a GeoTIFF raster with a density value per pixel. It supports kernel selection, weighted points, dynamic radius per point, and produces a persistent output file. Use this for any actual analysis, cartographic output, or publication.
The heatmap renderer is for exploration. The Processing algorithm is for output.
Running the KDE
In the Processing Toolbox, search “Heatmap” or navigate to Interpolation → Heatmap (Kernel Density Estimation). Key parameters:
Input point layer: Your geocoded point shapefile. This layer must be in a projected CRS — see the CRS section below before running.
Radius: The search distance around each point, in the units of the layer’s CRS (meters if you reprojected to UTM). All points within this radius contribute to the density value at any output pixel. Too small and you get isolated spikes at each point. Too large and the entire extent blurs to a uniform wash. For a city-scale analysis, start at 1,000–5,000 m and adjust visually.
Pixel size: Output raster resolution. Smaller pixels = finer resolution, larger output file, slower rendering. For a city-scale map, 100 m per pixel is sufficient for publication. For a national-scale map, 1,000 m or more.
Kernel function: Controls how density decays with distance from each point.
| Kernel | Character |
|---|---|
| Quartic (biweight) | Default; smooth, balanced decay |
| Epanechnikov | Smooth, parabolic; slightly sharper centre than Quartic |
| Triangular | Linear decay; sharp peaks, distinct boundaries |
| Gaussian | Smooth, long tail; hotspots appear larger |
| Uniform | Constant density within radius; creates blocky output |
For most mailing list analyses — where you want to see clusters without knowing their exact extent — Quartic or Epanechnikov is the right choice. Gaussian makes diffuse clusters look more prominent, which can be misleading in sparse data.
Weight from field: Optional. If your data includes a count column — number of purchases, number of household members, a response score — use it here to weight each point’s contribution to the density calculation. A customer with ten orders contributes ten times the density of one with a single order.
Run the algorithm. The output is a temporary raster layer. Right-click and export to GeoTIFF to save it permanently.
For symbology: apply a pseudocolor ramp on the raster value range. Natural breaks (Jenks) classification works well for analysis; quantile classification makes low-density areas visible in sparse data where natural breaks would collapse them to a single class.
Coordinate systems — the silent failure
This is the most common failure mode in the pipeline, and it produces no error message.
If your point layer is in WGS84 (EPSG:4326) — latitude and longitude in decimal degrees — and you run the KDE with a radius of 1000, the algorithm interprets that as 1000 degrees. The output raster covers an area the size of a continent, or more commonly produces a blank raster because the 1000-degree radius exceeds the bounding box of your data. The algorithm completes successfully. No warning appears. You just get an empty or nonsensical output.
Fix: reproject the point layer to a metric projected CRS before running the KDE. Use Vector → Data Management Tools → Reproject Layer:
- UTM: Find your zone at
epsg.ioby clicking the map. A dataset centred on Berlin uses EPSG:32633 (UTM zone 33N). A dataset centred on Chicago uses EPSG:32616 (UTM zone 16N). Units are meters. - State Plane (USA): More accurate for small areas, breaks across zone boundaries.
- National grid systems: OSGB36 for UK (EPSG:27700), RD New for Netherlands (EPSG:28992), GDA2020 for Australia (EPSG:7844).
If you are not doing geodetic survey work, any UTM zone covering your data’s extent is accurate enough for density mapping.
After reprojection, re-run the KDE on the reprojected layer with the radius in meters. The output raster will be in the projected CRS. If you need to overlay it on a web basemap (EPSG:3857 or EPSG:4326), reproject the raster afterwards using Raster → Projections → Warp (Reproject).
If the CRS issue surfaces in other parts of your GIS pipeline — when validating inbound data from contractors or converting between formats before analysis — the shapefile validation checklist covers CRS (.prj) checks specifically, and GeoJSON as an interchange format covers why format conversions in pipelines involving multiple tools need explicit CRS handling at each step.
The complete pipeline
In order:
- Clean CSV to Nominatim-compatible columns, save UTF-8
- MMQGIS → Geocode CSV with Web Service → Nominatim → get point layer + failures CSV
- Review failures CSV; fix correctable ones and re-run failed rows
- Reproject the point layer to local projected CRS (UTM or national grid)
- Processing → Heatmap (KDE) → set radius in meters, kernel = Quartic, pixel size appropriate to scale
- Export raster to GeoTIFF
- Apply pseudocolor symbology (natural breaks or quantile); overlay on basemap
The pipeline works at the scale of a local nonprofit’s mailing list (a few hundred addresses) and at the scale of a national survey (tens of thousands of addresses, with the time caveat for the public Nominatim server). The rate limit is the only real constraint on volume. For data that needs to move between this point layer output and other spatial formats — GeoPackage for offline use, GeoParquet for analytics pipelines, Shapefile for legacy GIS tools — GeoConvert handles the format layer with geometry validation and CRS preservation intact.