Our First Natural Language Map Query: What It Felt Like When It Actually Worked
We typed a question about coffee shop density near public transport in Melbourne, hit enter, and a map appeared. No code. No SQL. No GIS tool. Just the answer.
The query was not remarkable. We typed it as a test case, not as something we expected to produce anything interesting: "Show coffee shop density near public transport stops in Melbourne CBD." Standard proximity aggregation, publicly available data, straightforward geographic scope. We had run hundreds of test queries by that point, most of which had produced partial results, error messages, or outputs that were technically correct but geographically wrong in ways that were embarrassing to explain.
This one worked. A map appeared. The coffee shop density was visualised correctly against the CBD transit network, the clustering around the Flinders Street interchange was accurate, the falloff toward the fringes of the defined area matched what anyone who commutes through that part of Melbourne would recognise intuitively.
Mei Lin was sitting across the desk. We both looked at the output for a moment without saying anything. Then we tried a follow-up query, and that worked too.
What actually went right
The thing that had been failing in previous queries was not the spatial execution layer. The geometry operations, the buffer calculations, the point-in-polygon counts, those were working reliably by the time we reached this test. What had been failing was the parsing layer: the component that converts natural language geographic intent into a structured spatial query that the execution layer can process.
Geographic language is ambiguous in ways that are hard to resolve without context. "Near public transport" could mean within 200 metres walking distance, within 500 metres, or within a named catchment zone depending on the context. "Melbourne CBD" has at least three different possible geographic definitions in standard Australian boundary datasets: the City of Melbourne LGA, the Statistical Area 1 boundaries within the CBD, or the colloquially understood street boundary that most Melbourne residents would describe. "Coffee shop density" requires a decision about what density means: point count, point count per unit area, or kernel density estimation.
The version of the parsing layer that worked had learned to make defensible defaults for each of those ambiguities, communicate what default it had chosen, and surface the output in a way that made it easy for the user to recognise if the default was wrong and correct it. "Near public transport" defaulted to 400 metres walking radius. "Melbourne CBD" resolved to the SA1 boundaries matching the colloquial city centre definition. "Density" was rendered as kernel density estimation with a 250-metre bandwidth.
All three defaults were right for the query as we intended it. The map that appeared reflected what we had meant.
Why the correct default matters
Getting the default right is not just about producing a correct first output. It is about building appropriate trust in the system.
If a natural language query produces an output that looks wrong to the person who asked it, the first instinct is not to check whether the default parameters were appropriate. The first instinct is to doubt the whole system. And that doubt is rational: a system that gets defaults wrong often produces outputs that are technically valid but geographically misleading, which is worse than an error because it is not immediately recognisable as wrong.
We had seen that failure mode multiple times. An early version of the system defaulted "near a transit stop" to 1-kilometre radius, which is defensible as a pedestrian access measure but much larger than most users expect. The density outputs looked plausible but were systematically overestimating clustering. Users accepted the outputs as correct and we caught the error only when one of our pilot users noticed that a suburb they knew well was appearing high-density when it clearly was not.
That incident drove a significant amount of the calibration work on geographic defaults. The goal is not to produce technically correct outputs. It is to produce outputs that match the user's geographic intent, which requires knowing what the most likely intent is for a given phrasing.
The follow-up query that confirmed it
The first query worked. The follow-up was what convinced us the parsing layer was genuinely functioning.
We typed: "Now show only the areas within 300 metres of a heavy rail station." This required the system to hold the previous query context (Melbourne CBD, coffee shop density), filter the transit network to the heavy rail subset (Flinders Street, Melbourne Central, Flagstaff, Spencer Street), and apply a new, tighter proximity threshold to the existing output.
The result was a refined map showing the concentrated density zone around the four CBD rail stations. The system had correctly interpreted "now show only" as a filter on the previous output, "heavy rail" as a subset of the transit network, and "300 metres" as a replacement for the previous 400-metre default.
That is three separate pieces of natural language context resolution in a single follow-up query. In our earlier versions, any one of those three would have produced an error or a misinterpretation. In this version, all three resolved correctly.
What it felt like to the people building it
The honest answer is relief more than excitement. We had been working on the parsing layer for about four months at that point. We had a clear picture of what the system needed to do, and we also had a clear picture of how many ways it could go wrong. Watching it go right felt like the confirmation that the approach was sound, not just in theory but in execution.
There was also something harder to describe: the recognition that the output was genuinely useful. It was not a demo output engineered to look impressive. It was a real geographic answer to a real geographic question, produced in under 30 seconds from a conversational query, with no GIS expertise on the user's side. The product we had described to ourselves was, in that moment, working.
Mei Lin pulled up the same query in QGIS to cross-check the result. The manual process took her about eight minutes, and she is fast. The outputs were equivalent. That comparison was not part of our test protocol; it was a spontaneous sanity check. When it confirmed that the natural language output matched the expert-produced output, we called it a working system.
What came next
That session in October 2025 was not the end of the calibration work. It was the point at which we shifted from "is the approach viable?" to "how do we make it reliable across the full range of query types?" The months between October and our early-access launch in early 2026 were primarily about expanding the query coverage, adding the Australian data layers that were most requested by the planning and logistics teams we were working with, and fixing the cases where defaults that worked for Melbourne did not work correctly for other Australian cities.
The first query working did not mean every subsequent query would work. But it was evidence that the core mechanism was sound: a natural language interface that resolves geographic intent correctly enough that a non-specialist can use it and trust the output. That is what we set out to build, and that October session was the first time we knew it was actually there.
Try MapAI
Ask your own location question
Free plan includes 50 queries per month. No credit card, no GIS background needed.
Start Free