How to get started in geospatial data science

Most people meet geospatial the hard way, jumping right into the middle. A client request, a technical tutorial, a new Python library, an overly complicated and technical error message. And even if the code runs and the map looks plausible, something may be quietly three metres off — and then the question comes: how to make it right?

I have seen this from both sides. Plenty of people arrive from QGIS or ArcGIS, where every analysis is a sequence of clicks that works beautifully once and is painful to repeat, share, or audit. Others, including me, arrive from data science, with pandas and scikit-learn already in hand, and no idea that the problems they have been solving — clustering customers, forecasting demand, comparing regions — were spatial all along. Both groups want the same thing: analyses that run again tomorrow, on a different city, with the same result, and that a colleague can read and check.

That is what this journey is for. It takes you from an empty computer to what I call geospatial-ready Python — the point where you can load, reproject, join, and map spatial data in code, and know which of the numbers to trust. Along the way it also settles a question that has become unavoidable: how to work with an AI assistant without working blind.

Step 1–2: see the field, then see the destination

Before touching a tool, it helps to know where the tools live. Step one is the article The Map of Geospatial Data Science, which draws the field as a map: where GIS, geospatial analytics, and data science overlap, and why so many data people walk past the middle of it without noticing. Reading it first means that when you later hit GeoPandas or a CRS error, you know which part of the map you are standing on.

Step two is a thirteen-minute video in which I build a full 3D Manhattan — 46,000 buildings at their real heights — in one live, unedited session with an AI assistant - actually, my first proper interaction with Claude Fable 5, a flagsip model in 2026. It sits this early on purpose. It shows what the end of the road looks like, and it shows the moments where the assistant needed a driver who knew the territory. Keep that in mind; the whole journey is about becoming that driver.

Step 3–5: Python, and the habit of checking

The ground floor is my book Data Science Essentials: 101 Practical Steps in Python. It starts with installing Python and ends with a clustered world map, and everything in between is organised around one idea: almost everything that goes wrong in real data work goes wrong quietly. A column of dollar amounts that is secretly text. An average dragged down by three countries recorded as having no people. A bar chart whose baseline was moved. A model whose impressive score was measured on rows it had already seen. None of these raise an error; all of them change a decision. So the book teaches pandas, NumPy, and matplotlib by teaching you to catch them — every chart drawn twice, the honest version beside the misleading one — and it ends with a one-line rule that beats a trained classifier, so you also learn when not to model.

Python Foundations for Geospatial Data Science covers the same ground in video, with three additions a book cannot deliver as well. It's a pre-recorded online course, which makes it a lot more interactive and focused than a book. Then, it highlights the habit of workflow craft — the config block, the staged save, and the habit of hunting down errors. And third, it extends the book by having a whole chapter on coding with an AI assistant: what a good prompt contains, what the failure modes look like, and how to check the machine rather than believe it.

Step five puts that chapter's thoughts to work. In I Mined My Entire Claude History I export 1,281 of my own conversations and build a pipeline to find out what I actually keep doing. The rule that came out of it — Python counts, the model judges — is the one sentence I would highlight on every AI-assisted project. AI is going to be part of your workflow whether you plan for it or not; this is where you learn to use it to your advantage without handing over the judgment.

Step 6–7: spatial thinking, then core spatial data toolkit

Only now does the route turn spatial, and it does so without code first. How to Think Spatially is about an hour of slides on the concepts every line of geospatial code is expressing: what makes data spatial, why nearby things break the independence assumption behind most statistics, how a globe gets flattened onto a screen and what that costs, why you cannot measure distance in latitude and longitude, and what "near" means when a radius, a walking time, and a nearest neighbor disagree. Every concept points to the exact step in the next book where you build it.

That book is Geospatial Data Science Essentials: 101 Practical Python Tips and Tricks, and it is where vector data starts for real. Points, lines, and polygons in Shapely; GeoDataFrames, spatial joins, overlays, and dissolves in GeoPandas; projections and EPSG codes; spatial indexing with RTree and H3; geocoding; a first pass at rasters; OpenStreetMap through OSMnx; spatial networks; and an introduction to raster data in Python, and spatial statistics and machine learning — 101 techniques, on the current stack. Every later course on this site names it as its companion.

What you will have, and what comes next

By the end of this journey, you can write clean, checkable Python for real data; explain coordinate systems, vector and raster, and Tobler's law before you touch a library; load, reproject, join, and map spatial data; and read a map or a model output critically. You will also have a working relationship with an AI assistant in which you are the one doing the checking — and a Claude skill, the Spatial Data Interrogator, that catches the few-meters-off problem before review does.

What this journey does not cover is any specific city, model, or the trending topic of GeoAI. That is what the Urban Data Science and GeoAI journeys are for, and both assume you have done this one. If you already work with data, you are closer to this than you think.

Extra steps

Once the core route is done, a few things are worth picking up on the side. Connecting the Dots is the reading for the ideas underneath everything here — networks, data, algorithms, how the systems around you decide what you see next — with no code at all. Two short tutorials extend the projections chapter of How to Think Spatially: The World Map with Many Faces shows, in Python, exactly where every flat world map lies, and Map Projections in Cartopy covers the one projection tool the core stack leaves out. Geospatial File Formats is the reference to keep open when Shapefile, GeoPackage, GeoJSON, and GeoParquet start showing up in the same project. And The map is wrong and nothing threw an error — a video guide plus a Claude skill, the Spatial Data Interrogator — is the "inspect before you trust" habit applied to spatial data specifically: how analyses fail quietly, and the tool that catches it before review does.

Next
Next

The New Science of Maps