Spatial Data Science across Languages - Jena edition

Notes on attending the 2026 Spatial Data Science across Languages (SDSL) workshop in Jena, Germany 🇩🇪 from 16-18 Sep 2026.

Photo of the workshop venue at the Max Planck Institute for Biogeochemistry.

Main working space for the 2026 Spatial Data Science across Languages workshop

Overall impressions

This is the fourth installment of the SDSL workshop series and my first time attending. Less than 20 people in-person and a handful joining online, predominantly European, mostly people from universities or research institutes, a sprinkling of freelancers and research software engineer types, practically all nerds working at the heart of innovative and core open source libraries in geospatial. It feel like a fairly tight-knit group, a very knights of the round table feel to it if I must say.

Main working space for the 2026 SDSL workshop

What we lacked in gender diversity was made up by multilingualness I guess, in both programming and spoken languages. Topics varied quite a bit, from LLM/Agentic coding (unfortunately) to knowledge graphs, COGs and large (1PB scale) data processing, GeoZarr and DGGS, spatial statistical learning and correctness, community building, etc. There were representative from high-level programming languages (Julia, R, Python), some who know Rust and maybe some C/C++, all sharing ideas across their respective ecosystems and/or figuring out how to reduce duplication.

Overall, I personally found it amazing that the organizers managed to bundle discussions covering a good breadth of spatial data science, rather than focusing on a single narrow topic you usually get for workshops with a couple dozen attendees. It wasn't just about one tabular/n-dimensional file format, nor was it about a single category of spatial statistical methods, but really about stepping back and seeing the bigger picture of how to do (hard) things accurately, questioning assumptions, and what needs to be done to evolve the spatial community. Granted, this type of gathering might not be suited for those preferring to stay in their bubbles or uncomfortable with debating the inevitable nuances of why there are X ways to do one thing (each method might be correct under specific conditions).

But I appreciate that there are still gatherings like this that are tackling complexity head-on in a grassroots manner.

We're all grappling with the changing nature of open source, and it's especially important to meet in person and build connections, because code might be cheap now, trust is not.


Before I go on, I do want to express my thanks to DevSeed, both for funding my out-of-proportion travel costs, and also for being such a respectable part of the open source geospatial community. I can't overstate just how much goodwill DevSeed has in this SDSL community, and I felt privilleged to be able to represent there to keep on top of cloud-native geospatial and spatial data science happenings!

Now on with the blog post.

Are robots smart enough yet?

LLM/Agentic coding by Maarten Pronk

The schedule had some last minute changes because someone had a flight cancellation, so instead of a keynote, we started with the LLM/Agentic session, and I thought it worked out fantastically. There was some discussion around AI contribution policies mostly, the GDAL one was brought up, geopandas just merged their policy the day before, and many people chimed in with their opinions or experiences. Everyone praised or vented, and we agreed as a group to not bring it up (too much) for the rest of the week.

Knowledge graphs by Anita Graser and Edzer Pebesma

Anita demo-ed a geo-knowledge graph (loosely inspired? by KnowWhere Graph) she built for some municipalities in Austria, totalling to about 800k triples in a Knowledge Graph, and using that graph as the basis for a chatbot to answer questions in a verifiable way. Not very scalable, but it definitely felt more grounded in facts than Retrieval Augmented Generation (RAG).

While powerful, building the graph seemed to require having access to rather structured data, and if you have structured data already ... using a query language seemed sufficient? There were some comparison of using different agents to perform SPARQL queries or something, I wish I could share more details but the code is still private. I got lost in some details towards the end, but caught some fuzzy mention that [Knowledge graphs] are a [perpetual] solution looking for a [killer] use-case, so... watch this space?

From file formats to equal-area projections and back again

Cloud-optimized GeoTIFF by me / GeoZarr by Felix Cremer

I was meant to present my Cloud-optimized GeoTIFF (COG) talk on Thursday, but they shifted it last minute to Wednesday morning. Luckily I had a reasonable version prepared!

2026 Ecosystem of GeoTiff readers

This was one of my attempts to map out the (Geo)TIFF reader library ecosystem for 2026. On the top are the low-level C++ (mostly GDAL, or nvTIFF) and Rust (a mess of image-tiff, async-tiff, bindings of C++ stuff...) implementations. The surrounding stuff along the bottom and right show language-specific libraries that either rely on the C++/Rust implementations (mostly the case for R and Python ecosystems), or independent implementations (as in the Julia and Javascript/Typescript). I'll also call out the xarray-related libraries in the Python box that includes no less than 4 backend/engines!

The GeoZarr talk Felix gave was mainly just browsing the https://geozarr.org/ website. It's probably on the way to becoming just as rich as the GeoTIFF one (jk). Though zarr-python recently announced in release v3.4.0 that they'll likely switch from a Pure Python to a Rust-backed backend (via zarrs and zarrista), so maybe things won't get as messy, at least on the I/O level.

That said, the data structures built on top of the seemingly infinitely flexible Zarr format can make anyone dizzy, and the conversation rightfully steered towards...

Discrete Global Grid Systems by Benoit Bovy & Anshul Singhvi

For anyone subscribed to mapping related news, you might have heard about the UN resolution to use an Equal-area projection system. Well, if you want to nerd out more, DGGS allows even more flexibility to better approximate the surface of a spherical or ellipsoidal Earth. Much like there isn't a single perfect coordinate system, there isn't a perfect DGGS either that satisfies multiple properties (hence, endless debates).

From a library maintainer's perspective, it's tricky to design a proper API around different DGGS types that use different shapes, indexing schemes, parent-child relationships, etc. If you thought lonlat vs latlon was bad, wait till you think about triangular, quadrilateral, pentagonal, hexagonal, or even irregular shapes... Most of the discussion that week seems to be around how to do efficient/accurate conservative regridding into a DGGS, optimal ways of storing a DGGS in Zarr, and so on.

Scaling data processing and communities

Large data processing - Yu-Feng Ho and Sajed Sarabandi at OpenGeoHub

It's easy to dismiss 'big data' nowadays, but having spoken to these guys yesterday, and heard them say 1PB+ volumes, boy I knew I should listen in. They've made a copy of the entire Landsat archive in Europe, rightfully by copying it on to physical hard drives in the US (taking ~1 week), shipping it over the Atlantic, and then plugging it into their server infrastructure. Because... it would have taken them 6 months to do that over the network, whereas this one took them 3 months (including organizing the logistics).

Fun fact: A decade ago, I used to do that for my first job for ~100GB volumes, running maybe a few hundred metres, so can totally relate!

Their flagship data product is this Landsat bi-monthly cloud-free mosaic, apparently 100TB worth of Cloud-optimized GeoTIFFs. They went into some details about how the data processing was optimized, by overlapping I/O and compute (nice), and writing many small files to local storage first before moving to a slower but bigger disk (cool). Considering that this is a first-year (?) PhD student and a software research engineer doing this, that's pretty impressive already.

Global Landsat Composite processing pipeline

There was some debate on why not do this on the commercial cloud, and counter arguments on data sovereignty and a good point on keeping it a non-US controlled backup of Landsat (knowing how some NOAA datasets were taken down before). I went into a further heated take (as a non-European) that what OpenGeoHub is doing with on-prem (instead of cloud) totally makes sense, because they are incentivized to make it efficient to access for researchers and can innovate on better network I/O hardware topologies, compared to commercial cloud providers who are not incentivized to make it cheap (see: egress) or efficient at all. Just the day before, I was chatting with Aaron how I've given up on big cloud for GPU training, and am really in favour of on-prem HPCs, because again, incentive structures.

Alas, I could spend many more paragraphs talking about how to configure GPUs to speed things up, but let's talk about another form of scaling - open source communities.

User Community, Teaching and Outreach by Claudiu Forgaci

This was the last talk, and a case study of various communities (ROpenSci, Carpentries) and how they operate. Claudiu started with a point about (maintainer) bus factors, and I guess ended by advocating for sustainable communities by growing the user-/contributor-base? It feels very trickle-down economics-like, getting more contributors who might eventually step up to become maintainers, and personally I'm not sure if it always works (knowing some projects with thousands of users but more or less one key maintainer).

Yet it did get me thinking, if better API/tutorial documentation can optimize user experience, and good CONTIBUTING.md files optimizes contributor experience, could there be a way to optimize (onboarding) maintainer experience somehow? It probably needs to be more than just release checklist templates and automated CI/CD though I feel, I've also tried developer-focused workshops before, and have even seen the hire-a-postdoc/RSE-to-maintain-this method many times which works well for a year or two, but doesn't work long-term when the postdoc/RSE moves on. There was a mention of the Netherlands (eScience?) evolving from mere data management planning to considering software lifecycles too (e.g. this Software Management Plan), which I guess is a step towards taking software more seriously. But maybe I'm just cynical now, having seen so many sides to software maintenance...

I reckon there needs to be more fun. What starts out feeling fun (when it's new) can slowly feel like a chore. How we keep things fun for longer, I don't know, I wish I knew.

Correctness and knowing what you're doing

Spatial statistical learning - by Levi Wolf (inaugural Lorena Abad award winner)

Read the nomination excerpt here. Levi is a co-host of the Geography, Life + Data (GLaD) podcast, co-author of the Geographic Data Science with Python book, and is writing a Causal Inference in Spatial Analysis book too. All sorts of content that I need to check out at some point! He drew a nice schematic of the Spatial ML landscape across languages along the axis of model training and model critique tools.

Schematic of the SpatialML landscape

I'm almost ashamed to admit that I haven't heard of, let alone used any of these libraries. At most, I'm aware of model critique techniques like SHAP, and some geostatistical methods from my GIS courses over a decade ago, but haven't applied them much in practice. But these are the very tools (especially those on the right side) people ought to reach for when your fancy machine learning model isn't performing well, and you want to find out whether it's something missing/lacking in your inputs or a bad model, or something.

And this ecosystem of libraries isn't as mature as the whole neural network landscape (based on trial and error, see: xkcd comic). Levi rightly identified that while there are some 'fast common frameworks for estimation' on the left, mostly C libraries (with differing levels of bus factor risk), there are practically none on the right that are 'fast common frameworks for (model) criticism'. Opportunity here? Or burden for someone to step in?

Before anyone gets the wrong idea, I want to point out that these spatialML tools need to be implemented correctly, and ideally by someone deep in the spatial statistics domain. It's also one thing to have a correct library implementation, but having someone blindly using these without caring about the assumptions necessary for the statistics to be valid is bad. So understand before using.

Correctness on Sphere by Maarten Pronk

This session mostly goes over the gaps identified last year in this "Spatial Data Science Languages: Commonalities and Needs" report. It covers the nuances of AREA_OR_POINT in GeoTIFF/raster cells, extensive (sum preservation) / intensive (average preservation) when aggregrating/disaggregating data, measurements on spherical/geodesic geometries, and maybe a few other nitty gritty details people aren't aware of.

On the raster AREA_OR_POINT issue, or gridline/pixel registration as we call in GMT/PyGMT, it comes down to having the data structure storing that piece of metadata, and ensuring operations (like resampling) are aware of that distinction to avoid half-pixel offsets. For the vector extensive/intensive case, I guess a similar metadata attribute could work, and operations use that to calculate stats correctly?

The spherical/geodesic-aware measurements for things like distance, area, perimeters seems to be going strong with libraries built-on top of s2geometry like s2geography, spherely and so on. It's just a matter of exposing that sanely to higher-level libraries like geopandas, duckdb-spatial, etc. Going from Cartesian to Geodesic to Vincenty distances for example.

All these, easier said that done.

Hackday, hallway track, and other tidbits

The final day (Friday) was the hackday, where less than half of the group remained, I think all Julia folks (except me 😹)? There were some hacking on DGGS/Zarr stuff, someone getting some GPU-accelerated code working on Apple Sillicon/NVIDIA in a hardware agnostic manner (via webgpu I think). I was just doing some TIFF stuff in Rust, including a TIFF TUI hobby project (navifd) I've been playing with to get used to Ratatui.

The interesting chats were of course in between sessions, during walks along the river (including one where a few of us nearly got hit by a tram, hence a joke about 'tram factor'). I had some fun chats with Aaron about icechunk and GPU stuff (mentioned above), got to know about the behind the scenes of OpenGeoHub a bit more, had coffee with an MPI guy that went into agents and security.

There were also lots of interesting chats during the dinners that were all Italian pasta/pizza based for some reason. Best one was on Friday night when I sat with an R, Python and Julia guy, and we complained about... packaging 📦

I think this is a good place to end. See you (maybe) next year at Delft!