首页 > 随笔 > State of the Map 2026: OpenStreetMap conference

State of the Map 2026: OpenStreetMap conference

Hacker News 2026-08-31 16:57 2 阅读 查看原文
On 29-30 August, I attended the State of the Map (SotM) conference, in particular the scientific part. It’s been 15 years since the last time that I attended the SotM conference (last time 2011!), and it’s an opportunity to fill in a knowledge gap that I developed over this period. Unlike other conference reports that I’ve written, I am not summarising sessions, but capturing my impressions and aspects that I note through the renewed engagement with OpenStreetMap (OSM). The scientific part of the conference was particularly interesting for me, because it expresses the type of researchers that selected to present their work back to the community. Although the academic track is peer-reviewed and operates more like a scientific conference, the aim of the conference as a whole is more towards the community of OSM than the usual academic conference. OSM is used extensively in research – in 2025, OpenAlex suggests over 1250 papers, so I don’t expect that attending the session will be completely a review of what is going on. But the 20 or so papers do provide a notion of what is researched by the researchers who are closer to the community. It is fortunate that I could attend SotM this year, considering that in June, I received the Test of Time award from the IEEE Pervasive Computing journal for the publication of the article on OpenStreetMap in 2008 (it is a top-cited paper in the journal), it is nice to get a sense of the papers that are citing it. In general, the academic/scientific track of the conference is doing well, with studies about OpenStreetMap and studies that use OSM data that filled the schedule for two days. My first takeaway from SotM is that it was nice to see many familiar faces – there is a core group of people in OpenStreetMap that have been around now for about 20 years. For some, it is part of their career and what they do. Other people are doing it as a hobby in addition to their work. Either way, it is valuable to see how engagement can continue over such a long time. Secondly, unlike citizen science, there is much more presence of commercial actors – as sponsors of the conference, as presenters, and there was even an area for professional geospatial people who use OSM in France. This does provide resources, places of work for people who are in between enthusiasts and professionals (or professionalising their enthusiasm), and an engagement with the changing needs of the data. Turning to the scientific track, from the start of scientific use of OSM, there were several characteristics that make it particularly attractive. It is an accessible, open, and hackable (in the sense that it is mutable and easy to understand) dataset. This makes OSM a site for experimentation in developing solutions to challenges such as routing, map generalisation, cartography, etc. However, it’s more messy data that needs to be examined, cleaned, and organised in order to use it for a specific investigation. And while the geometry might be complete, the attribute information continues to be hidden and variable. This messiness creates challenges for topology and routing – which makes it a persistent issue in the nature of the data produced. It is valuable to note how routing remains an area of experimentation and challenges. But there are plenty of things to map: for example, attribute completeness for dams in rural spaces is very low – below 1%. There is also interest in indoor mapping and completing details of public buildings. Of course, the world continues to change, and you need to understand where and how changes are happening. The emergence of multiple open geospatial data sources is making it possible to keep the OSM approach to the use of the data. Satellite imagery continues to play an important part as a source of information. Another area for improvement in mapping can be the indication of building entrances instead of centroids – which is very relevant for navigation applications. An example of the ongoing research on data quality is the exploration of completeness – using extrinsic data comparison, intrinsic attribute analysis, or statistical estimation. The methodology that was developed combines intrinsic attributes and statistical methods to evaluate completeness – assuming that there are features that will be captured first (say roads) and things that will be saturated at the end (say addresses) it is possible to check over time to see how features are being added until the map stabilises. The analysis of completeness in this way is relevant for a specific class of feature (or attribute) – so the question can be: is this a complete set of buildings? etc. In terms of the application areas that OSM data is being used on, there were examples from public health studies. OSM is considered relevant enough to explore if it can provide information on rural spaces (which wasn’t the case in the past). A similar example is applications in monitoring mining activities across the world. There are also new problems that need addressing – such as mapping the electricity grid (my very first large project in GIS was on digitising the mapping of the Israel Electric Company in 1991). Interestingly, the quality assurance of OSM is seen as valuable – with tools such as osmose. For the grid, consistency, completeness, and up-to-dateness are core parameters (in mapmygrid.org/quality). What is also interesting is that because of the good level of completeness – especially in large urban areas- there are increasing large scale studies that use OSM as a basis for analysis. This can be included in the issue of data quality – a persistence issue. The role of OSM as a humanitarian source of mapping in places where information is missing continues. On the practical side, it is an effective and efficient way to produce maps, and there are even evaluations of the low costs that such mapping involves. I was somewhat surprised to see that most of the examples that were shown didn’t use other open data projects and merged the data for evaluation and analysis. One of the only examples was the use of the Colouring Cities project that is running from the Touring Institute. Since my early days in geospatial research, I am baffled, and continue to be, about different analyses that are in the form “we try to solve problem X only with data from source Y”. For example, research on road lanes, or the characteristics of an urban park that only uses OSM without using other sources. I think that one of the major reasons that it continues to be the case, almost 30 years later, is the learning costs of getting familiar with a data source and knowing how to use it. PhDs, postdoc fellowships, or research projects are always limited in time, and it probably feels like the effort of learning all the ways in which you should use a dataset is time-consuming and complex. There is a lot of trial and error, so you stick to one source. Yet, maybe the thing that people should do is to reduce the geographical scope of their question while trying to explore multiple sources of information. There is so much open data of high quality out there – from Wikipedia, OpenStreetMap, Satellite data, Citizen science data, etc. Creative approaches to merging and using different data sources might be a more effective way to answer the question… There is also a space for a social and theoretical critique, such as noticing the limitations of digital humanitarian efforts, such as tendency towards solutions. For example, the way that remote mapping might override local contexts and the codifying of local knowledge through the standards that OSM offers. One topic that was discussed is the voting on tagging proposals. Also, consideration of inclusion and diversity through statistical analysis was covered (by Carlos Cámara). The critical literature from Crampton, Harley, Wood and others appears – even Ground Truth (Pickles 1995) made an appearance. Arguing that maps are not by/for elites. There is still inherent bias in participation, and therefore in representation. The analysis that looked at tagging proposals from 2006. They defined certain proposals as feminised or masculinised – based on gender performativity. The need for diversity in the OSMF board came up in the dedicated session – there is a need for increased wider participation. Only 982 users created a proposal, and most didn’t engage. Participation inequality appears in them. There are fewer feminised proposals that are receiving less attention. This research, however, assumes that official proposals for tagging matter – which is not. It’s a case, for me, of research that uses OSM without proper understanding of how the culture of do-ocracy is impacting the outputs. In this context, and maybe because of the tagging aspects, I was surprised that two people mentioned to me the community issues with Monica Stephens’ paper from 2013 “Gender and the GeoWeb: divisions in the production of user-generated cartographic information” (https://link.springer.com/article/10.1007/s10708-013-9492-z) about the misinterpretation of the governance power of tagging proposals and practices, although, as I pointed to the two, while the example might be wrong, the general problem was very real and well documented. It even came up in the discussion by OSM Foundation board… An interesting talk covered the governance implications of the task manager of the Humanitarian OSM Team (HOT), which covered concepts such as how geodatafication can be considered as a side impact of the humanitarian efforts. As expected, AI in its different forms: machine learning, image recognition and analysis, and LLMs appears in the work on OpenStreetMap – such as the automatic identification of road changes that should be updated in OSM. It is impacting the infrastructure with false accounts, requests for data, or badly written software applications that abuse the infrastructure. Hannah Boetcher also looked at AI ethics and OSM – positioning it within digital commons and digital capitalism, and the concept of tragedy of the commons. Looked at how AI systems extract data from OSM – creating infrastructure strain and licensing issues. Corporate mapping in OSM was seen as a threat, and the level of it is declining. OSM data is used for training AI – with automated queries by bots, large-scale, continuous crawling. There is no reciprocity or engagement (unlike the corporate edits). The scraping does not respect attribution. The fact that AI doesn’t contribute makes it justified to do a defensive closure. The burden of AI is known by the board, and the abuse by AI is clearly felt by them. With regard to corporate editing of OSM, it was recorded in 2019 (Anderson et al.), and there are normative questions and also impact questions: data quality, editing patterns, influencing disengagement of volunteers. Recent research suggests a reduction in corporate editing and engagement. Using collective intelligence framework by Grinberger at al, was looking at four basic conditions: independence, diversity, decentralisation, and aggregation mechanism (Surowiecki 2004) – used a measure for the reach of these elements by looking at tags, entities, and geometric complexity. Countries like Thailand and Malaysia are places where Facebook carried out AI-assisted road tracing, and in addition, the Philippines, Papua New Guinea, and Myanmar are used as examples of places with different degrees of active local communities. Corporate editing can be from Facebook, Grab, and other companies. There are two waves: Facebook between 2017 and 2020, while Grab 2023-2024 are editing and improving. The impacts of corporate mapping are complex; there is no clear direct impact of an intervention. The large-scale effort of Facebook naturally made a bigger impact. In summary: the context matters, and how corporates exist within the OSM community and need to be integrated. There is resilience by the community of mappers, so corporate mapping is not overpowering local mapping. Engagement with commercial organisations also came up in a discussion with the board. The comments from the OSM Foundation also reflected the increase in geopolitical tensions, with an increase in complaints about boundaries of countries. They are something that the board is getting regular complaints about, and these are being dealt with all the time, and it is a risk of being sued by a state entity about the position of the border. OSMF is using areas of control and not trying to record boundaries. There was also an indication of the ongoing impact of Brexit is the move of the OSM Foundation to Belgium, in order to be in a place that is within the EU. Explanation on why to do that is the protection of the IP of the database, due to the lack of protection in the UK (see https://osmfoundation.org/wiki/Board/Minutes/2026-03). Overall, it’s interesting to see the evolution and response to the changing internet and commercial environment over a period of 20 years, and how the increase in data coverage and quantity is opening up applications. For a project that is managed in a chaotic do-ochracy way, with all the problems and challenges that this creates, it is very impressive. It is slowly (very very slowly) maturing into an organisation with some staff (well 1.something now), and a turnover of the foundation of about €1m. In comparison, the European Citizen Science Association (ECSA), which is doing far less than OSM and started in 2013, is already on €1.5m with about 20 members of staff. It might be that the AI age will force such changes. There are also persistence problems, with a lot of overlap with citizen science: data quality and how you assess the quality; how you encourage and guide volunteers to share data of high quality and fitness for purpose; how you address inclusiveness, diversity and why you are supposed to do that in the first place; how to engage and work with commercial actors when you are a volunteer and human-focused project; how to deal with socio-technical assemblage that is the project (and both parts are critical); where and how the governance of the project plays out? There are, naturally, domain questions: routing and route planning is a challenge that provides a rich space for exploration, which is also visualisation, cartography, representation, and automatic generalisation. The behaviour of technology companies in the AI area was one of the lasting impressions, with the awareness on how their irresponsible behaviour towards shared open data, with lack of care towards who pays for the data traffic that they are creating and the servers that provide the data is something that is likely to harm open knowledge projects. As much as I don’t believe in the tragedy of the commons (https://news.cnrs.fr/opinions/debunking-the-tragedy-of-the-commons) and the problem with Hardin’s thinking and beliefs, there are clearly bad actors that need to be punished and stopped from abusing shared resources. While copyright does receive attention, I think that it is worth paying attention to this aspect too. Share this: Share on X (Opens in new window) X Share on Facebook (Opens in new window) Facebook Share on LinkedIn (Opens in new window) LinkedIn Share on Reddit (Opens in new window) Reddit Share on Pinterest (Opens in new window) Pinterest Print (Opens in new window) Print Email a link to a friend (Opens in new window) Email Share on Telegram (Opens in new window) Telegram Share on Tumblr (Opens in new window) Tumblr Share on WhatsApp (Opens in new window) WhatsApp Related Tagged ai artificial-intelligence chatgpt citizen science Commons Critical Cartography critical GIScience Education Environmental information humanitarian OpenStreetMap OSM SOTM sotm26 technology VGI Volunteered Geographic Information