GeoData Scout: Dataset Research & Readiness for Geospatial AI Base Models
Project Kickoff: October 1, 2026

Identifying and validating the open Very High Resolution (VHR) datasets that will train reusable AI base models for humanitarian mapping — so the building, solar panel, and tree-detection models published on HOT’s fAIr platform are built on data that actually transfers to the field, not just the datasets that were easiest to find. In this 12-week challenge, you’ll join a collaborative team of GIS specialists, data engineers, and geospatial researchers from around the world.
The problem
fAIr is the Humanitarian OpenStreetMap Team’s open mapping AI platform. A base model published there isn’t a finished map layer — it’s a blueprint that mapping teams discover, fine-tune on their own OpenAerialMap (OAM) imagery, and run to produce features for humanitarian mapping. That means a base model has to work well outside the exact region it was trained on, and the data it learns from is what decides whether that’s even possible.
Geospatial data doesn’t generalise easily. Building styles, vegetation, terrain, imagery sensors, resolution, season, and even local mapping conventions all shift from one region to the next, so a model trained on a narrow or unrepresentative dataset won’t transfer — no matter how well it scores on its own validation set.
Impact of the Problem:
The first dataset that turns up in a search is rarely the right one to build on. Some widely used benchmarks carry restrictive or expired licences; others cover only a handful of well-mapped regions, or are annotated at a level that doesn’t match the task — pixel masks where the model needs instance boundaries, bounding boxes where it needs polygons.
Discovering a licensing problem or an annotation mismatch after a model has already been trained and packaged is expensive: weeks of engineering effort have to be redone, or the model ships with a quiet blind spot in exactly the regions — often the least-mapped, most humanitarian-relevant ones — it was meant to serve.
This challenge exists to catch that problem before it’s expensive. For Buildings, Solar Panels, and Trees, it maps what open VHR data actually exists, checks what can legally and practically be used, and hands HOT’s engineering team exactly what to build on and how to prepare it — so model development starts from evidence, not from whatever downloads first.
The goals
The objective is to identify, assess, and prioritise open VHR datasets for the three committed feature types, and deliver a prioritised recommendation, a verified licence, and concrete data-preparation guidance for each one: output HOT’s engineering team can act on directly, not research they have to reinterpret.
Project Goals:
- Confirm criteria and review HOT’s starting datasets: Agree the suitability criteria and feature-type priorities with HOT, then assess the datasets HOT has already identified — the HOT VHR Building Segmentation Dataset and SpaceNet for Buildings, the Zenodo and DeepSolar sources for Solar Panels, and the Million Trees, OAM Tree Canopy, and WNA Hub datasets for Trees — as a starting point, not a final answer.
- Discover and assess additional open datasets: Search for open VHR alternatives that improve geographic diversity, imagery-source diversity, and representation of underrepresented regions, and score every candidate against resolution, annotation quality, size, and task fit.
- Verify licensing and deliver the first recommendation: Confirm every recommended dataset carries a licence compatible with AGPL-3.0-only, MIT, Apache-2.0, or BSD-3-Clause, and deliver a prioritised, licence-verified recommendation and data-preparation guidance per feature type.
- Validate datasets in practice: Check that the recommended datasets hold up once engineers start working with them — structure, label quality, real usability — and flag where coverage or annotation falls short.
- Fill gaps and test alternatives: Investigate complementary or alternative datasets wherever the first choice doesn’t fully cover a region or use case, and update the recommendations as new evidence comes in.
- Review geographic representation and hand off documentation: Establish which regions each recommended dataset does and doesn’t cover, close the gaps that open alternatives allow, and compile the licensing, provenance, and limitation evidence each model’s documentation needs.
Every one of those decisions feeds directly into whether a base model published on fAIr can be fine-tuned successfully by a mapping team working in a region it has never seen before.
Timeline
Sprint 1 (Weeks 1-2): Criteria & Landscape Review. Confirm suitability criteria and feature-type priorities with HOT, then review the starting datasets for Buildings, Solar Panels, and Trees while beginning the search for additional open VHR alternatives.
Sprint 2 (Weeks 3-4): Assessment & First Recommendation. Score every candidate dataset against resolution, annotation quality, and geographic diversity, verify licensing, and deliver a prioritised, licence-verified recommendation and data-preparation guidance per feature type.
Sprint 3 (Weeks 5-6): Practical Validation. Put the recommended datasets to work — checking structure, label quality, and real usability — and begin identifying where coverage or annotation falls short.
Sprint 4 (Weeks 7-8): Gap Filling & Alternatives. Test complementary or alternative datasets wherever the first choice doesn’t fully cover a region or use case, and update the recommendations with what validation reveals.
Sprint 5 (Weeks 9-10): Geographic Representation. Map which regions each recommended dataset does and doesn’t cover, and close the gaps that open alternatives allow.
Sprint 6 (Weeks 11-12): Documentation & Handoff. Compile the licensing, provenance, and limitation evidence each model’s documentation needs, and walk HOT through the final recommendations.
More details will be shared with the designated team.
First Omdena Project?
- Join the Omdena community to make a real-world impact and develop your career
- Build a global network and get mentoring support
- Earn money through paid gigs and access many more opportunities
Your Benefits
- Address a significant real-world problem with your skills
- Get hired at top companies by building your Omdena project portfolio (via certificates, references, etc.)
- Access paid projects, speaking gigs, and writing opportunities
Requirements
- Good English
- A very good grasp in computer science and/or mathematics
- Understanding of Machine Learning, Web Scraping and/or GIS Analysis
