Skip to content
AI Innovation Project

GeoBench: Foundational Model Research & Benchmarking for Geospatial AI

Project Kickoff: October 1, 2026

GeoBench: Foundational Model Research & Benchmarking for Geospatial AI

Establishing what already exists before HOT’s engineering team builds anything new — surveying the open base and foundation models available for building, solar panel, and tree detection, registering the benchmarks the community treats as standard, and recording the scores leading approaches actually report, so any new model has a real baseline to beat. In this 12-week challenge, you’ll join a collaborative team of machine learning researchers, computer vision engineers, and geospatial specialists from around the world.

The problem

HOT is moving from notebook-and-weights deliverables to reusable base models on its fAIr platform — models that mapping teams fine-tune on their own OpenAerialMap (OAM) imagery and deploy directly. That shift raises the bar for model selection: a strong validation score alone isn’t enough if the architecture can’t be exported to ONNX, served in a minimal environment, turned into GeoJSON, or transferred to regions outside its training data.

Building without first surveying the field risks two expensive mistakes: rebuilding something the community has already published and released openly, or building something that scores well in isolation but can’t be fairly compared with anything else — and therefore can’t be judged.

Impact of the Problem:

Reported benchmark figures are easy to misread. Two papers covering the same feature type may use different data splits, metrics, imagery resolutions, or regions, so their headline numbers often aren’t comparable at all — a model that looks state-of-the-art on paper may simply have been measured more generously.

Random data splits compound the problem: for mapping tasks, a random chip shuffle leaks neighbouring tiles between training and validation, inflating reported scores in a way that doesn’t hold up once the model meets a genuinely new region — exactly the situation a fAIr base model has to handle.

Without a documented, source-traceable baseline, HOT has no honest reference point for judging whether a proposed model is genuinely state of the art, merely competitive, or quietly duplicating existing work.

The goals

The objective is to survey the open base and foundation models available for Buildings, Solar Panels, and Trees, register the benchmarks the community accepts as standard for each task, and record what leading approaches actually score — so HOT has a documented, source-traceable baseline and a fAIr-compatible architecture shortlist for each feature type before any new model gets built.

Project Goals:
  • Confirm task formulations and review criteria with HOT: Agree the task framing for each feature type — instance segmentation for Buildings, object detection for Solar Panels, segmentation and/or object detection for Trees — and the criteria a candidate architecture must meet.
  • Survey the open model landscape: Identify the open pretrained and foundation models already published for each feature type, with their source, licence, and reuse conditions.
  • Register community-accepted benchmarks: Document the benchmarks, metrics, and evaluation protocols the geospatial and computer vision communities treat as standard for each task.
  • Collect and trace reported scores: Record what leading approaches actually report on those benchmarks, with every figure traceable to its published source and the evaluation conditions — split, metric, imagery, region — stated alongside it.
  • Assess comparability and fAIr readiness: Work out which reported scores can fairly be compared with one another, and review each candidate for transferability, compute needs, ONNX export readiness, and inference/post-processing requirements.
  • Deliver the architecture shortlist and baseline: Produce a prioritised, reasoned shortlist and a documented baseline per feature type — the reference score a new model should be measured against.

A model that can’t be exported to ONNX, can’t be fairly compared to what’s already published, or was only ever validated on a split that leaks neighbouring tiles isn’t a foundation HOT can build on — this challenge makes sure the field is actually surveyed before that happens.

Timeline

Sprint 1 (Weeks 1-2): Task Scoping & Model Survey. Confirm task formulations and review criteria with HOT, then begin surveying the open pretrained and foundation models already published for Buildings, Solar Panels, and Trees.

Sprint 2 (Weeks 3-4): Benchmark Registration. Identify and document the benchmarks, metrics, and evaluation protocols the community treats as standard for each task.

Sprint 3 (Weeks 5-6): Score Collection. Collect the reported scores of leading approaches, tracing every figure to its published source and recording the evaluation conditions behind it.

Sprint 4 (Weeks 7-8): Comparability & Readiness Review. Assess which reported scores can fairly be compared, and review each candidate for transferability, compute needs, ONNX export readiness, and inference/post-processing requirements.

Sprint 5 (Weeks 9-10): Shortlist & Baseline. Finalise a prioritised architecture shortlist per feature type and define the baseline and evaluation protocol a new model should be measured against.

Sprint 6 (Weeks 11-12): Documentation & Handoff. Compile the methodology, evaluation, and limitation documentation, and walk HOT through the final findings.

More details will be shared with the designated team.

Your Benefits

  • Address a significant real-world problem with your skills
  • Get hired at top companies by building your Omdena project portfolio (via certificates, references, etc.)
  • Access paid projects, speaking gigs, and writing opportunities

Requirements

  • Good English
  • A very good grasp in computer science and/or mathematics
  • Understanding of Machine Learning, Web Scraping and/or GIS Analysis

This challenge is hosted with our friends at

Omdena collaboratorsVisit the Collaborator DashboardOpen dashboard