A claim that an AI “beat human players” needs more detail than a final score. Which model? Which players? Could either side search the web? Did the images show famous landmarks or anonymous stretches of farmland?
For a satellite-location test, those choices define the task. The following is a proposed comparison you can reproduce, with a blank result sheet rather than claimed findings.
Choose the images before collecting guesses
Create a fixed set of images and keep the answer coordinates in a separate file. Give each image an identifier. Record how you selected the set, including any exclusions for clouds, missing tiles, or unreadable images.
Include different kinds of locations if you want to compare performance across them. Decide those categories in advance. A set made entirely of landmarks answers a narrower question than one that also includes rural landscapes.
Use imagery you have permission to share if you intend to publish the test. Remove place labels and location-bearing filenames from the material given to participants.
Give both sides the same information
Use the same image crop and resolution for every participant. Set the rules for time limits, zooming, map access, and outside searches before starting. Record any difference between the human and model interfaces.
For the AI, save the exact model identifier, test date, prompt, and available tools. Ask for one latitude-longitude pair in a fixed format. Decide in advance how to handle invalid responses and retries; quietly discarding failed attempts would distort the comparison.
For humans, record their relevant experience. A result from one beginner and one experienced player cannot establish an average for all geography players.
Keep a row for every guess
- Record the image identifier, participant, guessed coordinates, and answer coordinates.
- Calculate the distance error using the same method for every guess.
- Keep invalid answers and timeouts in the record, with a clear scoring rule.
- Save explanations separately from scores so a convincing explanation cannot compensate for a wrong location.
Report the misses as well as the average
Show median and mean distance error together. A few guesses on the wrong side of the planet can move the mean substantially. Publish the per-image results so readers can see whether a small number of rounds determined the outcome.
If you also use a game score, publish its formula. A scoring curve can compress large distance errors or reward very close guesses, changing how the comparison looks.
Keep the conclusion within the test
A model naming a landmark correctly does not reveal how it arrived at that answer. A wrong rural guess does not, by itself, prove that the model cannot reason. Describe the observed performance without inventing an explanation for its internal process.
The useful conclusion is specific: these participants, using these rules, produced these errors on this image set. Publish the materials and results before making a broader claim about AI or human geography skills.