Geolocating an image solely from its visual content is a critical capability for multiple cybersecurity and OSINT applications, yet current multimodal Large Language Models (LLMs) are rarely evaluated on closed-world data (i.e., images that do not appear in any public dataset). This work presents a systematic benchmark of nine state-of-theart multimodal LLMs (including five versions of Gemini, two of LLaMA, Qwen 3, and Gemma 3) on a closed-world geolocation task, using four key prompting strategies: Zero-Shot, Chain-of-Thought (CoT), Zero-Shot CoT, and ReAct. Results reveal a complex relationship between prompting and model architecture. While standard CoT often degrades performance due to noisy output, the Zero-Shot CoT strategy achieves peak accuracy by enforcing an internalized reasoning process while suppressing textual output. This demonstrates that performance degradation in standard CoT stems from output generation rather than the reasoning mechanism itself. Topperforming models, Gemini 3 Pro and Gemini 2.5 Pro, achieved a peak of 66% exact-location accuracy. These findings underscore both the potential and current limitations of multimodal LLMs for real-world geolocation, highlighting their applicability in content verification tasks (e.g.,mis- and dis- information detection) as well as in Open Source Intelligence (OSINT) contexts.

Benchmarking multimodal LLMs on closed-world image geolocation / Di Gisi, M., Fenza, G., Gallo, M., Palomba, M.. - 4198:(2026). (ITASEC & SERICS 2026 - Joint National Conference on Cybersecurity 2026 Cagliari, Italy 09-13/02/2026).

Benchmarking multimodal LLMs on closed-world image geolocation

Di Gisi Maria;
2026

Abstract

Geolocating an image solely from its visual content is a critical capability for multiple cybersecurity and OSINT applications, yet current multimodal Large Language Models (LLMs) are rarely evaluated on closed-world data (i.e., images that do not appear in any public dataset). This work presents a systematic benchmark of nine state-of-theart multimodal LLMs (including five versions of Gemini, two of LLaMA, Qwen 3, and Gemma 3) on a closed-world geolocation task, using four key prompting strategies: Zero-Shot, Chain-of-Thought (CoT), Zero-Shot CoT, and ReAct. Results reveal a complex relationship between prompting and model architecture. While standard CoT often degrades performance due to noisy output, the Zero-Shot CoT strategy achieves peak accuracy by enforcing an internalized reasoning process while suppressing textual output. This demonstrates that performance degradation in standard CoT stems from output generation rather than the reasoning mechanism itself. Topperforming models, Gemini 3 Pro and Gemini 2.5 Pro, achieved a peak of 66% exact-location accuracy. These findings underscore both the potential and current limitations of multimodal LLMs for real-world geolocation, highlighting their applicability in content verification tasks (e.g.,mis- and dis- information detection) as well as in Open Source Intelligence (OSINT) contexts.
2026
Large Language Models (LLMs), Geolocation, OSINT
File in questo prodotto:
File Dimensione Formato  
paper26.pdf

Accesso aperto

Descrizione: Benchmarking multimodal LLMs on closed-world image geolocation
Tipologia: Versione Editoriale (PDF)
Licenza: Creative commons
Dimensione 2.95 MB
Formato Adobe PDF
2.95 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.11771/44379
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • OpenAlex ND
social impact