← Back to Innovations
Geospatial & Earth Science Model Vision Experimental

FMV Grounding

Vision-Language Detection for Full-Motion Video

Try FMV Grounding on Microsoft Foundry → Try on Microsoft Foundry →
FMV Grounding

About FMV Grounding

FMV Grounding is a closed-set object detection model for full-motion video and overhead imagery that identifies six mission-critical categories — person, vehicle, bike, aircraft, ship, and maritime objects — across satellite, aerial, drone, and ground-based scenarios. Built on a vision-language transformer architecture, it pairs image understanding with language-guided object queries, and an optional text prompt can constrain detection to a specific subset of the supported classes — for example, only ship and maritime objects. Trained on roughly 3 million annotations and validated on 700,000, the model reaches 71.55% mAP and 83.43% average recall at IoU 0.5 on its internal test set.

Operational and geospatial intelligence workflows demand detectors tuned to a fixed, mission-relevant ontology rather than general-purpose recognition — objects outside the six categories are ignored by design, keeping outputs focused and auditable. On public benchmarks the model is especially strong on maritime and airborne targets, scoring 97.86% AP on Singapore Maritime ships and 98.76% on Mobdrone, with average inference latency near 200 ms on a single NVIDIA V100. Those characteristics suit maritime vessel monitoring, port and harbor analytics, drone-based perimeter surveillance, airfield detection, and border and coastal situational-awareness systems.

Key capabilities

  • Detects six land and maritime classes — person, vehicle, bike, aircraft, ship, maritime objects
  • Vision-language transformer pairs image understanding with language-guided object queries
  • Optional text prompts constrain detection to a chosen subset of supported classes
  • 71.55% mAP and 83.43% average recall at IoU 0.5 on internal test set
  • ~200 ms inference latency on a single NVIDIA V100 GPU

Ready to Explore?

Dive into platform integrations, source code, research papers, and announcements.