Benchmarking the Alignment of Data-Quality Metrics, Human Judgment and Land-Cover Segmentation Performance for Earth Observation

TL;DR

This study systematically evaluates automatic quality metrics versus human perception and downstream land-cover segmentation performance for Earth observation synthetic data.

eess.IV 🔴 Advanced 2026-06-24 65 views
Ümit Mert Çağlar Alptekin Temizel
remote sensing data quality generative models evaluation metrics human perception

Key Findings

Methodology

The authors generate synthetic Earth observation images using diffusion and GAN models, applying perturbations such as rotation, noise, and combined transformations. They evaluate automatic metrics (FID, KID, LPIPS, SSIM, PSNR) under these perturbations, conduct large-scale human perception surveys assessing recognition, realism, and preferences, and test the impact of mixed real-synthetic datasets on land-cover segmentation models. Correlations and discrepancies among these evaluation sources are analyzed to reveal biases and limitations of current metrics in geospatial contexts.

Key Results

  • Automatic metrics like FID are highly sensitive to semantic-preserving transformations such as rotation, with scores fluctuating significantly despite human recognition remaining stable. Synthetic samples with poor automatic scores still achieve high perceptual realism and can improve downstream segmentation performance when combined with real data. Metrics based on ImageNet features show poor correlation with human perception and task utility in Earth observation datasets.
  • In land-cover segmentation experiments, incorporating synthetic data improved model F1 scores by 1-3%, yet automatic metrics did not reliably predict these improvements. Disturbance experiments demonstrated that low-frequency smoothing and structural memorization inflate metric scores, while minor geometric changes cause large score variations, exposing metric fragility.
  • Overall, the study highlights that current automatic quality metrics are insufficient for geospatial data evaluation, emphasizing the need for task-oriented, human-aligned assessment frameworks.

Significance

This research exposes critical shortcomings in prevalent automatic quality metrics for Earth observation synthetic data, underlining the importance of human perception and task performance in evaluation. It challenges the reliance on distributional distances like FID in geospatial domains, advocating for more holistic, application-specific assessment methods. The findings have significant implications for remote sensing data augmentation, model training, and operational deployment, guiding future development of robust, meaningful quality metrics aligned with real-world utility.

Technical Contribution

The paper introduces a comprehensive evaluation framework combining perturbation tests, human perception surveys, and downstream task analysis. It systematically uncovers biases in distribution-based metrics, especially those relying on ImageNet features, and demonstrates their poor correlation with human judgment and task performance in Earth observation contexts. This work pioneers the integration of multi-source evidence for synthetic data quality assessment, providing a foundation for developing more reliable, domain-specific metrics.

Novelty

This is the first systematic study to compare automatic metrics, human perception, and task utility specifically for Earth observation synthetic data. It reveals the mechanisms behind metric biases under semantic-preserving transformations and emphasizes the importance of task-aligned evaluation. The approach of combining perturbation experiments, large-scale human surveys, and downstream performance analysis offers a novel, comprehensive perspective that advances the field beyond traditional fidelity measures.

Limitations

  • The experiments focus primarily on land-cover segmentation, limiting generalization to other remote sensing applications such as change detection or multi-modal fusion. Further validation across diverse tasks is needed.
  • Synthetic data generation models are limited to certain architectures (e.g., Stable Diffusion, StyleGAN3), and results may vary with other models or training regimes.
  • Perturbation types are constrained, not covering all real-world variations like illumination changes or cloud occlusion, which may influence metric robustness.

Future Work

Future research should develop multi-modal, task-specific quality metrics that incorporate geospatial features and semantic consistency. Extending evaluations to other remote sensing tasks and data modalities will improve robustness. Additionally, integrating perceptual models trained explicitly on geospatial data could bridge the gap between automatic metrics and human judgment, fostering more reliable synthetic data assessment frameworks.

AI Executive Summary

Remote sensing plays a vital role in environmental monitoring, urban planning, and disaster management, but the high costs of data acquisition limit its widespread application. To address this, deep generative models like Diffusion and GANs have been employed to synthesize realistic satellite imagery, promising to augment training datasets and improve model robustness. However, evaluating the quality of these synthetic images remains a challenge. Traditional metrics such as FID, KID, and LPIPS, originally designed for natural images, often fail to capture the nuances of geospatial data, especially under transformations like rotation or scale changes.

This study systematically investigates the reliability of automatic quality metrics in the context of Earth observation data. By applying controlled perturbations—including rotation, noise, and combined transformations—the authors measure how metrics respond, revealing their high sensitivity to semantic-preserving changes that humans hardly notice. Large-scale human perception surveys further demonstrate that people can recognize and judge the realism of synthetic images accurately, even when automatic scores suggest poor quality.

The core finding is a stark misalignment: metrics rooted in ImageNet features, such as FID, do not reliably reflect human perception or downstream task performance. In land-cover segmentation experiments, synthetic data improved model accuracy despite low fidelity scores, emphasizing that visual fidelity alone is insufficient. These insights challenge the current paradigm of quality assessment, advocating for evaluation frameworks grounded in task utility and human judgment.

Overall, this research underscores the need for more domain-specific, human-aligned metrics in remote sensing. It paves the way for developing evaluation tools that better predict real-world performance, ultimately enhancing the deployment of synthetic data in critical geospatial applications. The findings have broad implications for the future of data augmentation, model training, and operational remote sensing, highlighting a shift toward more holistic, application-aware quality assessment methods.

Deep Dive

Plain Language Accessible to non-experts

想象你在厨房里准备一道菜,菜的好坏不仅仅看它的颜色或外形,还要尝一尝味道。传统的评价方法就像只看菜的颜色,觉得颜色鲜亮就算好菜,但其实味道才是关键。自动指标就像只看菜的外表,容易被一些小变化骗过,比如把菜旋转一下,评分就会变差,但味道其实没变。人类就像品尝厨师做的菜,能很快判断出菜是否好吃。研究发现,自动指标在面对旋转或噪声等变化时,表现得很不稳定,而人类的感知能力却很强。这就像你用鼻子和味蕾去判断菜的好坏,而自动指标只看外表,不能完全反映真实情况。更有趣的是,有些虚拟的菜(用电脑模拟出来的)虽然在评分上不如真正的菜,但看起来还挺像真的,还能帮厨师改进菜肴。这个比喻告诉我们,要用更聪明的方法,结合人类的感觉和实际用途,才能真正评价一份菜是不是好吃。

ELI14 Explained like you're 14

想象你在学校的食堂吃饭,老师用一种特别的评分方法来判断菜的好坏。可是,这个评分方法只看菜的颜色和形状,而不真正尝味道。有时候,菜看起来很漂亮,但其实味道很差。你会不会觉得这个评分不太公平?其实,人们(像你和我)吃菜时,主要是靠味觉和感觉,而不是只看外表。研究发现,这些自动评分方法(就像只看菜的外表)经常被一些小变化骗过,比如把菜旋转一下,评分就变得很差,但其实菜的味道没变。反而,我们人类能很快判断菜的好坏。更妙的是,有些虚拟的菜(用电脑做出来的)虽然在评分上不如真正的菜,但其实看起来还挺像真的,还能帮厨师做出更好的菜。这告诉我们,要用更聪明的方法,结合我们自己的感觉和实际用途,才能真正判断一份菜是不是好吃。

Abstract

Volume and quality of datasets are crucial for deep learning model training, yet they are often constrained by availability and data acquisition costs. Synthetic data augmentation can extend existing datasets with realistic images, and the quality of these images is generally assessed through fidelity metrics such as FID, KID, IS, LPIPS and SSIM that measure structural or distributional similarity. However, such metrics, including the widely used FID, focus on visual fidelity without reflecting downstream utility, and can diverge from human perception under perturbations that are imperceptible to human observers. In this work, we systematically evaluate Earth observation datasets alongside synthetic counterparts generated by deep generative models, comparing automatic metrics against human perception and downstream tasks. Our results reveal a stark misalignment: semantics-preserving perturbations such as rotation drastically alter metric scores while leaving human recognition unaffected, and synthetic samples that score poorly on automatic metrics achieve comparable or higher perceived realism, and can improve downstream performance when combined with real data. By benchmarking semantic segmentation models trained on mixed real-synthetic datasets, we demonstrate that quality metrics rooted in ImageNet-pretrained feature spaces are unreliable indicators for geospatial data. Our findings underscore that automatic quality evaluation of synthetic datasets should be grounded in downstream task performance and human evaluation.

eess.IV cs.AI cs.CV cs.LG