Role
First author
Venue
ICECCME 2025
Status
Published
Context
Université Sorbonne Paris Nord

The problem

Object detection in aerial imagery is hard for two compounding reasons: the objects in a single scene span a very wide range of scales, and the images themselves vary a lot in quality.

A detector trained on that mixture spends capacity compensating for low-quality inputs instead of learning scale-invariant features.

The approach

  1. SRGAN as a dataset-level intervention

    Rather than changing the detector, a Super-Resolution Generative Adversarial Network generates higher-quality versions of the weakest images in the dataset, which then replace the originals during training and evaluation.

  2. Ten YOLO architectures, three scenarios

    The benchmark spans ten YOLO architectures across three evaluation scenarios, so the claim is about the strategy rather than about one lucky model.

  3. Two input resolutions

    Everything is measured at 416×416 and 640×640 pixels, since input resolution interacts directly with the small-object problem the method is trying to solve.

  4. DOTA v1.5

    A standard aerial detection benchmark with the scale heterogeneity the method targets.

Results

DOTA v1.5.

YOLOv5s-transformer Best configuration
SRGAN ×2 Upscaling factor
10 Architectures compared
416² and 640² Resolutions

The SRGAN-upscaled training strategy improves detection across the model family, not only for the best configuration.

Publication & code

Publication

da Rocha, W. F., Azzag, H., Lebbah, M., Mokraoui, A. Benchmarking SRGAN-Upscaled YOLO for Enhanced Object Detection in Aerial Imagery. ICECCME 2025, pp. 1–8.

Related work

Other parts of the same research line.