Publication
RG-DermNet: A Multimodal Attention-Based Model with Residual Block Usage for Skin Lesion Classification. International Joint Conference on Neural Networks (IJCNN), IEEE WCCI 2026.
Residual gated attention for multimodal skin lesion classification
Skin cancer accounts for nearly a third of all diagnosed tumours, and early recognition is what changes outcomes. Multimodal CAD systems — image plus clinical metadata — consistently beat image-only models, because a dermatologist does not diagnose from a photograph either.
The catch is that the metadata a real clinic collects is heterogeneous and frequently incomplete. Fusion mechanisms that treat metadata as a mandatory second input degrade badly when fields are missing, which is exactly the situation in the remote settings where CAD would help most.
A CNN or Transformer visual backbone (ResNet, Caformer-B36 and others) handles the lesion image. The metadata goes through a deliberately lightweight one-hot encoding pipeline — the point is that the expensive capacity should sit in the visual branch.
The fusion block gates the visual features with the metadata representation, and a residual connection carries the ungated visual signal forward. The gate can attenuate to nothing without the image information being lost, which is what makes the model degrade gracefully.
Splits are made per patient, not per image, so a lesion photographed several times cannot leak between train and test. The four dermatological datasets have heterogeneous metadata schemas, which is the point of the exercise.
Beyond the accuracy numbers, a SHAP-based analysis quantifies how much each clinical attribute contributed, giving the fusion block an interpretable account of what it used.
PAD-UFES-20, Caformer-B36 backbone, patient-wise cross-validation.
Outperforms the existing multimodal baselines under the same evaluation setting.
RG-DermNet: A Multimodal Attention-Based Model with Residual Block Usage for Skin Lesion Classification. International Joint Conference on Neural Networks (IJCNN), IEEE WCCI 2026.
The implementation is not public yet. If you would like to reproduce the experiments or discuss the setup, get in touch and I will share what I can.
Other parts of the same research line.