Role
First author
Venue
IJCNN 2026 · IEEE WCCI
Status
Preprint

The problem

Skin cancer accounts for nearly a third of all diagnosed tumours, and early recognition is what changes outcomes. Multimodal CAD systems — image plus clinical metadata — consistently beat image-only models, because a dermatologist does not diagnose from a photograph either.

The catch is that the metadata a real clinic collects is heterogeneous and frequently incomplete. Fusion mechanisms that treat metadata as a mandatory second input degrade badly when fields are missing, which is exactly the situation in the remote settings where CAD would help most.

Multimodal skin lesion classification pipeline A clinical image goes through a CNN or Transformer backbone while structured clinical metadata goes through a one-hot or sentence-embedding encoder. Both feature vectors meet in a gated attention fusion block, which produces a calibrated diagnosis together with a per-feature attribution explaining it. INPUT Clinical image Smartphone or dermatoscope Clinical metadata Age, region, itch, bleed, elevation… ENCODE Visual backbone CNN or Transformer (ResNet, Caformer…) Metadata encoder One-hot, or SBERT over LLM sentences FUSE Gated attention fusion RG-ATT · MetaBlock-SE Residual path keeps the image signal when the metadata is missing DECIDE Diagnosis Six lesion classes, calibrated confidence Explanation SHAP attribution per clinical feature residual path
The shape most of my work takes: two modalities, one fusion block, and an output that has to be both accurate and explainable. The papers differ in how that fusion block is built — gated attention in RG-DermNet, sentence embeddings in MetaBlock-SE, Bayesian updating in PRISM.

The approach

  1. Two encoders, deliberately asymmetric

    A CNN or Transformer visual backbone (ResNet, Caformer-B36 and others) handles the lesion image. The metadata goes through a deliberately lightweight one-hot encoding pipeline — the point is that the expensive capacity should sit in the visual branch.

  2. Residual gated attention (RG-ATT)

    The fusion block gates the visual features with the metadata representation, and a residual connection carries the ungated visual signal forward. The gate can attenuate to nothing without the image information being lost, which is what makes the model degrade gracefully.

  3. Patient-wise evaluation across four datasets

    Splits are made per patient, not per image, so a lesion photographed several times cannot leak between train and test. The four dermatological datasets have heterogeneous metadata schemas, which is the point of the exercise.

  4. SHAP analysis over clinical features

    Beyond the accuracy numbers, a SHAP-based analysis quantifies how much each clinical attribute contributed, giving the fusion block an interpretable account of what it used.

Results

PAD-UFES-20, Caformer-B36 backbone, patient-wise cross-validation.

0.75 ± 0.05 Accuracy
0.78 ± 0.03 Balanced accuracy
0.77 ± 0.04 F1-score
0.95 ± 0.01 AUC

Outperforms the existing multimodal baselines under the same evaluation setting.

Publication & code

Publication

RG-DermNet: A Multimodal Attention-Based Model with Residual Block Usage for Skin Lesion Classification. International Joint Conference on Neural Networks (IJCNN), IEEE WCCI 2026.

Code

The implementation is not public yet. If you would like to reproduce the experiments or discuss the setup, get in touch and I will share what I can.

Related work

Other parts of the same research line.