The problem

Two modalities, one decision

Skin cancer accounts for roughly a third of all diagnosed tumours. Computer-aided diagnosis can extend specialist reach into remote regions, but only if it survives the messiness of real clinical data: metadata that is incomplete, images captured on whatever device was available, and clinicians who need to understand why a model said what it said.

Multimodal skin lesion classification pipeline A clinical image goes through a CNN or Transformer backbone while structured clinical metadata goes through a one-hot or sentence-embedding encoder. Both feature vectors meet in a gated attention fusion block, which produces a calibrated diagnosis together with a per-feature attribution explaining it. INPUT Clinical image Smartphone or dermatoscope Clinical metadata Age, region, itch, bleed, elevation… ENCODE Visual backbone CNN or Transformer (ResNet, Caformer…) Metadata encoder One-hot, or SBERT over LLM sentences FUSE Gated attention fusion RG-ATT · MetaBlock-SE Residual path keeps the image signal when the metadata is missing DECIDE Diagnosis Six lesion classes, calibrated confidence Explanation SHAP attribution per clinical feature residual path
The shape most of my work takes: two modalities, one fusion block, and an output that has to be both accurate and explainable. The papers differ in how that fusion block is built — gated attention in RG-DermNet, sentence embeddings in MetaBlock-SE, Bayesian updating in PRISM.
Interests

Research interests

Multimodal learning

How images and structured clinical context should meet. Gated attention, residual fusion and cross-modal interaction — with the constraint that the image branch has to keep working when the other modality is absent.

Medical AI

Computer-aided diagnosis evaluated the way it will be used: patient-wise cross-validation, heterogeneous datasets, and metrics that survive class imbalance.

Computer vision

CNN and Transformer backbones for classification and detection, and the training strategies around them — including super-resolution as a way to handle objects that span a wide range of scales.

Large language models

LLMs as semantic transducers: turning sparse categorical metadata into verbose, anamnesis-like descriptions that a sentence encoder can represent far more richly than a one-hot vector.

Efficient architectures

Lightweight transformers and fusion blocks cheap enough in inference time and parameter count to be embedded in a real CAD system, not just a benchmark table.

Explainable AI

SHAP attribution over clinical features, stepwise Bayesian reasoning, and calibration protocols that keep confidence estimates trustworthy as evidence accumulates.

Current work

What I'm working on now

  1. Fusion that tolerates missing metadata

    Clinical records are incomplete in practice, and most multimodal models degrade sharply when a field is absent. I work on fusion blocks — residual gated attention, and sentence-embedding extensions of MetaBlock — that keep the visual pathway intact and treat metadata as evidence rather than as a requirement.

    RG-DermNet MetaBlock-SE LiwTERM-r

  2. LLMs as semantic transducers for clinical data

    Instead of encoding "age = 62, region = forearm, itch = yes" as a sparse vector, an LLM writes it out as an anamnesis-like description which SBERT then encodes. On PAD-UFES-20 and ISIC-2019 this is competitive with — and sometimes statistically better than — the traditional encoding, and it degrades more gracefully when the metadata schema is sparse.

    The more the merrier — SBCAS 2026

  3. Interpretability and calibration for clinical use

    A model that is confidently wrong is worse than one that abstains. This thread combines stepwise Bayesian updating — so a clinician can watch the diagnosis move as each piece of evidence arrives — with a calibration protocol that counteracts the overconfidence that sequential updating compounds.

    PRISM

  4. LLM-driven neural architecture search

    My MSc work at PPGI/Ufes: instead of hand-designing the fusion architecture, a local LLM acts as the search controller. It receives the search space and the history of what has already been evaluated, proposes the next configuration as JSON, and gets the resulting balanced accuracy back as the reward that shapes its next proposal. The same question is now being asked of vision-language models on PAD-UFES-20.

    LLM as NAS controller NAS for multimodal VLMs

  5. Training strategies for wide-scale-range detection

    Outside the medical domain, the same question shows up as scale: aerial scenes mix large and very small objects at uneven image quality. Using SRGAN to upscale the weakest images during training raises detection performance across the YOLO family.

    SRGAN-upscaled YOLO

Future work

Where I want to take this

Open directions, in rough order of how soon I expect to get to them.

  • Search that optimises for deployability, not only accuracy

    The LLM-driven search currently rewards balanced accuracy. The natural extension is an explicit efficiency budget in the reward — parameters, latency, memory — so the architectures it converges on are ones that can actually be deployed.

  • Multiple images per lesion, and longitudinal records

    A lesion is often photographed more than once, and a patient has a history. Models that consume a set of images together with previous evaluations should be more trustworthy than ones that see a single frame.

  • Prospective validation with clinicians in the loop

    Retrospective cross-validation is necessary but not sufficient. The natural next step is measuring whether the explanations actually change clinical decisions.

  • Edge deployment for remote care

    The clinical motivation for this work is reach. That argues for models that run on the device that took the photograph, which puts quantisation and distillation of multimodal fusion blocks squarely on the path.

Collaborations

Who I work with

This research is collaborative. The groups and co-authors below appear across the publications listed on this site.

Multimodal skin lesion group

André G. C. Pacheco · Luis A. de Souza Jr. · Thiago Oliveira dos Santos · Pedro H. G. Bouzon · Ana T. R. S. Pereira

The work around PAD-UFES-20 and its extensions: MetaBlock-SE, LiwTERM-r, PRISM, RG-DermNet and the LLM-metadata study.

Université Sorbonne Paris Nord

Hanene Azzag · Mustapha Lebbah · Anissa Mokraoui

Where I did my Master's research internship in France, and where the super-resolution work on aerial object detection came from.

International co-authors

Christoph Palm · João Paulo Papa

Co-authors on the lightweight transformer work for multimodal skin lesion detection published in the Journal of the Brazilian Computer Society.

VR & intelligent space group

Douglas Almonfrey · Mariana Rampinelli · Rafhael Milanezi de Andrade · Antônio Bento · Pablo Pereira e Silva · Luiza Mazzoni · Clebeson Canuto

The gait-recognition and virtual-reality rehabilitation research that started during my scientific initiation at Ifes.

Datasets I work with
  • PAD-UFES-20
  • PAD-UFES-20 extended
  • ISIC-2019
  • DOTA v1.5

Open to collaboration

Reviewing, co-authorship, PhD positions and research partnerships — get in touch.