Multimodal learning
How images and structured clinical context should meet. Gated attention, residual fusion and cross-modal interaction — with the constraint that the image branch has to keep working when the other modality is absent.
A diagnosis is never made from a photograph alone. My research is about building models that combine what a lesion looks like with what the clinician already knows about the patient — and that stay honest about how confident they are.
Skin cancer accounts for roughly a third of all diagnosed tumours. Computer-aided diagnosis can extend specialist reach into remote regions, but only if it survives the messiness of real clinical data: metadata that is incomplete, images captured on whatever device was available, and clinicians who need to understand why a model said what it said.
How images and structured clinical context should meet. Gated attention, residual fusion and cross-modal interaction — with the constraint that the image branch has to keep working when the other modality is absent.
Computer-aided diagnosis evaluated the way it will be used: patient-wise cross-validation, heterogeneous datasets, and metrics that survive class imbalance.
CNN and Transformer backbones for classification and detection, and the training strategies around them — including super-resolution as a way to handle objects that span a wide range of scales.
LLMs as semantic transducers: turning sparse categorical metadata into verbose, anamnesis-like descriptions that a sentence encoder can represent far more richly than a one-hot vector.
Lightweight transformers and fusion blocks cheap enough in inference time and parameter count to be embedded in a real CAD system, not just a benchmark table.
SHAP attribution over clinical features, stepwise Bayesian reasoning, and calibration protocols that keep confidence estimates trustworthy as evidence accumulates.
Clinical records are incomplete in practice, and most multimodal models degrade sharply when a field is absent. I work on fusion blocks — residual gated attention, and sentence-embedding extensions of MetaBlock — that keep the visual pathway intact and treat metadata as evidence rather than as a requirement.
Instead of encoding "age = 62, region = forearm, itch = yes" as a sparse vector, an LLM writes it out as an anamnesis-like description which SBERT then encodes. On PAD-UFES-20 and ISIC-2019 this is competitive with — and sometimes statistically better than — the traditional encoding, and it degrades more gracefully when the metadata schema is sparse.
A model that is confidently wrong is worse than one that abstains. This thread combines stepwise Bayesian updating — so a clinician can watch the diagnosis move as each piece of evidence arrives — with a calibration protocol that counteracts the overconfidence that sequential updating compounds.
My MSc work at PPGI/Ufes: instead of hand-designing the fusion architecture, a local LLM acts as the search controller. It receives the search space and the history of what has already been evaluated, proposes the next configuration as JSON, and gets the resulting balanced accuracy back as the reward that shapes its next proposal. The same question is now being asked of vision-language models on PAD-UFES-20.
Outside the medical domain, the same question shows up as scale: aerial scenes mix large and very small objects at uneven image quality. Using SRGAN to upscale the weakest images during training raises detection performance across the YOLO family.
Open directions, in rough order of how soon I expect to get to them.
The LLM-driven search currently rewards balanced accuracy. The natural extension is an explicit efficiency budget in the reward — parameters, latency, memory — so the architectures it converges on are ones that can actually be deployed.
A lesion is often photographed more than once, and a patient has a history. Models that consume a set of images together with previous evaluations should be more trustworthy than ones that see a single frame.
Retrospective cross-validation is necessary but not sufficient. The natural next step is measuring whether the explanations actually change clinical decisions.
The clinical motivation for this work is reach. That argues for models that run on the device that took the photograph, which puts quantisation and distillation of multimodal fusion blocks squarely on the path.
This research is collaborative. The groups and co-authors below appear across the publications listed on this site.
André G. C. Pacheco · Luis A. de Souza Jr. · Thiago Oliveira dos Santos · Pedro H. G. Bouzon · Ana T. R. S. Pereira
The work around PAD-UFES-20 and its extensions: MetaBlock-SE, LiwTERM-r, PRISM, RG-DermNet and the LLM-metadata study.
Hanene Azzag · Mustapha Lebbah · Anissa Mokraoui
Where I did my Master's research internship in France, and where the super-resolution work on aerial object detection came from.
Christoph Palm · João Paulo Papa
Co-authors on the lightweight transformer work for multimodal skin lesion detection published in the Journal of the Brazilian Computer Society.
Douglas Almonfrey · Mariana Rampinelli · Rafhael Milanezi de Andrade · Antônio Bento · Pablo Pereira e Silva · Luiza Mazzoni · Clebeson Canuto
The gait-recognition and virtual-reality rehabilitation research that started during my scientific initiation at Ifes.
Reviewing, co-authorship, PhD positions and research partnerships — get in touch.