Peer-reviewed work
10 journal and conference papers, 4 books and book chapters, spanning multimodal learning, medical AI, computer vision and embedded systems. Every entry links to the publisher, and carries its abstract and BibTeX.
-
Conference paper Preprint 2026
LLM-Driven Neural Architecture Search for Multimodal Skin Lesion Classification under Deployment Constraints
The Conference on Graphics, Patterns and Images (SIBGRAPI) - 2026 Main Track
Abstract
Neural Architecture Search (NAS) can automate multimodal model design, but its computational cost remains a major limitation in medical imaging. In this work, we investigate Large Language Models (LLMs) as NAS controllers for multimodal skin lesion classification under deployment-oriented constraints. We define a search space comprising 55,296 candidate architectures and sample 500 configurations, evaluating the generated models through patient-level stratified cross-validation on the PAD-UFES-20 dataset. Final model selection is performed with A-TOPSIS to balance predictive performance and stability. Our results suggest that Chain-of-Thought (CoT) prompting improves search stability, enabling even small LLMs to identify competitive architectures under constrained search budgets. The selected architecture reduces model size by 52% relative to MetaBlock and by 77% relative to concatenation-based baselines, while achieving statistically comparable performance to Cross-Attention, although remaining below MetaBlock in balanced accuracy. Transfer experiments on ISIC-2019 and MILK10K, without repeating the search on either dataset, indicate that the selected architecture retains useful binary classification performance, whereas multiclass settings remain challenging under severe class imbalance.
BibTeX
This paper is not in the proceedings yet — the citation will be added once it is published.
-
Conference paper Preprint 2026
RG-DermNet: A Multimodal Attention-Based Model with Residual Block Usage for Skin Lesion Classification
International Joint Conference on Neural Networks (IJCNN), IEEE WCCI 2026
Abstract
Skin cancer accounts for nearly one-third of all diagnosed tumors worldwide, making early and accurate recognition critical for improving patient outcomes. In this work, we propose RG-DermNet, a multimodal deep learning framework that integrates skin lesion images with structured clinical metadata through a residual gated-attention (RG-ATT) fusion mechanism. The architecture combines CNN- and Transformer-based visual backbones with a lightweight one-hot encoding pipeline for metadata, enabling effective cross-modal interaction. The proposed model is evaluated using a patient-wise cross-validation protocol across four dermatological datasets with heterogeneous metadata. On PAD-UFES-20, using Caformer-B36 as the visual backbone, RG-DermNet achieves an accuracy of 0.75±0.05, balanced accuracy of 0.78±0.03, F1-score of 0.77±0.04, and AUC of 0.95±0.01, outperforming existing multimodal baselines under the same evaluation setting. In addition, a SHAP-based analysis provides insights into the contribution of clinical metadata to the model's predictions, supporting both performance gains and interpretability.
BibTeX
This paper is not in the proceedings yet — the citation will be added once it is published.
-
Conference paper 2026
The more the merrier: the use of verbose metadata description in the multimodal classification of skin lesions
Anais do XXVI Simpósio Brasileiro de Computação Aplicada à Saúde (SBCAS), pp. 121–132
Abstract
CAD systems for skin cancer often rely on structured metadata with limited semantic depth. This study investigates replacing numeric/categorical encodings with verbose, anamnesis-like descriptions generated by Large Language Models (LLMs) as semantic transducers to convert structured metadata into verbose anamnesis-like descriptions for multimodal skin lesion classification. Clinical attributes from PAD-UFES-20 and ISIC-2019 were converted into natural language, encoded via SBERT, and fused with CNN visual features through the MetaBlock-SE architecture. Patient-wise 5-fold cross-validation shows that LLM-based models achieve competitive or statistically superior Accuracy and AUC compared to traditional baselines. These findings indicate that verbose metadata descriptions constitute a flexible and semantically rich alternative for multimodal skin lesion classification, particularly in scenarios with limited or semantically sparse metadata.
BibTeX
@inproceedings{pereira2026more, author = {Ana T. R. S. Pereira and Wyctor F. da Rocha and Pedro H. Bouzon and André G. C. Pacheco and Luis A. Souza Jr.}, title = {The more the merrier: the use of verbose metadata description in the multimodal classification of skin lesions}, year = {2026}, booktitle = {Anais do XXVI Simpósio Brasileiro de Computação Aplicada à Saúde}, pages = {121--132}, publisher = {SBC}, address = {Porto Alegre, RS, Brasil}, doi = {10.5753/sbcas.2026.20400} } -
Journal paper 2026
PRISM: a clinically interpretable stepwise framework for multimodal skin cancer diagnosis
Scientific Reports, vol. 16
Abstract
Skin cancer is the most common type of cancer worldwide, placing a substantial burden on healthcare systems. While numerous computer-assisted diagnostic systems have been developed, most offer limited transparency, making it difficult for healthcare professionals to understand how individual pieces of clinical information influence the diagnostic outcome. In this work, we propose PRISM: Probabilistic Reasoning Interpretable Stepwise Model, an interpretable multimodal skin cancer classification framework that integrates image-based predictions with clinical data using a Bayesian framework. Our approach allows the model to be evaluated with incrementally available metadata, enabling clinicians to interpret how each clinical feature contributes to the diagnostic decision. To address the compounding overconfidence inherent in sequential Bayesian updating, we implement a Stepwise Calibration protocol that dynamically scales with the volume of available evidence, ensuring statistically reliable confidence estimates at every intermediate diagnostic step. We validated the framework with four competitive vision backbones on three datasets. Compared to state-of-the-art attention-based methods, our framework achieves competitive or superior performance, reaching a peak balanced accuracy of 77.2±3.6% on the PAD-UFES-20 dataset. Finally, qualitative case studies indicate that PRISM creates a reasoning process that is consistent with clinical logic, suggesting it is a viable alternative for clinical decision support.
BibTeX
@article{bouzon2026prism, author = {Pedro H. G. Bouzon and Wyctor F. da Rocha and Luis A. de Souza and André G. C. Pacheco}, title = {PRISM: a clinically interpretable stepwise framework for multimodal skin cancer diagnosis}, year = {2026}, journal = {Scientific Reports}, volume = {16}, publisher = {Nature Publishing Group}, doi = {10.1038/s41598-026-47756-4} } -
Journal paper 2026
LiwTERM-r: A Revised Lightweight Transformer-based Model for Multimodal Skin Lesion Detection Robust to Incomplete Input
Journal of the Brazilian Computer Society, vol. 32, no. 1, pp. 305–315
Abstract
As the most common type of cancer in the world, skin cancer accounts for approximately 30% of all diagnosed tumor-based lesions. Early diagnosis can reduce mortality and prevent disfiguring in different skin regions. With the application of machine learning techniques in recent years, especially deep learning, promising results in this task could be achieved, presenting studies demonstrating that the combination of patients' clinical anamneses and images of the injured lesion is essential for improving the correct classification of skin lesions. Despite that, meaningful use of anamneses with multiple collected images of the same skin lesion is mandatory, requiring further investigation. Thus, this project aims to contribute to developing multimodal machine learning-based models to solve the skin lesion classification problem by employing a lightweight transformer model that is robust to missing clinical information input. As a main hypothesis, models can be fed by multiple images from different sources as input along with clinical anamneses from the patient's historical evaluations, leading to a more factual and trustworthy diagnosis. Our model deals with the not-trivial task of combining images and clinical information concerning the skin lesions in a lightweight transformer architecture that does not demand high computation resources or even all the information from the anamneses but still presents competitive classification results.
BibTeX
@article{souza2026liwterm, author = {Luis A. de Souza Júnior and André G. C. Pacheco and Thiago Oliveira dos Santos and Wyctor Fogos da Rocha and Pedro H. Bouzon and Christoph Palm and João Paulo Papa}, title = {LiwTERM-r: A Revised Lightweight Transformer-based Model for Multimodal Skin Lesion Detection Robust to Incomplete Input}, year = {2026}, journal = {Journal of the Brazilian Computer Society}, volume = {32}, number = {1}, pages = {305--315}, doi = {10.5753/jbcs.2026.5871} } -
Conference paper 2025
Benchmarking SRGAN-Upscaled YOLO for Enhanced Object Detection in Aerial Imagery
2025 5th International Conference on Electrical, Computer, Communications and Mechatronics Engineering (ICECCME), pp. 1–8
Abstract
Object detection in aerial imagery is challenging because of the wide variability in image quality and in the size of the objects within each scene. In this work, we propose a training and evaluation strategy based on a Super-Resolution Generative Adversarial Network (SRGAN) that generates higher-quality versions of the lower-quality images available in the dataset and uses them in their place. The SRGAN-based strategy improves how well the model learns to detect objects across a broad range of scales, raising the performance of the YOLO model family. We benchmark ten YOLO architectures across three evaluation scenarios at input resolutions of 416×416 and 640×640 pixels on the DOTA v1.5 dataset. The experimental results show that YOLOv5s-transformer trained on SRGAN-upscaled images with a scale factor of 2 achieves the best detection performance.
BibTeX
@inproceedings{rocha2025srgan, author = {Wyctor Fogos da Rocha and Hanene Azzag and Mustapha Lebbah and Anissa Mokraoui}, title = {Benchmarking SRGAN-Upscaled YOLO for Enhanced Object Detection in Aerial Imagery}, year = {2025}, booktitle = {2025 5th International Conference on Electrical, Computer, Communications and Mechatronics Engineering (ICECCME)}, pages = {1--8}, organization = {IEEE}, doi = {10.1109/ICECCME64568.2025.11277691} } -
Journal paper 2025
MetaBlock-SE: A Method to Deal With Missing Metadata in Multimodal Skin Cancer Classification
IEEE Journal of Biomedical and Health Informatics, vol. 29, pp. 8855–8862
Abstract
Skin cancer is the most common type of cancer worldwide. Its high incidence rate is a challenge for healthcare systems. Computer-aided diagnosis (CAD) systems have emerged as a promising tool to support skin cancer detection, particularly in remote regions. Recent advances in deep learning have enabled the development of multimodal models that integrate patient metadata with imaging data to improve diagnostic accuracy. However, these models may struggle to maintain high performance when dealing with missing or incomplete metadata. In this work, we propose an extension of the Metadata Processing Block (MetaBlock) architecture to address missing metadata in multimodal skin cancer classification. Our approach integrates a sentence embedding algorithm into the original method to generate dense vectors that capture the semantic relationships between the metadata features. We evaluate the proposed method on two different datasets: the PAD-UFES-20 and a new extended version of this dataset that is seven times larger. The proposed method outperforms the original MetaBlock architecture in all experiment scenarios, and obtains a balanced accuracy of up to 70.2±2.8 on the PAD-UFES-20 and 68.2±1.0 on the extended version with the ResNet-50 backbone. These results represent a step towards more robust multimodal models for skin lesion classification in real-world clinical scenarios. Finally, the new method is simple to implement and computationally efficient in terms of inference time and size, which makes it suitable to embed in real-world CAD systems.
BibTeX
@article{bouzon2025metablock, author = {Pedro H. G. Bouzon and Wyctor F. da Rocha and Luis A. Souza and André G. C. Pacheco}, title = {MetaBlock-SE: A Method to Deal With Missing Metadata in Multimodal Skin Cancer Classification}, year = {2025}, journal = {IEEE Journal of Biomedical and Health Informatics}, volume = {29}, pages = {8855--8862}, publisher = {IEEE}, doi = {10.1109/JBHI.2025.3612837} } -
Book 2024
Análise de sistemas: introdução ao funcionamento de sistemas
Atena Editora
Abstract
Fundamentals of system analysis — from the definition of a system through to the role of modelling and documentation in software and industrial contexts.
BibTeX
@book{amorim2024analise, author = {Vitor dos Santos Amorim and Wyctor Fogos da Rocha}, title = {Análise de sistemas: introdução ao funcionamento de sistemas}, year = {2024}, publisher = {Atena Editora}, doi = {10.22533/at.ed.602242808} } -
Book 2024
Introdução ao planejamento e controle da produção: conceitos e ferramentas
Atena Editora
Abstract
Support material for the Production Planning and Control discipline of a technical course in Logistics, used as a guide for classroom activities.
BibTeX
@book{amorim2024pcp, author = {Vitor dos Santos Amorim and Wyctor Fogos da Rocha and Ciro Xavier Maretto}, title = {Introdução ao planejamento e controle da produção: conceitos e ferramentas}, year = {2024}, publisher = {Atena Editora}, doi = {10.22533/at.ed.664242808} } -
Conference paper 2023
Detection of Anomalies in Civil Structures Using Machine Learning Models: Benchmark Study
2023 International Conference on Machine Learning and Applications (ICMLA), pp. 1913–1919
Abstract
This research introduces a unique dataset that includes four distinct categories of anomalies: Exposure Rebard, Infiltration, Normal, and Concrete Crack. The primary objective of this study is to evaluate the quality and robustness of the proposed dataset through a comprehensive benchmark study involving various machine learning models. The models evaluated in this study include VGG16, VGG19, MobileNet, MobileNetV2, ResNet50v2, ResNet101v2, DenseNetl21, and InceptionV3.The performance of each model evaluation was based on four key metrics: precision, recall, f1-score, and accuracy. These benchmark results provide valuable insights into the performance of different models on the proposed dataset, thereby guiding researchers and practitioners in selecting suitable models for their specific anomaly detection tasks in civil construction.
BibTeX
@inproceedings{amorim2023anomalies, author = {Vitor dos Santos Amorim and Wyctor Fogos da Rocha and Flávio Tongo da Silva and Aalah Adam}, title = {Detection of Anomalies in Civil Structures Using Machine Learning Models: Benchmark Study}, year = {2023}, booktitle = {2023 International Conference on Machine Learning and Applications (ICMLA)}, pages = {1913--1919}, organization = {IEEE}, doi = {10.1109/ICMLA58977.2023.00290} } -
Journal paper 2023
A diagnostic room for lower limb amputee based on virtual reality and an intelligent space
Artificial Intelligence in Medicine, vol. 143, art. 102612
Abstract
This article proposes a virtual reality (VR) system for diagnosing and rehabilitating lower limb amputees. A virtual environment and an intelligent space are the basis of the proposed solution. The target audiences are physiotherapists and doctors, and the aim is to provide a VR-based system to allow visualization and analysis of gait parameters and conformity. The multi-camera system from the intelligent space acquires images from patients during gait, making it possible to generate tridimensional information for the VR-based system. Among the provided functionalities, the user can explore the virtual environment and manage several features, such as gait reproduction and parameters displayed, using a head-mounted display and hand controllers. Besides, the system presents an automatic classifier that can assist physiotherapists and doctors in assessing abnormalities from conventional human gait. We evaluate the system through two quantitative experiments: a performance evaluation of the automatic classifier, and a Likert scale questionnaire submitted to a group of physiotherapists. The results from the questionnaire showed that the virtual environment is suitable for helping track patients' rehabilitation. Also, the neural network-based classifier results are promising, averaging higher than 91% for all evaluation metrics.
BibTeX
@article{silva2023diagnostic, author = {Pablo P. e Silva and Wyctor F. da Rocha and Luiza E. V. N. Mazzoni and Rafhael M. de Andrade and Antônio Bento and Mariana Rampinelli and Douglas Almonfrey}, title = {A diagnostic room for lower limb amputee based on virtual reality and an intelligent space}, year = {2023}, journal = {Artificial Intelligence in Medicine}, volume = {143}, pages = {102612}, publisher = {Elsevier}, doi = {10.1016/j.artmed.2023.102612} } -
Book 2023
Gestão de estoques
Atena Editora
Abstract
An approach to inventory optimization, covering demand forecasting, replenishment policies and the trade-offs behind stock levels.
BibTeX
@book{amorim2023estoques, author = {Vitor dos Santos Amorim and Wyctor Fogos da Rocha}, title = {Gestão de estoques}, year = {2023}, publisher = {Atena Editora}, doi = {10.22533/at.ed.323232411} } -
Conference paper 2021
The use of Convolutional Neural Network and StarRGB technique for gait movements recognition in remote physiotherapy
2021 International Conference on Electrical, Computer, Communications and Mechatronics Engineering (ICECCME), pp. 1–6
Abstract
Using the StarRGB technique to condense the spatial-temporal information of a gait sequence into a single image, and a convolutional neural network to classify the movements of lower limb amputees in a remote physiotherapy setting.
BibTeX
@inproceedings{rocha2021starrgb, author = {Wyctor Fogos da Rocha and Pablo Pereira e Silva and Luiza Emília V. N. Mazzoni and Clebeson Canuto and Rafhael Milanezi de Andrade and Antônio Bento and Mariana Rampinelli and Douglas Almonfrey}, title = {The use of Convolutional Neural Network and StarRGB technique for gait movements recognition in remote physiotherapy}, year = {2021}, booktitle = {2021 International Conference on Electrical, Computer, Communications and Mechatronics Engineering (ICECCME)}, pages = {1--6}, organization = {IEEE}, doi = {10.1109/ICECCME52200.2021.9590936} } -
Book chapter 2021
Sistema automático de controle de iluminação pública com base em sensores de presença e Bluetooth
Engenharia Elétrica: Comunicação Integrada no Universo da Energia, Atena Editora, pp. 36–49
Abstract
An extended treatment of the presence-sensor and Bluetooth-based lighting control system, applied to public street lighting.
BibTeX
@incollection{rocha2021iluminacaopublica, author = {Wyctor Fogos da Rocha and Mário Mestria}, title = {Sistema automático de controle de iluminação pública com base em sensores de presença e Bluetooth}, year = {2021}, booktitle = {Engenharia Elétrica: Comunicação Integrada no Universo da Energia}, publisher = {Atena Editora}, pages = {36--49}, doi = {10.22533/at.ed.3732123024} } -
Undergraduate thesis 2020
O uso da inteligência artificial na indústria brasileira
Instituto Federal do Espírito Santo (Ifes)
Abstract
A literature review on the adoption of artificial intelligence in Brazilian industry, mapping application areas, reported gains and the barriers to wider deployment.
BibTeX
@mastersthesis{rocha2020iaindustria, author = {Wyctor Fogos da Rocha}, title = {O uso da inteligência artificial na indústria brasileira}, year = {2020}, school = {Instituto Federal do Espírito Santo}, type = {Literature review} } -
Journal paper 2019
Sistema automático de controle de iluminação baseado em sensores de presença e Bluetooth
Revista SODEBRAS, vol. 14, no. 163, pp. 32–38
Abstract
An embedded automation system for lighting control based on presence sensors and Bluetooth communication, designed to reduce energy consumption in intermittently occupied environments.
BibTeX
@article{rocha2019iluminacao, author = {Wyctor Fogos da Rocha and Mário Mestria}, title = {Sistema automático de controle de iluminação baseado em sensores de presença e Bluetooth}, year = {2019}, journal = {Revista SODEBRAS}, volume = {14}, number = {163}, pages = {32--38}, doi = {10.29367/issn.1809-3957.14.2019.163.32} }
Looking for the methods behind the papers?
The project pages break down the architectures, the pipelines and the results.