ChartAttack studies the vulnerability of large language models to malicious prompting in chart generation tasks, showing how adversarial instructions can manipulate LLM-generated charts to misrepresent data.
@inproceedings{ortiz2026chartattack,title={ChartAttack: Testing the Vulnerability of LLMs to Malicious Prompting in Chart Generation},author={Ortiz-Barajas, Jesus-German and Tonglet, Jonathan and Gupta, Vivek and Gurevych, Iryna},booktitle={Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},year={2026},doi={10.48550/arXiv.2601.12983}}
2025
Statement-Tuning Enables Efficient Cross-lingual Generalization in Encoder-only Models
Ahmed Elshabrawy, Thanh-Nhi Nguyen, Yeeun Kang, and 8 more authors
In Findings of the Association for Computational Linguistics: ACL 2025, 2025
This paper shows that statement-tuning, a lightweight fine-tuning approach, enables efficient cross-lingual generalization in encoder-only models, achieving strong performance across languages without the compute costs of large decoder-only LLMs.
@inproceedings{elshabrawy2025statement,author={Elshabrawy, Ahmed and Nguyen, Thanh-Nhi and Kang, Yeeun and Feng, Li and Jain, Annant and Shaikh, Faadil Abdullah and Mansurov, Jonibek and Imam, Mohamed Fazli Mohamed and Ortiz-Barajas, Jesus-German and Chevi, Rendi and Aji, Alham Fikri},booktitle={Findings of the Association for Computational Linguistics: ACL 2025},year={2025},doi={10.48550/arXiv.2506.01592}}
CVQA is a culturally-diverse multilingual Visual Question Answering benchmark covering many languages and cultural contexts, designed to evaluate multimodal models on culturally grounded visual understanding.
@inproceedings{romero2024cvqa,title={CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark},author={Romero, David and Lyu, Chenyang and Wibowo, Haryo Akbarianto and Lynn, Teresa and Hamed, Injy and Kishore, Aditya Nanda and Mandal, Aishik and Dragonetti, Alina and Abzaliev, Artem and Tonja, Atnafu and Balcha, Bontu Fufa and Whitehouse, Chenxi and Salamea, Christian and Velasco, Dan John and Adelani, David Ifeoluwa and Meur, David and Villa-Cueva, Emilio and Koto, Fajri and Farooqui, Fauzan and Belcavello, Frederico and Batnasan, Ganzorig and Vallejo, Gisela and Caulfield, Grainne and Ivetta, Guido and Song, Haiyue and Ademtew, Henok Biadglign and Maina, Hern{\'a}n and Lovenia, Holy and Azime, Israel Abebe and Cruz, Jan Christian Blaise and Gala, Jay and Geng, Jiahui and Ortiz-Barajas, Jesus-German and Baek, Jinheon and Dunstan, Jocelyn and Alemany, Laura Alonso and Nagasinghe, Kumaranage Ravindu Yasas and Benotti, Luciana and D'Haro, Luis Fernando and Viridiano, Marcelo and Estecha-Garitagoitia, Marcos and Cabrera, Maximiliano and Rodr{\'\i}guez-Cantelar, Mario and Jouitteau, M{\'e}lanie and Mihaylov, Momchil and Imam, Mohamed Fazli Mohamed and Adilazuarda, Farid and Gochoo, Munkhjargal and Otgonbold, Munkh-Erdene and Etori, Naome and Niyomugisha, Olivier and Silva, Paula and Chitale, Pranjal A. and Dabre, Raj and Chevi, Rendi and Zhang, Ruochen and Diandaru, Ryandito and Cahyawijaya, Samuel and G{\'o}ngora, Santiago and Jeong, Soyeong and Purkayastha, Sukannya and Kuribayashi, Tatsuki and Jayakumar, Thanmay and Torrent, Tiago and Ehsan, Toqeer and Araujo, Vladimir and Kementchedjhieva, Yova and Burzo, Zara and Lim, Zheng and Yong, Zheng-Xin and Ignat, Oana and Nwatu, Joan and Mihalcea, Rada and Solorio, Thamar and Aji, Alham Fikri},booktitle={Advances in Neural Information Processing Systems (NeurIPS)},year={2024},doi={10.52202/079017-0366}}
HyperLoader: Integrating Hypernetwork-Based LoRA and Adapter Layers into Multi-Task Transformers for Sequence Labelling
Jesus-German Ortiz-Barajas, Helena Gómez-Adorno, and Thamar Solorio
HyperLoader combines hypernetwork-based LoRA and adapter layers into multi-task transformers for sequence labelling, allowing efficient parameter sharing across tasks while retaining task-specific adaptation.
@article{ortiz2024hyperloader,title={HyperLoader: Integrating Hypernetwork-Based LoRA and Adapter Layers into Multi-Task Transformers for Sequence Labelling},author={Ortiz-Barajas, Jesus-German and G{\'o}mez-Adorno, Helena and Solorio, Thamar},journal={arXiv preprint arXiv:2407.01411},year={2024},doi={10.48550/arXiv.2407.01411}}
MBZUAI-UNAM at SemEval-2024 Task 1: Sentence-CROBI, a Simple Cross-Bi-Encoder-Based Neural Network Architecture for Semantic Textual Relatedness
Jesus-German Ortiz-Barajas, Gemma Bel Enguix, and Helena Gómez-Adorno
In Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval-2024), 2024
This paper presents the MBZUAI-UNAM team’s submission to SemEval-2024 Task 1, applying the Sentence-CROBI cross-bi-encoder architecture to the Semantic Textual Relatedness task.
@inproceedings{ortiz2024mbzuai,author={Ortiz-Barajas, Jesus-German and Bel Enguix, Gemma and G{\'o}mez-Adorno, Helena},booktitle={Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval-2024)},year={2024},doi={10.18653/v1/2024.semeval-1.155}}
2023
Job offers classifier using neural networks and oversampling methods
Germán Ortiz, Gemma Bel Enguix, Helena Gómez-Adorno, and 2 more authors
In Recent Developments and the New Directions of Research, Foundations, and Applications: Selected Papers of the 8th World Conference on Soft Computing, February 03–05, 2022, Baku, Azerbaijan, Vol. I, 2023
Both policy and research benefit from a better understanding of individuals’ jobs. However, as large-scale administrative records are increasingly employed to represent labour market activity, new automatic methods to classify jobs will become necessary. We developed an automatic job offers classifier using a dataset collected from the largest job bank in Mexico known as Bumeran5. We applied machine learning algorithms such as Support Vector Machines, Naive-Bayes, Logistic Regression, Random Forest, and deep learning Long-Short Term Memory (LSTM). Using these algorithms, we trained multi-class models to classify job offers in one of the 23 classes (not uniformly distributed): Sales, Administration, Call Center, Technology, Trades, Human Resources, Logistics, Marketing, Health, Gastronomy, Financing, Secretary, Production, Engineering, Education, Design, Legal, Construction, Insurance, Communication, Management, Foreign Trade, and Mining. We used the SMOTE, Geometric-SMOTE, and ADASYN synthetic oversampling algorithms to handle imbalanced classes. The proposed convolutional neural network architecture achieved the best results when applied the Geometric-SMOTE algorithm.
@inproceedings{ortiz2023job,title={Job offers classifier using neural networks and oversampling methods},author={Ortiz, Germ{\'a}n and Enguix, Gemma Bel and G{\'o}mez-Adorno, Helena and Ameer, Iqra and Sidorov, Grigori},booktitle={Recent Developments and the New Directions of Research, Foundations, and Applications: Selected Papers of the 8th World Conference on Soft Computing, February 03--05, 2022, Baku, Azerbaijan, Vol. I},pages={235--248},year={2023},organization={Springer},}
2022
Overview of PAR-MEX at Iberlef 2022: Paraphrase Detection in Spanish Shared Task
Gemma Bel-Enguix, Gerardo Sierra, Helena Gómez-Adorno, and 3 more authors
Paraphrase detection is an important unresolved task in natural language processing; especially in the Spanish language. In order to address this issue, and contribute to the creation of high-performance paraphrase detection automated systems, we propose a shared task called PAR-MEX. For this task, we created a corpus, in Spanish, with topics in the domain of Mexican gastronomy. Afterwards, the participants in this task submitted their classification results on our corpus. In this paper, we explain the steps followed for the creation of the corpus, we summarize the results obtained by the various participants, and propose some conclusions regarding the paraphrase-detection task in Spanish.
@article{bel2022overview,title={Overview of PAR-MEX at Iberlef 2022: Paraphrase Detection in Spanish Shared Task},author={Bel-Enguix, Gemma and Sierra, Gerardo and G{\'o}mez-Adorno, Helena and Torres-Moreno, Juan-Manuel and Ortiz-Barajas, Jesus-German and V{\'a}squez, Juan},journal={Procesamiento del Lenguaje Natural},volume={69},pages={255--263},year={2022},}
Sentence-CROBI: A Simple Cross-Bi-Encoder-Based Neural Network Architecture for Paraphrase Identification
Jesus-German Ortiz-Barajas, Gemma Bel-Enguix, and Helena Gómez-Adorno
Since the rise of Transformer networks and large language models, cross-encoders have become the dominant architecture for various Natural Language Processing tasks. When dealing with sentence pairs, they can exploit the relationships between those pairs. On the other hand, bi-encoders can obtain a vector given a single sentence and are used in tasks such as textual similarity or information retrieval due to their low computational cost; however, their performance is inferior to that of cross-encoders. In this paper, we present Sentence-CROBI, an architecture that combines cross-encoders and bi-encoders to obtain a global representation of sentence pairs. We evaluated the proposed architecture in the paraphrase identification task using the Microsoft Research Paraphrase Corpus, the Quora Question Pairs dataset, and the PAWS-Wiki dataset. Our model obtains competitive results compared with the state-of-the-art by using model ensembles and a simple model configuration. These results demonstrate that a simple architecture that combines sentence pair and single-sentence representations without using complex pre-training or fine-tuning algorithms is a viable alternative for sentence pair tasks.
@article{ortiz2022sentence,author={Ortiz-Barajas, Jesus-German and Bel-Enguix, Gemma and G{\'o}mez-Adorno, Helena},journal={Mathematics},volume={10},number={19},pages={3578},year={2022},publisher={MDPI},doi={10.3390/math10193578}}
2020
Enhancing Job Searches in Mexico City with Language Technologies
Gerardo Sierra Martı́nez, Gemma Bel-Enguix, Helena Gómez-Adorno, and 7 more authors
In Proceedings of the 1st Workshop on Language Technologies for Government and Public Administration (LT4Gov), 2020
In this paper, we show the enhancing of the Demanded Skills Diagnosis (DiCoDe: Diagnostico de Competencias Demandadas), a system developed by Mexico City’s Ministry of Labor and Employment Promotion (STyFE: Secretaria de Trabajo y Fomento del Empleo de la Ciudad de Mexico) that seeks to reduce information asymmetries between job seekers and employers. The project uses webscraping techniques to retrieve job vacancies posted on private job portals on a daily basis and with the purpose of informing training and individual case management policies as well as labor market monitoring. For this purpose, a collaboration project between STyFE and the Language Engineering Group (GIL: Grupo de Ingenieria Linguistica) was established in order to enhance DiCoDe by applying NLP models and semantic analysis. By this collaboration, DiCoDe’s job vacancies system’s macro-structure and its geographic referencing at the city hall (municipality) level were improved. More specifically, dictionaries were created to identify demanded competencies, skills and abilities (CSA) and algorithms were developed for dynamic classifying of vacancies and identifying terms for searches on free text, in order to improve the results and processing time of queries.
@inproceedings{martinez2020enhancing,title={Enhancing Job Searches in Mexico City with Language Technologies},author={Mart{\'\i}nez, Gerardo Sierra and Bel-Enguix, Gemma and G{\'o}mez-Adorno, Helena and Torres-Moreno, Juan-Manuel and Hern{\'a}ndez-Garc{\'\i}a, Tonatiuh and Guadarrama-Olvera, Julio V and Ortiz-Barajas, Jes{\'u}s-Germ{\'a}n and Rojas, Angela Mar{\'\i}a and Damerau, Tomas and Mart{\'\i}nez, Soledad Arag{\'o}n},booktitle={Proceedings of the 1st Workshop on Language Technologies for Government and Public Administration (LT4Gov)},pages={15--21},year={2020},}
2019
Detection of Aggressive Tweets in Mexican Spanish Using Multiple Features with Parameter Optimization.
Germán Ortiz, Helena Gómez-Adorno, Jorge Reyes-Magaña, and 2 more authors
This paper explains our approach to Aggressiveness Identification in the MEX-A3T shared task, whose aim is the detection of aggressive tweets. The task proposes a binary classification for every tweet: aggressive and non-aggressive. We approached the problem using linguistically motivated features and several types of n-grams (words, characters, functional words, punctuation symbols, among others). We trained a Support Vector Machine using a combinatorial framework that optimizes the results of the classifier. Our best run achieved an F1-score of 0,4549, which is the 5th best among the twenty-six runs.
@inproceedings{ortiz2019detection,title={Detection of Aggressive Tweets in Mexican Spanish Using Multiple Features with Parameter Optimization.},author={Ortiz, Germ{\'a}n and G{\'o}mez-Adorno, Helena and Reyes-Maga{\~n}a, Jorge and Bel-Enguix, Gemma and Sierra, Gerardo},booktitle={IberLEF@ SEPLN},pages={520--525},year={2019},}