Akbar Karimi

Postdoctoral Researcher · Saarland University

LSV Lab — Spoken Language Systems

Akbar Karimi

I study AI and language models and their applications — from social media analysis to health and medical data, and from data augmentation to model robustness under adversarial perturbations. My goal is to find ways to improve AI models and their benefits to society.

Currently I am a postdoctoral researcher at Saarland University, working on chemical language models and trustworthy AI for drug discovery at the Spoken Language Systems (LSV) Lab led by Prof. Dietrich Klakow.

I completed my PhD at the University of Parma under Prof. Andrea Prati at IMP Lab, where I developed adversarial learning and data augmentation techniques for more robust language models.


News

Sep 2026 Attending German Conference in Bioinformatics (GCB) in Saarland Informatics Campus to present our poster: Structural Patterns in Drugs Guide Model Design for Glioblastoma.
Aug 2026 Attending Ellis Summer School for Trustworthy & Responsible AI in Drug Discovery in Saarland Informatics Campus.
Jun 2026 Joining LSV Lab at Saarland University to work on chemical/protein language models.
Apr 2026 Paper accepted to ACL 2026 Findings: More Agents Improve Math Problem Solving but Adversarial Robustness Gap Persists.
Mar 2026 Two papers accepted to DialRes Workshop @ LREC 2026: Can LLM Agents Identify Spoken Dialects like a Linguist? and Speaker Normalization via Voice Conversion Reveals a Human–Machine Dissociation in Dialect Classification.
Jan 2026 Paper accepted to WASSA Workshop @ EACL 2026 in Rabat, Morocco: Label-Consistent Data Generation for Aspect-Based Sentiment Analysis Using LLM Agents.
Dec 2025 Attending the Eurips Conference and presenting ArithmAttack and Multi-hop Reasoning with Hyperbolic Representations at the Ellis UnConference.
Sep 2025 Attending the ECMLPKDD Conference in Porto to present the results from the Colliding with Adversaries Challenge.
Aug 2025 Attending Interspeech Conference in Rotterdam to present our recent work on speech data augmentation for German dialects.
Aug 2025 Attending ACL Conference in Vienna to present ArithmAttack and Multi-hop Reasoning with Hyperbolic Representations.

Publications

2026
K Alavi, Z Yeltay, L Flek, A Karimi
Findings of ACL 2026
We evaluate multi-agent sampling-and-voting on adversarially perturbed math questions across six open-source models and four benchmarks, finding that more agents reliably improve accuracy while the robustness gap to noise, especially human-like typos, persists regardless of agent count.
PDF
2026
L Flek, O Janik, PA Jung, A Karimi, T Saala, A Schmidt, M Schott, P Soldin, M Thiesmeyer, C Wiebusch, U Willemsen
The European Physical Journal C
We present MiniFool, a physics-inspired adversarial attack that minimizes a χ²-based test statistic combined with a target-score deviation, testing the robustness of neural network classifiers on IceCube, CMS, and MNIST data, including unlabeled experimental data.
PDF
2026
T Bystrich, L Hamm, MH Akhter, L Fischbach, L Flek, A Karimi
DialRes Workshop @ LREC 2026
We explore whether LLM agents can classify Swiss German dialects from ASR-generated phonetic transcriptions combined with linguistic resources such as dialect feature maps, vowel history, and rules, comparing them against HuBERT, an LLM baseline, and a human linguist.
PDF
2026
C Kleen, L Fischbach, A Karimi, L Flek, A Lameli
DialRes Workshop @ LREC 2026
In perception experiments on nine German dialect regions, we show that voice conversion to a single target speaker leaves human dialect recognition unchanged while significantly improving a deep learning model, revealing a divergence between human and machine speech processing.
PDF
2026
P Bechtle, L Flek, PA Jung, A Karimi, T Saala, A Schmidt, M Schott, P Soldin, C Wiebusch, U Willemsen
arXiv preprint
We propose CONSERVAttack, an adversarial attack whose perturbations stay within simulation-versus-data uncertainty bounds, evading standard validation checks in high energy physics while fooling the model, and discuss strategies to mitigate such vulnerabilities.
PDF
2026
MHA Monfared, L Flek, A Karimi
WASSA Workshop @ EACL 2026
We propose an agentic data augmentation method for Aspect-Based Sentiment Analysis (ABSA) that uses iterative generation and verification to produce high-quality synthetic training examples.
PDF
2025
S Rawat, L Flek, A Karimi
WASP Workshop @ IJCNLP-AACL 2025
We build a multi-task SciBERT system for classifying telescope references, semantic attributes, and instrument mentions in astronomy papers, using stochastic segment sampling and majority voting, which significantly outperforms an open-weight GPT baseline.
PDF
2025
L Flek, PA Jung, A Karimi, T Saala, A Schmidt, M Schott, P Soldin, C Wiebusch
Computing and Software for Big Science
We present the Random Distribution Shuffle Attack (RDSA), which targets correlations between observables rather than individual features, and show that adversarial training with it improves classification on particle physics and five other tasks.
PDF
2025
T Saala, L Flek, A Karimi, PA Jung, A Schmidt, P Soldin, D Stefanopoulos, A Voskou, U Willemsen, C Wiebusch, M Schott
ECML PKDD 2025
We describe the Colliding with Adversaries challenge, with tasks on generating adversarial examples against a jet-classification model and on building models robust to unseen attacks, using simulated CMS collision data.
Paper
2025
L Fischbach, A Karimi, A Lameli, L Flek
RANLP 2025
We evaluate lightweight audio augmentation techniques on recordings from 20 German dialects, finding that frequency-based methods, especially frequency masking, consistently help while time masking or speaker-based insertion can hurt.
PDF
2025
L Fischbach, A Karimi, C Kleen, A Lameli, L Flek
Interspeech 2025
We use Retrieval-based Voice Conversion to map recordings to a single target speaker, reducing speaker variability so models focus on dialectal features, improving low-resource German dialect classification alone and combined with other augmentations.
PDF
2025
S Welz, L Flek, A Karimi
Findings of ACL
Through a simple integration of hyperbolic representations with an encoder-decoder model, we perform a controlled and comprehensive set of experiments to compare the capacity of hyperbolic space versus Euclidean space in multi-hop reasoning.
PDF
2025
WF Chen, Z Zhao, A Karimi, L Flek
Findings of ACL
We introduce HaluMap, a training-free, model-agnostic framework that detects hallucinations by mapping NLI entailment and contradiction relations between inputs and outputs, outperforming other training-free NLI-based methods by five points with interpretable explanations.
PDF
2025
Z Ul Abedin, S Qamar, L Flek, A Karimi
LLMSEC Workshop @ ACL
We propose ArithmAttack to examine how robust LLMs are when they encounter noisy prompts containing extra punctuation marks — an attack that causes no information loss yet consistently degrades performance across eight models.
PDF
2024
A Aliakbarzadeh, L Flek, A Karimi
Eighth Widening NLP Workshop (WiNLP 2024) Phase II
We study how real-world spelling mistakes from Wikipedia edit history affect 9 multilingual language models across NLI, NER, and intent classification in 6 languages, finding a 2.3 to 4.3 point performance gap, with mT5 models the most robust.
PDF
2024
S Nie, M Fromm, C Welch, R Görge, A Karimi, J Plepi, N Mowmita, N Flores-Herr, M Ali, L Flek
C3NLP Workshop @ ACL
We investigate the effect of multilingual training on bias mitigation by systematically training six LLMs of identical size (2.6B parameters): five monolingual and one multilingual model, showing that multilingual training consistently reduces bias while improving prediction accuracy.
PDF
2023
A Karimi, L Flek
SemEval 2023
We introduce a counterfactual data augmentation method based on verb replacement for identifying medical claims, yielding significant relative improvement on the minority class compared to three other augmentation techniques.
PDF
2022
A Karimi, L Flek
SMM4H 2022
We apply adversarial data augmentation in the input and embedding spaces to BioBERT for detecting disease mentions in Spanish tweets, outperforming a vocabulary-based baseline, with augmentation especially helpful in low-data settings.
PDF
2022
L Rossi, A Karimi, A Prati
JVCIR
We propose SBR-CNN, an evolution of HTC with loop mechanisms for box and mask refinement and an improved GRoIE, addressing IoU and feature-level imbalances and reaching 45.3% / 41.5% AP on COCO with a ResNet-50 backbone.
PDF
2022
L De Bruyne, A Karimi, O De Clercq, A Prati, V Hoste
LREC 2022
We present a multimodal dataset of 4,900 comments on 175 Instagram images annotated for aspect-based emotion analysis, and find that aspect and emotion classification benefit little from multimodal coreference resolution.
PDF
2022
L Rossi, A Karimi, A Prati
ICIAP 2022
We add a bounding box localization classification task to the Mean Teacher framework to better filter pseudo-labels, and show that box regression on unlabeled data helps as much as classification, improving SSOD on COCO by 1.14% AP.
PDF
2021
A Karimi, L Rossi, A Prati
ICNLSP 2021
We propose two simple modules, Parallel Aggregation and Hierarchical Aggregation, on top of BERT for Aspect Extraction and Aspect Sentiment Classification, improving performance without further training of the BERT model.
PDF
2021
A Karimi, L Rossi, A Prati
Findings of EMNLP
AEDA inserts punctuation marks randomly into text — simpler than EDA, lossless, and consistently superior across five classification datasets.
PDF
2021
L Rossi, A Karimi, A Prati
CAIP 2021
We propose R³-CNN, which replaces cascade architectures with a loop mechanism and recursive IoU-based re-sampling, surpassing the HTC model on COCO while significantly reducing the number of parameters.
PDF
2021
A Karimi, L Rossi, A Prati
SemEval 2021
We detect toxic spans by combining CharacterBERT, which handles misspelled toxic words through character-level embeddings, with a bag-of-words method that ensures frequently used toxic words are labeled.
PDF
2020
L Rossi, A Karimi, A Prati
ICPR 2020
We propose GRoIE, a Generic RoI Extractor that uses all FPN layers with non-local blocks and attention, integrating seamlessly into two-stage architectures and improving detection by up to 1.1% AP and instance segmentation by 1.7% AP.
PDF
2020
A Karimi, L Rossi, A Prati
ICPR
We propose BERT Adversarial Training (BAT), a novel architecture that applies adversarial training to Aspect Extraction and Aspect Sentiment Classification, outperforming both general and domain post-trained BERT — the first study of its kind in ABSA.
PDF
2018
A Karimi, E Ansari, BS Bigham
LREC 2018
We propose a bidirectional method to extract parallel sentences from document-aligned English and Persian Wikipedia, producing a corpus of about 200,000 sentences that improves statistical machine translation quality.
PDF

Teaching

2025
Dialog Systems
Undergraduate · University of Bonn
Lab sessions on LLM agents, agentic frameworks including SmolAgents, LangChain, and LlamaIndex, and building personal chatbots using small models with internet search and tool-use capabilities.
2024
Dialog Systems
Undergraduate · University of Bonn
Core components of a dialog system: ASR, NLU, dialog manager, dialog state tracking, and NLG, and the tasks involved in each module.
2023
Introduction to Natural Language Processing
Undergraduate · University of Marburg
Core NLP concepts including TF-IDF, word embeddings, RNNs, and Transformers, as well as evaluation methodologies and applications in conversational systems and computational social science.
2022
Dialog Systems
Undergraduate · University of Marburg

CV

Experience
Postdoctoral Researcher
2026 – Present · Saarland University
LSV Lab — Chemical/Protein Language Models
Postdoctoral Researcher & Scientific Coordinator
Oct 2023 – Dec 2025 · University of Bonn
CAISA Lab — Conversational AI and Social Analytics. Research on LLM robustness, adversarial NLP, and data augmentation; co-coordination of lab activities and student supervision.
Postdoctoral Researcher
2022 – 2023 · University of Marburg
Research on data augmentation and entity extraction for BioNLP; teaching Dialog Systems and Introduction to Natural Language Processing.
Education
Ph.D. in Information Technology
2022 · University of Parma
Thesis on adversarial learning and data augmentation for robust NLP models, under Prof. Andrea Prati at IMP Lab.
M.S. in Computer Science — Intelligent Systems
2017 · IASBS
B.S. in Computer Engineering — Software
2012 · Azad University of Zanjan