Multimodal AI
Vision-language models for meme understanding, visual question answering, and multimodal content analysis across multiple languages.
Research Assistant · QCRI
I'm a Research Assistant at the Qatar Computing Research Institute (QCRI), where I work on multilingual NLP, multimodal AI, content moderation, and media literacy. My research focuses on multimodal and vision-language approaches to meme and image understanding, and on developing explainable systems for detecting harmful content such as hate speech and propaganda.
I'm currently pursuing my MSc in Data Science and Engineering at Hamad Bin Khalifa University (HBKU), after earning a First Class Honours BSc in Computer Science and Mathematics from Liverpool John Moores University. Having lived in China for 15 years, I bring a global perspective to my research. My work has appeared at NAACL, EMNLP, WWW, and WISE.
Vision-language models for meme understanding, visual question answering, and multimodal content analysis across multiple languages.
Explainable detection of propaganda and hateful content in memes and text, using LLMs and chain-of-thought reasoning.
Building specialized LLMs and benchmarks for Arabic and other languages, bridging the gap in low-resource language technology.
AI systems for learning, media literacy, document understanding, and student analytics, including EduLens and CritiSense.
Papers grouped by Google Scholar year. Full list on my Google Scholar profile.
No publications match this filter yet.
ACL Demo 2026 Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics, System Demonstrations, pages 481-492.
A multilingual media-literacy app for prebunking misinformation through short interactive challenges, instant feedback, and modular lessons. The usability study reports 93 users, 83.9% overall satisfaction, 90.1% ease-of-use ratings, and 500+ active users over six months.
@inproceedings{alam-etal-2026-critisense,
title = {{C}riti{S}ense: Critical Digital Literacy and Resilience Against Misinformation},
author = {Alam, Firoj and Ahmad, Fatema and Shahroor, Ali Ezzat and Kmainasi, Mohamed Bayan and Sartori, Elisa and Da San Martino, Giovanni and Hasnat, Abul and Ali, Raian},
booktitle = {Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations)},
year = {2026},
pages = {481--492},
doi = {10.18653/v1/2026.acl-demo.48},
url = {https://aclanthology.org/2026.acl-demo.48/}
}
arXiv 2026 1 citation arXiv:2601.12539.
A unified multilingual, multitask VLM for meme understanding, built from 38 public meme datasets mapped into a shared taxonomy across 20 tasks and 9 languages.
WWW Companion 2026 Companion Proceedings of the ACM Web Conference 2026, pages 935-944.
An empirical study of thinking-based multimodal LLMs for hateful meme analysis, showing how reinforcement learning and rationale supervision improve classification and explanation quality.
arXiv 2026 arXiv:2606.15307.
A reinforcement-learning approach for explainable hateful and propagandistic meme detection, using chain-of-thought supervision, task-specific rewards, and GRPO across English and Arabic benchmarks.
EMNLP 2025 Findings 13 citations Explanation-enhanced propaganda detection in Arabic and English.
A multilingual Arabic-English dataset and model setup for propaganda detection with explanations, designed to make classification decisions more transparent and useful for media-literacy contexts.
@inproceedings{hasanain2025propxplain,
title = {PropXplain: Can LLMs Enable Explainable Propaganda Detection?},
author = {Hasanain, Maram and Hasan, Md. Arid and Kmainasi, Mohamed Bayan and Sartori, Elisa and Shahroor, Ali Ezzat and Da San Martino, Giovanni and Alam, Firoj},
booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2025},
year = {2025},
url = {https://aclanthology.org/2025.findings-emnlp.1296/}
}
EMNLP 2025 12 citations Explainable multimodal detection with the MemeXplain dataset.
An explanation-enhanced approach for Arabic propagandistic memes and English hateful memes, pairing label detection with natural-language rationales.
@inproceedings{kmainasi2025memeintel,
title = {MemeIntel: Explainable Detection of Propagandistic and Hateful Memes},
author = {Kmainasi, Mohamed Bayan and Hasnat, Abul and Hasan, Md. Arid and Shahroor, Ali Ezzat and Alam, Firoj},
booktitle = {Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing},
year = {2025},
url = {https://aclanthology.org/2025.emnlp-main.1539/}
}
arXiv 2025 6 citations arXiv:2501.09768.
An Arabic legal judgment prediction study using Saudi commercial court judgments, benchmarking open-source LLMs under zero-shot, one-shot, and LoRA fine-tuning settings.
arXiv 2025 1 citation arXiv:2510.06371.
A culturally grounded spoken visual QA framework and resource, also released as OASIS, covering image, text, speech, and speech+image settings for Arabic and English contexts.
WISE 2024 20 citations International Conference on Web Information Systems Engineering, pages 406-420.
A study of native, non-native, and mixed-language prompting strategies across 11 Arabic NLP tasks and 198 experiments, showing how prompt language affects LLM performance.
@inproceedings{kmainasi2024native,
title = {Native vs Non-Native Language Prompting: A Comparative Analysis},
author = {Kmainasi, Mohamed Bayan and Khan, Rakif and Shahroor, Ali Ezzat and Bendou, Boushra and Hasanain, Maram and Alam, Firoj},
booktitle = {Web Information Systems Engineering -- WISE 2024},
pages = {406--420},
year = {2024},
url = {https://doi.org/10.1007/978-981-96-0576-7_30}
}
NAACL 2025 Findings 17 citations Findings of ACL: NAACL 2025, pages 5627-5649.
A specialized multilingual LLM for analyzing news and social media content across Arabic, English, and Hindi, covering 18 tasks and 52 datasets.
@inproceedings{kmainasi2025llamalens,
title = {LlamaLens: Specialized Multilingual LLM for Analyzing News and Social Media Content},
author = {Kmainasi, Mohamed Bayan and Shahroor, Ali Ezzat and Hasanain, Maram and Laskar, Sahinur Rahman and Hassan, Naeemul and Alam, Firoj},
booktitle = {Findings of the Association for Computational Linguistics: NAACL 2025},
year = {2025},
url = {https://aclanthology.org/2025.findings-naacl.313/}
}
Co-developing EduLens, an AI-driven platform for HBKU enabling LLM-assisted assignment submissions and analytics. Conducting research on LLMs with emphasis on multilingual capabilities, with papers at EMNLP, NAACL, WWW, and WISE.
Achieved Second Place in both the QCRI Programming Contest 2024 and Project award. Contributed to research on multilingual LLMs and content moderation.
Mastered ML foundations including linear algebra, probability, and statistics. Built proficiency in data preprocessing, supervised/unsupervised learning, NLP, and deep neural networks with TensorFlow and Keras.
Assisted the front-end ASU larval commentary unit in controlling the audio system inside a CCR room during the FIFA World Cup 2022, ensuring flawless broadcast quality to millions of viewers worldwide.
Apps, datasets, models, and systems connected to my research interests in media literacy, multimodal AI, and multilingual NLP.
A multilingual media-literacy app for prebunking misinformation through short interactive challenges, instant feedback, and modular learning content. It supports nine languages and has reached 500+ active users.
A browser-based party quiz game with Jeopardy boards, Family Feud bonus rounds, host controls, presets, final recap, and share-card export.
An AI-driven platform for HBKU enabling LLM-assisted assignment submissions, document management, and student-model interaction analytics.
Specialized multilingual LLM with instruction datasets in Arabic, English, and Hindi for analyzing news and social media.
Explainable detection of propagandistic and hateful memes, released with the MemeXplain dataset.
Explainable propaganda detection with LLMs, released with the PropXplain dataset.
The fastest way to reach me is by email. I'm open to research collaborations, especially around multimodal NLP, Arabic language technology, and AI for education.