AI News HubLIVE
Original source2 min read

Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes

Internet memes intertwine visual cues, textual content, and cultural context, making them particularly challenging to interpret in scenarios where humor, sarcasm, and harmful intent coexist. Existing multimodal classifiers either overlook these interdependencies or provide only limited interpretability. This paper introduces MAR-12, a framework using Vision Language Models (VLMs) to interpret memes through twelve structured perspectives, a role-aware soft-gated attention mechanism, and a prototype-based classifier. It achieves 80.3% accuracy for humor detection and 75.9% for hate detection on PrideMM and Memotion datasets, outperforming state-of-the-art approaches, and produces coherent, persuasive explanations.

SourcearXiv AIAuthor: Shanhong Liu, Pai Chet Ng, De Wen Soh, Malika Meghjani, Konstantinos N. Plataniotis

-->

[Submitted on 16 Jul 2026]

Title:Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes

View a PDF of the paper titled Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes, by Shanhong Liu and 4 other authors

View PDF HTML (experimental)

Abstract:Internet memes intertwine visual cues, textual content, and cultural context, making them particularly challenging to interpret in scenarios where humor, sarcasm, and harmful intent coexist. These complexities highlight the need for explainable meme understanding systems that can provide reliable and structured reasoning to support both accurate classification and human interpretability. However, existing multimodal classifiers either overlook these interdependencies or provide only limited interpretability. In this paper, we introduce MAR-12, a novel framework that leverages Vision Language Models (VLMs) for meme detection and understanding in settings where humorous and hateful elements may coexist. The framework first interprets each meme through twelve structured perspectives derived from humor and hate theories. It then applies a role-aware soft-gated attention mechanism to learn how much each perspective should contribute, followed by a prototype-based classifier for the final prediction. Finally, explanations are synthesized using both perspective-specific reasoning and learned attention weights, ensuring transparent and context-grounded justifications. We evaluate MAR-12 on the PrideMM and Memotion datasets, where it achieves up to 80.3% accuracy for humor detection and 75.9% accuracy for hate detection, outperforming state-of-the-art approaches. Furthermore, both human and GPT-4-based evaluations confirm that MAR-12 produces coherent and persuasive explanations, particularly for memes in which humorous and harmful cues co-occur.

Comments: Accepted for Publication at AAAI-ICWSM 2027

Subjects:

Artificial Intelligence (cs.AI)

Cite as: arXiv:2607.15442 [cs.AI]

(or arXiv:2607.15442v1 [cs.AI] for this version)

https://doi.org/10.48550/arXiv.2607.15442

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Shanhong Liu [view email] [v1] Thu, 16 Jul 2026 20:27:31 UTC (8,225 KB)

Full-text links:

Access Paper:

View a PDF of the paper titled Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes, by Shanhong Liu and 4 other authors

View PDF

HTML (experimental)

TeX Source

view license

Current browse context:

cs.AI

new | recent | 2026-07

Change to browse by:

cs

References & Citations

NASA ADS

Google Scholar

Semantic Scholar

Loading...

Data provided by:

Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)

Related Papers

Recommenders and Search Tools

Link to Influence Flower

Influence Flower (What are Influence Flowers?)

Core recommender toggle

CORE Recommender (What is CORE?)

Author

Venue

Institution

Topic

About arXivLabs

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)