Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed
State-of-the-art toxicity detectors for text-to-image generation are found to fail marginalized communities: about 35% of images labeled safe are considered harmful by disability communities. This paper advocates for community-specific toxicity detection (CTD) and demonstrates its feasibility with safety guidelines for dwarfism and blind/low vision communities. Experiments reveal that general-purpose detectors achieve F1 scores lower than random (0.32, 0.37) in zero-shot settings, while prompt-based methods improve up to 0.78 (GPT-4o). However, CTD performance still lags far behind general-purpose detection (F1≈0.9), highlighting the need for sustained research.
-->
[Submitted on 27 Jul 2026]
Title:Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed
View a PDF of the paper titled Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed, by Xinnuo Xu and 8 other authors
View PDF HTML (experimental)
Abstract:State-of-the-art toxicity detectors for text-to-image generation adopt a one-size-fits-all approach: a single universal model applying fixed safety guidelines to all users. Our empirical evidence shows that these detectors fail to shield marginalized communities: approximately 35% of generated images labeled safe are considered harmful by disability communities. In this position paper, we argue for community-specific toxicity detection (CTD). To demonstrate its feasibility, we collaborate with disability experts to develop safety guidelines for two communities: dwarfism and blind/low vision. Using a dataset of 2,400 annotated T2I-generated images we demonstrate that both large vision-language models and existing general-purpose toxicity detectors catastrophically fail to recognize harmful content under these guidelines in zero-shot settings with F1 score lower than random guessing (F1 0.32 and 0.37). Promisingly, prompt-based adaptation methods (ICL, VQA) substantially improve harm detection performance (GPT-4o: F1 0.50 and 0.78), while parameter-efficient fine-tuning improves smaller models (0.5b-7b with best F1 0.48 and 0.59) with less than 100 demonstrations, but remains sensitive to evolving guidelines. Despite these gains, CTD performance remains far below F1 $\approx 0.9$ achieved for general-purpose toxicity detection, highlighting the challenge and the need for sustained research effort.
Comments: 18 pages, under review
Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2607.24898 [cs.CV]
(or arXiv:2607.24898v1 [cs.CV] for this version)
https://doi.org/10.48550/arXiv.2607.24898
arXiv-issued DOI via DataCite (pending registration)
Submission history
From: Xinnuo Xu [view email] [v1] Mon, 27 Jul 2026 16:42:36 UTC (20,546 KB)
Full-text links:
Access Paper:
View a PDF of the paper titled Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed, by Xinnuo Xu and 8 other authors
View PDF
HTML (experimental)
TeX Source
view license
Current browse context:
cs.CV
new | recent | 2026-07
Change to browse by:
cs cs.AI
References & Citations
NASA ADS
Google Scholar
Semantic Scholar
Loading...
Data provided by:
Bibliographic Tools
Bibliographic and Citation Tools
Bibliographic Explorer Toggle
Bibliographic Explorer (What is the Explorer?)
Connected Papers Toggle
Connected Papers (What is Connected Papers?)
Litmaps Toggle
Litmaps (What is Litmaps?)
scite.ai Toggle
scite Smart Citations (What are Smart Citations?)
Code, Data, Media
Code, Data and Media Associated with this Article
alphaXiv Toggle
alphaXiv (What is alphaXiv?)
Links to Code Toggle
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
DagsHub Toggle
DagsHub (What is DagsHub?)
GotitPub Toggle
Gotit.pub (What is GotitPub?)
Huggingface Toggle
Hugging Face (What is Huggingface?)
ScienceCast Toggle
ScienceCast (What is ScienceCast?)
Demos
Demos
Replicate Toggle
Replicate (What is Replicate?)
Spaces Toggle
Hugging Face Spaces (What is Spaces?)
Spaces Toggle
TXYZ.AI (What is TXYZ.AI?)
Related Papers
Recommenders and Search Tools
Link to Influence Flower
Influence Flower (What are Influence Flowers?)
Core recommender toggle
CORE Recommender (What is CORE?)
Author
Venue
Institution
Topic
About arXivLabs
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)