Skip to content
AI News HubLIVE
Original source2 min read

Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibratione

Summary

Safety-aligned large language models often suffer from over-refusal, incorrectly rejecting benign instructions that merely appear safety-related. This paper analyzes over-refusal through dynamic routing conflicts inside transformer attention, identifying a sparse subset of “Hypersensitive Safety Heads” that misfire on Hard-Safe prompts and entangle harmless target entities with refusal semantics. To mitigate this, the authors propose Semantic Routing Calibration (SRC), a lightweight, training-free inference framework that localizes and suppresses these heads at inference time and uses dual-branch logits fusion as a safety regularizer during decoding. Experiments show reduced over-refusal while preserving intrinsic safety performance as much as feasible. The paper was accepted to the EMNLP…

SourcearXiv Computational LinguisticsAuthor: Zixuan Wang, Bingjie Zhang, He Zhao, Dandan Guo
Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibratione
Report an error

The correction channel is not available yet. You can copy the article reference below for later.

Correction instructions
Read article

[Submitted on 6 Sep 2026]

Title:Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibratione

View a PDF of the paper titled Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibratione, by Zixuan Wang and 2 other authors

View PDF HTML (experimental)

Abstract:Large language models (LLMs) aligned for safety often suffer from over-refusal, incorrectly rejecting benign yet safety-related instructions. Prior studies primarily attribute this to static representation overlap, largely overlooking the underlying dynamic mechanisms. In this paper, we present the mechanistic analysis of over-refusal through the lens of internal routing conflicts within transformer attention. We discover that a sparse subset of Hypersensitive Safety Heads misfires on Hard-Safe prompts, exhibiting abnormal attention entanglement that forcefully binds harmless target entities to refusal semantics. This triggers a severe, high-entropy routing conflict that deprives target entities of necessary attention. To counteract this, we propose Semantic Routing Calibration (SRC), a lightweight, training-free inference framework. SRC precisely localizes and dynamically suppresses these hypersensitive safety heads at the inference stage. Coupled with a dual-branch logits fusion that acts as a safety regularizer during subsequent decoding, SRC seamlessly restores trustworthy reasoning. Extensive experiments demonstrate that SRC alleviates over-refusal, with intrinsic safety performance preserved as much as feasible.

Comments: 33 pages, 13 figures, accepted to the EMNLP 2026 Main Conference

Subjects:

Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

Cite as: arXiv:2609.25049 [cs.CL]

(or arXiv:2609.25049v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2609.25049

arXiv-issued DOI via DataCite

Submission history

From: Zixuan Wang [view email] [v1] Sun, 6 Sep 2026 09:37:56 UTC (10,075 KB)

Full-text links:

Access Paper:

View a PDF of the paper titled Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibratione, by Zixuan Wang and 2 other authors

View PDF

HTML (experimental)

TeX Source

view license

Current browse context:

cs.CL

new | recent | 2026-09

Change to browse by:

cs cs.AI

References & Citations

NASA ADS

Google Scholar

Semantic Scholar

Loading...

Data provided by:

Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)

Related Papers

Recommenders and Search Tools

Link to Influence Flower

Influence Flower (What are Influence Flowers?)

Core recommender toggle

CORE Recommender (What is CORE?)

Author

Venue

Institution

Topic

About arXivLabs

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Key points and analysis

Article intelligence

EngineersAdvanced

Key points

  • Prior work often blames static representation overlap; this paper focuses on dynamic routing conflicts within transformer attention.
  • A sparse subset of hypersensitive safety heads misfires on Hard-Safe prompts, binding harmless entities to refusal semantics and causing high-entropy routing conflicts.
  • SRC is training-free: it localizes and dynamically suppresses hypersensitive safety heads, with dual-branch logits fusion acting as a safety regularizer.
  • Experiments report alleviated over-refusal with intrinsic safety preserved as much as feasible; accepted to EMNLP 2026 Main Conference.

Highlights and analysis are generated automatically and may contain errors. Check the original source.