AI News HubLIVE
Original source2 min read

Enabling Fully Integer-Only Inference for Lightweight Detection Transformers

This paper introduces I-LW-DETR, the first fully integer-only lightweight DETR, addressing deployment challenges of Vision Transformer detectors on NPUs and microcontrollers. It achieves 3.6x model size reduction and over an order of magnitude computational cost reduction with moderate accuracy loss.

SourcearXiv Computer VisionAuthor: Thanh Cong Le, Michal Szczepanski, Martyna Poreba

-->

[Submitted on 27 Jul 2026]

Title:Enabling Fully Integer-Only Inference for Lightweight Detection Transformers

View a PDF of the paper titled Enabling Fully Integer-Only Inference for Lightweight Detection Transformers, by Thanh Cong Le and 1 other authors

View PDF HTML (experimental)

Abstract:Vision Transformer detectors now approach the accuracy of CNNs but remain difficult to deploy on NPUs and microcontrollers because key components, including deformable attention, feature fusion, and nonlinear activation functions, are not natively compatible with integer arithmetic. Existing quantized detectors either retain operators such as Softmax, GELU, and LayerNorm or focus on heavyweight backbones, leaving lightweight detection transformers without an end-to-end integer implementation. We address this gap with I-LW-DETR, the first fully integer-only lightweight DETR, in which every operation in the forward pass, including transformer nonlinearities, is executed in integer arithmetic. I-LW-DETR is built upon three key components: a scale-preserving split convolution that assigns independent activation scale to each branch of the multi-scale projector; SD-ShiftGELU, a sign-dependent GELU approximation that preserves element-wise behavior while avoiding the accuracy degradation; and a constrained Shiftmax that maintains stable Softmax normalization. Experimental results demonstrate that the proposed quantization pipeline consistently produces efficient fully integer-only models across different model scales. Across all model scales, the proposed pipeline incurs only a moderate accuracy degradation while reducing the model size by approximately $3.6\times$ and the computational cost by more than one order of magnitude.

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as: arXiv:2607.24981 [cs.CV]

(or arXiv:2607.24981v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2607.24981

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Martyna Poreba [view email] [v1] Mon, 27 Jul 2026 18:33:18 UTC (6,254 KB)

Full-text links:

Access Paper:

View a PDF of the paper titled Enabling Fully Integer-Only Inference for Lightweight Detection Transformers, by Thanh Cong Le and 1 other authors

View PDF

HTML (experimental)

TeX Source

view license

Current browse context:

cs.CV

new | recent | 2026-07

Change to browse by:

cs

References & Citations

NASA ADS

Google Scholar

Semantic Scholar

Loading...

Data provided by:

Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)

Related Papers

Recommenders and Search Tools

Link to Influence Flower

Influence Flower (What are Influence Flowers?)

Core recommender toggle

CORE Recommender (What is CORE?)

Author

Venue

Institution

Topic

About arXivLabs

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)