AI News HubLIVE
Original source3 min read

A Systems Engineering Framework for Vision-Language-Enabled UAV Triage and Disaster Response

This paper proposes a systems engineering framework that embeds Vision Language Models (VLMs) as coordination agents in human-UAV disaster response loops, moving beyond decision support. Developed with Model-Based Systems Engineering (MBSE), the architecture integrates natural language interaction, mission-level task coordination, software-in-the-loop implementation, and Incident Command System-aligned communications. A preliminary human-factors study with seven participants found reduced perceived workload and high ratings for AI trust and communication clarity.

SourcearXiv RoboticsAuthor: Swapnil Saha, Bhuvan Rajanasiriyur Jagadeesha, Karishma Patnaik, Neelakshi Majumdar

-->

[Submitted on 30 Jul 2026]

Title:A Systems Engineering Framework for Vision-Language-Enabled UAV Triage and Disaster Response

View a PDF of the paper titled A Systems Engineering Framework for Vision-Language-Enabled UAV Triage and Disaster Response, by Swapnil Saha and 3 other authors

View PDF HTML (experimental)

Abstract:Recent advances in Vision Language Models (VLMs) have created new opportunities for disaster response, where responders must interpret large volumes of sensor data under time pressure. Current VLM applications include social media monitoring for situational awareness, generation of draft action plans, and translation of technical alerts into public-facing messages. While these efforts can accelerate information flow, they remain largely limited to decision-support roles. Such approaches can increase operator burden because humans must still translate outputs into coordinated actions across teams and robotic assets. This study explores the viability of embedding VLMs as coordination agents within the human-UAV loop. The proposed architecture integrates natural language interaction, mission-level task coordination, software-in-the-loop implementation, and communication aligned with the Incident Command System (ICS). Rather than functioning solely as advisory tools, VLMs facilitate communication between human operators, mission control logic, and UAV task execution. The framework was developed using a Model-Based Systems Engineering (MBSE) approach, with use case and block definition diagrams representing system roles, internal structure, and component interactions. Three key elements, the VLM Coordinator Agent, UAV Mission Control, and Task Allocator, were implemented within an integrated simulation and control environment. A preliminary human-factors evaluation with seven participants showed reduced perceived workload across mental demand, effort, and frustration, along with high ratings for AI trust and communication clarity. By integrating MBSE, software-in-the-loop testing, and human-factors evaluation, this work advances scalable human-autonomy teaming for high-stakes disaster response, with broader implications for aerospace autonomy and civil safety.

Comments: 10 pages, 8 figures. Author accepted manuscript of AIAA Paper 2026-4010, published in the AIAA AVIATION 2026 Forum

Subjects:

Robotics (cs.RO); Artificial Intelligence (cs.AI); Systems and Control (eess.SY)

Cite as: arXiv:2607.27597 [cs.RO]

(or arXiv:2607.27597v1 [cs.RO] for this version)

https://doi.org/10.48550/arXiv.2607.27597

arXiv-issued DOI via DataCite (pending registration)

Journal reference: AIAA AVIATION 2026 Forum, AIAA Paper 2026-4010, 2026

Related DOI:

https://doi.org/10.2514/6.2026-4010

DOI(s) linking to related resources

Submission history

From: Swapnil Saha [view email] [v1] Thu, 30 Jul 2026 02:35:39 UTC (4,090 KB)

Full-text links:

Access Paper:

View a PDF of the paper titled A Systems Engineering Framework for Vision-Language-Enabled UAV Triage and Disaster Response, by Swapnil Saha and 3 other authors

View PDF

HTML (experimental)

TeX Source

view license

Current browse context:

cs.RO

new | recent | 2026-07

Change to browse by:

cs cs.AI cs.SY eess eess.SY

References & Citations

NASA ADS

Google Scholar

Semantic Scholar

Loading...

Data provided by:

Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)

Related Papers

Recommenders and Search Tools

Link to Influence Flower

Influence Flower (What are Influence Flowers?)

Core recommender toggle

CORE Recommender (What is CORE?)

Author

Venue

Institution

Topic

About arXivLabs

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)