AI for multimodal and multisensor situational awareness
Project Overview
Research scope & objectives
Efficient Vision-Language Inference for Real-Time Situational Awareness
The project focused on developing a real-time situational awareness framework that enables efficient vision-language inference over video streams. The work addresses the computational and payload overhead of continuously processing visual data with cloud-based Vision-Language Models by selectively processing relevant scene changes and maintaining compact representations of active objects. The framework also supports semantic visualization for immersive XR environments, with evaluation focused on inference efficiency, latency, communication overhead, and quality of scene understanding.
Military Object Detection Dataset Development and Image Curation
This project focused on building and curating a dataset for military object detection. The work involved collecting images from multiple sources, organizing them into relevant object categories, and reviewing the data for consistency and quality. Additional images were collected to improve the representation of less common military objects and to include challenging conditions such as occlusion, camouflage, and cluttered backgrounds.