AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Gengyuan, Hannan, Tanveer, Kleiner, Hermine, Aydemir, Beste, Xie, Xinyu, Lan, Jian, Seidl, Thomas, Tresp, Volker, Gu, Jindong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unveiling the "Fairness Seesaw": Discovering and Mitigating Gender and Race Bias in Vision-Language Models
von: Lan, Jian, et al.
Veröffentlicht: (2025)
von: Lan, Jian, et al.
Veröffentlicht: (2025)
Multi-event Video-Text Retrieval
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2023)
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2023)
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos
von: Hannan, Tanveer, et al.
Veröffentlicht: (2024)
von: Hannan, Tanveer, et al.
Veröffentlicht: (2024)
Can Vision-Language Models be a Good Guesser? Exploring VLMs for Times and Location Reasoning
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2023)
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2023)
Memory Helps, but Confabulation Misleads: Understanding Streaming Events in Videos with MLLMs
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2025)
Localizing Events in Videos with Multimodal Queries
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2024)
Multimodal Pragmatic Jailbreak on Text-to-image Models
von: Liu, Tong, et al.
Veröffentlicht: (2024)
von: Liu, Tong, et al.
Veröffentlicht: (2024)
FedPop: Federated Population-based Hyperparameter Tuning
von: Chen, Haokun, et al.
Veröffentlicht: (2023)
von: Chen, Haokun, et al.
Veröffentlicht: (2023)
DocSLM: A Small Vision-Language Model for Long Multimodal Document Understanding
von: Hannan, Tanveer, et al.
Veröffentlicht: (2025)
von: Hannan, Tanveer, et al.
Veröffentlicht: (2025)
RGNet: A Unified Clip Retrieval and Grounding Network for Long Videos
von: Hannan, Tanveer, et al.
Veröffentlicht: (2023)
von: Hannan, Tanveer, et al.
Veröffentlicht: (2023)
ReEXplore: Improving MLLMs for Embodied Exploration with Contextualized Retrospective Experience Replay
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2025)
Self-Discovering Interpretable Diffusion Latent Directions for Responsible Text-to-Image Generation
von: Li, Hang, et al.
Veröffentlicht: (2023)
von: Li, Hang, et al.
Veröffentlicht: (2023)
FocalPO: Enhancing Preference Optimizing by Focusing on Correct Preference Rankings
von: Liu, Tong, et al.
Veröffentlicht: (2025)
von: Liu, Tong, et al.
Veröffentlicht: (2025)
SVAG-Bench: A Large-Scale Benchmark for Multi-Instance Spatio-temporal Video Action Grounding
von: Hannan, Tanveer, et al.
Veröffentlicht: (2025)
von: Hannan, Tanveer, et al.
Veröffentlicht: (2025)
Provably Better Explanations with Optimized Aggregation of Feature Attributions
von: Decker, Thomas, et al.
Veröffentlicht: (2024)
von: Decker, Thomas, et al.
Veröffentlicht: (2024)
FedBiP: Heterogeneous One-Shot Federated Learning with Personalized Latent Diffusion Models
von: Chen, Haokun, et al.
Veröffentlicht: (2024)
von: Chen, Haokun, et al.
Veröffentlicht: (2024)
Visual Question Decomposition on Multimodal Large Language Models
von: Zhang, Haowei, et al.
Veröffentlicht: (2024)
von: Zhang, Haowei, et al.
Veröffentlicht: (2024)
True Multimodal In-Context Learning Needs Attention to the Visual Context
von: Chen, Shuo, et al.
Veröffentlicht: (2025)
von: Chen, Shuo, et al.
Veröffentlicht: (2025)
Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?
von: Chen, Shuo, et al.
Veröffentlicht: (2023)
von: Chen, Shuo, et al.
Veröffentlicht: (2023)
Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries
von: Amoroso, Roberto, et al.
Veröffentlicht: (2024)
von: Amoroso, Roberto, et al.
Veröffentlicht: (2024)
AViT: Adapting Vision Transformers for Small Skin Lesion Segmentation Datasets
von: Du, Siyi, et al.
Veröffentlicht: (2023)
von: Du, Siyi, et al.
Veröffentlicht: (2023)
Stop Reasoning! When Multimodal LLM with Chain-of-Thought Reasoning Meets Adversarial Image
von: Wang, Zefeng, et al.
Veröffentlicht: (2024)
von: Wang, Zefeng, et al.
Veröffentlicht: (2024)
VideoINSTA: Zero-shot Long Video Understanding via Informative Spatial-Temporal Reasoning with LLMs
von: Liao, Ruotong, et al.
Veröffentlicht: (2024)
von: Liao, Ruotong, et al.
Veröffentlicht: (2024)
AUVIC: Adversarial Unlearning of Visual Concepts for Multi-modal Large Language Models
von: Chen, Haokun, et al.
Veröffentlicht: (2025)
von: Chen, Haokun, et al.
Veröffentlicht: (2025)
How the (Tensor-) Brain uses Embeddings and Embodiment to Encode Senses and Symbols
von: Tresp, Volker, et al.
Veröffentlicht: (2024)
von: Tresp, Volker, et al.
Veröffentlicht: (2024)
When and Where do Events Switch in Multi-Event Video Generation?
von: Liao, Ruotong, et al.
Veröffentlicht: (2025)
von: Liao, Ruotong, et al.
Veröffentlicht: (2025)
Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?
von: Chen, Shuo, et al.
Veröffentlicht: (2024)
von: Chen, Shuo, et al.
Veröffentlicht: (2024)
Distributionally Robust Optimization with Multimodal Decision-Dependent Ambiguity Sets
von: Yu, Xian, et al.
Veröffentlicht: (2024)
von: Yu, Xian, et al.
Veröffentlicht: (2024)
Does Machine Unlearning Truly Remove Knowledge?
von: Chen, Haokun, et al.
Veröffentlicht: (2025)
von: Chen, Haokun, et al.
Veröffentlicht: (2025)
A Comparative Study on How Data Normalization Affects Zero-Shot Generalization in Time Series Foundation Models
von: Ahmed, Ihab, et al.
Veröffentlicht: (2025)
von: Ahmed, Ihab, et al.
Veröffentlicht: (2025)
Improving Perturbation-based Explanations by Understanding the Role of Uncertainty Calibration
von: Decker, Thomas, et al.
Veröffentlicht: (2025)
von: Decker, Thomas, et al.
Veröffentlicht: (2025)
Why Uncertainty Calibration Matters for Reliable Perturbation-based Explanations
von: Decker, Thomas, et al.
Veröffentlicht: (2025)
von: Decker, Thomas, et al.
Veröffentlicht: (2025)
Context Matters: Leveraging Spatiotemporal Metadata for Semi-Supervised Learning on Remote Sensing Images
von: Bernhard, Maximilian, et al.
Veröffentlicht: (2024)
von: Bernhard, Maximilian, et al.
Veröffentlicht: (2024)
Bag of Tricks for Subverting Reasoning-based Safety Guardrails
von: Chen, Shuo, et al.
Veröffentlicht: (2025)
von: Chen, Shuo, et al.
Veröffentlicht: (2025)
WebArbiter: A Principle-Guided Reasoning Process Reward Model for Web Agents
von: Zhang, Yao, et al.
Veröffentlicht: (2026)
von: Zhang, Yao, et al.
Veröffentlicht: (2026)
Agentic Neural Networks: Self-Evolving Multi-Agent Systems via Textual Backpropagation
von: Ma, Xiaowen, et al.
Veröffentlicht: (2025)
von: Ma, Xiaowen, et al.
Veröffentlicht: (2025)
ClarAVy: A Tool for Scalable and Accurate Malware Family Labeling
von: Joyce, Robert J., et al.
Veröffentlicht: (2025)
von: Joyce, Robert J., et al.
Veröffentlicht: (2025)
PrAViC: Probabilistic Adaptation Framework for Real-Time Video Classification
von: Trędowicz, Magdalena, et al.
Veröffentlicht: (2024)
von: Trędowicz, Magdalena, et al.
Veröffentlicht: (2024)
TexAVi: Generating Stereoscopic VR Video Clips from Text Descriptions
von: Srihari, Vriksha, et al.
Veröffentlicht: (2025)
von: Srihari, Vriksha, et al.
Veröffentlicht: (2025)
AViTMP: A Tracking-Specific Transformer for Single-Branch Visual Tracking
von: Tang, Chuanming, et al.
Veröffentlicht: (2023)
von: Tang, Chuanming, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Unveiling the "Fairness Seesaw": Discovering and Mitigating Gender and Race Bias in Vision-Language Models
von: Lan, Jian, et al.
Veröffentlicht: (2025) -
Multi-event Video-Text Retrieval
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2023) -
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos
von: Hannan, Tanveer, et al.
Veröffentlicht: (2024) -
Can Vision-Language Models be a Good Guesser? Exploring VLMs for Times and Location Reasoning
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2023) -
Memory Helps, but Confabulation Misleads: Understanding Streaming Events in Videos with MLLMs
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2025)