Chain-of-Anomaly Thoughts with Large Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Domingos, Pedro, Pereira, João, Lopes, Vasco, Neves, João, Semedo, David |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-ReS: Self-Reflection in Large Vision-Language Models for Long Video Understanding
by: Pereira, Joao, et al.
Published: (2025)
by: Pereira, Joao, et al.
Published: (2025)
FineVAU: A Novel Human-Aligned Benchmark for Fine-Grained Video Anomaly Understanding
by: Pereira, João, et al.
Published: (2026)
by: Pereira, João, et al.
Published: (2026)
Zero-Shot Action Recognition in Surveillance Videos
by: Pereira, Joao, et al.
Published: (2024)
by: Pereira, Joao, et al.
Published: (2024)
ReCCur: A Recursive Corner-Case Curation Framework for Robust Vision-Language Understanding in Open and Edge Scenarios
by: Wei, Yihan, et al.
Published: (2026)
by: Wei, Yihan, et al.
Published: (2026)
TRACE: A Self-Improving Framework for Robot Behavior Forecasting with Vision-Language Models
by: Puthumanaillam, Gokul, et al.
Published: (2025)
by: Puthumanaillam, Gokul, et al.
Published: (2025)
ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models
by: Yu, Chung-En Johnny, et al.
Published: (2025)
by: Yu, Chung-En Johnny, et al.
Published: (2025)
Learning Collective Dynamics of Multi-Agent Systems using Event-based Vision
by: Lee, Minah, et al.
Published: (2024)
by: Lee, Minah, et al.
Published: (2024)
Concept-RuleNet: Grounded Multi-Agent Neurosymbolic Reasoning in Vision Language Models
by: Sinha, Sanchit, et al.
Published: (2025)
by: Sinha, Sanchit, et al.
Published: (2025)
Hydra: An Agentic Reasoning Approach for Enhancing Adversarial Robustness and Mitigating Hallucinations in Vision-Language Models
by: Chung-En, et al.
Published: (2025)
by: Chung-En, et al.
Published: (2025)
MaCTG: Multi-Agent Collaborative Thought Graph for Automatic Programming
by: Zhao, Zixiao, et al.
Published: (2024)
by: Zhao, Zixiao, et al.
Published: (2024)
Enhancing Agentic Autonomous Scientific Discovery with Vision-Language Model Capabilities
by: Gandhi, Kahaan, et al.
Published: (2025)
by: Gandhi, Kahaan, et al.
Published: (2025)
Visual Sensor Pose Optimisation Using Visibility Models for Smart Cities
by: Arnold, Eduardo, et al.
Published: (2021)
by: Arnold, Eduardo, et al.
Published: (2021)
SPAgent: Adaptive Task Decomposition and Model Selection for General Video Generation and Editing
by: Tu, Rong-Cheng, et al.
Published: (2024)
by: Tu, Rong-Cheng, et al.
Published: (2024)
AdaptFly: Prompt-Guided Adaptation of Foundation Models for Low-Altitude UAV Networks
by: Chen, Jiao, et al.
Published: (2025)
by: Chen, Jiao, et al.
Published: (2025)
CollaMamba: Efficient Collaborative Perception with Cross-Agent Spatial-Temporal State Space Model
by: Li, Yang, et al.
Published: (2024)
by: Li, Yang, et al.
Published: (2024)
PreGSU-A Generalized Traffic Scene Understanding Model for Autonomous Driving based on Pre-trained Graph Attention Network
by: Wang, Yuning, et al.
Published: (2024)
by: Wang, Yuning, et al.
Published: (2024)
Metropolis-Hastings Captioning Game: Knowledge Fusion of Vision Language Models via Decentralized Bayesian Inference
by: Matsui, Yuta, et al.
Published: (2025)
by: Matsui, Yuta, et al.
Published: (2025)
Chain of Questions: Guiding Multimodal Curiosity in Language Models
by: Iji, Nima, et al.
Published: (2025)
by: Iji, Nima, et al.
Published: (2025)
Autonomous Computer Vision Development with Agentic AI
by: Kim, Jin, et al.
Published: (2025)
by: Kim, Jin, et al.
Published: (2025)
AIDE: Agentically Improve Visual Language Model with Domain Experts
by: Chiu, Ming-Chang, et al.
Published: (2025)
by: Chiu, Ming-Chang, et al.
Published: (2025)
V-Agent: An Interactive Video Search System Using Vision-Language Models
by: Park, SunYoung, et al.
Published: (2025)
by: Park, SunYoung, et al.
Published: (2025)
Visual Multi-Agent System: Mitigating Hallucination Snowballing via Visual Flow
by: Yu, Xinlei, et al.
Published: (2025)
by: Yu, Xinlei, et al.
Published: (2025)
V2X-DGPE: Addressing Domain Gaps and Pose Errors for Robust Collaborative 3D Object Detection
by: Wang, Sichao, et al.
Published: (2025)
by: Wang, Sichao, et al.
Published: (2025)
Fast2comm:Collaborative perception combined with prior knowledge
by: Zhang, Zhengbin, et al.
Published: (2025)
by: Zhang, Zhengbin, et al.
Published: (2025)
Hollywood Town: Long-Video Generation via Cross-Modal Multi-Agent Orchestration
by: Wei, Zheng, et al.
Published: (2025)
by: Wei, Zheng, et al.
Published: (2025)
AniMaker: Multi-Agent Animated Storytelling with MCTS-Driven Clip Generation
by: Shi, Haoyuan, et al.
Published: (2025)
by: Shi, Haoyuan, et al.
Published: (2025)
Enhancing CLIP Robustness via Cross-Modality Alignment
by: Zhu, Xingyu, et al.
Published: (2025)
by: Zhu, Xingyu, et al.
Published: (2025)
Med-GRIM: Enhanced Zero-Shot Medical VQA using prompt-embedded Multimodal Graph RAG
by: Madavan, Rakesh Raj, et al.
Published: (2025)
by: Madavan, Rakesh Raj, et al.
Published: (2025)
A Case Study of Counting the Number of Unique Users in Linear and Non-Linear Trails -- A Multi-Agent System Approach
by: Rahman, Tanvir
Published: (2025)
by: Rahman, Tanvir
Published: (2025)
TraF-Align: Trajectory-aware Feature Alignment for Asynchronous Multi-agent Perception
by: Song, Zhiying, et al.
Published: (2025)
by: Song, Zhiying, et al.
Published: (2025)
VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning
by: Chen, Boyu, et al.
Published: (2025)
by: Chen, Boyu, et al.
Published: (2025)
Allen: Rethinking MAS Design through Step-Level Policy Autonomy
by: Zhou, Qiangong, et al.
Published: (2025)
by: Zhou, Qiangong, et al.
Published: (2025)
Multi-Agent Amodal Completion: Direct Synthesis with Fine-Grained Semantic Guidance
by: Fan, Hongxing, et al.
Published: (2025)
by: Fan, Hongxing, et al.
Published: (2025)
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
by: Kugo, Noriyuki, et al.
Published: (2025)
by: Kugo, Noriyuki, et al.
Published: (2025)
Active Scout: Multi-Target Tracking Using Neural Radiance Fields in Dense Urban Environments
by: Hsu, Christopher D., et al.
Published: (2024)
by: Hsu, Christopher D., et al.
Published: (2024)
LogiStory: A Logic-Aware Framework for Multi-Image Story Visualization
by: Meng, Chutian, et al.
Published: (2026)
by: Meng, Chutian, et al.
Published: (2026)
Sentinel: Embodied Cooperative Spatial Reasoning and Planning
by: Lin, Xiangye, et al.
Published: (2026)
by: Lin, Xiangye, et al.
Published: (2026)
Cascading multi-agent anomaly detection in surveillance systems via vision-language models and embedding-based classification
by: Rehman, Tayyab, et al.
Published: (2026)
by: Rehman, Tayyab, et al.
Published: (2026)
FetalAgents: A Multi-Agent System for Fetal Ultrasound Image and Video Analysis
by: Hu, Xiaotian, et al.
Published: (2026)
by: Hu, Xiaotian, et al.
Published: (2026)
FootBots: A Transformer-based Architecture for Motion Prediction in Soccer
by: Capellera, Guillem, et al.
Published: (2024)
by: Capellera, Guillem, et al.
Published: (2024)
Similar Items
-
Self-ReS: Self-Reflection in Large Vision-Language Models for Long Video Understanding
by: Pereira, Joao, et al.
Published: (2025) -
FineVAU: A Novel Human-Aligned Benchmark for Fine-Grained Video Anomaly Understanding
by: Pereira, João, et al.
Published: (2026) -
Zero-Shot Action Recognition in Surveillance Videos
by: Pereira, Joao, et al.
Published: (2024) -
ReCCur: A Recursive Corner-Case Curation Framework for Robust Vision-Language Understanding in Open and Edge Scenarios
by: Wei, Yihan, et al.
Published: (2026) -
TRACE: A Self-Improving Framework for Robot Behavior Forecasting with Vision-Language Models
by: Puthumanaillam, Gokul, et al.
Published: (2025)